What an AI Consulting Company Actually Does (and When Not to Hire One)

Where an AI assessment earns its fee, how it is billed, and the cases where AI is the wrong answer

September 17, 2026
7 min read

An AI consulting company is not a model vendor and not a staffing agency. It is a short, bounded engagement that answers one question: where does AI change your economics, and what would it cost to prove that before you commit a build budget? A useful engagement ends with something you can act on — a ranked list of use cases, one measured test on your own data, and a costed plan for the parts worth building.

What an AI consulting company actually does

The work is assessment and de-risking, not model research. A consulting engagement should map the places where a model beats the rule you are running today, check whether your data can support them, and put a number on both the value and the running cost.

In practice that is five deliverables:

  • An opportunity map. Every candidate use case ranked by value per decision and by how much data it needs — not a list of technologies.
  • A data readiness read. What you already store, what would have to be labelled, and where the examples of a correct decision actually come from. Most AI projects die here, not at the model.
  • One measured proof. A narrow test on your own data with an agreed accuracy bar, evaluated the same way each run, so the result is a number rather than an impression.
  • A cost-to-run estimate. Inference per request, storage, the human review step, and monitoring — the recurring bill that never appears in a demo.
  • A build, buy, or stop call. Including the honest case that an existing product already solves it, or that the process should be fixed before any model is added.

Why the assessment has to run on your data

A benchmark score tells you almost nothing about your own workload. Public benchmarks measure a model against someone else's examples, and their definition of correct. Your catalogue, your ticket history, and your document formats are what the system will actually see.

That is why the first test should be a small evaluation set built from real cases you already handled — twenty to a few hundred of them, with the correct outcome written down. It is unglamorous work and it is the whole point: it tells you whether the idea clears your accuracy bar before anyone writes production code, and it becomes the regression test that protects the feature afterwards.

The three ways AI consulting is billed

ModelTypical shapeBest when
Fixed-fee assessmentOne price, two to six weeks, defined deliverablesYou need a decision and a budget number before committing further
Embedded sprintA day rate or monthly fee for a named team working inside your productA proof is already accepted and you are moving into the build
Advisory retainerA set number of hours per month, no delivery commitmentYou have engineers and need a second opinion on architecture and evaluation

Watch for the fourth model, which is not on any pricing page: an open-ended discovery phase billed by the hour with no defined end state. If the engagement has no deliverable list, the meter keeps running after the useful answer has been given.

When AI is the wrong answer

Some problems that arrive as AI projects are not AI problems, and the fastest win is saying so.

  • The decision is already written down as rules and those rules are followed. A model adds variance to a process that needed consistency.
  • You have no examples of a correct decision. Without them there is no evaluation set and no way to know whether the output improved.
  • Nobody owns the errors. A model that is right most of the time needs a person accountable for the rest, plus a path for them to correct it.
  • The accuracy bar is effectively total. If a single wrong answer is unacceptable — pricing a contract, releasing a payment — use AI to assist the human, never to decide.
  • The cost per request is larger than the value per request. High-volume, low-margin steps rarely clear that bar once review time is included.
  • The real problem is the workflow. If the data exists but no one can find it, search and process design beat a model.

What a credible assessment hands over

Ask for the artefact, not the meeting. At the end you should hold a written pack that includes the ranked use cases with a value estimate each, the data each one needs and who owns it, the accuracy bar you agreed, the measured result of the proof, a per-request running cost, a build-versus-buy recommendation, and the specific risks that would sink it. If any of those is missing, the engagement is not finished.

Red flags in an AI proposal

  • Accuracy figures quoted before anyone has seen your data.
  • No mention of an evaluation set, a test harness, or how the result will be measured.
  • No estimate of inference or review cost per request.
  • A model or vendor chosen in the proposal, before the problem was defined.
  • No failure plan: what the product does when confidence is low.
  • Scope written as hours rather than as deliverables.

What happens after the assessment

The useful outcome of an assessment is a build that starts with the risky part. If the proof held, the next step is production work: retrieval over your real documents, evaluation wired into the deploy, human review where the accuracy bar demands it, and monitoring that reports the error rate rather than the uptime. Custom app development covers that stage, and what an AI development company builds walks through the components in detail.

If the numbers did not hold, you have spent weeks instead of quarters, and you know which use case to revisit when the data or the models change.

Frequently asked questions

What does an AI consulting company deliver?

A ranked set of use cases with value estimates, a data readiness read, one measured test on your own examples against an agreed accuracy bar, a cost-per-request estimate including human review, and a build, buy, or stop recommendation. The deliverable is a decision plus the evidence for it, not a strategy deck.

How long does an AI consulting engagement take?

A scoped assessment usually runs two to six weeks. The part that sets the schedule is building the evaluation set from real cases, because that work depends on your data and your people, not on the model. Anything longer than that without a written deliverable is scope drifting.

How much does AI consulting cost?

Fixed-fee assessments are priced by scope — number of use cases, how much labelling is needed, and how much integration surface the proof touches. The honest answer is that the price is set by data work, not by model work, which is why a vendor who quotes before looking at your data is guessing.

Can we just build without an assessment?

You can, and it is the most expensive way to learn. The failure pattern is a demo built on clean sample data that cannot be reproduced on production input, discovered after the integration budget is spent. A two-week evaluation set and a held-out test removes most of that risk.

Should we hire AI consultants or use an existing product?

If a product you can buy already solves the problem at an acceptable accuracy and price, buy it and spend the consulting budget on integration. Custom work earns its place when the decision is specific to your workflow, your data is the advantage, or no product handles your documents and rules.

If you are deciding whether an AI engagement is worth the budget, start with scope: we map what you would build, what it integrates with, and what it costs to run, then quote it in writing with the assumptions listed. Request a scope and quote, or talk it through first. You can also see what we build and host for clients running these systems in production.

Share this post