Cost & Delivery

What an enterprise agentic AI pilot actually costs

13 min read · September 2026

Every enterprise AI budget conversation we join starts in the same place: someone has a number in their head, and it came from a vendor deck. The number is usually either far too low, because it counts only model tokens, or far too high, because it prices a transformation program when what is needed is a pilot. This piece is the breakdown we give clients before they commit anything, using the shape of engagements we actually run.

First, definitions, because the word pilot does too much work. A pilot here means a bounded agentic system — multi-step, tool-using, running against real data in a non-production or restricted-production environment — evaluated against a metric someone will sign their name to, delivered in about 30 days, and ending in a documented go or no-go. It is not a proof of concept on synthetic data, and it is not production. It is the artifact that tells you whether production is worth funding.

The largest line is always engineering time, and it is not close. A typical pilot team is two senior engineers for four weeks, plus roughly a quarter of an architect's time for design and review. At blended senior consulting rates in the United States, that lands somewhere in the range of 75,000 to 120,000 dollars for the delivery team. If a proposal for a genuine agentic pilot comes in materially under that, the difference is being made up somewhere: junior staffing, synthetic data, no evaluation, or no security work.

Model spend during a pilot is almost always trivial, and people are consistently surprised by this. A pilot processing 5,000 documents with a frontier model, including the retries and the evaluation runs, typically costs between 200 and 2,000 dollars over the month. Evaluation runs can exceed inference for the pilot itself, because you re-run the whole suite on every prompt change. We budget 2,500 dollars and are usually well under.

The trap is projecting that number forward. Pilot volume is not production volume, and pilot prompts are not production prompts after the context grows. The useful pilot output is not the total spend — it is the measured cost per case, which is what you multiply by real volume later. We instrument that number from day one and put it in the same report as the quality scores.

Infrastructure for a month is modest: a vector store or search index, a container environment, a database, and observability. Two to five thousand dollars covers it in almost every engagement, and a good deal of that is observability tooling that the organization already pays for. Where this line inflates is when the pilot requires a new isolated account, a new landing zone, or a private network path that does not exist yet — then you are paying for platform work, not pilot work, and it should be a separate line.

Now the costs that do not appear in vendor proposals, which in aggregate frequently exceed the delivery fee.

Security review and architecture approval. In a bank, insurer, or health system, expect 40 to 120 hours of internal effort across security architecture, identity, data governance, and privacy — plus calendar time, which is the more expensive resource. This is the line that turns a 30-day pilot into a 90-day pilot when it is not started on day one. Our standing advice: open the security workstream in week one, in parallel with the build, and treat the review as a deliverable with an owner rather than a gate you arrive at.

Data access. Getting a senior engineer credentialed and pointed at real data in a regulated environment routinely takes two to four weeks and involves three teams. It is not billable to anyone, it is invisible on the plan, and it is the most common reason a pilot underdelivers — the team spends its first fortnight on access and its second on the actual problem. Budget it as a task with a named owner and a date.

Subject-matter expert time. An agentic pilot needs a real underwriter, adjuster, analyst, or clinician for roughly four to six hours a week: defining what a good outcome is, labeling the evaluation set, and reviewing output. This is the highest-leverage cost in the whole program and the most frequently under-provided. A pilot with no committed expert produces a system that impresses executives and fails the people who would have to use it.

Legal and vendor onboarding. New model provider, new data processing agreement, new sub-processor disclosure, sometimes a customer notification review. Fifteen to forty hours of internal legal time, mostly front-loaded. It costs nothing if your organization already has a Bedrock or Vertex arrangement in place, which is one practical reason we default to deploying models through a cloud provider the enterprise has already contracted.

Put together, a realistic all-in figure for a first agentic pilot in a large regulated organization is 120,000 to 200,000 dollars, of which roughly 60 to 70 percent is external delivery fee and the rest is internal effort that is real whether or not anyone books it to the project. A second pilot in the same organization costs meaningfully less — typically 30 to 40 percent less — because the access paths, the security pattern, the evaluation harness, and the deployment topology already exist.

What production adds is a different conversation, and the multiplier is smaller than people fear but larger than they hope. Hardening, integration with systems of record, on-call and runbooks, drift monitoring, and the residual compliance work generally run one and a half to three times the pilot delivery cost, spread across the following two to four months. Ongoing run cost is dominated by inference at real volume plus roughly a quarter to a half of an engineer for maintenance — model versions move, prompts drift, upstream schemas change.

The honest cost question is not what the pilot costs. It is what a wrong go decision costs. Funding an eighteen-month program off a demo that was never evaluated is how organizations spend seven figures discovering the thing a 150,000-dollar pilot would have told them in a month. The value of the pilot is largely in its ability to return a credible no.

So we structure pilots to be able to fail informatively. There is a written metric and a threshold agreed before the build starts. There is an evaluation harness, so the result is a number rather than an impression. There is a documented decision at the end with three options — proceed, iterate on a named weakness, or stop — and stopping is a legitimate outcome that we have recommended and been paid for.

Two ways to spend less without weakening the result. Pick a workload where the data is already accessible; the fastest pilots we have run were fast almost entirely because access was solved before the engagement started. And scope to one workflow with one clear owner, rather than a platform that serves three departments. Platform ambition is the right instinct at the wrong stage, and it is where 30-day pilots become 6-month programs with nothing decided at the end.

One closing note on rates, since it is the question we get asked privately. Senior engineering time is expensive and the alternative is more expensive. The pattern we see most often in salvage engagements is a cheap pilot delivered by a large team of junior staff, which produced a demo, no evaluation, no security artifact, and no reusable infrastructure — and then a second, properly staffed pilot funded six months later to answer the original question. Paying once is cheaper.

— Related services

Builders Newsletter

Get our field notes in your inbox.

One thoughtful read a month on what's shipping in commerce, AI, cloud, and security — from the engineers building it.

No spam. Unsubscribe anytime.