AI Agents

From AI Pilot to Production: Why Agents Stall and How to Ship

Why AI pilots stall on the way to production, the eight failure modes behind it, a ten-gate readiness scorecard, the reliability math and real cost ranges.

Cover image for an article about moving an AI pilot into production
In this article
  1. What the data actually says about AI pilot to production rates
  2. Why AI pilots stall: eight failure modes, in the order they bite
  3. The production readiness scorecard: ten gates to clear before launch
  4. Why your pilot's accuracy number is lying to you
  5. How to choose an autonomy level you can actually govern
  6. The integration and operations work nobody budgets for
  7. What will change under you: model retirements, quotas and rate limits
  8. What does it cost to take an AI pilot to production?
  9. A 60-day plan to unstick a stalled pilot
  10. When not to put a pilot into production
  11. Next steps
  12. Frequently asked questions
  13. Sources

Most AI pilots do not fail because the model is not good enough. They stall on the way from AI pilot to production because nobody resolved three things before the demo: who owns the agent's mistakes, how its data gets into production safely, and what number has to be true for it to go live. The model was never the hard part.

This article is the practical version of that argument. It covers what the credible data actually says about pilot-to-production rates, the eight failure modes that stop projects in roughly the order they bite, a ten-gate readiness scorecard you can run against your own project this week, the arithmetic that explains why a 95% accurate agent is unusable over twenty steps, the autonomy levels worth governing differently, the platform changes that will break your agent whether you plan for them or not, and honest cost ranges for the production push.

It is written for the person who has a working demo, an impatient executive, and a nagging feeling that the last 20% is going to take longer than the first 80%. It will.

Key takeaway: The distance between a pilot and production is measured in integration, evaluation and ownership, not in model quality. Budget accordingly.

What the data actually says about AI pilot to production rates

Start by distrusting the headline number. The most quoted statistic in this field — the claim that 95% of generative AI pilots produce no measurable return — comes from a 2025 MIT Media Lab NANDA report. As of October 2026, the URL that report was published at redirects to the NANDA project overview page, which does not list the report among its publications. A statistic whose primary source has quietly disappeared is not a statistic to put in a board deck.

Here is what is documented and checkable.

Abandonment is real and it got worse. S&P Global Market Intelligence surveyed 1,006 midlevel and senior IT and line-of-business professionals across North America and Europe for its Voice of the Enterprise: AI & Machine Learning, Use Cases 2025 study. It found that the percentage of companies abandoning the majority of their AI initiatives before they reach production "has surged from 17% to 42% year over year," with organizations reporting that an average of 46% of projects are scrapped between proof of concept and broad adoption.

Success rates inside IT operations are modest but not catastrophic. A Gartner survey of 782 infrastructure and operations leaders, fielded in November and December 2025, found that only 28% of AI use cases in I&O fully succeed and meet ROI expectations, while 20% fail outright. The rest land somewhere in between: live, partially useful, not obviously worth the money.

The projects that survive look different from the start. In a fourth-quarter 2024 Gartner survey of 432 respondents across six countries, 45% of leaders at organizations with high AI maturity said their AI initiatives stay in production for three years or more, compared with 20% at low-maturity organizations. The same survey found that business units trust and are ready to use new AI solutions at 57% of high-maturity organizations versus 14% of low-maturity ones, and that 63% of high-maturity leaders run financial analysis on risk factors, conduct ROI analysis and concretely measure customer impact.

Forecasters expect a lot of cancellations. Gartner predicted in June 2025 that "over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls." In May 2026 it added a second prediction: by 2027, 40% of enterprises will demote or decommission autonomous AI agents because of governance gaps identified only after production incidents occur.

Two things follow. First, the base rate for getting a pilot live and keeping it there is genuinely poor, so a cautious plan is not pessimism. Second, the differences between the survivors and the casualties are organizational and operational, and they are visible before you write the first line of production code.

Why AI pilots stall: eight failure modes, in the order they bite

These are the eight patterns that actually stop projects. They are ordered roughly by when they show up, not by how much damage they do.

1. No one agreed what "working" means

A pilot is judged by whether the demo impressed people. Production needs a threshold. Anthropic's guidance on defining success criteria is blunt about what a usable target looks like: specific ("accurate sentiment classification" rather than "good performance"), measurable, achievable against current model capability, and relevant to the application's purpose. It also notes that "most use cases need multidimensional evaluation along several success criteria," naming task fidelity, consistency, relevance and coherence, tone and style, privacy preservation, context utilization, latency and price.

If your project has one number ("accuracy"), or no number, you do not have a go-live decision. You have an argument waiting to happen.

2. The pilot ran on data the production system cannot have

Pilots run on an exported spreadsheet, a sanitized sample, or a copy of last quarter's tickets. Production needs live, permissioned, current data. Gartner's February 2025 research found that 63% of organizations either do not have, or are unsure whether they have, the right data management practices for AI, based on a July 2024 survey of 1,203 data management leaders, and predicted that through 2026 organizations would abandon 60% of AI projects unsupported by AI-ready data.

Its April 2026 I&O survey puts a sharper number on it: 38% of leaders who faced setbacks said poor data quality or limited data availability was a direct cause of AI project failure, the same share who blamed persistent skills gaps.

3. The use case was never a good fit for an agent

Gartner's June 2025 release described vendors engaging in "agent washing" — rebranding existing assistants, robotic process automation and chatbots without substantial agentic capability — and estimated that only about 130 of the thousands of agentic AI vendors are real. Its analyst put the buyer-side version plainly: "Many use cases positioned as agentic today don't require agentic implementations."

The S&P Global data points the same direction from the opposite end. Organizations reported higher enterprise impact from internal system interaction, complex multi-step processes such as design and prototyping, and sales outreach than from the popular choices of content creation and information summarization. If your pilot summarizes things, the summary is probably nice and the return is probably not measurable.

4. Per-step reliability does not survive multi-step work

This one has arithmetic behind it, and it gets its own section below. The short version: a tool-calling agent's end-to-end success rate is roughly its per-step success rate raised to the power of the number of steps, and that curve is brutal.

Research supports the pessimism. A June 2025 study, More Vulnerable than You Think: On the Stability of Tool-Integrated LLM Agents, found agents are highly susceptible to errors at every stage of tool invocation — reading documentation, selecting the tool, generating parameters and processing the response — and that scaling the model up does not substantially improve tool invocation reasoning and can increase vulnerability to adversarial inputs that look like ordinary instructions.

5. Governance was treated as a switch instead of a dial

Gartner's May 2026 analysis names this directly: "Enterprises are treating AI agent governance as binary, either locked down or fully trusted, and that is the root cause of failure." Over-restricting a simple read-only agent slows delivery and pushes people into shadow builds. Under-restricting an autonomous one creates operational, security and compliance exposure that only becomes visible after an incident.

6. Nobody owns it after launch

An agent is not a report. It drifts, its tools change, its upstream data changes, and it needs a person whose job includes noticing. Gartner's September 2026 prediction that 70% of enterprises will abandon agentic AI built by vendor forward-deployed engineering by 2028 is fundamentally an ownership failure: the observation is that "FDE engagements often fail structurally before they fail technically," because scope, incentives, governance, IP ownership and exit were never settled.

7. The unit economics only worked at pilot volume

A pilot runs a few hundred requests. Production runs a few hundred thousand, with longer contexts, retries, and tool calls that each cost tokens. Costs do not scale linearly with usefulness, and the finance conversation arrives late. We cover the levers in detail in LLM API cost optimization; the point here is that cost per completed task belongs in the go-live criteria, not in the first invoice.

8. Trust never transferred to the people who have to use it

The Gartner maturity survey's trust gap — 57% versus 14% — is the quietest failure mode and often the final one. An agent that users route around is indistinguishable from an agent that does not exist. Adoption is not a launch-day activity; it is a design input.

The production readiness scorecard: ten gates to clear before launch

This is the checklist we use to decide whether a pilot is ready to be promoted. Score each gate 0 (not started), 1 (partial) or 2 (done and evidenced). Anything below 16 out of 20 means you are shipping risk, not a product. Any gate at 0 in the Control or Operate group should block launch outright.

A four-stage production readiness scorecard with ten gates covering decision, evidence, control and operations
The ten gates between a working AI pilot and a production system you can defend
# Group Gate What "done" looks like
1 Decide Named business owner One person, not a committee, is accountable for the outcome and the budget after launch
2 Decide Quantified success threshold A written number per dimension: task success, escalation rate, latency, cost per completed task
3 Decide Failure budget and kill switch An agreed error rate that triggers rollback, and a tested way to turn the agent off in minutes
4 Evidence Held-out evaluation set from real traffic At least 100 real cases the agent has never seen, labeled by someone who does the job
5 Evidence Trajectory evaluation, not just final answers Tool-call sequences compared against a reference, so you know how it reached the answer
6 Control Autonomy level chosen and justified The agent's write permissions match its measured reliability, not its potential
7 Control Least-privilege access and audited write paths Scoped credentials per tool, every write logged with the prompt, inputs and approver
8 Operate Tracing and per-task cost attribution Every run traceable end to end, with token and tool-call counts attributed to a workflow
9 Operate Model migration plan A documented replacement candidate and a re-runnable evaluation set
10 Operate Human escalation path that someone staffs A real queue with a real owner and a response-time target, not an unmonitored inbox

Key takeaway: Gates 1, 3 and 10 are organizational and cost nothing but a decision. They are also the three most commonly missing when a pilot stalls.

The structure mirrors how the major cloud providers frame the same problem. The AWS Well-Architected Generative AI Lens, published in November 2025, walks the generative AI lifecycle through scoping, model selection, customization, development, deployment and continuous improvement across six pillars, including operational excellence ("achieve consistent model output quality, monitor and manage operational health, maintain traceability") and reliability. If you want a vendor-neutral governance scaffold to hang this on, the NIST AI Risk Management Framework (NIST AI 100-1, January 2023) organizes the work into Govern, Map, Measure and Manage, and NIST published a Generative AI Profile (NIST AI 600-1) in July 2024. Both are voluntary, and neither is a substitute for your compliance team's judgment.

Why your pilot's accuracy number is lying to you

Here is the single most useful piece of arithmetic in agent development, and the reason experienced teams get nervous when someone says "it works 95% of the time."

If an agent completes a task by taking a number of independent steps, and each step succeeds with the same probability, the probability that the whole task succeeds is approximately the per-step rate raised to the power of the step count. Steps are not perfectly independent in practice, and well-built agents recover from some errors, so treat this as a floor rather than a forecast. It is still the right mental model, and it explains almost every "but it worked in the demo" conversation.

A chart showing end-to-end agent success collapsing as the number of steps grows, for per-step reliability of 95, 99 and 99.9 percent
End-to-end success rate by number of steps, calculated as per-step reliability raised to the power of the step count
Per-step success 5 steps 10 steps 20 steps 50 steps
95% 77.4% 59.9% 35.8% 7.7%
99% 95.1% 90.4% 81.8% 60.5%
99.9% 99.5% 99.0% 98.0% 95.1%

Read the first row again. A 95% reliable step is a good-looking demo and an unusable twenty-step workflow. Flip it around: to keep end-to-end success at or above 95%, a 95% step rate buys you one step, a 99% rate buys you five, and a 99.9% rate buys you fifty-one.

Three practical consequences:

Shorten the chain before you improve the model. Collapsing four tool calls into one well-designed API endpoint does more for reliability than a model upgrade. This is a large part of why teams invest in purpose-built tool layers; we go into the economics of that in MCP server development cost.

Measure the path, not just the destination. Google Cloud's Gen AI evaluation service distinguishes "final response evaluation: evaluate the final output of an agent (whether or not the agent achieved its goal)" from "trajectory evaluation: evaluate the path (sequence of tool calls) the agent took to reach the final response," and offers trajectory metrics including exact match, in-order match, any-order match, precision, recall and single tool use. A final-answer-only evaluation will happily pass an agent that got the right answer by accident after three wrong tool calls, and that agent will hurt you in production. The feature is in preview at the time of writing, so check its current status before you build a release gate on it.

Add checkpoints where the chain is long. If a workflow genuinely needs twenty steps, break it into stages with validation between them, so a failure costs one stage rather than the whole task.

How to choose an autonomy level you can actually govern

Gartner's May 2026 guidance proposes classifying agents by autonomy level and applying proportional governance to each. The four levels it describes are Observe (read-only, output visible only to the requesting user), Advise (generates recommendations or drafts, humans execute), Act with Approval (can write data, send communications or change configuration, but only after explicit human approval for every action), and Act Autonomously (executes within defined guardrails, with humans reviewing exceptions and aggregate outcomes).

The useful move is to tie the level you deploy at to the reliability you have actually measured, and to promote only when the evidence supports it.

Autonomy level Deploy when your measured end-to-end success is Minimum controls Typical first use case
Observe Any level, if outputs are clearly labeled Scoped read access, authentication, usage logging Policy and document lookup
Advise Reliable enough that reviewers are not rubber-stamping Hallucination and accuracy testing, user training on how much to rely on it Draft responses, first-pass triage
Act with Approval High on the specific action, with reviewers who have time Approval workflow with audit trail, agent-specific incident response Ticket updates, scheduling, order changes
Act Autonomously High and stable over weeks of production traffic Continuous monitoring, enforced guardrails, rapid rollback, circuit breakers, named owner Narrow, high-volume, low-blast-radius tasks

Two warnings are worth repeating from that research. At the approval level, "human review is effective only if it remains a meaningful control" — without clear workflows and audit trails, approvals degrade under time pressure into a false sense of safety. At the autonomous level, actions execute at a speed that can outpace human oversight, which is why circuit breakers that halt operation on threshold violations belong in the first release, not the second.

The common mistake is launching at Act with Approval because it feels safe, then discovering that the approval queue is the bottleneck and quietly loosening it. Decide the promotion criteria in advance and write them down.

If your agent touches regulated data or decisions, the autonomy question is also a supervisory question. For banks and credit unions, our guide to SR 26-2 model risk management covers how examiners think about model inventories and validation; for healthcare, see the HIPAA-compliant AI development checklist. Confirm the specifics with your counsel or compliance team before you set an autonomy level.

The integration and operations work nobody budgets for

This is where the schedule actually goes. In the Gartner I&O survey, the leaders who succeeded attributed it primarily to integrating AI into existing workflows and systems and securing full executive support, and 33% of successful leaders said they embed AI into the systems and processes people already use rather than running it as a side project.

Here is the work that sits behind that sentence.

Identity and permissions. The agent needs its own service identity, scoped per tool, with credentials that rotate. "The agent uses the developer's API key" is a pilot pattern and a production incident. The OWASP Top 10 for LLM Applications 2025 lists this directly as LLM06: Excessive Agency.

Write-path safety. Every action that changes state needs idempotency keys so a retry does not double-charge a customer, a dry-run mode, and a reversal path. Design the undo before the do.

Input and output handling. LLM01 (Prompt Injection) and LLM05 (Improper Output Handling) in the OWASP list are the two that most often turn a helpful agent into an attack surface, particularly when the agent reads untrusted content — email, web pages, uploaded documents — and then calls a tool.

Tracing and cost attribution. You cannot operate what you cannot see. The OpenTelemetry generative AI semantic conventions define metrics purpose-built for this, including gen_ai.client.operation.duration, gen_ai.server.time_to_first_token, gen_ai.invoke_agent.duration, gen_ai.invoke_agent.inference_calls, gen_ai.invoke_agent.tool_calls and gen_ai.execute_tool.duration. Those last three are exactly what tells you an agent has started looping before the bill does. The conventions are still marked as in development, so pin your versions and expect churn, but the attribute names are the right vocabulary to standardize on now.

Latency budgets. OpenAI's production guidance notes that "the bulk of the latency typically arises from the token generation step," which means the biggest levers are model choice and output length rather than infrastructure. Set a 95th-percentile latency target per workflow and measure it with real context lengths, not with your test prompts.

Rate limits and capacity. This surprises almost everyone. Microsoft's Azure OpenAI quotas and limits documentation warns that when you exceed your usage tier, "latency can vary and, in some cases, may be more than two times higher than when operating within your usage tier," and that you might receive 429 responses "even when token usage metrics appear below your quota." It also notes that quota increase requests are prioritized for customers already using their existing allocation, and that requests which do not meet that condition "might be denied." Plan capacity before launch week, not during it.

Retrieval quality. If the agent reads from a knowledge base, retrieval quality usually matters more than model quality. The cost and effort profile of doing that properly is covered in RAG chatbot development cost.

What will change under you: model retirements, quotas and rate limits

A pilot is a snapshot. Production is a subscription to someone else's roadmap, and the platform will change beneath you on a published schedule most teams never read.

Models retire, on a clock, with no extensions. Microsoft's Foundry Models lifecycle and support policy sets a generally available model's retirement date 18 months out at launch. At 12 months the model becomes Deprecated, which blocks new customers. At 18 months it is Retired and "all inference requests return 410 Gone." Customers with active deployments get at least 60 days of notice before a GA retirement and at least 30 days for a preview model. The documentation is explicit that retirement dates are not extendable, and that generally available models from Anthropic, DeepSeek, Fireworks and Mistral AI on that platform follow a 12-month lifecycle rather than 18.

Anthropic's own model deprecations page commits to "at least 60 days' notice before model retirement for publicly released models." The recent history shows what that means in practice: Claude Sonnet 4.5 was deprecated on September 30, 2026 with a retirement date of November 30, 2026, a 61-day window, and Claude Sonnet 4, released in May 2025, was retired on June 15, 2026, about thirteen months later.

Preview models are not production models. Microsoft states it directly: "Preview models aren't recommended for production workloads." Preview deployments can be force-upgraded to a replacement with 30 days' notice, and there is no option to stay on a retiring preview model. If your pilot was built on a preview model, that is a migration you have already committed to.

Provisioned capacity does not migrate itself. On Azure, Standard-tier deployments are auto-upgraded at retirement, but provisioned deployments are not, and those customers must migrate manually. The team that chose provisioned throughput for predictable latency is the team that owns the migration.

The practical response is not to fight this. It is to make the model a replaceable part:

  • Keep the evaluation set from gate 4 of the scorecard re-runnable on demand, so swapping models is a measurement rather than an argument.
  • Keep model identifiers in configuration, never in code.
  • Put a quarterly reminder in the calendar to check each provider's deprecation page, and subscribe to the service health alerts your platform offers.
  • Assume one model migration per year per workload and put it in the run budget.

Microsoft's own guidance is worth taking at face value here: do not wait for a replacement to be named, evaluate candidates "against your own application, prompts, and representative data," and "compare quality, latency, and cost together rather than relying on public benchmarks alone."

What does it cost to take an AI pilot to production?

Costs here are ranges built from stated assumptions, not quotes. Anyone who gives you a fixed number before seeing your systems is guessing.

The labor input. The U.S. Bureau of Labor Statistics reports that "the median annual wage for software developers was $135,980 in May 2025." Loading that at 1.3 to 1.4 times for payroll taxes, benefits and overhead gives roughly $177,000 to $190,000 a year, or about $85 to $92 an hour at 2,080 hours. US-based senior AI engineers and MLOps specialists typically sit above that median, and offshore or nearshore teams below it. Use your own blended rate if you have one.

The effort. The table below covers the production-hardening work that comes after you already have a working pilot. It assumes the pilot's core logic is sound, the target systems have usable APIs, and no new data platform is required.

Workstream Typical effort What drives the range
Evaluation set, harness and trajectory metrics 2-5 person-weeks How much labeling real cases requires, and whether domain experts are available
Data access, permissions and pipeline hardening 3-10 person-weeks Number of source systems, API maturity, how much of the data lives in someone's spreadsheet
Identity, write-path safety, idempotency, audit logging 2-6 person-weeks How many write actions the agent has, and whether the target systems support reversal
Observability, tracing and cost attribution 1-3 person-weeks Whether you already run distributed tracing
Security review and threat modeling against the OWASP LLM risks 1-4 person-weeks Internal review only, versus a third-party penetration test
Compliance review and documentation 0-8 person-weeks Zero for internal tooling; the top of the range for regulated decisions needing model validation
Change management, training and rollout 1-4 person-weeks Number of users and how much the workflow changes for them

At the low end that is 10 person-weeks of work, and at the high end 40. At $85 to $92 an hour and 40 hours a week, that puts the production push in a band of roughly $34,000 to $147,000 in labor for a single-workflow agent, before run costs. Regulated deployments and multi-system agents land above it.

The run costs. Budget three line items that pilots never have: inference, which scales with traffic and context length; the human escalation queue, a real salaried cost and often the largest of the three; and maintenance, which should assume one model migration per year plus ongoing evaluation. For a starting point on the first of those, see AI agent development cost.

Key takeaway: The production push commonly costs more than the pilot did, because the pilot bought a capability and production buys trust, safety and operability. Teams that budget for one phase get stuck halfway through the second.

A 60-day plan to unstick a stalled pilot

If you have a demo that impressed people three months ago and has not moved since, this sequence tends to break the logjam. It assumes one workflow, one owner and no new data platform.

Days 1 to 10: decide. Name the business owner. Write the success thresholds — task success rate, escalation rate, 95th-percentile latency, cost per completed task — as numbers with the owner's signature next to them. Define the failure budget that triggers rollback. Pick the autonomy level you will launch at, which should usually be one level below where you expect to end up. None of this is technical and all of it blocks everything else.

Days 11 to 25: prove. Build the held-out evaluation set from real traffic, labeled by people who do the job. Run both final-response and trajectory evaluation. Publish the result even if it is bad, especially if it is bad, because a measured gap is a plan and an unmeasured one is an argument. If the numbers miss the threshold, shorten the chain before you reach for a bigger model.

Days 26 to 45: control. Give the agent its own identity with least-privilege, per-tool credentials. Make every write idempotent, logged and reversible. Stand up tracing with per-task token and tool-call attribution. Threat-model against the OWASP LLM risks, with particular attention to prompt injection if the agent reads anything a user or a third party can influence. Build and test the kill switch.

Days 46 to 60: ship small and watch. Launch to a limited population — one team, one region, one customer segment — with the escalation queue staffed and a named person watching the dashboards daily. Compare live metrics to the thresholds weekly. Expand when the live numbers hold, not when the calendar says so.

Two notes on sequencing. Do not run these phases in parallel to save time, because phase 2 regularly changes the design that phase 3 would have hardened. And if phase 1 takes more than ten days, the problem is not the technology, and no amount of engineering will fix it.

For a fuller view of how this fits into an overall timeline, including the pilot phase that precedes it, see how long it takes to build an AI agent.

When not to put a pilot into production

Honest advice includes the cases where the answer is no.

When the workflow is rare. If the task happens fifty times a month and takes a person ten minutes, the agent saves about eight hours a month and costs more than that to operate and govern. Automate the thing that happens ten thousand times.

When nobody will own it. An agent without an owner degrades silently. If you cannot name the person, do not launch; the question is organizational, and shipping will not answer it.

When the measured reliability does not support any safe autonomy level. If the agent is not good enough to act and not trusted enough to advise, production will only move the failure somewhere more expensive.

When the real problem is the process. A surprising share of stalled AI pilots are automating a workflow that should be deleted or simplified. Gartner's I&O research found failures concentrated where leaders expected AI to "fix long-standing operational issues." AI is poor at fixing processes and good at executing clear ones. Sometimes the right answer is not an agent at all but better software, a trade-off we wrote about in when to replace spreadsheets with custom software.

When the compliance question is unresolved. Launching into a regulated process while the control question is open is the most expensive sequencing error available. Resolve it first, with your compliance function and counsel, and see our AI compliance overview for the shape of that work.

Next steps

Getting from AI pilot to production is a management problem wearing a technical costume. The data says the organizations that succeed choose use cases on business value and feasibility, weigh compliance, risk and data availability before they start, measure results with more than one metric, and give the thing an owner. The ones that stall usually skipped a decision, not a sprint.

If you have a stalled pilot, do three things this week. Score it honestly against the ten gates above. Run the step-reliability math on your actual workflow length. And name the person who will own it in production, out loud and in writing.

If you would rather not do the production push alone, Fleurant AI builds and hardens production agents through a vetted network of specialist teams with one accountable partner, including a risk-free proof of value for AI agent development projects and a free discovery call. You can also talk to a specialist about where your pilot is stuck, or read how to evaluate outside help in how to hire an AI agent development company.

Frequently asked questions

What percentage of AI pilots make it to production?

There is no single credible number, and you should be suspicious of articles that give you one. The best-documented figures come from S&P Global Market Intelligence, whose May 2025 survey of 1,006 IT and line-of-business professionals in North America and Europe found that the share of organizations abandoning the majority of their AI initiatives before production rose from 17% to 42% year over year, with an average of 46% of projects scrapped between proof of concept and broad adoption.

Why do AI agents work in a demo but fail in production?

A demo runs one happy path once, with a human steering it. Production runs thousands of messy paths unattended. Three things change: per-step errors compound across tool calls, real data is dirtier and less accessible than the pilot's sample, and every write action now needs permissions, audit logs, rollback and an owner. Most of the work between pilot and production is integration and control, not modeling.

How long does it take to move an AI pilot into production?

For a scoped single-workflow agent with existing APIs and a clear owner, plan on 6 to 12 weeks after a working pilot. Projects that need new data pipelines, identity work, a formal model risk review or vendor security approval routinely run 4 to 9 months. The variable that moves the schedule most is not the model; it is how long it takes to get production data access and a signed-off decision on who owns the agent's mistakes.

What should I measure before letting an AI agent act on its own?

Measure both the final answer and the path taken to it. Google Cloud's Gen AI evaluation service splits these into final response evaluation and trajectory evaluation, where trajectory metrics compare the agent's actual sequence of tool calls against a reference sequence. Add a per-step success rate, an escalation rate, a cost per completed task, and the rate at which humans override the agent. Track those on real traffic, not on your original test set.

Do I need a governance framework before going live?

You need proportional controls, not a uniform policy. Gartner predicts that by 2027, 40% of enterprises will demote or decommission autonomous AI agents because of governance gaps found only after a production incident, and warns that applying the same controls to every agent causes both over-restriction and under-restriction. The NIST AI Risk Management Framework is a reasonable, voluntary scaffold. Confirm anything that touches regulated data with your compliance team or counsel.

What happens when the model my pilot was built on is retired?

It stops working. Microsoft sets an Azure OpenAI GA model's retirement date 18 months out at launch and returns HTTP 410 Gone afterwards, with at least 60 days of notice and no extensions. Anthropic commits to at least 60 days' notice before retiring a publicly released model. Treat the model as a replaceable component: keep an evaluation set you can re-run, and budget for a migration roughly once a year.

Should we build the agent ourselves or have a vendor's engineers build it on site?

Either can work, but decide the exit before you start. Gartner predicts that by 2028, 70% of enterprises will abandon agentic AI built by vendor forward-deployed engineering, and notes that these engagements often fail structurally before they fail technically. Contract for knowledge transfer, IP ownership and a transition plan on day one, and measure success by whether your team can change the agent without the vendor.

Sources

  1. Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, Gartner
  2. Gartner Says Applying Uniform Governance Across AI Agents Will Lead to Enterprise AI Agent Failure, Gartner
  3. Gartner Predicts 70% of Enterprises Will Abandon Agentic AI Built by Vendor Forward-Deployed Engineering by 2028, Gartner
  4. Gartner Survey Finds 45% of Organizations With High AI Maturity Keep AI Projects Operational for at Least Three Years, Gartner
  5. Gartner Says AI Projects in I&O Stall Ahead of Meaningful ROI Returns, Gartner
  6. Lack of AI-Ready Data Puts AI Projects at Risk, Gartner
  7. Generative AI experiences rapid adoption, but with mixed outcomes - Highlights from VotE: AI & Machine Learning, S&P Global Market Intelligence
  8. NANDA project overview, MIT Media Lab
  9. Foundry Models lifecycle and support policy, Microsoft Learn
  10. Azure OpenAI in Microsoft Foundry Models quotas and limits, Microsoft Learn
  11. Model deprecations, Anthropic
  12. Evaluate Gen AI agents, Google Cloud
  13. Define your success criteria, Anthropic
  14. Production best practices, OpenAI
  15. OWASP Top 10 for LLM Applications 2025, OWASP GenAI Security Project
  16. AI Risk Management Framework, National Institute of Standards and Technology
  17. AWS Well-Architected Generative AI Lens, Amazon Web Services
  18. OpenTelemetry semantic conventions for generative AI, OpenTelemetry
  19. More Vulnerable than You Think: On the Stability of Tool-Integrated LLM Agents, arXiv
  20. Occupational Outlook Handbook: Software Developers, Quality Assurance Analysts, and Testers, U.S. Bureau of Labor Statistics

Free, no-obligation consultation

Have a question about getting AI pilots into production?

Tell us what you're working on or what you'd like to know. A specialist will get back to you with practical next steps, whether or not we end up working together.

  1. 1Send your question or project details (takes 2 minutes)
  2. 2A specialist reviews it and replies within 1 business day
  3. 3Get clear, practical next steps, free