AI Agents

How Long Does It Take to Build an AI Agent? Week-by-Week Timeline

How long does it take to build an AI agent? A week-by-week timeline from discovery to production, the seven things that cause delays, and a scoring worksheet.

Cover image for an article about how long it takes to build, test and launch a production AI agent
In this article
  1. Why published AI agent timelines are close to useless
  2. The three clocks: person-weeks, calendar weeks and decision weeks
  3. How long does it take to build an AI agent? Four scopes, four answers
  4. The week-by-week timeline for a production agent
  5. Seven things that make AI agent timelines slip
  6. Estimate your own timeline: a seven-factor worksheet
  7. How to compress the timeline without breaking things
  8. When the honest answer is "not yet"
  9. What "done" actually means: a production-readiness checklist
  10. Timelines in regulated industries
  11. Putting a date on the calendar
  12. Frequently asked questions
  13. Sources

Short answer: how long does it take to build an AI agent depends far less on the model than on your own organization. A working prototype on your real data takes 1–2 weeks. One agent handling one business workflow, running in production with real users, takes 8–16 weeks at most mid-sized US companies. A multi-agent system inside a bank, insurer or health system takes 6–12 months. Swapping frontier models changes those numbers by days. Data access, the number of systems the agent writes to, and the length of your approval queues change them by months.

Most timeline articles give you a table of ranges and stop. This one shows you where the weeks actually go, so you can estimate your own project instead of borrowing someone else's average. You will get a week-by-week schedule for a typical 12-week build, the seven delays that cause almost every slip, a seven-factor worksheet for scoring your own project, and an honest list of the cases where the right answer is "not yet."

Why published AI agent timelines are close to useless

Search this question and you will find confident numbers: 2–4 weeks here, 8 months there, 12 weeks as a "graduation window." Almost none of them cite anything. There is no credible public survey of how long AI agent builds take, because the category is too new and too inconsistently defined for one to exist.

What is verifiable is less flattering and more useful.

Gartner predicts that over 40% of agentic AI projects will be canceled by the end of 2027, "due to escalating costs, unclear business value or inadequate risk controls." In the same June 2025 release, Gartner analyst Anushree Verma says most agentic projects are "early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied," which "can blind organizations to the real cost and complexity of deploying AI agents at scale, stalling projects from moving into production." Gartner also names the mechanism directly: "Integrating agents into legacy systems can be technically complex, often disrupting workflows and requiring costly modifications."

That is the honest starting point. The technology is not the constraint. Stanford's 2026 AI Index reports that AI agents jumped "from 12% to ~66% task success on OSWorld," a benchmark of real computer tasks — a large leap, and still roughly one failure in three on structured benchmarks. Models are good enough to be useful and not good enough to be trusted unsupervised, which means every serious project has to budget time for evaluation, guardrails and human review. That time is the timeline.

Key takeaway: If a vendor quotes you a delivery date before asking who owns your data access and who signs off on the agent's actions, the date is decoration.

Testing that before you sign is a solvable problem: our guide to how to hire an AI agent development company has the screening questions, the scorecard and the proof-of-value design that surface it early.

The three clocks: person-weeks, calendar weeks and decision weeks

Estimating an agent build goes wrong when people conflate three different measurements.

Person-weeks are engineering effort: one engineer working one week. This is what quotes are priced on. The US Bureau of Labor Statistics reports a median annual wage of $135,980 for software developers as of May 2025, and engineers with production agent experience sit above the median. We break the money side down in our guide to AI agent development cost; this article is about the clock.

Calendar weeks are elapsed time. Two engineers do not halve a 12-week schedule, because much of the schedule is not engineering.

Decision weeks are the stretches where your team is blocked on someone else: a security questionnaire, an identity approval, a legal review, a vendor's support queue, a committee that meets on the third Tuesday. These are invisible in every estimate and they are usually the difference between a project that lands in three months and one that lands in seven.

A useful rule from planning real builds: for a first agent at a company that has never shipped one, engineering effort is roughly 50–60% of elapsed time. The rest is waiting and deciding. For a second or third agent at the same company, that ratio improves sharply, because the access, the platform and the review path already exist. This is why the first agent is expensive in time and the fourth one is not.

How long does it take to build an AI agent? Four scopes, four answers

Agents are not one thing. Here are four scopes with distinct timelines. "Person-weeks" is engineering effort; "calendar" assumes one or two engineers working steadily and a responsive organization.

Scope What it does Person-weeks Calendar Main constraint
Prototype / proof of value Answers real questions on real data, in a test environment, for a handful of internal reviewers 2–4 1–2 weeks Getting a data extract
Read-only assistant Searches your documents or records and drafts answers; a human sends everything 6–10 6–10 weeks Content quality, evaluation
Production agent that acts Takes an action in one system — creates a ticket, updates a record, sends a draft for approval 12–24 10–16 weeks Write access, permissions, audit logging
Multi-agent or regulated Several coordinated agents, or any agent touching regulated data or decisions 40–100+ 6–12 months Governance, review cycles, integration count

Three things to notice.

The prototype is dramatically faster than everything else, and that is the point. A prototype that works on real data in two weeks tells you whether the workflow is automatable at all — before anyone commits a six-figure budget. Fleurant AI builds working AI agent prototypes in 1–2 weeks for exactly this reason: it is the cheapest way to find out you were wrong.

Calendar time and person-weeks diverge as scope grows. At prototype scale they are nearly identical. At production scale, calendar time is roughly 1.5–2x the engineering effort. At regulated scale the multiple can be 3x or more, because review cycles dominate.

The jump from "read-only" to "acts" is the expensive one. The moment an agent can change something, you inherit permissions, approvals, audit trails, rollback and a security review. Many teams discover this at week 9 of a 10-week plan.

The week-by-week timeline for a production agent

Here is a schedule for the most common real project: one agent, one workflow, two or three integrations, at a mid-sized US company with no unusual regulatory burden.

Schedule chart showing eight overlapping phases of an AI agent project across sixteen weeks, from discovery and data access through prototype, pilot, integrations, evaluation, security review and launch
Phases overlap on purpose. The blue bars are calendar time you do not control, which is why they start in week 0.

Week 0: discovery and scoping

One week, and it decides most of what follows. The output is not a document; it is four specific answers.

  • Which single workflow? Named, bounded, with a volume number. "Triage inbound support email" beats "help the support team."
  • What does the agent do at the end? Reads and summarizes? Drafts for a human? Acts directly? This answer sets your timeline more than anything else on the list.
  • What does "correct" mean? If nobody can describe a right answer, you cannot evaluate, and without evaluation you cannot launch.
  • Who says yes? One named person who can approve production access and accept residual risk.

In the same week, file every access request you will eventually need. Not when the code works — now. More on why below.

Weeks 1–2: prototype on real data

Build the narrowest possible version that touches real data. No integrations, no permissions model, no interface beyond something a reviewer can click.

Anthropic's engineering guidance on building effective agents is worth following here: find "the simplest solution possible, and only increasing complexity when needed," since "optimizing single LLM calls with retrieval and in-context examples is usually enough." Agentic systems "often trade latency and cost for better task performance," and their autonomy "means higher costs, and the potential for compounding errors." A large share of projects labeled "agent" are workflows, and workflows ship faster and break less.

Deliverable at the end of week 2: 20–50 real cases run end to end, with a documented success rate and a list of failure patterns. That list is the plan for the next ten weeks.

Weeks 3–6: pilot with real users

Five to fifteen people doing their actual jobs. This is where you learn the things no test set predicts: the agent is right but too slow to be useful, the answer format does not fit the workflow, staff do not trust it, or the 15% of cases that fail are the 15% that matter most.

Run the pilot for four to six weeks. Less than three and you only see the happy path. More than eight without a decision usually means nobody agreed in advance what score the agent had to hit to graduate — set that number in week 0.

Weeks 5–12: integrations and hardening

This overlaps the pilot deliberately, and it is the longest engineering stretch. It covers real authentication and per-user permissions, error handling and retries for every external call, audit logging of every action the agent takes, rate limits and cost controls, and rollback for anything the agent can change.

The integration count drives this phase more than the model does. A standard protocol helps: the Model Context Protocol is an open standard for connecting AI applications to external systems, described by its maintainers as "like a USB-C port for AI applications." Where a vendor already publishes an MCP server, you save real time. Where you need a custom one, add one to three weeks per non-trivial system — and see our line-by-line breakdown of MCP server development cost for what that work involves.

Weeks 7–13: security and compliance review

Prompt injection is the reason this phase exists. OWASP's LLM01:2025 Prompt Injection entry is blunt that the problem is structural: "Given the stochastic influence at the heart of the way models work, it is unclear if there are fool-proof methods of prevention." Its mitigations are design decisions, not a checklist you bolt on at the end — constrain model behavior, define and validate output formats, filter input and output, enforce privilege control so the agent has minimum necessary access, require human approval for high-risk operations, segregate untrusted external content, and conduct adversarial testing.

Two of those — privilege control and human approval — change your architecture. Discover them in week 11 and you rebuild. Decide them in week 1 and the review is a formality.

Week 12 onward: launch and operate

Launching is not finishing. Microsoft's guidance on observability in generative AI splits the lifecycle into three stages — "base model selection," "pre-production evaluation" and "post-production monitoring" — and the third one never ends. It covers continuous evaluation of sampled production traffic, scheduled evaluation against test datasets to detect drift, scheduled adversarial testing, and alerts when outputs fail quality thresholds.

The AWS Well-Architected Generative AI Lens, published November 19, 2025, names the same shape across "scoping, model selection, customization, development, deployment, and continuous improvement." Budget ongoing engineering time from launch day. An agent nobody maintains degrades as your business, your documents and your models change underneath it.

Seven things that make AI agent timelines slip

Almost every overrun traces back to one of these. Each has a fix that costs nothing if you apply it early.

1. Data access takes longer than the build. The agent needs production data. Getting it means a data owner, a classification review, possibly a contract amendment, and an environment where the data can legally sit. Fix: file every request in week 0 and track them like engineering tickets.

2. Write access needs someone else's approval. Identity and access governance is designed to be a gate. Microsoft Entra ID's admin consent workflow routes application permission requests to designated reviewers by email, with a configurable expiry, and only Global Administrators can approve apps requesting Microsoft Graph application permissions. Even the mechanics have lag: Microsoft notes it "can take up to an hour for the workflow to become enabled." That is one system. Now multiply it. Fix: identify every permission the agent needs in week 0 and name the approver for each.

3. Vendor environments have their own runway. In healthcare, Epic's developer documentation shows the shape clearly: changes "may take up to 1 hour to sync with the Sandbox," and after marking an app ready for production, requests for client secrets "typically appear within 5 minutes" but you should "allow an hour." Then each Epic community member must sign the open.epic API Subscription Agreement and request your app before you can provision credentials for them. Those customer-side steps are not on your calendar and not under your control. Fix: start vendor onboarding the week you decide the integration is in scope.

4. Nobody defined "correct." Without a labeled set of examples, every release is an argument about vibes. Microsoft's pre-production stage exists precisely to validate "performance through evaluation datasets," identify edge cases and measure "task adherence, groundedness, relevance, and safety." Fix: build a 50–200 case evaluation set in weeks 1–3, from real historical work, with a human-agreed right answer for each.

5. You built an agent where a workflow would do. Autonomy is the most expensive design choice available. Gartner is direct that "many use cases positioned as agentic today don't require agentic implementations." Fix: ship the deterministic workflow first. Add autonomy only where the workflow demonstrably cannot cope.

6. Governance arrives late and unowned. The NIST AI Risk Management Framework, released January 26, 2023, organizes work into four functions — Govern, Map, Measure and Manage — and Govern comes first for a reason. Bolting governance on at the end means redoing decisions. Fix: in week 0, write down who owns the agent, what it may and may not do, and what happens when it is wrong.

7. "While we're in there" scope creep. The agent works, so someone asks for three more use cases before launch. Fix: freeze scope at the end of the pilot. Log every request as version two. Ship version one.

Key takeaway: Six of these seven delays are organizational, and all seven are visible in week 0 if you look. The teams that ship in twelve weeks are not faster engineers. They started the waiting earlier.

Estimate your own timeline: a seven-factor worksheet

Borrowed averages are worthless. Score your own project instead. Each factor is worth 0, 1 or 2 points; add them up and read the range.

Scoring worksheet with seven factors — process clarity, data access, systems written to, action approval, evaluation baseline, regulatory review and decision owner — each scored zero to two points, with totals mapped to calendar ranges
Score the project, not the technology. Nearly every point on this card is something your organization controls.
Factor 0 points 1 point 2 points
Process clarity Documented and agreed Lives in people's heads Teams disagree on the rules
Data access Your team can read it today One ticket away Needs a contract or new environment
Systems written to None — read only One Two or more
Action approval A human approves everything Human reviews exceptions Fully autonomous
Evaluation baseline Labeled examples exist Past tickets you can mine Nothing yet
Regulatory review None Internal policy only Formal committee sign-off
Decision owner One person, named, available Named but hard to reach A committee

0–3 points: 8–12 weeks to production. You have a well-scoped project. The main risk is scope creep.

4–8 points: 3–6 months. Normal for a first agent at a company that has not shipped one. Attack the 2-point rows first; each one you can move to 1 typically returns two to four weeks.

9+ points: 6 months or more. This is a real project, but it is the wrong first project. Find a narrower version — usually the read-only, human-approves variant of the same workflow — ship that in ten weeks, and use the access, evaluation harness and trust you build to do the hard version second.

That last move is the single highest-leverage decision on this page. A 2-point score on "systems written to" and "action approval" can often be reduced to 0 by changing what the agent does at the end, not by changing what it understands.

How to compress the timeline without breaking things

There are five levers that genuinely shorten a build, and a lot of popular advice that does not.

Start the waiting in week 0. Access requests, security questionnaires, vendor onboarding and committee agenda slots all have lead times measured in weeks. Filed in week 0, they resolve while you build. Filed in week 9, they become the critical path. This is the cheapest month you will ever save.

Ship a thin vertical slice, not a broad horizontal one. One workflow, end to end, in production, beats five workflows at 80%. A slice in production generates the operational feedback that makes everything after it faster.

Stage the autonomy. Read → draft for a human → act with approval → act within limits. Each stage is a shippable product and each earns the trust needed for the next. Teams that try to land at the final stage on day one usually spend the extra time in review, not in code.

Use a managed agent runtime. Google Cloud's Vertex AI Agent Engine handles deployment, scaling, session state, memory, tracing and monitoring so your team does not build that layer. Managed services compress infrastructure work, not discovery, evaluation or review — so expect them to shorten the build, not the project.

Run evaluation continuously, not as a phase. Microsoft's guidance is to integrate "automated quality gates into CI/CD pipelines." A test suite that runs on every change catches regressions in minutes instead of in a pilot review three weeks later.

What does not work: adding engineers to a project blocked on approvals, skipping the pilot to "save four weeks" (you spend them after launch, more expensively), and choosing a more capable model to compensate for an undefined problem.

When the honest answer is "not yet"

Some projects should not start, and saying so early is worth more than any schedule.

The process is genuinely undefined. If two experienced staff handle the same case differently and both are right, an agent will not resolve that ambiguity; it will amplify it. Document the process first. That is a business project, not an AI project.

The volume does not justify it. Ten cases a week handled in five minutes each is about four hours a month. A 12-week build will not pay that back in any reasonable horizon.

The data does not exist yet. If the knowledge lives in people's heads and nowhere else, you are budgeting a knowledge-capture project with an AI phase at the end. Document-heavy processes usually start as a custom AI retrieval project and only become an agent once the knowledge is written down.

Nobody owns the outcome. If you cannot name the person who will accept the risk and make the launch call, the project will stall at week 10 regardless of how well it works.

You need a decision this quarter. If the answer is needed in six weeks and the honest estimate is sixteen, a prototype that informs the decision is a better use of six weeks than a build that misses.

What "done" actually means: a production-readiness checklist

Use this as the graduation gate from pilot to production. An agent should not launch until every line has an owner and an answer.

  • Accuracy. A documented score on the evaluation set, at a threshold agreed in week 0 — not a threshold chosen afterward to match the result.
  • Failure behavior. The agent says "I don't know" and escalates rather than guessing. Verified, not assumed.
  • Permissions. The agent acts with the least privilege that works, and where possible as the requesting user rather than as a superuser.
  • Human approval. Required for every high-risk action, per OWASP's guidance on privileged operations.
  • Audit trail. Every action logged with inputs, tool calls, outputs, the model version and the user. In regulated settings, retained to your record-retention policy.
  • Cost controls. Per-user and per-day limits, with alerts, so a loop cannot produce a surprise invoice.
  • Monitoring. Quality and safety evaluated on sampled production traffic, with alerts on threshold breaches.
  • Rollback. A documented way to turn the agent off and reverse what it changed, tested at least once.
  • Ownership. A named owner, a maintenance budget, and a review date.

If your vendor's plan does not include these, the plan is shorter than the project. Our AI agent development work treats this checklist as the definition of delivery rather than a follow-on phase.

Timelines in regulated industries

Banks, insurers, healthcare organizations and law firms should add 6–16 weeks to any range above. The extra time is mostly calendar, not engineering, which means it can run in parallel if you start it in week 0.

What typically gets added: a documented risk assessment and entry in a model or AI inventory; third-party due diligence on every vendor in the chain, including model providers; contractual data terms such as a business associate agreement in healthcare; a privacy review of what leaves your environment; and sign-off by a committee whose meeting schedule you do not control. Committee cadence alone frequently adds a month, because missing one meeting costs four weeks.

The NIST AI RMF's four functions — Govern, Map, Measure, Manage — map onto this work cleanly, and using a recognized framework shortens review because your reviewers recognize the structure. Deployment architecture matters here too; our comparison of private LLM options versus ChatGPT Enterprise covers the data-residency decisions that regulated buyers usually have to settle before an agent project can start. For the governance side specifically, see our AI compliance work. None of this is legal advice — confirm the specifics with your counsel or compliance team.

One compensation: regulated organizations that build the review path once move much faster on agents two through five. The first project pays for the road.

Putting a date on the calendar

If you need a single planning number, use this: for one workflow, one agent, two or three integrations, at a company that has not shipped an agent before, plan twelve weeks from kickoff to production, and expect the critical path to run through access and approvals rather than code.

Then do four things this week. Pick one workflow with real volume and a describable right answer. Name the person who can approve production access. File every access request you will eventually need. And score the project on the seven-factor worksheet above, so the date you commit to is derived from your own constraints rather than someone else's blog post.

If you want a second opinion on the scope and the schedule before you commit a budget, talk to a specialist. A short prototype on your real data is usually the fastest way to replace an estimate with evidence — and it is a much cheaper way to discover a project is not ready than finding out in month five.

Frequently asked questions

How long does it take to build an AI agent?

A working prototype on your real data takes one to two weeks. One agent that handles one business workflow in production takes eight to sixteen weeks for most mid-sized US companies. A multi-agent system in a regulated environment takes six to twelve months. The model you choose barely changes these numbers. Data access, the number of systems the agent writes to, and your own approval queues change them a lot.

Why do AI agent projects take longer than the demo suggests?

A demo runs on data someone hand-picked, with no permissions, no error handling and no audit trail. Production requires the agent to work on the messy 20% of cases, to authenticate as someone, to log what it did, and to pass a security review. Gartner notes that integrating agents into legacy systems is technically complex and often requires costly modifications. That gap, not model quality, is where the weeks go.

Can you build an AI agent in a week?

You can build a credible prototype in a week if the process is already documented, the data is already accessible to your team, and the agent only reads rather than writes. That prototype is genuinely useful: it tells you whether the workflow is automatable before you commit a real budget. What you cannot do in a week is production access, evaluation, security review and change management.

What is the longest part of an AI agent project?

Waiting. In most projects the single longest stretch is getting production access to data and systems: identity approvals, vendor security reviews, API credentials and contract changes. These are other teams' queues, not your engineering backlog. Teams that start these requests in week zero rather than after the code works routinely save a month or more of elapsed time.

How long should a pilot run before going to production?

Four to six weeks with a small group of real users is usually enough to see the failure patterns that matter. Shorter than three weeks and you only see the happy path. Longer than about eight weeks without a decision is a warning sign: it usually means nobody agreed in advance what score the agent had to hit to graduate, so the pilot has no end condition.

Does using a platform like Microsoft Foundry, Amazon Bedrock or Vertex AI make it faster?

Yes, for infrastructure. Managed agent runtimes remove the work of hosting, scaling, tracing and session storage, which can save several weeks. They do not shorten discovery, data access, evaluation, security review or change management, and those are usually the bigger share of the calendar. Expect a managed platform to compress the build, not the project.

How long does an AI agent take to build in a regulated industry?

Add six to sixteen weeks to any of the ranges in this article. Banks, insurers and healthcare organizations typically need a documented risk assessment, model inventory entry, vendor due diligence, a business associate agreement or equivalent, and sign-off from a committee that may meet monthly. Most of that time is calendar time rather than engineering time, so it can run in parallel if you start it early.

Sources

  1. Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, Gartner
  2. Building effective agents, Anthropic
  3. Observability in Generative AI, Microsoft Learn
  4. AWS Well-Architected Generative AI Lens, Amazon Web Services
  5. Vertex AI Agent Engine overview, Google Cloud
  6. LLM01:2025 Prompt Injection, OWASP Gen AI Security Project
  7. AI Risk Management Framework, National Institute of Standards and Technology
  8. Configure the admin consent workflow, Microsoft Learn
  9. Epic on FHIR documentation: testing and production, Epic
  10. What is the Model Context Protocol?, Model Context Protocol
  11. Software Developers, Quality Assurance Analysts, and Testers, U.S. Bureau of Labor Statistics
  12. The 2026 AI Index Report, Stanford Institute for Human-Centered Artificial Intelligence

Free, no-obligation consultation

Have a question about AI agent timelines?

Tell us what you're working on or what you'd like to know. A specialist will get back to you with practical next steps, whether or not we end up working together.

  1. 1Send your question or project details (takes 2 minutes)
  2. 2A specialist reviews it and replies within 1 business day
  3. 3Get clear, practical next steps, free