AI Agents

How to Hire an AI Agent Development Company: Scorecard and RFP

How to hire an AI agent development company: a 30-point scorecard, a 10-part RFP outline, the contract terms to settle first, and the red flags that matter.

Cover image for an article about how to hire an AI agent development company, including a vendor scorecard and RFP outline
In this article
  1. Why hiring an AI agent development company is not like hiring a web shop
  2. Before you contact anyone: the one-page requirement brief
  3. The 30-point vendor scorecard
  4. Ten screening questions, and the answers that should worry you
  5. The ten-part RFP outline that produces comparable proposals
  6. Run a paid proof of value on your own data
  7. Contract terms to settle before the statement of work
  8. Red flags, and how to verify instead of trust
  9. When not to hire an outside firm at all
  10. What good looks like, and what to do next
  11. Frequently asked questions
  12. Sources

Short answer: how to hire an AI agent development company comes down to four moves, and only one of them involves a sales call. Write a one-page requirement before you contact anyone. Screen on written answers instead of decks. Pay two finalists to build the same thing on your real data for two weeks, scored against a test set you never show them. Settle intellectual property, data use and exit terms before you sign a statement of work.

That sequence exists because the usual sequence fails. Most buyers start with a shortlist of impressive-looking firms, sit through demos, pick the one that felt most confident, and discover in month four that nobody agreed what "working" meant.

This article gives you the tools to run it the other way: a 30-point weighted scorecard, ten screening questions with the answers that should worry you, a ten-part request for proposal outline, a design for a paid proof of value, a reference-call script, and a contract checklist adapted from how the US federal government is now required to buy AI. It also tells you when the right answer is not to hire anyone.

Why hiring an AI agent development company is not like hiring a web shop

Three things make this purchase different, and every mistake below traces back to one of them.

The output is probabilistic, so "done" needs a number. A checkout page either works or it does not. An agent that classifies and routes inbound requests is right some percentage of the time, and that percentage depends on which cases you feed it. If the contract does not define correctness, measured on a specific set of cases, there is nothing to accept or reject at the end.

Most of the work is not AI work. The model is a commodity you rent. The expensive part is access to your systems, permissions, audit logging, error handling, human review and the plumbing between them. A vendor who talks mostly about models and barely about your identity provider has not built one of these in production.

The vendor's leverage grows after launch. The prompts, tool definitions, retrieval configuration, evaluation set and deployment scripts are the actual asset. If those live in the vendor's account, under their license, in their framework, your switching cost is not the code. It is starting over.

Add a hype cycle on top. In a June 25, 2025 press release, Gartner predicted that over 40% of agentic AI projects will be canceled by the end of 2027, "due to escalating costs, unclear business value or inadequate risk controls." The same release names the supply-side problem: many vendors are "engaging in 'agent washing' – the rebranding of existing products, such as AI assistants, robotic process automation (RPA) and chatbots, without substantial agentic capabilities," and Gartner "estimates only about 130 of the thousands of agentic AI vendors are real."

Gartner analyst Anushree Verma is blunt about the demand side too: "Most agentic AI propositions lack significant value or return on investment (ROI), as current models don't have the maturity and agency to autonomously achieve complex business goals or follow nuanced instructions over time," and "Many use cases positioned as agentic today don't require agentic implementations."

That last sentence is worth holding onto during vendor calls. Anthropic's engineering guidance draws a clean line: "Workflows are systems where LLMs and tools are orchestrated through predefined code paths," while "Agents are systems where LLMs dynamically direct their own processes and tool usage." Its advice is to find "the simplest solution possible, and only increasing complexity when needed," because agentic systems "often trade latency and cost for better task performance." A vendor who proposes a deterministic workflow for a deterministic problem is not underselling you. They are the ones who have done this before.

Key takeaway: You are not buying a model. You are buying integration work, an evaluation method and a maintenance relationship. Score vendors on those three things and the field narrows fast.

Before you contact anyone: the one-page requirement brief

The single highest-leverage hour in this process happens before any vendor knows you exist. Write one page that answers seven questions. Send the same page to everyone.

  1. Which one workflow? Named and bounded. "Triage and route inbound service email" beats "improve customer support."
  2. What volume? Items per day or per month today, and the peak. Run costs and payback both depend on it.
  3. Which systems, and read or write? List them. Mark each one read-only or write. The read-to-write jump is where timelines and security reviews double.
  4. What does "correct" mean, and who judges it? One named person who can look at 50 outputs and say right or wrong. If nobody can, you have a research project, not a build.
  5. How sensitive is the data? Personal data, protected health information, customer financial data, or none of the above. This decides half the contract.
  6. Who decides and who signs off on risk? One name each.
  7. What is the budget band and the date that actually matters? A range is fine. A refusal to give one produces proposals you cannot compare.

Two benefits. Vendors who cannot respond to a one-pager with substance disqualify themselves cheaply. And the page becomes the core of your request for proposal later, so you write it once.

How to build a shortlist of eight to twelve

Cast wider than you think you need, because the screen is fast and the hit rate is low. Mix vendor types deliberately, because they fail in different ways.

Vendor type Best fit Real strength How it goes wrong
Independent contractors Prototypes, internal read-only assistants Speed, low overhead, direct access to the builder Single point of failure; little security or compliance muscle; nobody on call after handover
Boutique AI specialist Production agents, evaluation-heavy work, regulated data Has shipped and measured LLM systems before Capacity limits; may be thin on your industry's domain rules
Full-service dev agency Agents inside a larger app or portal build Front end, DevOps and project management under one roof Often has no one who has run an agent in production; learns on your budget
Systems integrator Multi-system programs at larger companies Enterprise change management, existing access to your stack Slow, expensive, and the named seniors in the pitch rarely write the code
Platform or cloud partner Builds standardized on one cloud's AI services Certified engineers, aligned with your existing cloud commitments Architecture bends toward that platform whether or not it fits

Where to look: your cloud provider's partner directory, the maintainers and contributors of the open-source tools in your intended stack, engineering leaders in your network who have shipped something similar, and the vendors your peers in the same industry actually pay. Ignore ranking sites that charge for placement.

The 30-point vendor scorecard

Impressions do not survive comparison. Scores do. Fill this in the same day as each call, one scorer per vendor, from evidence you saw or read rather than from how the conversation felt.

A thirty-point weighted scorecard for AI agent vendors, showing seven categories with their point weights, three hard gates that fail a vendor regardless of score, and how to read the total
Weights reflect what actually predicts a successful build: evidence of production work, judgment, and terms you can live with.
Category Points Full marks looks like Zero looks like
Evidence of production agents 6 Walks you through a live system: daily task volume, success rate, cost per task, and the last three failures with fixes A demo video, a logo wall, or a pilot they call production
Technical depth and judgment 5 Proposes the simplest architecture that could work, and can say why not fine-tuning, why not multi-agent Every problem gets the same framework; fine-tuning proposed in week one
Data handling, security, compliance 5 Names where data sits, who can see it, subprocessors, SOC 2 scope, and what they will not do with your data "Security is handled" with no artifact behind it
Evaluation: how they prove it works 4 Brings a method: test set construction, per-case scoring, regression checks before each release "We test it thoroughly" or "the model is 95% accurate" with no basis
Commercial terms: IP, portability, exit 4 Agrees in writing to IP assignment, source delivery, your cloud account, and transition assistance Proprietary runtime you cannot self-host; licensing you need a lawyer to parse
Delivery model and named people 3 Names the engineers, puts them on the call, states who is on the project full time and for how long Senior names in the pitch, unnamed "delivery team" in the contract
Run and support after launch 3 A concrete support model: hours, response times, who fixes a broken integration, what it costs monthly Support discussed as an afterthought, priced later

Three hard gates fail a vendor regardless of score. No named engineer on the technical call. Refusal to run a paid proof of value on your data. Refusal to assign intellectual property or accept portability terms. Each one predicts a specific kind of pain later, and none of them is negotiable for a first project.

Reading the total: 24 to 30 is a finalist, send them to the paid proof of value. 18 to 23 is conditional, and the gaps must be closed in writing before they advance. Below 18 is a no, whatever the price says.

Ten screening questions, and the answers that should worry you

Send these in writing. Written answers are comparable, they take less of everyone's time than four discovery calls, and the quality of writing is itself a signal.

  1. Describe one agent you have running in production today. What does it do, how many tasks a day, what does a task cost, and what percentage needs human intervention? Worrying answer: only pilots, or an NDA excuse for every detail including volume.
  2. On a project like ours, what share of the effort is model work versus integration, permissions and error handling? Good answer: integration dominates, usually by a lot. Worrying answer: a long description of model selection.
  3. How would you measure correctness on our workflow, and what score would you refuse to launch below? Worrying answer: any accuracy figure quoted before they have seen your data.
  4. What would you tell us not to build? Good answer: something specific, ideally a part of your own brief. Worrying answer: enthusiasm for all of it.
  5. Who exactly would work on this, and will they be on the next call? Worrying answer: roles instead of names.
  6. What happens to our data? Is any of it used to train, fine-tune or improve anything of yours or a third party's? Good answer: a flat no, and an offer to put it in the contract.
  7. Which components run in our cloud account and which run in yours? Worrying answer: everything in theirs, with no self-host path.
  8. What exactly do we receive at the end? Good answer: a list — source, infrastructure as code, prompts, tool definitions, evaluation set, runbook, architecture notes. Worrying answer: "access to the platform."
  9. What went wrong on your last agent project, and what did you change? Worrying answer: nothing has ever gone wrong.
  10. What is the monthly run cost at our volume, and what drives it? Good answer: a per-task estimate with the assumptions stated. Worrying answer: a flat monthly number with no volume attached.

Key takeaway: Every question above asks for an artifact or a number. That is the point. A vendor's willingness to be specific in writing is the cheapest predictor of how the project will run.

Reference calls: six questions that work

Ask to speak with the customer's engineer or operations lead, not only the executive sponsor. The sponsor remembers the outcome. The engineer remembers the project.

  1. What did the first version get wrong, and how long did it take to fix?
  2. How long from signature to something real users actually used? Compare it to the plan they were given.
  3. What did you end up doing yourselves that you expected the vendor to do? Data cleanup and access are the usual answers, and they are the usual budget overrun.
  4. What happened when a model or API changed under you? Who noticed, who paid, how long it took.
  5. Who fixes it at 2am now, and what does that cost per month?
  6. Would you let them build something that writes to a system of record? The hesitation is the answer.

If a vendor can only offer sponsor-level references, or cites confidentiality for every operational detail, weigh that as evidence. Plenty of firms can produce an engineer who will talk about volumes and failure modes without naming a client.

The ten-part RFP outline that produces comparable proposals

The US federal government rewrote its own AI buying rules in April 2025, and the result is the best free procurement template available to a private buyer. OMB Memorandum M-25-22, Driving Efficient Acquisition of Artificial Intelligence in Government, issued April 3, 2025, rescinds and replaces M-24-18 and is organized around competition, performance tracking and cross-functional engagement. Its mechanics translate directly to a mid-sized company buying one agent.

The memo urges "performance-based techniques" because they "allow agencies to understand and assess vendor claims about their proposed use of AI systems or services prior to contract award." That phrase is the whole reason to write an RFP rather than take three proposals in three different shapes.

Ten sections, two to six pages total:

  1. Objective and outcome. The business result, not the technology. One paragraph.
  2. Scope: the workflow. Your one-pager, expanded. Include volumes, seasonality and the cases you know are hard.
  3. Systems and data. Each system, the access method, read or write, who owns the credential, and the current state of the data.
  4. Performance requirements and acceptance criteria. The numbers that define done: task success rate on your test set, maximum latency, maximum escalation rate, cost per completed task. State that acceptance is measured, not demonstrated.
  5. Evaluation method. Tell them you will score against a set of real cases they will not see. M-25-22 is explicit about this for federal buyers: evaluation data "should not be accessible to the vendor, and should be as similar as possible to the data used when the system is deployed."
  6. Security and compliance requirements. Where data may reside, encryption, logging, subprocessor disclosure, SOC 2 expectations, and any regulated-data terms. Ask them to map their controls to the OWASP Top 10 for LLM Applications, whose 2025 list includes LLM01 Prompt Injection, LLM02 Sensitive Information Disclosure, LLM06 Excessive Agency and LLM05 Improper Output Handling. Those four are the ones that bite agents with write access.
  7. Deliverables and handover artifacts. The list from question 8 above, named individually.
  8. Run and support. Expected monthly cost, monitoring, on-call expectations, response times, and how model or API changes are handled.
  9. Commercial terms. Your IP, data-use, portability, liability and exit requirements, stated up front so pricing reflects them.
  10. Pricing format. Mandate the shape: engineer-days by role and phase, third-party and cloud costs separately, support priced monthly, and an explicit list of exclusions.

Section 10 is the one buyers skip and then regret. If every proposal expresses effort in engineer-days by role, three quotes become comparable in ten minutes instead of unanswerable.

How to decode a quote and compare pricing models

Divide the build price by the engineer-days to get an implied day rate. Then compare it to what the same work costs you internally.

The US Bureau of Labor Statistics reports a median annual wage of $135,980 for software developers as of May 2025, which is about $65 an hour across 2,080 hours. In the BLS Employer Costs for Employee Compensation release for June 2026, "Wages and salaries averaged $32.82 and accounted for 70.0 percent of employer costs, while benefit costs averaged $14.07 and accounted for the remaining 30.0 percent" in private industry. Dividing $65 by 0.70 puts a loaded in-house cost near $93 an hour, or roughly $745 a day. The 70% share covers all private-industry workers, so treat it as an approximation, and note that engineers with production LLM experience typically sit above the median.

A vendor's day rate is legitimately higher than that, because it also covers management, quality assurance, bench time, tooling and margin. What matters is that you can see the multiple and ask what it buys. An implied rate several times your loaded internal cost needs a reason, and "senior AI expertise" is not one unless the named seniors are the ones writing the code.

Pricing model Works when How it fails What to require
Time and materials Scope is genuinely unknown; you have a technical owner watching burn Open-ended spend, no incentive to finish Weekly burn reports, a not-to-exceed cap per phase
Capped time and materials The common case for a first production agent Cap becomes the estimate, and quality is trimmed to reach it Acceptance criteria in the contract, plus a written change process
Fixed price per phase Discovery and proof of value, where scope is small and bounded Fixed price for the whole build hides padding or produces change orders One fixed price per phase, re-quoted at each gate
Monthly retainer Run, support and continuous improvement after launch Drifts into paying for availability rather than output A defined scope of work per month and a quarterly review
Outcome or per-task pricing Mature, high-volume, well-instrumented workflows Attribution disputes; you cannot audit their cost base Agreed measurement source, a floor and a ceiling

For most first projects the sequence that behaves best is a fixed-price discovery, a fixed-price proof of value, then capped time and materials for the build with a gate at each phase. Our AI agent development cost guide breaks down build and run budgets by complexity, and the timeline guide shows where the weeks actually go, so you can sanity-check both halves of a quote.

Always ask what a quote excludes. The usual omissions are model and API spend, cloud infrastructure, third-party connector licenses, the security review, data cleanup, and anything called a change request.

Run a paid proof of value on your own data

This is the stage that changes decisions, and the stage most buyers skip because it looks like a delay. It is the opposite: two weeks here routinely prevents four wasted months.

A six-stage vendor selection process across seven weeks, showing what you do at each stage and how many vendors remain, from eight to twelve down to one
Stage 5 is the only stage that tests work instead of claims. Budget for it deliberately.

The design:

  • Two finalists, the same brief, the same two weeks. Pay both a small fixed fee. You own whatever is produced, in writing, before day one.
  • Real data in a controlled environment. M-25-22 tells federal buyers to seek "detailed demonstrations and tests of potentially useful AI systems or services in scenarios that closely reflect the intended real-world operating environment." Synthetic data hides exactly the problems you are shopping for.
  • A test set you build and never share. Fifty to a hundred real cases with agreed correct answers, including the ugly ones. Assemble it before the vendors start. This is your acceptance instrument for the whole project, not just the trial.
  • Fixed measurements for both. Task success on the held-back set, cost per completed task, median and worst-case latency, escalation rate, and how a failure surfaces. Anthropic's guidance recommends "extensive testing in sandboxed environments, along with the appropriate guardrails" for autonomous agents, and a two-week trial is the cheapest place to find out whether a vendor works that way by habit.
  • One shared debrief. Each finalist presents what they learned, what they would change, and what they now think it costs. The gap between the original proposal and this number tells you how the vendor estimates.

A working prototype in one to two weeks is a realistic bar in 2026, not an aggressive one. Fleurant AI builds prototypes on real data in that window and offers a risk-free proof of value for AI agent projects, precisely because it is the cheapest way for both sides to find out the truth early.

Key takeaway: The held-back test set is the most valuable artifact you will create during vendor selection. It survives the trial, becomes your acceptance criteria, and later becomes your regression suite.

Contract terms to settle before the statement of work

M-25-22 requires federal contracts for AI systems and services to address a specific list: intellectual property rights and use of government data, privacy, vendor lock-in protections, risk-management compliance, ongoing testing and monitoring, vendor performance requirements, and notification of new AI features. That list is a gift to commercial buyers, because it was written by people who have been burned. Here it is translated.

Term What to require Why it matters
IP assignment Written assignment of source code, system prompts, tool definitions, retrieval configuration, infrastructure as code and the evaluation set; delivery into your repository These are the asset. Vendors sometimes classify prompts and eval sets as proprietary
Use of your data No use of your inputs or the system's outputs to train or improve any model or product without your explicit written consent M-25-22 has agencies "permanently prohibit the use of non-public inputted agency data and outputted results to further train publicly or commercially available AI algorithms ... absent explicit agency consent"
Third-party components A list of open-source and third-party components with licenses, and disclosure of AI coding tools used Licenses ride along with the code you now own
Portability and exit Knowledge transfer, data and model portability, clear licensing, pricing transparency, a defined transition assistance period, export formats, and a data deletion certificate M-25-22 names "knowledge transfer, clear data and model portability practices, clear licensing terms, and pricing transparency" as lock-in protections
Where it runs Production in your cloud account, under your identity provider, with no proprietary runtime you cannot self-host Determines whether switching vendors is a project or a rebuild
Testing and monitoring rights Your right to evaluate the system on your own data on a stated cadence, and to see the vendor's test procedures and results M-25-22 requires contracts to "detail the examination, testing, and validation procedures of the vendor" and not to prohibit internal disclosure of results
Performance and rollback Performance standards met before a new version deploys, and rollback if a new version fails them M-25-22 has agencies require vendors "to meet performance standards before deploying a new version ... or to roll-back to a previous version"
Change notification Advance written notice before changing the model, the provider, or the agent's level of autonomy A silent model swap can change behavior overnight
Security artifacts SOC 2 report with its scope and period, subprocessor list, penetration test summary, and mapped controls for prompt injection and excessive agency A SOC 2 covers controls "relevant to Security, Availability, Processing Integrity, Confidentiality, or Privacy" per AICPA; a Type 2 covers operating effectiveness over a period, a Type 1 only design at a point in time
Liability for wrong outputs Named responsibility and a cap for losses caused by an incorrect automated action, plus required human checkpoints Without explicit language, the cost of an agent's mistake tends to land on you

Two specifics deserve their own paragraphs.

The AI-generated-code wrinkle. If your vendor's engineers generate a large share of the code with AI assistants, "you own all the IP" is a slightly more complicated promise than it sounds. The US Copyright Office's Part 2 report on copyrightability, released January 29, 2025, concluded that outputs of generative AI can be protected by copyright "only where a human author has determined sufficient expressive elements," that protection does not extend to "the mere provision of prompts," and that "the use of AI to assist in the process of creation or the inclusion of AI-generated material in a larger human-generated work does not bar copyrightability." The practical response is not to ban AI-assisted development, which would be unenforceable and unwise. It is to stop relying on copyright alone: require assignment of all rights that exist, plus confidentiality and trade-secret protection, plus actual delivery of the source and the right to modify it. Confirm the wording with your own counsel.

Regulated data changes the shortlist, not just the paperwork. If protected health information is involved, your vendor is a business associate and needs a business associate agreement. HHS lists a "Third-party vendor Artificial Intelligence (AI) chatbot on a provider's patient portal that provides services involving the patient's PHI" among its examples of business associates, and notes that a business associate "is also directly liable for complying with certain provisions of the HIPAA Rules." Banks and credit unions buying AI have their own governance expectations; our guide to SR 26-2 model risk management covers what changed in April 2026. For a vendor-neutral structure to ask about, the NIST AI Risk Management Framework (NIST AI 100-1, released January 26, 2023) organizes AI risk work into Govern, Map, Measure and Manage, and a vendor who can describe their practice in those terms is usually one who has faced an auditor. Our AI compliance work starts from the same framework.

Red flags, and how to verify instead of trust

Regulators have already made the point that AI claims are claims like any other. Announcing Operation AI Comply on September 25, 2024, the FTC said through then-Chair Lina Khan that "Using AI tools to trick, mislead, or defraud people is illegal" and that "there is no AI exemption from the laws on the books." In March 2024 the SEC charged two investment advisers with false and misleading statements about their use of AI, settling for $400,000 in combined penalties; then-Chair Gary Gensler's statement was that advisers "should not mislead the public by saying they are using an AI model when they are not."

You are not going to bring an enforcement action against a development shop. But you can borrow the posture: ask for the artifact, not the assertion.

Claim Artifact to request
"We have built dozens of production agents" One live system, screen-shared, with its dashboard and last week's error log
"It is 95% accurate" The test set, the scoring method, and the date it was last run
"Enterprise-grade security" SOC 2 report with scope and period, subprocessor list, pen test summary
"You own everything" The draft IP clause, plus the deliverables list, before the SOW
"Our senior team will handle this" Names, and those names on the next call
"We are partners with [platform]" The partner listing in that platform's own directory
"We reduced their handle time 40%" The baseline, the measurement window, and who measured it

The red flags that most reliably predict trouble:

  • A fixed price before anyone has seen your data or systems. Either a guess or a plan for change orders.
  • A demo video instead of a live run. Ask for a real input you supply, live, with logs visible.
  • A proprietary runtime you cannot self-host. Portability is not a feature request; it is the exit door.
  • Fine-tuning proposed before retrieval or plain prompting has been tried. Anthropic's own guidance is to add complexity "only when it demonstrably improves outcomes."
  • No evaluation plan. If they cannot say how they will measure it, they cannot say when it is done.
  • Case-study percentages with no baseline. A number without a measurement method is decoration.
  • Senior names in the pitch, unnamed team in the SOW. Fix it in the contract or walk.
  • Reluctance to write anything about data use. This is a five-minute clause. Hesitation is the answer.

When not to hire an outside firm at all

The honest version of this advice includes the cases where the answer is no.

The workflow is not documented and the data does not exist yet. Spend six weeks writing down how the work is actually done and getting the data into one place. Any vendor would bill you to do this, and you will do it better.

Nobody owns the decision. A project with a committee instead of an owner will stall at the security review regardless of who builds it.

A product you already pay for does this. Check your existing help desk, CRM, ERP and document platform first. Buying a feature beats building an agent when the feature is close enough.

The volume does not support the payback. Twenty items a day at four minutes each is about 1.3 hours of work. Build costs do not shrink to match. Do the arithmetic before the RFP, not after.

The regulatory decision is not made. If your compliance team has not decided whether this class of decision may be automated, a build will sit finished and unused.

You have good engineers and just need focus. Two of your own developers, given two weeks and a narrow scope, often beat a procurement cycle. Consider buying three days of architecture review instead of a project, then decide.

If any of these apply, a short paid assessment is usually the better purchase. It is also a legitimate way to test a vendor: how good is their advice when they are not selling a build?

What good looks like, and what to do next

A vendor worth hiring will tell you your scope is too big, ask about your data before your budget, name the engineer who will do the work, propose measuring correctness before proposing an architecture, and agree in writing that you own the prompts. They will quote in engineer-days, list what is excluded, and treat a two-week proof of value as normal rather than insulting.

Run the process in this order and it takes about seven weeks: the one-page brief, written screening, a technical call with an engineer present, the 30-point scorecard, two reference calls per finalist, a paid proof of value against a held-back test set, then contract terms before the statement of work. The stages that feel like overhead are the stages that produce evidence.

If you want a second opinion on a proposal you already have, or a sounding board for your requirement brief, talk to a specialist. Our AI agent development team will tell you if what you are buying should be a workflow, a product subscription or nothing at all, and our custom AI and MCP integration guides cover the pieces most quotes leave out. A discovery call is free, and a specialist replies within one business day.

Frequently asked questions

What should I ask an AI agent development company before signing?

Ask for one agent they have running in production today, with its daily task volume and cost per task. Ask how they would measure correctness on your workflow and what score they would refuse to launch below. Ask who exactly will write the code, what happens to your data, and what they hand over at the end. Answers in writing, not a deck.

Should I hire a freelancer, a general dev agency, or an AI specialist firm?

Match the vendor type to the risk. A contractor is fine for a prototype or an internal read-only assistant. A specialist firm earns its premium when an agent writes to a system of record, touches regulated data, or needs an evaluation harness. A general web agency without anyone who has shipped and measured an LLM system is the most common expensive mistake.

Who owns the code, the prompts and the evaluation set?

You should, and the contract has to say so explicitly for each one. System prompts, tool definitions, retrieval configuration and the test set are the assets that make the agent work, and some vendors treat them as proprietary. Require written assignment, delivery of source into your repository, and a list of any third-party or open-source components with their licenses.

How long should it take to choose a vendor?

About seven weeks if you run it properly: one week to write the brief, one to screen written answers, one for technical calls and references, two to three for a paid proof of value with two finalists, and one to settle contract terms. Compressing it below three weeks usually means you are buying a deck instead of a capability.

Do I need a SOC 2 report from an AI development vendor?

Ask for one whenever the vendor will hold or process your data. A SOC 2 Type 2 report covers the design and operating effectiveness of controls over a period, which is more useful than a Type 1 snapshot. Read the scope and the exceptions, not just the cover page, and check that the systems in scope are the ones your project will actually use.

Is a paid proof of value worth the money?

In most cases yes, because it is the only stage that tests work rather than claims. Two finalists building the same thing for two weeks on your real data, scored against a test set you never showed them, costs a fraction of a production build and routinely changes the decision. Insist that you own whatever is produced.

What is the single biggest red flag?

A fixed price offered before anyone has looked at your data or your systems. It means the vendor is either guessing or planning to recover the difference through change orders. The second biggest is a demo video in place of a live run on real inputs, with logs visible.

Sources

  1. M-25-22, Driving Efficient Acquisition of Artificial Intelligence in Government, Office of Management and Budget
  2. Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, Gartner
  3. FTC Announces Crackdown on Deceptive AI Claims and Schemes, Federal Trade Commission
  4. SEC Charges Two Investment Advisers with Making False and Misleading Statements About Their Use of Artificial Intelligence, U.S. Securities and Exchange Commission
  5. Copyright Office Releases Part 2 of Artificial Intelligence Report, U.S. Copyright Office
  6. AI Risk Management Framework, National Institute of Standards and Technology
  7. OWASP Top 10 for LLM Applications 2025, OWASP Gen AI Security Project
  8. Building Effective AI Agents, Anthropic
  9. Software Developers, Quality Assurance Analysts, and Testers, U.S. Bureau of Labor Statistics
  10. Employer Costs for Employee Compensation, June 2026, U.S. Bureau of Labor Statistics
  11. System and Organization Controls: SOC Suite of Services, AICPA & CIMA
  12. Business Associates, U.S. Department of Health and Human Services

Free, no-obligation consultation

Have a question about hiring an AI agent development company?

Tell us what you're working on or what you'd like to know. A specialist will get back to you with practical next steps, whether or not we end up working together.

  1. 1Send your question or project details (takes 2 minutes)
  2. 2A specialist reviews it and replies within 1 business day
  3. 3Get clear, practical next steps, free