Custom AI
RAG Chatbot Development Cost: Build, Run and Accuracy
What does RAG chatbot development cost? A build model in engineer-days, monthly run costs from published API prices, and the line item most budgets miss.
In this article
- What you are actually paying for
- What does RAG chatbot development cost to build?
- Why knowledge base size barely moves your infrastructure bill
- What a RAG chatbot costs to run each month
- What the retrieval layer really costs
- The line item most budgets miss
- Build, buy, or use a managed RAG service?
- What makes one quote three times another
- When not to build one
- A budget you can defend
- Where this leaves you
- Frequently asked questions
- Sources
RAG chatbot development cost splits into three honest tiers, and none of them are mostly about infrastructure. An internal pilot that answers questions over one document source is roughly 15 to 33 engineer-days. A production assistant with several synced sources, permission-aware retrieval, citations and a working evaluation harness is roughly 65 to 150 engineer-days. A regulated or customer-facing build is roughly 100 to 235. Multiply by your own fully loaded day rate and you have a budget you can defend.
The running costs are smaller than almost every guide on this subject claims. At published prices, embedding a 10,000-document knowledge base costs under a dollar, and serving 5,000 conversations a month costs between $8 and $825 in model tokens depending entirely on which model you choose. The expensive part is the part nobody quotes: the engineering attention required to make the answers correct and keep them correct.
This article gives you the line items, the arithmetic behind every number, and the assumptions each one rests on. Every price links to the vendor's own pricing page and carries a date, because they all move.
What you are actually paying for
Retrieval-augmented generation, usually shortened to RAG, is a simple idea wrapped in intimidating vocabulary. Instead of hoping a language model already knows your refund policy, you search your own documents for the passages most likely to answer the question, put those passages into the prompt, and ask the model to answer using only that material.
That means a RAG chatbot is four systems, not one:
- Ingestion. Getting documents out of SharePoint, Confluence, Zendesk, a file share or a database, and keeping the copy current as they change.
- Indexing. Splitting documents into chunks, converting each chunk to a numeric vector with an embedding model, and storing those vectors somewhere searchable.
- Retrieval. Turning a user question into a search, filtering results by what that user is allowed to see, and ranking the survivors.
- Generation. Handing the top passages to a language model with instructions about citing sources, admitting uncertainty, and when to escalate to a human.
Key takeaway: Almost everyone budgets for steps 2 and 4, because those are the ones with public price lists. The overruns live in steps 1 and 3, where the work is integration and judgment rather than API calls.
What does RAG chatbot development cost to build?
The only durable way to price a build is in engineer-days, because dollar quotes bundle a day rate you cannot see. Convert to money with a rate you control.
Converting engineer-days to dollars
The U.S. Bureau of Labor Statistics reports a median annual wage of $135,980 for software developers as of May 2025. Wages are not the whole cost of an employee: the BLS Employer Costs for Employee Compensation release for June 2026, published on September 9, 2026, puts wages and salaries at 70.0 percent of total compensation for private industry workers, with benefits making up the other 30.0 percent.
That gives a fully loaded annual cost of about $194,000, and across roughly 260 working days a year, about $750 per engineer-day for a median US developer employed directly. A senior specialist costs more. A US agency or consultancy bills more again, because its rate has to carry recruiting, management, bench time and profit. An offshore team bills less and usually needs more days.
Use whichever rate matches how you will actually staff the work, and apply it consistently to every tier below.
The three tiers
| Tier | What it is | Engineer-days | At $750/day |
|---|---|---|---|
| 1. Internal pilot | One document source, read-only, small internal audience, no per-user permissions | 15–33 | $11,000–$25,000 |
| 2. Production assistant | Three to five synced sources, permission-aware retrieval, citations, evaluation harness, monitoring, human handoff | 65–150 | $49,000–$113,000 |
| 3. Regulated or customer-facing | Tier 2 plus PII handling, identity integration, audit logging, security review, adversarial testing, governance documentation | 100–235 | $75,000–$176,000 |
The line items behind tier 2, which is where most companies land:
| Work | Engineer-days |
|---|---|
| Discovery, question inventory, success criteria | 4–8 |
| Connectors for three to five sources with incremental sync | 10–25 |
| Permission-aware retrieval (per-user access filtering) | 8–20 |
| Chunking strategy, metadata, reranking, tuning | 8–18 |
| Answer generation, citations, refusal behavior, escalation | 6–12 |
| Evaluation harness, labeled test set, regression runs | 8–18 |
| Observability, feedback capture, analytics | 5–10 |
| User interface or integration into an existing tool | 8–18 |
| Security hardening and deployment | 5–12 |
| Content owner and admin workflows | 3–8 |
Two of those rows deserve attention, because they are the ones vendors quietly leave out. Permission-aware retrieval means the assistant must never surface a document the asker could not have opened themselves, which requires carrying access control lists through ingestion and applying them as search filters. The evaluation harness is what tells you whether a change made answers better or worse. Without it, every adjustment is a guess.
What this does not include
Elapsed time is longer than engineering time. A pilot is commonly three to six weeks of calendar time, because getting sanctioned access to the documents sits in someone else's queue. A tier 2 build commonly runs three to five months. Regulated builds add a security review and a governance cycle on top.
Why knowledge base size barely moves your infrastructure bill
Competing guides frequently claim that knowledge base size is the single largest cost variable. At published prices, that is close to backwards. Here is the arithmetic.
Assumptions: an average document of about 1,500 words, or roughly 2,000 tokens; chunks of 500 tokens with 15 percent overlap; embeddings of 1,536 dimensions stored as 4-byte floats; English text at roughly 4 characters per token.
| Documents | Tokens | Chunks | Raw text | Vector data | One-time embedding cost |
|---|---|---|---|---|---|
| 1,000 | 2M | 4,600 | 8 MB | 28 MB | $0.05 |
| 10,000 | 20M | 46,000 | 80 MB | 283 MB | $0.46 |
| 100,000 | 200M | 460,000 | 800 MB | 2.8 GB | $4.60 |
| 1,000,000 | 2,000M | 4,600,000 | 8 GB | 28 GB | $46 |
Embedding costs use OpenAI's text-embedding-3-small at $0.02 per 1M tokens as of September 2026, applied to the token count including chunk overlap. The larger text-embedding-3-large is $0.13 per 1M tokens, and the older text-embedding-ada-002 is $0.10, so even the more expensive options put a million-document corpus in the low hundreds of dollars, once.
Re-embedding is just as cheap. If a tenth of a 10,000-document corpus changes every month, that is about 2.3M tokens, or five cents.
Key takeaway: Multiplying your knowledge base by a hundred multiplies your embedding bill by a hundred, and a hundred times almost nothing is still almost nothing. What a large corpus genuinely costs you is engineer-days, because retrieval quality gets harder as the haystack grows.
That second sentence is the real relationship. In a 1,000-document corpus, straightforward search usually works. In a 500,000-document corpus with a decade of superseded policies in it, you will spend real time on metadata filters, date-based ranking, deduplication and reranking, and on deciding which documents should never have been indexed at all.
What a RAG chatbot costs to run each month
Take a reference workload: 10,000 documents, 5,000 conversations a month, three turns per conversation, so 15,000 model calls. Assume each call sends about 4,000 input tokens (six retrieved chunks of roughly 500 tokens, a system prompt, and recent conversation history) and produces about 300 output tokens.
At published prices as of September 2026:
| Model | Input / output per 1M tokens | Cost per turn | 15,000 turns |
|---|---|---|---|
| GPT-6 Luna | $0.10 / $0.50 | $0.00055 | $8 |
| GPT-4o-mini | $0.15 / $0.60 | $0.00078 | $12 |
| Claude Haiku 4.5 | $1.00 / $5.00 | $0.0055 | $83 |
| o4-mini | $1.10 / $4.40 | $0.0057 | $86 |
| Claude Sonnet 5 | $2.00 / $10.00 | $0.011 | $165 |
| GPT-6 Sol | $2.00 / $10.00 | $0.011 | $165 |
| Claude Opus 5.5 | $4.00 / $20.00 | $0.022 | $330 |
| GPT-6 Astra | $10.00 / $50.00 | $0.055 | $825 |
Other providers sit inside this range rather than outside it, and several publish promotional rates with an expiry date attached, so re-check the figure you budgeted against before your renewal.
Three practical notes on this table.
The spread is about a hundredfold, and model size is not the same as answer quality. In grounded question answering, where the correct answer is already in the prompt, a small model frequently matches a large one. Test both on fifty of your own real questions before you commit to a price point.
Prompt caching helps less in RAG than elsewhere. Anthropic prices cache reads at $0.20 per 1M tokens for Sonnet 5 against $2.00 for ordinary input, and OpenAI similarly discounts cached input. But caching only pays when a long prefix repeats, and in RAG the retrieved passages change with every question. You can cache the system prompt and any fixed instructions; you usually cannot cache the evidence.
Batch processing does not fit chat. Anthropic and OpenAI both discount asynchronous batch work by around half, which is excellent for re-indexing or bulk evaluation runs and useless for a person waiting for an answer.
What the retrieval layer really costs
For the same 10,000-document reference workload, about 283 MB of vectors, 80 MB of raw text and 15,000 retrievals a month, here is what the common options charge as of September 2026:
| Option | How it is priced | Reference workload |
|---|---|---|
| Pinecone Starter | Free up to 2 GB storage, 1M read units and 2M write units per month | $0 |
| Pinecone Builder | $20/month flat, up to 10 GB storage | $20 |
| Pinecone Standard | $50/month minimum, then $0.33/GB/month, $16–18 per 1M read units, $4–4.50 per 1M write units | $50 |
| Amazon Bedrock Knowledge Bases | $5.00 per GB of raw data per month, plus $1.00 per 1,000 Retrieve API calls | About $15 |
| OpenAI File Search | $0.10 per GB per day with 1 GB free, plus $2.50 per 1,000 tool calls | About $38 |
| Amazon OpenSearch Serverless | $0.24 per OCU-hour, minimum 2 OCUs for a classic collection, $0.02 per GB/month storage | About $350 |
| Self-hosted pgvector | A Linux virtual server you administer yourself | $24–$44 |
Some arithmetic worth showing. OpenSearch Serverless bills a minimum of 2 OCUs for the first classic collection, which is 2 × 730 hours × $0.24 = $350.40 a month before a single query. AWS documents a dev-test deployment mode without redundant standby nodes that halves that, and NextGen collections that have no minimum and scale to zero after ten minutes of inactivity. Which mode you pick matters more than how much data you store.
The self-hosted figure uses Amazon Lightsail's fixed bundles, $24 a month for 4 GB of RAM and 2 vCPUs or $44 for 8 GB, running PostgreSQL with the pgvector extension yourself. It is genuinely the cheapest line on the table and genuinely not free, because you now own patching, backups and the 2 a.m. page.
Key takeaway: Your retrieval bill at this scale is set by which product you choose, not by how much data you have. The gap between $0 and $350 a month is a procurement decision.
The line item most budgets miss
Add up the reference workload honestly. A mid-range model at $165, a managed retrieval layer at $15, application hosting and logs at perhaps $100. Call it $280 a month of infrastructure.
Now add the person. One day a week of a median US software developer, at the $750 fully loaded day rate derived above, is about $3,200 a month, more than ten times the infrastructure. That day goes on document owners changing things without telling you, a source system's API changing, a question category the assistant handles badly, a model version being deprecated, and reviewing what people actually asked last week.
This is the most useful thing to know when reading vendor quotes. A proposal that details token pricing to four decimal places and says nothing about who owns the system in month seven is describing the cheap part of the problem in expensive detail.
Accuracy is a budget line, not a phase
The work that makes answers trustworthy is measurable, so it can be planned. Microsoft's documentation for the RAG evaluators in Azure AI Foundry is a good public description of what "measured" means in practice, and it splits the problem in two:
- Process evaluation asks whether retrieval found the right passages. The Document Retrieval evaluator computes search-quality metrics including fidelity, NDCG and XDCG against human relevance labels.
- System evaluation asks whether the final answer was any good. The Groundedness evaluator measures whether the response sticks to the retrieved context without fabricating content, Relevance measures whether it addressed the question, and Response Completeness measures whether it left out something important. These score from 1 to 5, with a default pass threshold of 3.
You do not have to use that product to use that structure. The point is that "the answers seem good" is not a measurement, and without a labeled test set you cannot tell whether last Thursday's chunking change helped or hurt.
Retrieval improvements are cheap to buy and expensive to skip. Anthropic's published work on contextual retrieval reports that adding generated context to each chunk before embedding cut the top-20 retrieval failure rate by 35 percent, that combining contextual embeddings with contextual BM25 keyword search cut it by 49 percent, and that adding a reranking step brought the total reduction to 67 percent. The stated one-time cost to generate those contextualized chunks is $1.02 per million document tokens, about $20 for our 10,000-document corpus.
Twenty dollars and a few engineer-days for a two-thirds reduction in retrieval failures is the best trade in this entire article.
For regulated deployments, the governance layer is a separate budget again. The NIST AI Risk Management Framework, released on January 26, 2023, organizes that work into four functions — Govern, Map, Measure and Manage — and NIST published a Generative AI Profile, NIST-AI-600-1, on July 26, 2024. If your assistant touches customer financial data, protected health information or claims decisions, factor in the AI compliance work before scoping, and confirm specifics with your own counsel or compliance team. Nothing here is legal advice.
Build, buy, or use a managed RAG service?
There are three real options, and the cheapest one depends almost entirely on volume.
Buy a product. For customer support specifically, priced-per-outcome products exist. Intercom prices its Fin agent from $0.99 per resolution, counted when a customer confirms their issue is resolved, does not ask for more help, or Fin completes a workflow, plus seats from $29 to $132 per month depending on plan. At 5,000 resolutions a month that is about $4,950.
Use a managed RAG service. Amazon Bedrock Knowledge Bases, Azure AI Search, OpenAI File Search and similar services handle parsing, chunking, embedding and retrieval for you. You still write the application, the permissions and the evaluation, but you skip a meaningful chunk of tier 1 and some of tier 2.
Build custom. You own every layer, which matters when your permissions model is unusual, your documents are unusual, your data cannot leave a particular boundary, or the assistant needs to take actions rather than only answer.
The break-even math
Compare buying at $0.99 per resolution against a tier 2 custom build of 65 to 150 engineer-days ($49,000 to $113,000) with a realistic $3,500 a month to run, including the one day a week of engineering.
At 5,000 resolutions a month, buying costs about $4,950 monthly and building costs about $3,500 monthly. You save roughly $1,450 a month, so even the cheapest custom build takes about 34 months to pay back. Buying wins, and it is not close.
At 50,000 resolutions a month, buying costs about $49,500 monthly. A custom system serving the same volume needs about 150,000 model calls, which is roughly $1,650 in tokens on a mid-range model, plus retrieval, hosting and the same engineering attention. Call it $5,300 a month. You save roughly $44,000 a month, so even the expensive end of a custom build pays back in under three months.
Key takeaway: Per-resolution pricing scales linearly with volume and a custom system's cost does not. Find your own crossover point before you argue about architecture.
Two honest caveats. A product's price buys you continuous improvement you would otherwise fund yourself, so this comparison flatters custom builds slightly. And volume is not the only axis: a 500-conversation-a-month assistant that answers underwriting questions correctly can be worth more than a 50,000-conversation one that deflects password resets.
What makes one quote three times another
When two proposals for "a RAG chatbot" differ by a factor of three, it is almost always one of these six things. Ask about each explicitly.
- Permissions. Does every user see the same documents, or is retrieval filtered by that person's actual access rights? This single question can double a build.
- Number and quality of sources. One clean document library is not five systems with different authentication, different formats, and no reliable change feed. Scanned PDFs and spreadsheets cost more than clean HTML.
- Freshness. A one-time index is a fraction of the cost of incremental sync with deletions handled correctly.
- Evaluation. Is there a labeled test set and a regression run, or does "testing" mean someone tried twenty questions before the demo?
- Escalation and actions. Answering is cheaper than doing. The moment the assistant creates tickets or updates records, you are building an agent, and the cost model in our guide to AI agent development cost applies instead.
- Who owns month seven. Maintenance either appears as a retainer or appears as a surprise.
When not to build one
Three situations where the honest answer is to spend the money elsewhere.
Your documents are wrong. RAG retrieves what exists. If three versions of the expense policy are in circulation and nobody knows which is current, a chatbot will confidently cite the wrong one faster than a human would. Fix the source of truth first. That project is often cheaper and always more valuable.
The questions have a deterministic answer. "What is my PTO balance?" is a database query. Wrapping a language model around a lookup adds latency, cost and a small chance of a wrong number. A well-placed search box beats a chatbot more often than anyone likes to admit.
Nobody owns it. An assistant with no named owner degrades quietly. Documents change, retrieval drifts, trust erodes after two bad answers, and usage goes to zero months before anyone cancels the contract.
A budget you can defend
Bring these twelve lines to your next vendor conversation. If a proposal cannot fill them in, it is not a proposal.
- Engineer-days for ingestion, by source, with the format of each source stated.
- Engineer-days for permission-aware retrieval, or an explicit statement that all users see everything.
- Engineer-days for the evaluation harness, and the size of the labeled test set.
- The day rate being applied, and whether it is blended.
- Elapsed weeks, separating engineering from access approvals and reviews.
- Chosen model and the per-turn token cost at your expected context size.
- Expected conversations a month and turns per conversation, with the source of that estimate.
- Retrieval product and its pricing basis, including any monthly minimum.
- Re-indexing frequency and its cost.
- Monthly hours of ongoing engineering, and who supplies them.
- What happens when a model version is deprecated.
- Who owns the code, the index and the evaluation data.
The first three lines catch most underquoting. Line 10 catches most of the rest.
Where this leaves you
RAG chatbot development cost is dominated by engineer-days, not by tokens or vector storage. Embedding is nearly free, retrieval infrastructure ranges from nothing to a few hundred dollars a month at ordinary scale, and model spend is a choice you make with a hundredfold range in it. What costs real money is connecting to messy systems, respecting permissions, and proving the answers are right, then keeping them right.
If you are deciding between a pilot and a production build, run the pilot first and set a measurable bar for it. If you are weighing whether to keep company data inside your own boundary, our comparison of private LLM deployments and ChatGPT Enterprise covers the data-handling side of the same decision, and our guide to how long it takes to build an AI agent covers the timeline question this article only touches. If retrieval alone does not close the quality gap and someone suggests training a model instead, our line-by-line breakdown of custom AI model development cost prices that path before you commit to it.
Fleurant AI builds custom AI and generative AI systems, including retrieval assistants over company data, with a free discovery call and a working prototype in one to two weeks. If you would like a second opinion on a quote or a scope, talk to a specialist and you will hear back within one business day.
Frequently asked questions
How much does RAG chatbot development cost?
An internal pilot over one document source is roughly 15 to 33 engineer-days. A production assistant with several synced sources, permission-aware retrieval, citations and an evaluation harness is roughly 65 to 150. A regulated or customer-facing build is roughly 100 to 235. Multiply by your own fully loaded day rate, or by the day rate hidden inside a vendor quote, to convert engineer-days into money.
What does it cost per month to run a RAG chatbot?
For 10,000 documents and 5,000 conversations a month, the infrastructure is small: roughly $8 to $825 in model tokens depending on which model you pick, plus $0 to $350 for retrieval. The larger recurring cost is engineering attention. One day a week of a US software developer is about $3,200 a month using BLS wage data, which is more than the whole infrastructure bill.
Does a bigger knowledge base cost more?
Far less than most guides claim. Embedding 10,000 documents with OpenAI's text-embedding-3-small costs under $1 at published prices, and going to 1,000,000 documents costs about $46. What a larger corpus really costs you is engineering: retrieval quality degrades as the corpus grows, so you spend more days on chunking, metadata, filtering and reranking.
Is it cheaper to buy a chatbot product than to build one?
Usually yes at low volume. Intercom's Fin agent is priced from $0.99 per resolution, so 5,000 resolutions a month is about $4,950. A custom production build of 65 to 150 engineer-days would take years to pay that back at that volume. The math flips at high volume, because per-resolution pricing scales linearly and a custom system's run cost does not.
What is the most common reason a RAG budget overruns?
Accuracy work. Getting a demo to answer well takes a few days. Getting it to answer correctly, refuse when it should, and keep doing both as documents change requires a labeled test set, retrieval metrics and regression runs. Teams that skip the evaluation harness usually pay for it later in rework and lost trust.
Do I need a vector database?
Not always. For a corpus under a few hundred thousand chunks, PostgreSQL with the pgvector extension on a server you already run is often enough, and managed options such as Amazon Bedrock Knowledge Bases charge $5.00 per GB of raw data per month plus $1.00 per 1,000 retrieval calls. Dedicated vector databases earn their price at larger scale or when you need their specific filtering and hybrid search features.
How long does a RAG chatbot take to build?
A pilot is commonly three to six weeks of elapsed time, because getting access to the documents takes longer than the engineering. A production assistant is commonly three to five months. Regulated builds add a security review and a governance cycle, which are queue time more than work time.
Should we use a cheap model or an expensive one?
Test both on your own questions before deciding. At 4,000 input and 300 output tokens per turn, published prices put a small model near $0.0006 per turn and a flagship near $0.055, roughly a hundredfold spread. In grounded question answering, a well-retrieved context often matters more than model size, so the cheap model wins more often than teams expect.
Sources
- API Pricing, OpenAI
- Pricing, Anthropic
- Introducing Contextual Retrieval, Anthropic
- Pinecone Pricing, Pinecone
- Amazon OpenSearch Service Pricing, Amazon Web Services
- Amazon Bedrock Pricing, Amazon Web Services
- Amazon Lightsail Pricing, Amazon Web Services
- Retrieval-Augmented Generation (RAG) Evaluators, Microsoft Learn
- Software Developers, Quality Assurance Analysts, and Testers, U.S. Bureau of Labor Statistics
- Employer Costs for Employee Compensation, June 2026, U.S. Bureau of Labor Statistics
- Pricing, Intercom
- AI Risk Management Framework, National Institute of Standards and Technology