Custom AI

Custom AI Model Development Cost: A Line-by-Line Budget

What custom AI model development cost really looks like: engineer-days per tier, published training and hosting prices, and a three-year comparison.

Cover image for an article about the line-by-line cost of developing a custom AI model
In this article
  1. What people mean by "custom AI model" (and why it changes the price by 10x)
  2. What drives custom AI model development cost
  3. The day rate: turning engineer-days into dollars you can defend
  4. Line by line: what a fine-tuned hosted model actually costs
  5. The data line item: the 40 percent nobody budgets
  6. The evaluation line item: the part that decides whether you ship
  7. The line item that wrecks budgets: hosting a custom model
  8. Three-year total cost of ownership, compared
  9. The costs that start after launch
  10. When a custom model is worth it, and five signs it is not
  11. How to scope the project so the budget survives
  12. What to ask a vendor
  13. Putting it together
  14. Frequently asked questions
  15. Sources

Custom AI model development cost is a people number wearing a hardware costume. The GPUs get the attention, but in a typical mid-sized company project they account for less than one percent of the build. Priced in engineer-days, the honest tiers are 20 to 45 days to fine-tune a hosted model, 35 to 90 days to train a task-specific model on your own data, and 70 to 130 days to fine-tune open weights and run them yourself. At a fully loaded US day rate of about $750, derived below from federal wage data, that is roughly $15,000 to $98,000 of build effort.

The training compute inside those projects usually costs between $12 and $1,500. The monthly hosting can cost more than the entire training run, every month, indefinitely.

This article gives you each line item, the arithmetic behind it, and the published price it comes from. Every rate links to the vendor's own pricing page and carries a date, because they all move. By the end you should be able to read a $250,000 proposal and say exactly which line you are paying for.

What people mean by "custom AI model" (and why it changes the price by 10x)

Four very different projects hide behind the same phrase. Confusing them is the most expensive mistake in this category, because a company will budget for the fourth when it needs the first.

  1. Prompting plus retrieval. No training at all. You give a general model your documents and clear instructions. This is what most companies actually need, and it is covered in our guide to RAG chatbot development cost.
  2. Fine-tuning a hosted model. You upload examples to OpenAI, Azure or Amazon Bedrock, and they return a private version of a model that behaves the way your examples do. You never touch a GPU.
  3. Training a task-specific model. Fraud scoring, demand forecasting, defect detection, document routing. These are often gradient-boosted trees or small neural networks rather than language models, and they are frequently the cheapest way to get a large, measurable business result.
  4. Continued pre-training or training from scratch. Taking an open-weight model and teaching it a new domain at scale, or building a foundation model. This is a research program, not a project.
Decision tree showing which kind of custom AI model to build, from prompting with retrieval through fine-tuning a hosted model to self-hosted open weights
Work down the questions and stop at the first yes. Each step down the tree multiplies the budget, and most companies stop at box one or two.

The scale of the fourth option is worth stating plainly so you can rule it out with confidence. Stanford's 2026 AI Index reports that global AI compute capacity has grown 3.3 times per year since 2022, reaching 17.1 million H100-equivalents, and that the estimated training emissions of a single recent frontier model reached 72,816 tons of CO2 equivalent. That is the weight class you would be entering. Almost nothing in a mid-sized company's roadmap requires it.

Key takeaway: Decide which of the four projects you are funding before you accept any quote. A vendor quoting option four for a problem that option two solves is not overcharging. They are answering a different question.

What drives custom AI model development cost

Six things, in descending order of impact on the invoice.

  • How clearly the task is defined. A task with an agreed definition of "correct" can be evaluated automatically, which makes everything downstream fast. A task where three experts disagree about the right answer is an open-ended research bill.
  • Where the data lives and who owns it. Data already sitting in one system with clear usage rights is days of work. Data spread across a document management system, a mailbox and a vendor platform, governed by contracts that never mentioned AI, is weeks of work before anyone trains anything.
  • How much labeling you need. Labels created by a subject-matter expert are the most expensive data in the project, and usually the only labels that matter.
  • Where the model has to run. A hosted endpoint costs cents per thousand calls. A dedicated endpoint inside your own boundary costs thousands of dollars a month whether or not anyone uses it.
  • What has to be proved before launch. In regulated settings the evidence package can equal the engineering.
  • How often it has to be retrained. A model retrained quarterly needs a repeatable pipeline. A model trained once needs a script.

Notice what is not on that list: parameter count, GPU generation, and which model family you pick. Those decisions matter for quality. They barely move the budget.

The day rate: turning engineer-days into dollars you can defend

Every number in this article is expressed in engineer-days so you can apply your own rate. Here is the rate used throughout, built from federal data rather than assertion.

The U.S. Bureau of Labor Statistics reports a median annual wage of $135,980 for software developers and $120,230 for data scientists, both as of May 2025. Wages are not the whole cost of an employee. In the BLS Employer Costs for Employee Compensation release for June 2026, total compensation for private industry workers averaged $46.89 per hour worked, of which wages and salaries were 70.0 percent and benefits 30.0 percent.

Grossing the developer median up for benefits gives $135,980 divided by 0.70, or $194,257 a year. Divided by 260 working days, that is $747 a day, rounded here to $750. A data scientist on the same arithmetic is about $660 a day.

That is an internal cost. An agency or consultancy day rate in the US commonly lands between $900 and $2,000 depending on seniority and market, which is why the same 60-day project can be quoted at $45,000 or $120,000 without anyone being dishonest. When you compare quotes, divide by a day rate and compare day counts. That is the only apples-to-apples comparison available to you.

Line by line: what a fine-tuned hosted model actually costs

This is the most common custom model project in mid-sized companies, so it is worth walking through properly. Assume a document classification and extraction service: incoming PDFs are sorted into 14 categories, six fields are pulled out of each one, and the output has to be valid JSON every time.

Line item Low Mid High What moves it
Problem framing and evaluation design 4 8 14 How much experts disagree on correct output
Data collection, rights review, cleaning 6 14 30 Number of source systems, contract review
Labeling and expert review 3 10 28 Label count, ambiguity, review rounds
Training runs and hyperparameter work 3 6 12 Number of experiments, not model size
Evaluation harness and error analysis 4 9 20 Regulated review, edge-case hunting
Integration, deployment, monitoring 6 13 26 Systems touched, auth, rollback, logging
Total engineer-days 26 60 130
At $750 a day $19,500 $45,000 $97,500
Horizontal bar chart of a sixty engineer-day fine-tuning project by line item, with the GPU bill shown as a fraction of one percent
The mid-case build, drawn to scale. Training compute is $12 to $300 against about $45,000 of people time.

The bottom of that range is a realistic internal project with clean data and a tolerant use case. The top is a regulated build where every label is reviewed twice and the evaluation evidence has to survive an audit.

Where the compute actually lands

OpenAI publishes fine-tuning training prices per million tokens: $5.00 for gpt-4.1-mini, $25.00 for gpt-4.1, $1.50 for gpt-4.1-nano and $3.00 for gpt-4o-mini, as of September 2026.

Work the arithmetic for our example. One thousand training examples at roughly 800 tokens each is 800,000 tokens per pass. Three passes over the data is 2.4 million training tokens. On gpt-4.1-mini that is 2.4 multiplied by $5.00, or $12.00. Ten experiment runs while you tune the dataset still lands under $150.

Amazon Bedrock prices customization the same way. In US East (N. Virginia), as of September 2026, customization training is $0.001 per 1,000 tokens for Amazon Nova Micro, $0.002 for Nova Lite and $0.008 for Nova Pro, which is $1, $2 and $8 per million training tokens. Storing each resulting custom model costs $1.95 a month.

If you run the training yourself on rented GPUs, the numbers are still small for anything short of a full pre-train. At AWS on-demand rates in US East (N. Virginia) as of September 2026, a g6e.xlarge with one NVIDIA L40S costs $1.861 an hour, so a twelve-hour parameter-efficient fine-tune of an 8-billion-parameter model is about $22. A p5.48xlarge with eight H100 GPUs costs $55.04 an hour, so a twenty-hour full fine-tune of a 70-billion-parameter model on one of those is about $1,101. Spot capacity cuts both further, at the price of interruption.

Key takeaway: Across every one of these paths, training compute is between 0.03 percent and 3 percent of the build. If a proposal is organized around GPU cost, it is organized around the wrong thing.

The data line item: the 40 percent nobody budgets

Two thirds of the mid-case project above is data and evaluation work. Three things drive it.

Rights and provenance. Before a single example is used, someone has to confirm you are allowed to use it. Customer contracts signed before 2023 rarely contemplated model training. Vendor terms may prohibit using their outputs to train competing models. Personal data brings purpose-limitation questions. This is a lawyer-and-engineer exercise, and on regulated projects it is the most common cause of a three-week slip. Our overview of AI compliance work covers the governance side of that review.

Labeling. You have two options and most projects use both. Crowd labeling is cheap per item: on Amazon Mechanical Turk, requesters pay a 20 percent commission on worker rewards, with an additional 20 percent for tasks with 10 or more assignments and 5 percent more for the Masters qualification, at a minimum of $0.01 per assignment. But crowd workers cannot tell you whether a claim was correctly denied under your policy. Expert labeling is what most business tasks actually require, and at $660 to $750 a day it is the expensive half.

Volume, which matters less than people think. OpenAI's guidance states that "the minimum number of examples you can provide for fine-tuning is 10," recommends "starting with 50 well-crafted demonstrations and evaluating the results," and reports improvements from 50 to 100 examples. It also warns: "If 50 examples have no impact, rethink your task or prompt before adding training data."

That last sentence is worth a lot of money. The instinct when a fine-tune underperforms is to label more data. Frequently the real problem is that the task is underspecified, and 5,000 more examples of an ambiguous task produce an expensively confused model.

The evaluation line item: the part that decides whether you ship

OpenAI's fine-tuning guide opens with a blunt instruction: "Good evals first! Only invest in fine-tuning after setting up evals." That is not a formality. Without a test set you cannot tell whether the fine-tune helped, whether last week's change broke something, or whether the model has drifted three months from now. Teams that skip the harness do not save nine engineer-days. They move those days into rework later, with less information to work from.

A workable harness for the example project is 200 to 500 held-out documents labeled by experts, a script that runs every candidate model against them, accuracy and field-level error rates, and a stored record of every run. Budget four to twenty engineer-days depending on how much human review the results need.

In regulated settings this artifact is also your evidence. The NIST AI Risk Management Framework, released January 26, 2023, organizes the work into four functions, Govern, Map, Measure and Manage, and NIST published a Generative AI Profile, NIST AI 600-1, on July 26, 2024. Both are voluntary, but examiners and enterprise customers increasingly expect something shaped like them. Confirm specifics with your own counsel or compliance team. Nothing here is legal advice.

The line item that wrecks budgets: hosting a custom model

Training is a one-off. Hosting is permanent, and it is where custom models diverge sharply from general ones.

Hosted fine-tunes are billed per token, at a premium. OpenAI prices a fine-tuned gpt-4.1-mini at $0.80 per 1M input tokens and $3.20 per 1M output tokens, against $0.40 and $1.60 for the base model. Fine-tuning exactly doubles the unit price. On Bedrock, custom Amazon Nova Lite inference is priced at $0.00006 per 1,000 input tokens and $0.00024 per 1,000 output tokens, which is $0.06 and $0.24 per million.

Dedicated capacity is billed by the clock. This is the number that surprises people. Amazon Bedrock Provisioned Throughput for Nova Lite is $55.00 per model unit per hour on a one-month commitment in US East (N. Virginia), which is about $40,150 a month for a single unit that runs whether or not anyone sends it a request.

Imported open-weight models are billed by the minute. Bedrock Custom Model Import is priced at $0.05718 per Custom Model Unit per minute for Llama, Mistral, Mixtral, Flan and Qwen architectures in US East (N. Virginia), and $0.1433 for OpenAI open-weight architectures. AWS documents the formula as running model copies multiplied by Custom Model Units per copy, multiplied by the rate per unit per minute, multiplied by the number of five-minute billing windows divided by 60, charged from the first successful inference call. One unit running continuously is 0.05718 multiplied by 60 multiplied by 730, or about $2,504 a month.

Azure bills hosting separately from tokens. Microsoft states that "each customized (fine-tuned) model that's deployed incurs an hourly hosting cost regardless of whether chat completions or response API calls are made to the model," and that a deployment left inactive for more than 15 days is deleted, though the underlying custom model is retained and can be redeployed. Azure's Developer tier removes the hourly hosting fee but carries no availability SLA and is documented as being for candidate evaluation rather than production use.

Running it on your own instances is not free either. Two g6e.xlarge instances for basic redundancy at $1.861 an hour is about $2,717 a month, before load balancers, storage, logging, or the engineer who keeps them patched.

Hosting path Published rate (US East, September 2026) Cost if it runs continuously
Hosted fine-tune, pay per token $0.80 / $3.20 per 1M tokens (gpt-4.1-mini) About $144 a month at 100M in, 20M out
Bedrock custom model, on-demand $0.06 / $0.24 per 1M tokens (Nova Lite) About $17 a month at the same volume
Bedrock Custom Model Import $0.05718 per Custom Model Unit per minute About $2,504 a month per unit
Two EC2 g6e.xlarge instances $1.861 an hour each About $2,717 a month
Bedrock Provisioned Throughput $55.00 per model unit per hour, 1-month commitment About $40,150 a month

Key takeaway: The cheapest training path and the cheapest serving path are usually different decisions. Decide how the model will be served before you decide how it will be trained, because serving is the recurring number.

Three-year total cost of ownership, compared

Assume an internal service handling 200,000 requests a month at 500 input and 100 output tokens each, which is 100 million input and 20 million output tokens monthly. Assume $750 an engineer-day, and ongoing engineering attention of two, three and five days a month respectively, which is what these three architectures realistically demand.

Prompt and retrieval, base model Fine-tuned hosted model Fine-tuned open weights, self-hosted
Build effort 15 to 30 days 20 to 45 days 70 to 130 days
Build cost $11,250 to $22,500 $15,000 to $33,750 $52,500 to $97,500
Training compute $0 $12 to $300 $200 to $1,500
Monthly inference about $72 about $144 about $2,717
Monthly engineering $1,500 $2,250 $3,750
Three-year total $67,800 to $79,100 $101,200 to $120,000 $285,500 to $331,800

Three observations about that table.

First, the gap between the columns is almost entirely people and hosting. The training compute row is a rounding error in all three.

Second, self-hosting is roughly three times the three-year cost of the hosted fine-tune at this volume. There is a volume at which that flips. Holding the five-to-one input-to-output ratio, the fine-tuned hosted model costs $1.44 per million input tokens all in, so it takes about 1.9 billion input tokens a month, roughly 3.8 million requests, before the hosted bill matches $2,717 of GPU rental.

Third, that break-even is optimistic, and honest vendors will say so. Two small GPU instances will not serve 3.8 million requests a month at acceptable latency for most models, so reaching the crossover means buying more GPUs, which pushes the crossover further out. Self-hosting is usually the right answer for reasons other than price: data residency, air-gapped environments, contractual restrictions, or predictable latency. Our comparison of private LLM deployments and ChatGPT Enterprise works through those non-price reasons in detail, and if the GPU bill is the part that worries you, cloud cost optimization is a separate discipline worth applying before you commit to instances.

The costs that start after launch

A custom model is a perishable asset. Four recurring lines belong in the budget from day one.

Retraining. Anything trained on business data reflects the business at a moment in time. New product lines, new document formats and new policies all degrade a model quietly. A quarterly retrain on a maintained pipeline is two to five engineer-days per cycle. A retrain that means rediscovering how the original was built is two to four weeks.

Monitoring. You need to know when quality drops before your users tell you. That means logging inputs and outputs with retention rules, sampling for human review, and running the evaluation harness on a schedule. Budget one to three engineer-days a month, plus storage.

Base model deprecation. This risk is specific to fine-tuning. Your custom model is welded to a base model version, and base versions are retired. When the provider sunsets it, you retrain on the successor and revalidate. Plan for at least one forced migration in any three-year horizon, at roughly 20 to 40 percent of the original training and evaluation effort.

License obligations for open weights. "Open weights" is not "no legal work." The Llama 3.3 Community License requires a separate license from Meta if your products exceed 700 million monthly active users, requires you to "prominently display 'Built with Llama'" on a related website, user interface, blog post, about page or product documentation, and requires derivative model names to "include 'Llama' at the beginning." Most companies clear the user threshold easily, but the attribution and naming conditions are real obligations that somebody has to own.

When a custom model is worth it, and five signs it is not

Fine-tuning earns its cost in a narrower set of cases than the market suggests. OpenAI's documentation names the ideal cases as classification, nuanced translation, generating content in a specific format, and correcting instruction-following failures. That list is a useful filter: all four are about behavior, not knowledge.

It is probably worth it when:

  • A measured accuracy or format gap survives serious prompt engineering and retrieval work.
  • The task repeats at enough volume that a shorter prompt saves real money, since fine-tuning lets you drop lengthy instructions from every call.
  • The output has a rigid structure that a general model gets right 95 percent of the time when you need 99.
  • Latency or cost pressure justifies moving work from a large model to a small fine-tuned one.
  • The weights must stay inside your boundary for reasons that are not negotiable.

It is probably not worth it when:

  1. You have not built the evaluation set yet. Without it you cannot prove the model helped, so you cannot justify the next invoice.
  2. The real problem is missing knowledge. Fine-tuning teaches style and structure. It is a poor and expensive way to teach facts that change.
  3. The requirements are still moving. Every scope change invalidates part of the training data.
  4. The volume is low. A hundred documents a month does not repay a 60-day build, however satisfying the demo is.
  5. Nobody owns it after launch. An unowned model degrades into a liability within a year.

How to scope the project so the budget survives

A sequence that consistently keeps these projects honest, and that matches the phasing described in our guide to how long it takes to build an AI agent:

  1. Write the eval before the model. Fifty to two hundred real examples with expert-agreed correct answers. Two to five days. If you cannot produce them, stop here, because that is the finding.
  2. Set the baseline. Run a general model with a strong prompt and retrieval against that set, and record the score. Everything after this is measured as an improvement on that number.
  3. Do the rights review in parallel. Confirm you can use the data before you build anything that depends on it.
  4. Run the smallest possible fine-tune. Fifty examples, the cheapest model, one afternoon, about $1 of compute. Compare against the baseline.
  5. Decide with the numbers. If the small fine-tune closes a meaningful part of the gap, scale the data. If it does nothing, the problem is the task definition or the retrieval, and more data will not save it.
  6. Choose the serving path before scaling up. Per-token, dedicated endpoint, or your own instances. The answer changes the three-year cost by a factor of three.
  7. Budget the second year. Retraining, monitoring, one forced base-model migration, and a named owner.

Steps 1, 2 and 4 usually cost under ten engineer-days combined, and they routinely prevent a six-figure mistake. That is the highest-return work in the whole category.

What to ask a vendor

Four questions separate a scoped proposal from a number someone thought of.

  • "Restate this quote as engineer-days by phase." A real plan converts. A guess does not.
  • "What is the evaluation set, who labels it, and what score are we contracted to beat?" If quality is not measurable, it is not deliverable.
  • "What does this cost to run per month in year two, including hosting and retraining?" Compare the answer against the hosting table above.
  • "What happens when the base model is deprecated?" The answer tells you whether they have operated one of these before.

Vendors who answer all four cleanly are usually the ones who have shipped. For a version of this checklist covering agent projects, see our breakdown of AI agent development cost, which applies the same engineer-day method to a different architecture. If the model will end up calling APIs or taking actions rather than only producing text, that cost model applies instead of this one, and AI agent development is the right starting point.

Putting it together

Custom AI model development cost is best understood as three separate decisions, each with its own arithmetic. The build is engineer-days, dominated by data and evaluation, and runs roughly 20 to 45 days for a hosted fine-tune or 70 to 130 for a self-hosted open-weight model. The training compute is tens to low thousands of dollars and rarely deserves a meeting. The serving choice sets a recurring bill that ranges from under $200 a month to over $40,000 for the same model.

Start with the evaluation set, prove the gap, run the cheapest possible fine-tune, and only then commit. Most companies that do this in order discover they needed better retrieval and a better prompt, which is a far cheaper outcome than the one they budgeted for. The ones that genuinely need a custom model arrive at that decision with a number to defend it, and their projects tend to finish.

Fleurant AI builds custom AI and generative AI systems, including fine-tuned and self-hosted models for regulated environments, with a free discovery call and a working prototype in one to two weeks. If you want a second opinion on a quote, a scope, or a build-versus-buy decision, talk to a specialist and you will hear back within one business day.

Frequently asked questions

How much does custom AI model development cost?

Price the work in engineer-days, not dollars, then convert. Fine-tuning a hosted model is roughly 20 to 45 engineer-days. Training a task-specific model on your own data is roughly 35 to 90. Fine-tuning open weights and self-hosting them is roughly 70 to 130. At a fully loaded US day rate near $750, that is about $15,000 to $98,000 of build effort before any running costs.

What does the training compute actually cost?

Far less than most guides imply. Fine-tuning gpt-4.1-mini is priced at $5.00 per 1M training tokens, so a 1,000-example dataset run for three epochs costs about $12. Fine-tuning an 8-billion-parameter open-weight model on one L40S GPU at the $1.861 per hour on-demand rate for g6e.xlarge costs about $22 for a twelve-hour run. Compute is rarely the line that breaks a budget.

Why do agency quotes say $150,000 to $500,000 then?

Because they are quoting people, integration and risk, and usually bundling a product around the model. A quote is a day rate multiplied by a day count plus margin. Ask any vendor to restate the number as engineer-days by phase. If they cannot, the quote is a guess. If they can, you can compare it line by line with the estimates in this article.

Is fine-tuning cheaper than retrieval?

Usually not, and they solve different problems. Retrieval fixes missing knowledge; fine-tuning fixes behavior, format and tone. Fine-tuning also raises your per-token price: at published rates, a fine-tuned gpt-4.1-mini costs $0.80 per 1M input tokens against $0.40 for the base model, exactly double. Try prompting and retrieval first, and measure the gap that is left.

How much training data do we need?

Less than people expect for behavior, more than people expect for quality. OpenAI's guidance says the minimum is 10 examples, recommends starting with 50 well-crafted demonstrations, and reports improvements from 50 to 100 examples. The expensive part is not volume. It is getting subject-matter experts to agree on what a correct answer looks like.

What is the hidden cost of self-hosting a custom model?

Idle capacity. A dedicated endpoint bills whether or not anyone uses it. Amazon Bedrock prices Custom Model Import at $0.05718 per Custom Model Unit per minute, about $2,504 a month for one unit running continuously, and Microsoft states that each deployed fine-tuned Azure OpenAI model incurs an hourly hosting cost regardless of traffic. Two GPU instances for redundancy on EC2 run about $2,717 a month.

How long does a custom AI model project take?

For a fine-tuned hosted model, six to twelve weeks of elapsed time is common, and most of it is data access and evaluation rather than training. Self-hosted open-weight builds usually run three to six months, because deployment, security review and monitoring join the critical path. Getting permission to use the data is often the longest single step.

When is a custom model genuinely worth it?

When a measured gap survives serious prompt and retrieval work, when the task repeats at volume, when the output format is rigid, or when the weights must stay inside your own boundary. If you cannot state the gap as a number on a test set you already have, you are not ready to spend the budget.

Sources

  1. API Pricing, OpenAI
  2. Supervised Fine-Tuning, OpenAI
  3. Amazon Bedrock Pricing, Amazon Web Services
  4. Calculate the Cost of Running a Custom Model, Amazon Web Services
  5. Set Up Inference for a Custom Model, Amazon Web Services
  6. Amazon EC2 On-Demand Pricing, Amazon Web Services
  7. Deploy a Fine-Tuned Model, Microsoft Learn
  8. Customize a Model With Fine-Tuning, Microsoft Learn
  9. Software Developers, Quality Assurance Analysts, and Testers, U.S. Bureau of Labor Statistics
  10. Data Scientists, U.S. Bureau of Labor Statistics
  11. Employer Costs for Employee Compensation, June 2026, U.S. Bureau of Labor Statistics
  12. Amazon Mechanical Turk Pricing, Amazon Mechanical Turk
  13. Llama 3.3 Community License Agreement, Meta
  14. AI Risk Management Framework, National Institute of Standards and Technology
  15. The 2026 AI Index Report: Research and Development, Stanford Institute for Human-Centered AI

Free, no-obligation consultation

Have a question about custom AI model costs?

Tell us what you're working on or what you'd like to know. A specialist will get back to you with practical next steps, whether or not we end up working together.

  1. 1Send your question or project details (takes 2 minutes)
  2. 2A specialist reviews it and replies within 1 business day
  3. 3Get clear, practical next steps, free