Custom AI

AWS Bedrock vs Azure OpenAI vs Google Vertex AI: How to Choose

AWS Bedrock vs Azure OpenAI vs Google Vertex AI, compared on what decides real projects: data handling, residency controls, pricing mechanics, quotas and lock-in.

Cover image for an article comparing AWS Bedrock, Azure OpenAI in Microsoft Foundry and Google Vertex AI for mid-sized companies
In this article
  1. What these three platforms are actually called now
  2. What everyone gets wrong: model exclusivity is over
  3. How each platform handles your data
  4. Data residency: what "stays in region" actually means
  5. Pricing mechanics: four ways to buy the same tokens
  6. Throughput and quotas: where projects actually stall
  7. Compliance paperwork: BAAs, scope lists and the per-model trap
  8. Lock-in: how hard is it to leave?
  9. The decision framework
  10. When to use more than one
  11. What to do next
  12. Frequently asked questions
  13. Sources

If you are comparing AWS Bedrock vs Azure OpenAI vs Google Vertex AI, here is the short answer: pick the one that sits in the cloud where your data already lives, unless a hard requirement forces you elsewhere. All three now offer broad model catalogs, contractual commitments not to train on your data, a discounted batch tier, reserved-capacity options and controls over where inference runs. The differences that actually decide projects are narrower and less glamorous than most comparisons suggest: how each platform handles data for abuse monitoring, how precisely you can pin the processing location, how throughput is rationed before you commit money, and which models carry the platform's own contract rather than a third party's.

This article compares the three on those grounds, using each provider's own documentation as of October 2026. You will get a side-by-side on data handling and residency, the four different ways each platform sells the same tokens, where quota limits bite, how the compliance paperwork differs, an honest read on lock-in, and a scoring framework you can run in an afternoon with your own weights.

Key takeaway: The model you want is probably available on at least two of these platforms. The data path, the residency boundary and the contract behind that model are not interchangeable, and that is where the decision lives.

What these three platforms are actually called now

Start with the names, because two of the three changed and stale names cause real confusion in procurement documents.

Amazon Bedrock is still Amazon Bedrock. AWS describes it as a fully managed service providing "secure, enterprise-grade access to high-performing foundation models from leading AI companies," with 100+ foundation models from Amazon, Anthropic, DeepSeek, Moonshot AI, MiniMax, OpenAI and xAI.

Azure OpenAI is now a family of models inside Microsoft Foundry. Microsoft's documentation refers to "Foundry Models sold by Azure," a designation that "includes Azure OpenAI models," and the governing privacy article is titled Data, privacy, and security for Foundry Models sold by Azure. The catalog itself is enormous: Microsoft states it "includes over 10,000 models, with approximately 50 new models published each month."

Google Vertex AI's generative AI surface is now documented as the Gemini Enterprise Agent Platform. Google's data-residency and zero-data-retention pages carry that name, while the underlying API endpoints and aiplatform service names are unchanged.

This matters practically. If your security questionnaire asks a vendor about "Azure OpenAI data retention," the authoritative answer now lives under a Microsoft Foundry URL, and the policy differs depending on whether the model you picked is "sold by Azure" or sold by a partner. That turns out to be the sharpest trap in this comparison.

What everyone gets wrong: model exclusivity is over

The most common framing in platform comparisons is that Bedrock gives you Anthropic and open models, Azure gives you OpenAI, and Vertex gives you Gemini. That was roughly true two years ago. It is not true now.

As of October 2026, AWS documents OpenAI's GPT-6 Astra, GPT-6 Sol, GPT-6 Luna, GPT-6.1 Sol and the GPT-5.x family running on Bedrock, reachable through OpenAI's own Responses and Chat Completions APIs on the bedrock-runtime endpoint, with cross-Region inference support. Anthropic's Claude family runs on all three platforms: natively on Bedrock, as a partner offering in Microsoft Foundry, and in Google's Model Garden. Google's Gemini models remain the one genuinely exclusive family, available only on Google Cloud.

The buying implication is straightforward. "We need GPT" or "we need Claude" is no longer a platform decision for most teams. It becomes one only when you need a specific brand-new model on its launch day, or a specialized variant with limited availability. Pick the platform on data, operations and contract instead.

Key takeaway: Treat model availability as a tie-breaker, not a gate. Check it on each provider's region-and-model availability page for the exact model and region you need, because availability lags announcements, especially for geography-pinned deployments.

How each platform handles your data

This is the section regulated buyers should read first, and it is where the three genuinely diverge. Every provider makes a strong training commitment. None of them offers unconditional zero retention.

The training commitments are equivalent

AWS states that Amazon Bedrock "uses a zero operator access (ZOA) data security model," meaning "no operators of the service can access model input or output," and a "zero data retention (ZDR) data security model," meaning "by default, Amazon Bedrock does not store model inputs or outputs." AWS also runs each model provider's software in a dedicated Model Deployment Account the provider cannot reach, so "they don't have access to Amazon Bedrock logs or to customer prompts and completions."

Microsoft states that your prompts, completions, embeddings and training data "are NOT available to OpenAI or other providers of Models sold by Azure," "are NOT used by providers of Models sold by Azure to improve their models or services," and "are NOT used to train any generative AI foundation models without your permission or instruction." The models are stateless: "no prompts or completions are stored in the model."

Google points to the Training Restriction in its Service Specific Terms: Google "won't use your data to train or fine-tune any AI/ML models without your prior permission or instruction," and states that this "applies to all managed models on Gemini Enterprise Agent Platform, including GA and pre-GA models."

On training, there is nothing to choose between them. Stop using it as a differentiator.

The abuse-monitoring exceptions are not equivalent

Every platform logs something, sometimes, to detect abuse. The details differ enough to change architecture decisions.

Amazon Bedrock. The default is no storage, but AWS documents model-specific carve-outs. For OpenAI GPT models on Bedrock, "classifier-flagged traffic will be retained for up to 30 days for automated offline abuse detection," and "eligible customers may request full ZDR through their AWS account team." For Claude Fable 5 and Claude Fable 5.1, "all traffic will be retained for up to 30 days," with flagged traffic subject to potential human review by AWS. Retained data "is stored and processed by AWS and are not shared with third-party model providers." Critically for residency: "If cross-region inference is enabled for these models, retained inputs and outputs are stored in destination regions."

Microsoft Foundry. Abuse monitoring applies by default. When indicators are detected, "a sample of customer's prompts and completions may be selected for review," conducted "by automated means including by AI models such as LLMs by default, with additional reviews by human reviewers as necessary." For automated review, "customer's prompts and completions are not stored by the system." The human-review data store sits in each geography, is logically separated by customer resource, and reviewers use Secure Access Workstations with just-in-time approval. Managed customers "may apply to modify abuse monitoring"; once approved, "the data storage and human review process described above is not performed," and you can verify it by checking that the ContentLogging capability on the Foundry resource reads false.

Google. Prompt logging for abuse monitoring is conditional and bounded: if classifiers detect suspicious activity, Google "may log customer prompts solely for the purpose of examining whether a violation of the AUP or Prohibited Use Policy has occurred," stored "for up to 90 days in the same region or multi-region selected by the customer." Those logs "are not encrypted by Customer-managed encryption keys (CMEK)." Only customers governed by the Google Cloud Platform Terms of Service are in scope, which means "customers with a Google Cloud Master Agreement are exempt from prompt logging for this abuse monitoring by default." Separately, models designated "Advanced AI" under Google's Advanced AI Safety Addendum are logged more aggressively: "All prompts and responses will be logged and securely stored for up to 30 days," and "it may not be possible to opt-out." Google names the Claude Mythos and Claude Fable families in that scope, and consent to the addendum is required once per project before those models can be enabled.

There is also a quieter retention source on Google. By default, "Google's published Gemini models cache Customer Data (inputs, outputs, and derived data) in-memory to reduce latency," project-isolated, with a 24-hour TTL, honoring data residency. It can be disabled per project with a cacheConfig API call. Grounding with Google Search stores derived queries for up to three days with no way to disable it; Grounding with Google Maps stores prompts and outputs for 30 days.

Table comparing how Amazon Bedrock, Microsoft Foundry and Google Gemini Enterprise Agent Platform handle customer prompts, covering training commitments, default retention, abuse-monitoring exceptions and opt-out routes
Data handling on the three platforms, from each provider's own documentation, October 2026

The practical conclusion

If your workload cannot tolerate any third-party retention of prompts, you cannot answer that from the platform brochure. You have to pick a model whose documented abuse-monitoring profile matches your requirement, then pursue the opt-out route that platform offers: an AWS account team request, a Microsoft modified-abuse-monitoring application, or a Google Cloud Master Agreement plus avoiding Advanced AI models. That is a procurement task with a lead time, and it belongs in your project plan, not in a late-stage security review. We cover the surrounding control set in our HIPAA-compliant AI development checklist.

Data residency: what "stays in region" actually means

All three platforms now offer a three-tier residency model, and the vocabulary differs enough to cause mistakes in architecture documents.

Residency need Amazon Bedrock Microsoft Foundry Gemini Enterprise Agent Platform
Cheapest, routes anywhere Global cross-Region inference profile (about 10% savings) Global Standard / Global Provisioned (GlobalStandard) Global endpoints (aiplatform.googleapis.com)
Stays in a multi-country boundary Geographic cross-Region inference profile (US, EU, APAC) Data Zone Standard / Provisioned / Batch (US, EU, APAC) Jurisdictional multi-region endpoints (us, eu)
Stays in one region or geography Single-Region model invocation Standard / Regional Provisioned (ProvisionedManaged) Locational endpoints (us-central1, europe-west9)
Data at rest Customer-controlled, in your Region "Data stored at rest remains in the designated Azure geography" for all deployment types "Data stored at rest in the customer selected location remains at rest in that location"

Three details are easy to miss.

Global is cheaper on Bedrock, not just more available. AWS documents Global cross-Region inference as delivering "approximately 10% savings" versus Geographic, and notes that all cross-Region traffic "remains on the AWS network and does not traverse the public internet" and is "encrypted in transit between AWS Regions." CloudTrail records the processing Region in additionalEventData.inferenceRegion, which is how you evidence residency to an auditor. There is no extra routing charge, and the price is calculated based on the Region you call from.

Microsoft's Global tier gets models first. Microsoft's guidance is explicit: "For most workloads, start with Global Standard. It launches first when a new model releases, has the lowest price, and offers the broadest region coverage." New deployment types arrive in a set order — Global, then Data Zone, then geography-based — and "geography-based deployment types arrive last, have no guaranteed availability date, and depend on capacity that frees up as older models retire." If you have a strict single-geography requirement, assume a delay of months behind the headline launch.

Google's EU endpoint is narrower than "Europe." Google notes that "the European Union multi-region (eu) endpoint strictly covers data residency within EU member states" and that "geographies outside the European Union political boundary, including the United Kingdom and Switzerland, are excluded from this endpoint." Google also distinguishes residency tiers by compliance ceiling: locational endpoints "meet standard enterprise data governance and sovereign requirements (such as GDPR and HIPAA)," while workloads needing Department of Defense Impact Level 5 or ITAR isolation "should be deployed on jurisdictional endpoints as part of a broader compliant architecture."

Key takeaway: Write the residency requirement as a specific boundary — one region, one jurisdiction, or unrestricted — then confirm your chosen model supports that tier in that place. All three providers publish per-model, per-location availability tables. The gap between "the platform supports EU data zones" and "this model supports EU data zones today" is where schedules slip.

Pricing mechanics: four ways to buy the same tokens

Chasing per-token list prices across three providers is a poor use of time. Rates change, discounts get negotiated, and the published number for a given model is usually similar across clouds. What differs, durably, is the shape of the purchase. Each platform offers roughly four buying modes, and picking the wrong one is a much bigger cost error than picking the wrong cloud.

Pay-per-token, the default

All three default to pay-per-token with no commitment. This is correct for development, pilots and anything with unpredictable traffic. Microsoft's own guidance is that "standard deployments remain the better fit for development, testing, low-volume usage, or highly variable traffic."

Asynchronous batch, about half price

All three sell a batch tier at roughly a 50% discount with a 24-hour target turnaround and no real-time SLA.

  • Bedrock: batch inference is priced at a 50% discount to standard tier pricing. Note that prompt caching "is not supported with the batch inference API," so you cannot stack the two.
  • Microsoft Foundry: Global Batch and Data Zone Batch process "asynchronous groups of requests with separate quota and a 24-hour target turnaround, at 50% less cost than Global Standard." The separate enqueued-token quota is the underrated benefit: batch jobs do not steal capacity from live traffic.
  • Google: batch inference "is offered at a 50% discounted rate compared to real-time inference," with "24 hours turnaround time."

If you have document processing, classification, summarization or back-catalog enrichment, batch is the single easiest way to cut an AI bill in half. We go deeper on this and the other levers in LLM API cost optimization.

Caching, where the real savings hide

Caching discounts are large and the mechanics differ meaningfully.

Google's implicit caching is "enabled by default" for Gemini 2.5 and Gemini 3 models and "provides a 90% discount on cached tokens compared to standard input tokens." Importantly, "the discounts for cache and batch don't stack. The 90% cache hit discount takes precedence over the batch discount." On Provisioned Throughput the discount shows up as a reduced burndown rate instead: for Gemini 2.5 Pro, "1 input cached text token = 0.1 tokens."

Bedrock supports both implicit and explicit prompt caching. Cached reads bill at a model-specific cache-read rate, while cache writes can cost more than standard input tokens — for GPT-5.6 and later, "tokens written to cache are billed at 1.25× the uncached input token rate." Claude models on Bedrock support both a 5-minute and a 1-hour TTL, with minimums from 512 to 4,096 tokens depending on the model, plus a simplified single-breakpoint mode that looks back roughly 20 content blocks to find the longest match. Two operational details are worth designing around: cache hits "are not deducted against your rate limit," and on OpenAI models "cached input tokens read through prompt caching do not count against the input-tokens-per-minute quota." Caching therefore buys throughput headroom as well as money.

Reserved capacity, and the three units you have to learn

This is where the platforms are least alike, and where finance teams get surprised.

Amazon Bedrock Microsoft Foundry Gemini Enterprise Agent Platform
Unit Tokens per minute (Reserved tier); model units (Provisioned Throughput) Provisioned throughput unit (PTU) Generative AI scale unit (GSU)
Commitment Reserved tier: 1 or 3 months Hourly, or a 1-month or 1-year Azure Reservation Fixed-term subscription, weekly or monthly
Portable across models? Capacity is per model Yes — "the same PTU quota can be used to deploy any supported model" No — "specific to a project, region, model, and version"
Minimums Reserved tier: 100,000 input TPM and 10,000 output TPM Per-model minimum PTU counts Per-model minimum GSU purchase and increments
Overflow behavior "Automatically overflows to the Standard tier" Optional spillover to a standard deployment in the same resource On-demand requests are deprioritized behind Provisioned Throughput
Stated uptime Reserved tier "targets 99.5% uptime for model response" Latency targets per model; standard tiers have no latency SLA Assured capacity, with an SLA

Three warnings apply to all of them.

First, quota is not capacity. Microsoft says it plainly: "Having PTU quota doesn't guarantee that capacity is available. If capacity in the region is insufficient for the requested PTU count, the deployment fails." And: "Reservations don't guarantee capacity. First create deployments to confirm that capacity is available, then purchase the reservation to lock in the discounted rate." Google's equivalent warning is that "unused throughput doesn't accumulate or carry over to the next month."

Second, reserved capacity bills whether you use it or not. Microsoft: the meter "starts when the deployment is created and stops when it's deleted," regardless of tokens consumed. AWS: Reserved tier "billing continues until you delete the Reserved Tier reservation with the help of your AWS account manager." An idle provisioned deployment left over from a pilot is one of the more expensive mistakes in this category, and it is exactly the pattern we look for in cloud cost optimization work.

Third, do not plan to scale provisioned capacity up and down with traffic. Microsoft explicitly warns against it, because "capacity might not be available when you need to scale back up" and "continuous hourly billing at high utilization typically exceeds reservation pricing."

Bedrock adds a fourth mode the others handle differently: request-level service tiers. You set service_tier to priority, default or flex on the API call itself. Priority "delivers the fastest response times for a price premium over standard on-demand pricing" and is prioritized over the other tiers; Flex "offers cost-effective processing for a pricing discount" for workloads that tolerate longer processing, such as model evaluations and agentic workflows. Your on-demand quota is shared across those three tiers, while Reserved capacity is separate. Microsoft offers a comparable Priority processing tier with a defined latency target per model, plus Flex processing for delay-tolerant work. Google's equivalent ladder is Standard PayGo, Priority PayGo and Provisioned Throughput.

Throughput and quotas: where projects actually stall

Teams rarely fail to launch because the model was wrong. They fail because they hit an HTTP 429 in week three and discover the quota approval takes longer than the remaining sprint. The three platforms ration capacity on completely different principles, and this is arguably the most decision-relevant difference of all.

Microsoft Foundry: request it, per subscription. Default rate limits for most Foundry Models are 400,000 tokens per minute, 1,000 requests per minute and 300 concurrent requests, defined per region, per subscription and per model or deployment type. Azure OpenAI models vary per model and SKU. Structural limits matter too: 100 Foundry resources per region per subscription, 250 projects per resource, and 32 model deployments per resource. Since May 2026, Microsoft has been moving models to subscription-level quota pools, so "all Global Standard deployments of the same model and version under a subscription now draw from a single shared quota pool across all regions." Increases go through a request form, and Microsoft states that "priority goes to customers who actively use their existing quota allocation. Requests that don't meet this condition might be denied." Note also that models from partners and community, Anthropic excepted, "don't support quota increases."

Google: earn it by spending. Google's Standard PayGo "dynamically adjusts your organization's baseline throughput capacity, based on its total spend on eligible Agent Platform services over a rolling 30-day period." The published tiers, as of October 2026, are:

30-day organization spend Baseline TPM, Gemini Pro models Baseline TPM, Gemini Flash and Flash-Lite
$10 – $250 500,000 2,000,000
$250 – $2,000 1,000,000 4,000,000
$2,000 – $50,000 2,000,000 10,000,000
Over $50,000 10,000,000 50,000,000

The limit applies independently to each model in the family, there is "no separate requests-per-minute (RPM) limit for each tier," and traffic can "burst beyond this limit on a best-effort basis." Google's advice is to smooth traffic across the minute, because "high and instantaneous traffic can lead to throttling even if your average per-minute usage is below your limit."

This is a genuinely different model. A small team spending $300 a month starts with 1,000,000 TPM on Gemini Pro without filing a ticket. The same team on Microsoft Foundry starts at 400,000 TPM for most models and files a form to go higher. If your project has a hard launch date and an unpredictable ramp, that difference is worth more than a few cents per million tokens.

Bedrock: account-level, and partly discretionary. Bedrock controls inference with per-model token quotas, visible through Service Quotas, with separate allocations for the bedrock-runtime and bedrock-mantle endpoints — traffic to the two "is tracked against separate quotas, even when calling the same underlying model." AWS is unusually candid that the starting point varies: "the default quotas assigned to an account might be updated depending on regional factors, payment history, fraudulent usage, and/or approval of a quota increase request." A brand-new AWS account is not a good proxy for what an established one will get.

Key takeaway: Before you commit, run a 30-minute load test at your expected peak on each shortlisted platform and record the 429 rate. Then ask each account team, in writing, what quota you would get at launch and how long an increase takes. That answer predicts your timeline better than any benchmark.

Compliance paperwork: BAAs, scope lists and the per-model trap

For regulated buyers the contract matters as much as the architecture, and here the three differ in both substance and effort.

HIPAA. AWS lists Amazon Bedrock and Amazon Bedrock AgentCore in its HIPAA Eligible Services Reference, and the page states that covered entities and business associates "agree not to use these HIPAA Eligible Services for any purpose or in any manner involving Protected Health Information ... without first entering into an AWS business associate agreement." Unless specifically excluded, generally available features of a listed service are also eligible. Microsoft takes a different approach: the Microsoft HIPAA Business Associate Agreement "is available through the Microsoft Online Services Data Protection Addendum by default to all customers who are covered entities or business associates under HIPAA," with Azure and Azure Government in the in-scope list — there is no separate signature step. Google requires customers subject to HIPAA who want to use "any Google Cloud products in connection with PHI" to "review and accept Google's Business Associate Agreement," and states that the BAA "covers Google Cloud's entire infrastructure (all regions, all zones, all network paths, all points of presence)" plus the services it lists, with Gemini Enterprise Agent Platform among the AI and machine learning entries.

All three are quick to add the point that matters most: a BAA is necessary but never sufficient. Microsoft's own FAQ answers "Does having a Business Associate Agreement with Microsoft ensure my organization's compliance with HIPAA and the HITECH Act?" with a flat "No." Your access controls, audit logging, minimum-necessary design and risk analysis are still yours to build. Confirm the specifics with your counsel and compliance team.

The per-model trap. This is the most important compliance finding in the comparison, and it is specific to Microsoft Foundry. The catalog is split into two categories with materially different terms. "Models sold by Azure" are hosted and sold by Microsoft under Microsoft Product Terms, with Microsoft support, enterprise SLAs, and billing "via Azure meters as First Party Consumption Services." "Models from partners and community" — which Microsoft says "make up the vast majority of the Foundry Models" — are billed "through Azure Marketplace, in accordance with the Microsoft Commercial Marketplace Terms of Use," and are "supported by their providers."

Anthropic's Claude models sit in the partner category, and Microsoft's dedicated privacy page for them is explicit: "Anthropic is the seller and operator of Claude models in Microsoft Foundry and acts as an independent data processor for prompts and outputs." There are two hosting options. "Hosted on Azure" processes prompts and outputs on Azure infrastructure with data at rest in your selected Azure geography, but "automatic safeguards flag content that might be sent to Anthropic Trust & Safety for review," with Anthropic personnel reviewing "on an exceptions-only basis." "Hosted on Anthropic Infrastructure" is a different proposition entirely: "your prompts and outputs are processed on Anthropic hosted infrastructure. Data might be processed outside of Azure including outside of your selected Azure region."

A security reviewer who reads only the Azure OpenAI data-privacy page and then deploys Claude on Anthropic-hosted infrastructure has signed up for a data flow their documentation does not describe. Google has a parallel structure: partner models are subject to partner-specific terms, and the Advanced AI Safety Addendum requires explicit per-project consent that only an administrator holding the aiplatform.consents.update permission can grant. Build your AI compliance review around the specific model and hosting option, not the platform name.

Decision tree for choosing between Amazon Bedrock, Microsoft Foundry and Google Gemini Enterprise Agent Platform, starting from where company data already lives and branching on residency, retention and model requirements
A decision tree for mid-sized companies choosing a cloud AI platform

Lock-in: how hard is it to leave?

The honest answer is that the model call is the easy part to move, and nobody's lock-in lives there.

API surfaces are converging fast. Bedrock now accepts Anthropic's Messages API, OpenAI's Responses API and OpenAI's Chat Completions API alongside its own Converse and InvokeModel APIs, so you can point an existing OpenAI SDK at a Bedrock endpoint. Microsoft Foundry exposes the Azure OpenAI v1 APIs. Google exposes its own SDK plus OpenAI-compatible access. Swapping the inference call in a well-structured application is usually days of work, not months.

What does not move:

  • Retrieval. Your vector index, chunking pipeline, permission filters and document connectors are platform-specific if you used managed services. Build these on portable components if portability matters.
  • Identity and networking. IAM policies, Microsoft Entra ID app registrations, VPC endpoints, Private Link and VPC Service Controls are all cloud-specific, and typically represent more engineering hours than the model integration.
  • Guardrails and safety configuration. Bedrock Guardrails, Azure AI Content Safety and Google's safety filters have different categories, thresholds and failure modes. Re-tuning them is real work and requires re-testing.
  • Evaluation and observability. Your evaluation sets should be plain data you own. If they live inside a provider's evaluation product, you have quietly added lock-in to the one thing you most need in order to compare platforms.
  • Reserved capacity. A 1-year Azure Reservation or a 3-month Bedrock Reserved tier commitment is a contractual anchor. Match commitment length to your confidence, not to the size of the discount.

A reasonable mitigation, and the pattern we default to in custom AI development, is a thin provider adapter: one interface for chat, tool calling and embeddings; prompts, tool schemas and evaluation sets stored as data outside provider code; and a scheduled evaluation run against a second provider so you always know what switching would cost in quality. That is a few days of extra work at the start, and it converts a strategic risk into a scheduling decision.

The decision framework

Score each platform 1 to 5 on these nine criteria, multiply by your own weights, and the answer usually falls out. The weights below are a starting point for a US company of 50 to 1,000 employees. Change them to match your constraints.

Criterion Suggested weight What a 5 looks like Where to verify
Data gravity 25% The systems your AI must read already live in this cloud Your own system inventory
Residency fit 15% Your required boundary is supported for your required model, today Provider region-and-model availability tables
Retention fit 15% Documented abuse-monitoring behavior already meets your rule, or the opt-out is available to you Provider data-privacy and abuse-monitoring pages
Launch throughput 10% You get the tokens per minute you need at launch without a ticket Quota pages, plus your own load test
Team skills 10% Your engineers already hold certifications and operate in this cloud daily Your own team
Contract simplicity 10% One contract, one invoice, one support path for every model you need Model catalog category and billing terms
Model fit 5% The specific model that won your evaluation is available here Your evaluation results
Cost mechanics fit 5% Batch, caching and reserved options match your traffic shape Provider pricing pages
Exit cost 5% A provider swap is weeks, not quarters Your own architecture review

Two rules make this framework work. First, if you are regulated, treat residency fit and retention fit as gates rather than weights: a score of 1 on either is disqualifying regardless of the total. Second, run the scoring before the model evaluation, not after. Teams that benchmark models first tend to rationalize the platform afterwards.

Three worked scenarios

A 180-person regional insurance carrier, Microsoft 365 shop, claims data in Azure SQL. Data gravity and team skills point hard at Microsoft Foundry. The work is to pick a model "sold by Azure" so the Microsoft contract and SLA cover it, use Data Zone Standard to keep processing in the US, and apply for modified abuse monitoring if the claims narratives are sensitive. Expect the geography-pinned deployment type to lag the newest model by months, and plan the pilot on Global Standard with synthetic data while you wait.

A 60-person healthcare analytics company already on AWS with PHI in Amazon S3. Bedrock is the default. Sign the AWS BAA before any protected health information touches the service, use a Geographic (US) cross-Region inference profile rather than Global so any retained abuse-monitoring data stays in the US, and read the specific model's abuse-detection entry before choosing between Claude and GPT variants. Use the Flex tier for batch-like evaluation work, and reserve capacity only after a month of real traffic data.

A 400-person logistics company on Google Cloud with telemetry in BigQuery. Gemini Enterprise Agent Platform wins on data gravity, and the spend-based throughput tiers are an advantage: at a few thousand dollars a month of platform spend you reach 2,000,000 TPM on Gemini Pro models without filing anything. Disable in-memory caching if your policy requires it, avoid Grounding with Google Search where the three-day query retention is unacceptable, and use a jurisdictional us endpoint if a customer contract requires US-only processing. If the workflows become multi-step, the same reasoning carries into AI agent development.

When to use more than one

Multi-cloud AI is usually a bad default and occasionally the right answer. It is justified when:

  • A single model materially outperforms the alternatives on your own evaluation set and lives on only one platform.
  • You have a hard residency or retention requirement that one platform cannot meet for the model you need.
  • A specific customer contract names a cloud.
  • You are large enough that reserved-capacity negotiating leverage across two vendors is worth the operational overhead.

It is not justified by a desire to "avoid lock-in" in the abstract. Running two platforms roughly doubles the surrounding work — two identity models, two logging pipelines, two guardrail configurations, two cost reports, two sets of quota relationships — and that work is most of the project. If you want optionality, buy it with a provider adapter and a parallel evaluation harness, not with a second production platform.

Key takeaway: One platform, chosen on data gravity, with a thin adapter and portable evaluations, beats two platforms chosen on hedging. Add the second only when a named requirement forces it.

What to do next

Work through this in order and you will have a defensible decision in about two weeks.

  1. Write down where the data lives. List the five systems your AI must read, and note which cloud each sits in. This usually settles most of the decision on its own.
  2. Write the residency rule as a boundary. One region, one jurisdiction, or unrestricted. Then check your candidate model's availability for that tier in that place on the provider's own table.
  3. Write the retention rule. Decide whether any third-party retention of prompts is acceptable. If it is not, read the model-specific abuse-monitoring page and start the opt-out conversation now, because it has a lead time.
  4. Load-test for quota, not quality. Thirty minutes at expected peak on each shortlisted platform, recording 429s. Ask each account team in writing for launch quota and increase turnaround.
  5. Evaluate two or three models on your own data, with your own test set, stored as plain data you own.
  6. Model the cost in the right shape. Separate interactive traffic from batch-eligible traffic, apply the batch and caching discounts, and only then compare. Do not buy reserved capacity before you have a month of real traffic.
  7. Design the adapter before the first integration. One interface, prompts as data, evaluations outside provider code.

The platform decision is less consequential than it feels, and the decisions around it — residency boundary, retention rule, quota headroom, which contract covers which model — are more consequential than they look. Get those four right and any of the three will serve you.

If you want a second opinion on a shortlist, or a working prototype on your own documents before you commit to a platform, talk to a specialist. A free discovery call and a short prototype are usually cheaper than a year of a reservation chosen for the wrong reasons. You can also browse our other custom AI and generative AI work to see how we approach this.

Frequently asked questions

Is Azure still the only place to get OpenAI models?

No, and this is the most out-of-date assumption in platform comparisons. As of October 2026, OpenAI's GPT models run on Amazon Bedrock too: AWS documents GPT-6 Astra, GPT-6 Sol, GPT-6 Luna and the GPT-5.x family on the bedrock-runtime endpoint, reachable through OpenAI's own Responses and Chat Completions APIs. Model exclusivity is a weak reason to pick a platform now.

Do any of these platforms train on my prompts?

All three commit in writing that they do not. AWS states that Bedrock uses a zero data retention model and that model providers have no access to customer prompts and completions. Microsoft states that prompts and completions are not used to train any generative AI foundation models without your permission. Google's Service Specific Terms contain a training restriction covering all managed models. The real variable is abuse-monitoring retention, not training.

Which platform is best for HIPAA workloads?

All three can support protected health information with the right paperwork and controls, so the deciding factor is usually which cloud already holds your data. AWS lists Amazon Bedrock and Bedrock AgentCore as HIPAA eligible services and requires an AWS business associate agreement first. Microsoft offers its HIPAA BAA through the Online Services Data Protection Addendum by default, with Azure in scope. Google requires you to review and accept its BAA. Confirm the specifics with your counsel and compliance team.

What does zero data retention actually mean here?

It means the platform does not store your prompts and outputs by default, but every provider carves out exceptions for abuse monitoring that depend on the model you choose. Bedrock retains all traffic for up to 30 days for some models. Google logs prompts for up to 90 days when classifiers flag activity, and logs all prompts and responses for up to 30 days for models designated Advanced AI. Microsoft stores flagged prompts for human review unless you are approved for modified abuse monitoring. Read the model-specific page, not the platform headline.

How do the pricing models differ?

The per-token rate for a given model matters less than the buying mechanics. All three sell pay-per-token plus a roughly 50% discounted asynchronous batch tier. Beyond that, Bedrock adds request-level Flex and Priority tiers and a Reserved tier sold in tokens per minute. Microsoft sells provisioned throughput units with 1-month or 1-year Azure Reservations. Google sells Provisioned Throughput in generative AI scale units tied to a project, region, model and version.

Can I run the same application on more than one of these platforms?

Yes, and it is increasingly practical because the API surfaces are converging: Bedrock exposes OpenAI's Responses and Chat Completions APIs and Anthropic's Messages API alongside its own Converse API. The work that does not port is everything around the model, including retrieval, identity, logging, guardrail configuration and evaluation harnesses. Build a thin provider adapter early and keep prompts, tools and evaluations outside provider-specific code.

Which platform should a small or mid-sized company pick by default?

Start with the cloud that already holds the data your AI needs to read. Identity, networking, private connectivity, logging, cost reporting and your team's existing skills all come free in that cloud, and they are a larger share of the project than the model call. Only override that default when a hard requirement, such as a specific residency boundary or a specific model, is unavailable there.

Sources

  1. Overview of Amazon Bedrock, Amazon Web Services
  2. Data protection in Amazon Bedrock, Amazon Web Services
  3. Amazon Bedrock abuse detection, Amazon Web Services
  4. Service tiers for optimizing performance and cost, Amazon Web Services
  5. Route model inference requests across AWS Regions with cross-Region inference, Amazon Web Services
  6. Prompt caching for faster model inference, Amazon Web Services
  7. Quotas for Amazon Bedrock, Amazon Web Services
  8. HIPAA Eligible Services Reference, Amazon Web Services
  9. Amazon Bedrock Pricing, Amazon Web Services
  10. Data, privacy, and security for Foundry Models sold by Azure, Microsoft Learn
  11. Understanding deployment types in Microsoft Foundry Models, Microsoft Learn
  12. Provisioned throughput for Foundry Models, Microsoft Learn
  13. Microsoft Foundry Models overview, Microsoft Learn
  14. Data, privacy, and security for Anthropic Claude models in Microsoft Foundry, Microsoft Learn
  15. Microsoft Foundry Models quotas and limits, Microsoft Learn
  16. HIPAA and the HITECH Act, Microsoft Learn
  17. Gemini Enterprise Agent Platform and zero data retention, Google Cloud
  18. Abuse monitoring, Google Cloud
  19. Data residency, Google Cloud
  20. Standard PayGo usage tiers, Google Cloud
  21. Calculate Provisioned Throughput requirements, Google Cloud
  22. Batch inference with Gemini, Google Cloud
  23. Google Cloud compliance with HIPAA, Google Cloud

Free, no-obligation consultation

Have a question about cloud AI platform selection?

Tell us what you're working on or what you'd like to know. A specialist will get back to you with practical next steps, whether or not we end up working together.

  1. 1Send your question or project details (takes 2 minutes)
  2. 2A specialist reviews it and replies within 1 business day
  3. 3Get clear, practical next steps, free