Independent reporting on artificial intelligence.


The Frontier Wire

Comparisons

Top 5 AI gateways for automatic failover across Bedrock, Vertex AI and Azure OpenAI

Claude now runs on Amazon Bedrock, Google Cloud and Microsoft Foundry, and GPT models run on both OpenAI and Azure. That makes cross-cloud failover practical, but each cloud throttles, authenticates and names models differently. Five gateways, measured against seven published criteria.

TL;DR

  • Claude is sold through Anthropic, Amazon Bedrock, Google Cloud and Microsoft Foundry, so a Bedrock Claude workload can fail over to the same model on another cloud instead of to a different model.
  • Each cloud throttles differently: Bedrock sets quotas per Region and model, Google pools Claude quota by model lineage per endpoint, and Azure is moving quota to the subscription level with spillover from provisioned to standard deployments.
  • Bifrost ranks first: it speaks each cloud’s native authentication, treats 429, 401/403/402 and 5xx errors differently, gives every fallback its own retry budget and runs self-hosted under Apache 2.0.
  • LiteLLM, Kong AI Gateway, Vercel AI Gateway and Azure API Management complete the list; the right pick depends on whether the gateway must be self-hosted and which cloud carries most of the traffic.

Teams running Bedrock Claude in production have a reliability option that did not exist two years ago: the same Claude models are now hosted on Amazon Bedrock, Google Cloud, Microsoft Foundry and Anthropic’s own API, and GPT models are hosted on both OpenAI and Azure. When one host throttles or degrades, a request can move to the same model on another cloud. Doing that by hand means handling three authentication schemes, three ways of naming a model and three quota systems in every service. An AI gateway handles that translation once. This comparison ranks five gateways on how well they fail over across Bedrock, Vertex AI and Azure OpenAI, with the criteria stated before the ranking.

Cross-cloud failover for Claude and GPT models

Cross-cloud failover sends a model request to the same model on a different cloud when the first host returns throttling or server errors. It is different from switching to another model: the output style, context window and tool behavior stay the same, so the fallback path needs far less separate testing than a model swap does.

Three facts make it practical in 2026:

  • Claude runs on three clouds plus Anthropic. Anthropic documents Claude in Microsoft Foundry in Global Standard and US Data Zone Standard deployments, alongside Bedrock and Google Cloud. Google lists current Claude models for its Agent Platform, which is how Google’s documentation now refers to Vertex AI.
  • GPT models run on OpenAI and Azure. An Azure OpenAI deployment and an OpenAI project are two hosts for the same model family.
  • Gemini runs on Google’s two surfaces. Google AI Studio and Vertex AI both serve Gemini.

Failover is not free. Each host needs its own account, quota and credentials, and data-residency rules may forbid some routes. The trigger is usually visible on a status page; Bifrost publishes a live status monitor for the Anthropic API alongside the providers’ own pages. The Frontier Wire’s explainer on LLM failover and load balancing covers the general error taxonomy, retry budgets and circuit breakers. This piece covers the product question: which gateway handles the three clouds best.

How Bedrock, Vertex AI and Azure OpenAI throttle traffic

All three clouds return HTTP 429 when a request exceeds quota, but the quota sits in different places, and the place determines which failover move works. Rotating to another key helps only when the quota is per key; moving Region helps only when it is per Region.

CloudQuota scopeIn-cloud failover optionThrottling signal
Amazon BedrockPer account, Region and model; separate quotas for bedrock-runtime and bedrock-mantleCross-Region inference profiles (geographic or global)429 ThrottlingException; 503 ServiceUnavailable; 529 overloaded_error
Google Cloud (Vertex AI)Claude models launched after May 26, 2026 share one quota per model lineage per endpoint; older models use per-model QPM and TPMGlobal and multi-region endpoints, each its own quota bucket429 “Resource exhausted” (pay-as-you-go)
Azure OpenAI in FoundryMoving to subscription-level pools; Global Standard deployments of a model share one pool across RegionsGlobal and Data Zone deployment types; spillover from provisioned to standard deployments429, including while metrics show usage below quota

The details matter for routing:

  • Bedrock. The quotas page says the defaults applied to an account “might be lower than the default values listed” and that increases are not automatic. Cross-Region inference lets Bedrock pick a Region inside a geography or across all commercial Regions, with no extra routing cost; the global option is priced about 10 percent lower. Inference profiles do not support Provisioned Throughput.
  • Google Cloud. Under shared lineage quotas, a call to any Opus version on the global endpoint draws from one anthropic-claude-opus bucket. Quotas on the global endpoint and each multi-region endpoint are independent, so moving between endpoints is itself a failover move.
  • Azure. Microsoft’s spillover feature sends a request that a provisioned deployment cannot serve (429, 500 or 503) to a standard deployment in the same resource and marks it with an x-ms-spillover-from-deployment header.

Every in-cloud option stops at the cloud boundary. None of them moves a Claude request from Bedrock to Google, or a GPT request from Azure to OpenAI. That is the gap a gateway fills.

An application sends one model name to an AI gateway, which maps it to the same Claude model on Amazon Bedrock, Google Cloud Vertex AI, Microsoft Foundry and Anthropic's API Figure 1: The gateway owns the mapping from one model name to each cloud’s model ID, credentials and quota, so the application never changes when traffic moves.

How the gateways were evaluated

Seven criteria decide whether a gateway can fail over across these three clouds without extra glue code. Each product was checked against its own documentation in October 2026.

CriterionWhat was checkedWhy it matters
Native cloud authenticationSigV4 and IAM roles for Bedrock, service accounts or ADC for Vertex AI, Entra ID or managed identity for AzureStatic keys are the weakest option on all three clouds
Cross-provider fallbackOne request can move from Bedrock to Vertex AI to Azure with format translationThe core requirement
Error-aware retries429, auth, billing and 5xx errors handled differentlyA spend cap or revoked key should not burn the retry budget
Key and Region poolsSeveral keys, accounts or Regions per provider with weightsAbsorbs per-Region quotas before leaving a cloud
Health trackingCooldowns, circuit breakers or scoring that stop sending traffic to a failing targetAvoids paying for retries on a host already known to be down
Model mappingOne client-facing name mapped to each cloud’s model ID or deploymentBedrock inference profiles, Google model names and Azure deployment names all differ
Deployment modelSelf-hosted, managed or bothDecides where prompts and cloud credentials live

Cloudflare AI Gateway was considered and left out of this list. Its OpenAI-compatible endpoint, which dynamic routes use, supports Google Vertex AI but not Bedrock or Azure OpenAI, and the older Universal Endpoint that offered cross-provider fallbacks is deprecated. Bedrock and Azure remain available through Cloudflare’s provider-specific passthrough endpoints, without a shared fallback chain.

The five gateways compared at a glance

The table summarizes how each gateway covers the three clouds. “Paid” marks capabilities that need a commercial licence or plan.

RankGatewayDeploymentBedrock authVertex AI authAzure authCross-cloud fallbackHealth tracking
1BifrostSelf-hosted (Apache 2.0)Explicit keys, IAM role chain, API keyService account JSON, ADCAPI key, Entra ID, managed identityFallback chain, per-provider retry budgetAdaptive balancing and circuit breaker (paid)
2LiteLLMSelf-hosted (MIT core)YesYesYesModel-group fallbacks, weighted failoverCooldowns per deployment
3Kong AI GatewaySelf-hosted data planeYesYesYesPriority balancer, failover_criteria (paid)Circuit breaker, 3.13+ (paid)
4Vercel AI GatewayManagedAccess key and secret (BYOK)Service account fields (BYOK)API key and resource name (BYOK)Provider order per model slugManaged by Vercel
5Azure API ManagementAzure-managed; self-hosted gateway on Developer and PremiumImported passthrough APIImported APIManaged identityBackend pools; unified model API in previewBackend circuit breaker honoring Retry-After

The ranking rewards a gateway that covers all three clouds natively, makes error-specific decisions and can run inside the customer’s own network. A team that already runs Kong or Azure API Management for its other APIs would reasonably move those entries up.

1. Bifrost

Bifrost is an open-source AI gateway written in Go by Maxim AI and published under Apache 2.0. It ranks first because it is the only entry that covers every criterion in its free build: native authentication for all three clouds, error-specific retry logic, model mapping per provider and a cross-provider fallback chain. It gets a longer write-up here because its documentation describes each of those behaviors in enough detail to evaluate.

Coverage and speed. Bifrost exposes more than 25 providers and over 10,000 models through one OpenAI-compatible API. Bifrost’s published benchmark reports 11 µs of added overhead per request at 5,000 RPS on an AWS t3.xlarge. Teams whose primary host is AWS can start from the Bifrost overview for Amazon Bedrock, which covers governance and guardrails on Bedrock traffic.

Native cloud authentication. The Bedrock provider signs requests with SigV4 and accepts explicit credentials, an inherited IAM role or a Bedrock API key; an aliases map points a model name at an inference profile ID or ARN, such as a us.anthropic cross-Region profile. The Vertex AI provider accepts a service account JSON or Application Default Credentials and builds the endpoint from project and region. The Azure provider supports API keys, Entra ID service principals and managed identity through the default Azure credential chain, maps model names to deployment IDs, and serves both OpenAI and Anthropic models hosted on Azure. A walkthrough of calling Bedrock, Vertex AI, Gemini and Anthropic models through Bifrost shows the same setup end to end.

Error-specific retries. The retries and fallbacks documentation separates three cases:

  • Rate limits (429). The key is marked used for this cycle and Bifrost rotates to another key, still applying backoff because account-level quotas can be shared across keys.
  • Auth and billing failures (401, 403, 402). The key is marked dead for the rest of the request and Bifrost rotates immediately, without backoff. If every key is dead it returns 502 upstream_credentials_exhausted.
  • Server and network errors (5xx). Bifrost retries the same key with exponential backoff and 0.8 to 1.2 jitter, from 500 ms to a 5,000 ms cap by default.

Validation errors (400, 404, 422) are never retried. Retries are off by default (max_retries: 0), so they need to be set per provider.

Fallback chains. When retries are exhausted, the request moves to the next provider/model entry in a fallbacks list, for example a Bedrock Claude model, then the same model on Vertex AI, then Azure. Each fallback gets its own full retry budget and runs every plugin again, including caching and governance. When every entry fails, Bifrost returns the primary provider’s original error. The Bifrost team’s post on enabling automatic fallback when a primary provider fails walks through a configured chain. Each retry and fallback transition is written to the request’s routing log with the error type and status code, but never the upstream message, which can echo keys or user input.

Routing from virtual keys. Under provider routing, a virtual key can list several providers with weights, allowed models, budgets and rate limits. Bifrost removes any provider that is over budget or over its rate limit, picks one of the rest by weighted random selection, and adds the remaining providers as fallbacks in descending weight. A 70/30 split between Azure and OpenAI for the same GPT model is a single configuration block. Within a provider, key management spreads traffic across keys by weight and can restrict keys to specific models.

Azure streams. Azure can report an error inside an HTTP 200 stream after sending startup metadata. Bifrost buffers recognized startup events so those errors still reach retry and fallback logic, then replays the successful attempt’s events in order.

Bifrost resolves a virtual key to weighted Bedrock, Vertex AI and Azure providers, retries within each provider, then moves down the fallback chain Figure 2: Retries run inside each cloud and fallbacks run across clouds, and every cloud gets its own retry budget before the chain moves on.

Enterprise tier. Adaptive load balancing, which rescores providers and keys every 5 seconds on error rate and token-aware latency, and a header-driven circuit breaker are part of Bifrost Enterprise, along with clustering and in-VPC deployment. The circuit breaker opens on configured response headers, can read its cooldown from a header such as retry-after-ms, and its documented example policy redirects an Azure provisioned deployment to a pay-as-you-go deployment.

Best for: platform teams that want one self-hosted gateway to move Claude and GPT traffic across Bedrock, Google Cloud and Azure using each cloud’s native identity, with failover decisions they can audit request by request.

2. LiteLLM

LiteLLM is a Python SDK and proxy that calls more than 100 LLM APIs through the OpenAI format, including bedrock/, vertex_ai/ and azure/ model prefixes. It ranks second because its reliability features are broad and well documented, and every major cloud is supported in the open-source build.

Fallbacks. The fallbacks documentation moves a request to another model group after num_retries is exhausted, in the listed order. Three fallback types exist: general errors such as rate limits, context_window_fallbacks for prompts that exceed a model’s window and content_policy_fallbacks for provider content-filter rejections. The second and third are useful across clouds, because context limits and content filters can differ between hosts of the same model.

Load balancing and cooldowns. A model group can hold deployments on several clouds or Regions. The routing documentation recommends the default simple-shuffle weighted pick for production. A deployment that fails more than allowed_fails times in a minute is cooled down for cooldown_time, configurable per deployment. With enable_weighted_failover, a failed request first tries another deployment in the same group, such as another Azure Region, before moving to a different model group.

Encrypted reasoning. The documentation notes a cross-host detail that applies to any gateway: a fallback cannot decrypt encrypted reasoning produced by the failed deployment on earlier turns. LiteLLM drops what the target cannot read and keeps readable summaries, so the fallback reasons afresh.

Trade-offs. In production LiteLLM uses Redis to share cooldown and usage state across proxy instances, so a multi-instance deployment runs Redis alongside the proxy. Audit logs, SCIM, secret managers and IP allowlists are on the Enterprise plan, and SSO is free only up to five users. Teams moving off it can read the Frontier Wire’s LiteLLM alternatives comparison.

Best for: Python-centric teams that want the widest provider catalogue and fine-grained fallback types, and are prepared to run Redis alongside the proxy.

3. Kong AI Gateway

Kong AI Gateway is the AI layer of Kong Gateway. Its cross-cloud features live in the AI Proxy Advanced plugin, which is part of the AI Gateway Enterprise offering and needs Kong Gateway 3.8 or later. It ranks third because it covers all three clouds with strong balancing options, but only on the paid tier.

Providers and formats. The plugin lists Azure OpenAI, Amazon Bedrock, Gemini and Vertex AI among its targets. From 3.10, fallback works across targets in any supported format, so a Bedrock target can fall back to a Vertex AI or Azure target, and native Bedrock or Gemini formats can be proxied without conversion.

Balancing and failover. Seven algorithms are available: round robin, consistent hashing, least connections, lowest latency, lowest usage, semantic and priority. Priority gives tiered failover across groups of targets. Client errors do not trigger failover by default; failover_criteria can add codes such as http_429 and http_502. From 3.13, a circuit breaker stops routing to a target after max_fails failures until fail_timeout elapses.

Trade-offs. Every capability in this section is Enterprise-only. Kong is a strong choice when the API platform is already Kong; as a standalone purchase for model failover it is heavier than the alternatives.

Best for: organizations already running Kong for API traffic that want cross-cloud model failover under the same control plane and policies.

4. Vercel AI Gateway

Vercel AI Gateway is a managed gateway that routes a model slug such as anthropic/claude-sonnet-5 to whichever providers host that model. It ranks fourth because it offers the simplest cross-cloud setup in this list, at the cost of running only as a hosted service.

Provider ordering. The provider routing documentation sets order, only and sort under providerOptions.gateway. An order of Bedrock then Anthropic tries Bedrock first and falls back to Anthropic, with Vertex AI still available afterwards unless excluded with only. Sorting by cost, time to first token or throughput is also available. A separate models list adds fallback to different models.

Visibility. Responses carry provider metadata, including the final provider, the fallbacks available and each provider attempt with its credential type, so an application can see which cloud served a request.

Credentials. With bring your own key, Bedrock takes an access key, secret and Region, Vertex AI takes a project, location and service account fields, and Azure takes an API key and resource name, with model mappings for custom deployment names. If a request with the customer’s credentials fails, Vercel retries it with its own system credentials unless zero data retention rules prevent that.

Trade-offs. Every request passes through Vercel’s network, and the system-credential retry means a failed BYOK request can be served under Vercel’s provider agreements rather than the customer’s, which matters for teams that chose Bedrock or Azure for contractual reasons.

Best for: application teams that want cross-cloud Claude failover by setting a provider order, without operating a gateway.

5. Azure API Management

Azure API Management (APIM) is Microsoft’s API gateway, and its AI gateway capabilities extend it rather than ship as a separate product. It ranks fifth because it is the strongest option for failover inside Azure, while Bedrock and Vertex AI are not native backends.

Backend pools. A backend pool holds up to 30 backends and balances them by round robin, weight or priority, with optional session awareness. Lower-priority groups receive traffic only when every backend in the higher group has a tripped circuit breaker. That pattern suits Azure OpenAI deployments in several Regions, or a provisioned deployment ahead of standard ones.

Circuit breaker. A backend circuit breaker trips on a count or percentage of failure status codes in a time window and can accept the Retry-After header as its trip duration. Microsoft notes the rules are approximate across gateway instances, only one rule is allowed per backend, and the breaker is not available in the Consumption tier.

Other clouds. The unified model API, in preview, exposes several backends through one OpenAI Chat Completions endpoint, translates formats and supports model failover across providers. Its supported backend formats are the OpenAI Chat Completions API and the Anthropic Messages API, authenticated with headers or managed identity. Vertex AI and Bedrock endpoints can otherwise be imported as APIs rather than configured as native providers. Teams weighing APIM against standalone gateways for multi-cloud traffic can compare Azure AI gateway alternatives for multi-cloud LLM traffic.

Best for: Azure-first organizations that need Region and provisioned-throughput failover for Azure OpenAI under their existing API Management policies.

Designing failover across three clouds

Cross-cloud failover works when every route is decided before an incident: which errors trigger it, which hosts are allowed for which data, and how model names line up. A gateway enforces those decisions, but it does not make them. For the patterns underneath, see this production guide to retries, fallbacks and circuit breakers in LLM apps and a survey of failover routing strategies for enterprise LLM applications.

A decision flow sending 429 errors to another key or Region, 5xx errors to a retry then the next cloud, and validation or spend-cap errors back to the caller Figure 3: Only throttling and server errors justify leaving a cloud; validation errors and spend caps fail the same way everywhere.

Five design choices matter most:

  1. Exhaust the cheap moves first. Rotate keys, then use Bedrock cross-Region inference, Google’s global endpoint or Azure spillover, before leaving the cloud. In-cloud moves keep billing, logging and data residency unchanged.
  2. Keep one name per model. Map one client-facing alias to the Bedrock inference profile, the Google model name and the Azure deployment name. Version drift between clouds is a common cause of fallback output that differs from the primary. Bifrost’s provider guides list the model mappings and endpoints for Bedrock, Vertex AI and Azure.
  3. Respect residency. A geographic inference profile, a Data Zone deployment and an EU multi-region endpoint keep data inside a boundary; a global route may not. Fallback lists need the same review as primary routes.
  4. Account for caching. Prompt caches are per host, so the first requests after failover pay full input-token prices and run slower.
  5. Plan for streams. A stream that fails after output has started cannot continue on another cloud. Gateways can only retry streams that fail before the first token.

Cost control interacts with failover too: a fallback cloud may price the same model differently, and budgets should cover the fallback path. The Frontier Wire’s guide to LLM cost control with an AI gateway covers budgets and virtual keys in depth. For coding agents that already default to Bedrock or Vertex AI, the companion piece on AI gateways for Claude Code, Codex CLI and Cursor covers the client side, and a separate guide covers running Claude Code on Bedrock, Vertex AI or self-hosted models through a gateway.

How to choose

The choice follows from two questions: must the gateway run inside your own network, and which cloud carries most of the traffic?

SituationStrongest fitWhy
Self-hosted, Claude and GPT across all three cloudsBifrostNative auth for all three, error-specific retries, Apache 2.0
Python platform team, many providersLiteLLMBroadest catalogue, context-window and content-policy fallbacks
Kong already runs the API estateKong AI GatewayPriority failover and circuit breaker in the same control plane
No infrastructure to operateVercel AI GatewayProvider order per model slug, managed
Azure-first, PTU and Region failoverAzure API ManagementBackend pools with priority and Retry-After-aware circuit breaking

Teams that need the gateway in an isolated network should also read the Frontier Wire’s survey of AI gateways for air-gapped and in-VPC deployments, and the ranking of AI gateways for production covers the wider market beyond failover.

The verdict follows the criteria. Cross-cloud failover is a translation problem as much as a routing one: three identity systems, three model naming schemes and three quota models have to line up behind one name. Bifrost covers that translation in its open-source build and makes the most granular error decisions, which is why it leads this list. LiteLLM is the closest alternative for teams comfortable running Python and Redis, Kong and Azure API Management make sense where those platforms already exist, and Vercel is the shortest path for teams that would rather not run a gateway at all. Whichever gateway is chosen, test the fallback path deliberately: revoke a key, inject 429s and confirm in the logs that the second cloud served the request with the model you expected.

Sources

  1. Amazon Bedrock: cross-Region inference
  2. Amazon Bedrock: quotas
  3. Amazon Bedrock: Provisioned Throughput
  4. Amazon Bedrock: API error codes
  5. Google Cloud: quotas for Anthropic Claude models
  6. Google Cloud: request predictions with Claude models
  7. Google Cloud: error code 429
  8. Microsoft Foundry: Azure OpenAI quotas and limits
  9. Microsoft Foundry: spillover for provisioned deployments
  10. Microsoft Foundry: deploy and use Claude models
  11. Anthropic: Claude in Microsoft Foundry
  12. Bifrost docs: retries and fallbacks
  13. Bifrost docs: key management and weighted load balancing
  14. Bifrost docs: provider routing
  15. Bifrost docs: AWS Bedrock provider
  16. Bifrost docs: Vertex AI provider
  17. Bifrost docs: Azure provider
  18. Bifrost docs: circuit breaker (enterprise)
  19. Bifrost docs: enterprise overview
  20. Bifrost docs: benchmarking
  21. Bifrost source repository (GitHub, Apache 2.0)
  22. LiteLLM docs: fallbacks
  23. LiteLLM docs: routing, cooldowns and weighted failover
  24. Kong docs: AI Proxy Advanced plugin
  25. Vercel AI Gateway: provider routing and fallbacks
  26. Vercel AI Gateway: provider filtering and ordering
  27. Vercel AI Gateway: bring your own key
  28. Azure API Management: backends, pools and circuit breaker
  29. Azure API Management: unified model API (preview)
  30. Cloudflare AI Gateway: OpenAI-compatible endpoint

Questions readers ask

Is Claude a Bedrock model?

Claude is Anthropic's model family, and Amazon Bedrock is one of several places it is hosted. The same Claude models are also sold through Anthropic's own API, Google Cloud's Vertex AI (now documented as Gemini Enterprise Agent Platform) and Microsoft Foundry. Each host uses its own model IDs, authentication and quotas, which is why failing over between them needs a translation layer such as an AI gateway.

What happens to my application during an Anthropic outage?

If the application calls only Anthropic's API, requests fail or slow down until the incident is resolved. If the same Claude model is also configured on Bedrock or Vertex AI behind a gateway, the gateway can retry, exhaust its retry budget and send the request to the next host in the fallback chain. Streams that fail after output has started cannot be resumed on another host.

Is Bedrock cross-Region inference the same as gateway failover?

No. Cross-Region inference routes a request to another AWS Region inside Bedrock, using an inference profile tied to a geography or to all commercial Regions. It never leaves AWS and never changes provider. A gateway adds a second layer that can move a request from Bedrock to Vertex AI, Azure or Anthropic's API when the whole Bedrock path is throttled or unavailable.

Does Azure OpenAI spillover replace an AI gateway?

Spillover covers one case well. When a provisioned deployment is fully used and returns 429, 500 or 503, Azure sends the request to a standard deployment of the same model in the same resource. It does not move traffic to another region's resource, another model or another cloud, and it does not manage keys or budgets across providers.

Do Bedrock rate limits apply per Region?

Bedrock quotas are set per account, per Region and per model, and AWS adjusts the defaults by account history. The bedrock-runtime and bedrock-mantle endpoints are tracked against separate quotas even for the same model. Exceeding a quota returns a 429 ThrottlingException, which a gateway should treat as a reason to rotate keys, Regions or providers.

More comparisons