Independent reporting on artificial intelligence.


The Frontier Wire

Research

LLM failover and load balancing: keeping AI apps up when providers fail

Model APIs fail in documented, predictable ways. The status pages, error codes and rate-limit rules published by OpenAI and Anthropic show which failures a retry fixes, which need a different provider, and which no amount of rerouting will solve.

TL;DR

  • LLM failover moves a request, or a share of traffic, to another key, model or provider when the first one returns errors; load balancing spreads traffic across several of them before anything fails.
  • Provider incidents are routine: Anthropic’s status page listed 16 incidents touching the Claude API in August 2026, five rated major or critical, and OpenAI’s listed 25 incidents between September 10 and 25, 2026.
  • Not every error should trigger a retry. Spend-cap and credit-exhausted 429s, 400-class validation errors and mid-stream failures after output has started need different handling from 5xx and ordinary rate-limit errors.
  • Retries stacked in the SDK, the application and a gateway multiply load on a struggling provider; resilience logic works best in one layer, with a retry budget, jitter and a circuit breaker.
  • A fallback model changes output behavior and cost, so fallback paths need their own evaluation and budget controls.

LLM failover is the practice of sending a model request to a different API key, model or provider when the first choice returns errors, and LLM load balancing is the practice of spreading requests across several of those targets so that one rate limit or outage does not stop the application. Both exist because hosted model APIs fail often enough to plan for: provider status pages record incidents most weeks, and the providers’ own documentation lists the error codes, rate limits and spending caps that produce them. The sections below cover those failure signals, the approaches to handling them, and how one open-source gateway, Bifrost, implements retries, fallback chains, weighted key selection and health-aware routing. Bifrost is an open-source AI gateway written in Go and published by Maxim AI under the Apache 2.0 license; the configuration shown later comes from its documentation.

What LLM failover and load balancing mean

LLM failover and load balancing are two halves of one routing problem. Load balancing decides where a healthy request goes among equivalent targets; failover decides where it goes after a target fails, and for how long traffic avoids that target. The vocabulary is used loosely, so four terms are worth fixing:

  • Retry: the same request is resent to the same provider, usually after an exponential backoff delay.
  • Fallback: the request moves to the next entry in an ordered chain of models or providers after the current one fails.
  • Failover: a broader shift of traffic away from a degraded provider, key or region, often triggered by a health signal rather than a single error.
  • Load balancing: proportional distribution of requests across keys, deployments or providers, by static weight or by measured performance.

Each technique fixes a different failure: a retry fixes a transient network error, key rotation fixes a rate limit on one credential, and a fallback fixes an outage at one provider. None of them fixes an exhausted spending cap or a malformed request.

How LLM providers fail in production

Provider status pages show failures arriving weekly, not yearly. Between August 1 and August 31, 2026, Anthropic’s status page recorded 16 incidents that listed the Claude API as an affected component, five of them rated major or critical, including a critical “Service disruption on Claude services” on August 16. On the day this piece was published, the same page showed an open major incident titled “Elevated errors on claude.ai, Claude Code, Claude Cowork and the Claude API.”

OpenAI’s status page listed 25 incidents between September 10 and September 25, 2026. Most concerned ChatGPT surfaces, but several were API-specific, including “Elevated errors for GPT-5.6 Sol on the API” on September 11 and “We are seeing elevated error rates across API models” on September 17. Many incidents affect one model rather than a whole platform, so a fallback to another model at the same provider is sometimes enough.

The errors that reach an application are documented. The table below collects the codes that matter for routing decisions, taken from the Anthropic error reference and the OpenAI error code guide.

StatusProvider meaningRetry the same target?Reroute?
429 rate limitAnthropic rate_limit_error; OpenAI “Rate limit reached” or slow_downYes, after retry-after or backoffRotate to another key first
429 spend or credit limitAnthropic enforced_spend_limit_reached; OpenAI credit_balance_exhausted, spend-limit codesNoOnly to a different account or provider
500Anthropic api_error; OpenAI server errorYes, with backoffIf retries are exhausted
503OpenAI server_is_overloaded (model lacks capacity)Yes, after Retry-AfterYes, to another model or provider
504Anthropic timeout_errorSometimes; consider streamingYes
529Anthropic overloaded_error (high traffic across all users)Yes, with backoffYes
401, 402, 403Bad key, billing problem, missing permissionNoRotate to another key
400, 404, 422Malformed or invalid requestNoNo; the request itself is wrong

Two further details shape failover design. Anthropic notes that on a streamed response “an error can occur after the API returns a 200 response.” It also warns that a sharp increase in usage can trigger 429 errors from acceleration limits, so a sudden burst of failover traffic arriving at a backup account can itself be rate limited.

Rate limits and the errors retries cannot fix

Rate limits are the errors a healthy provider returns on purpose. Anthropic measures limits in requests, input tokens and output tokens per minute for each model class and uses a token bucket, so capacity refills continuously. Its rate-limit documentation adds that a limit of 60 requests per minute “might be enforced as 1 request per second,” which means short bursts can fail even when the per-minute total looks safe.

OpenAI’s rate-limit guide measures requests and tokens per minute and per day, sets limits at the organization and project level, and places accounts in usage tiers from Free to Tier 5. It also states that “unsuccessful requests contribute to your per-minute limit, so continuously resending a request won’t work.” A retry loop without backoff therefore consumes the same quota it is waiting on.

The failures that no retry or fallback can fix are billing states. Anthropic assigns each of its Start, Build and Scale tiers a monthly spend cap ($500, $1,000 and $200,000 respectively). Once an organization reaches it, requests return a 429 with the same rate_limit_error type as an ordinary rate limit, but with no retry-after header and an error_code of enforced_spend_limit_reached, and Anthropic warns that retrying, “including the SDKs’ automatic retries, fails until access resumes.” A spend limit that an organization sets for itself returns a 400 instead. OpenAI uses separate 429 codes for exhausted prepaid credit and for organization or project spend limits. A router that treats every 429 as transient will spend its whole retry budget on a request that cannot succeed at that account.

Approaches to LLM failover

Resilience logic can live in the application’s own code, the provider SDK, or a gateway between applications and providers. The main difference is how much of the fleet a single decision covers.

ApproachWhere it runsStrengthsWeaknesses
Application retries with backoffEach serviceFull control; no extra hopReimplemented per service and language; no shared view of provider health
SDK built-in retriesProvider client libraryZero code; honors retry-afterSame provider only; no cross-provider fallback; stacks with other layers
Gateway fallbacks and load balancingA proxy in front of all providersOne policy for every service; key pools; shared health and metricsExtra network hop; the gateway must itself be highly available
Weighted multi-key or multi-region poolsGateway or custom routerAbsorbs per-key rate limits; spreads loadAccount-level quotas may still be shared across keys
Circuit breaker and health scoringGateway or service meshStops sending traffic to a failing target; faster recoveryNeeds tuning or signals; can shift load abruptly onto the backup

Client-side and SDK retries

The official SDKs already retry. Anthropic’s SDKs retry connection errors, rate limits and 5xx errors with exponential backoff “twice by default,” honoring the retry-after header. The OpenAI Python library retries connection errors, 408, 409, 429 and 5xx responses two times by default and times out requests after 10 minutes. Both only ever retry the same provider.

Gateway-level fallbacks

A gateway centralizes key pools and fallback chains so that every service inherits them by changing a base URL. The trade-off is operational: the gateway sits on the request path and needs its own redundancy. Readers weighing specific products can start with this LLM gateway failover comparison, which scores failover and routing alongside governance and deployment model, and with the Frontier Wire explainer on running an AI gateway in front of every model call.

Circuit breakers and health checks

A circuit breaker, in Martin Fowler’s description of the pattern, wraps a call and, “once the failures reach a certain threshold,” trips so that further calls return an error immediately without reaching the failing service, then lets a trial call through to test recovery. For LLM traffic the breaker usually redirects rather than rejects: when the primary is open, requests go straight to the fallback instead of paying for retries on a target already known to be down.

Retry amplification

The strongest argument for keeping retries in one layer comes from Google’s SRE book chapter on addressing cascading failures. If three layers each issue three retries, “a single user action may create 64 attempts (4^3)” on the overloaded backend. The chapter recommends randomized exponential backoff, a per-request retry cap and a process-wide retry budget, such as 60 retries per minute. For LLM stacks, that means reducing SDK retries in services whose gateway already retries.

LLM load balancing across keys, regions and providers

LLM load balancing distributes requests across several targets that can serve the same model, so that no single rate-limit bucket or deployment saturates. The targets are usually multiple API keys at one provider, multiple deployments of the same model (for example, an OpenAI model on both OpenAI and Azure), or equivalent models at different providers.

Three strategies cover most deployments:

  • Static weights: each target receives a fixed share, such as 80 percent and 20 percent, which keeps cost and quota planning predictable.
  • Health-aware weights: shares are recalculated from observed error rates and latency, so a degraded target receives less traffic.
  • Sticky sessions: one conversation or agent run stays on the same provider and key, preserving provider prompt caches.

Weights have a limit that is easy to miss. Keys on the same account may share an account-level quota, so rotating keys helps with per-key limits but not with an organization limit. Anthropic sets limits “at the organization level,” and its inference_geo values share one rate-limit pool, so spreading requests across geographies adds no capacity there. Cross-provider balancing adds independent capacity, and it is also where model behavior starts to differ. The Frontier Wire guide to LLM cost control with an AI gateway covers the related cost trade-offs.

How Bifrost implements failover and load balancing

Bifrost handles resilience in two nested loops: retries inside a provider, and fallbacks across providers once retries are exhausted. Load balancing happens before either loop, when the gateway picks a provider and a key. Retries, fallbacks, weighted keys and routing rules are in the open-source gateway; adaptive load balancing and the circuit breaker are enterprise features.

Retries and key rotation

Bifrost’s retry and fallback engine classifies each failure before deciding what to do. Transient server failures (5xx, DNS, refused connections) are retried on the same key with exponential backoff and jitter. Per-key failures rotate to a different key from the provider’s pool: a 429 rotates with backoff still applied, because account-level quotas may be shared, while a 401, 402 or 403 marks the key dead for the rest of that request and rotates immediately without waiting. Validation errors (400, 404, 422), plugin-enforced blocks and cancelled requests are not retried.

Backoff is min(initial × 2^attempt, max) × jitter(0.8 to 1.2), with defaults of 500 and 5,000 milliseconds. Retries are off by default (max_retries: 0) and set per provider. If every key fails authentication or billing, Bifrost returns 502 upstream_credentials_exhausted rather than the raw 401, so callers know the upstream credentials failed, not their own.

Fallback chains

A request can carry a fallbacks array of provider/model strings. When the primary exhausts its retries on a retryable error, Bifrost tries each fallback in order, and each fallback gets its own full retry budget. With three retries on the primary and on each of two fallbacks, a request can make up to 12 attempts. The response’s extra_fields.provider field names the provider that served it. If every fallback fails, Bifrost returns the original primary error, and a plugin can set AllowFallbacks = false to halt the chain when a compliance rule should stop a request rather than reroute it. Each fallback runs as a fresh request, so caching, governance and logging run again.

Weighted keys and virtual keys

Within a provider, weighted key selection picks a key by weighted random choice, so a key with weight 0.7 receives about 70 percent of requests. Keys can be restricted to allowed models or given a blacklisted_models denylist. Across providers, virtual keys carry weighted provider configurations; the documented example sends 80 percent of gpt-4o traffic to Azure and 20 percent to OpenAI. If a request arrives with no fallbacks array, Bifrost builds one automatically from the virtual key’s providers, sorted by weight. Cross-provider routing is not automatic: a request for an OpenAI model will not reach Anthropic unless that model is explicitly allowed there.

Routing rules

For decisions that depend on runtime state, routing rules evaluate CEL expressions before governance provider selection. Expressions can read the model, headers, team and capacity metrics such as budget_used as a percentage; the documented example sends a team’s traffic to Groq, with an OpenAI fallback, when budget_used > 85. Each rule can carry its own ordered fallback list.

Adaptive load balancing and circuit breaking

In Bifrost Enterprise, adaptive load balancing recalculates route weights every 5 seconds from error rate (the primary signal), a latency score measured against peers and the route’s own baseline, and utilization. Routes move between Healthy, Degraded, Failed and Recovering states, and low-weight routes keep a small share of traffic so recovery is detected. The scoring is pre-tuned; operators get five switches, such as re-routing a pinned request whose provider is circuit-broken. Weights are computed per node, and the system is not cost-aware; cost policy belongs to governance routing.

The enterprise circuit breaker trips on provider response headers. Its documented example watches Azure’s X-Ms-Is-Spilled-Over header on a provisioned deployment and, when it appears, sends subsequent requests to a pay-as-you-go deployment for a 30-second default cooldown.

Session affinity

Session affinity keeps requests that share an x-bf-session-id header, or a session header sent by coding agents such as Claude Code and Codex CLI, on the provider and key that served them before. Provider prompt caches and rate-limit buckets are per key, and a conversation that moves between providers can change tool behavior mid-run.

A reference configuration from the Bifrost docs

Both snippets below come from Bifrost’s retries and fallbacks documentation. The first defines three equally weighted OpenAI keys and three retries per request, so Bifrost rotates keys on 429s and backs off on 5xx errors.

{
  "providers": {
    "openai": {
      "keys": [
        { "name": "openai-key-1", "value": "env.OPENAI_KEY_1", "models": ["*"], "weight": 1.0 },
        { "name": "openai-key-2", "value": "env.OPENAI_KEY_2", "models": ["*"], "weight": 1.0 },
        { "name": "openai-key-3", "value": "env.OPENAI_KEY_3", "models": ["*"], "weight": 1.0 }
      ],
      "network_config": {
        "max_retries": 3,
        "retry_backoff_initial": 500,
        "retry_backoff_max": 5000
      }
    }
  }
}

The second is a gateway request with a two-step fallback chain across providers, sent to Bifrost’s OpenAI-compatible endpoint.

curl -X POST http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-4o-mini",
    "messages": [
      { "role": "user", "content": "Explain quantum computing in simple terms" }
    ],
    "fallbacks": [
      "anthropic/claude-3-5-sonnet-20241022",
      "bedrock/anthropic.claude-3-sonnet-20240229-v1:0"
    ],
    "max_tokens": 1000
  }'

The model names are the documentation’s dated examples. Bifrost’s benchmarks report 11 microseconds of added latency per request at 5,000 requests per second on an AWS t3.xlarge, a vendor-reported figure worth checking against real traffic. Teams planning to self-host can compare it with other open-source LLM routers and gateways, a roundup that weighs external state dependencies, air-gapped viability and scaling mechanics.

Pitfalls when failing over between models

A fallback that returns a 200 is not automatically a success. The request reached a different model, with different behavior, price and limits, and several failure modes only appear after the switch.

Model behavior differences

Two models given the same prompt produce different lengths, formats and refusals, and they accept different parameters. Anthropic’s error reference lists model-specific 400 errors: some Claude models reject assistant-message prefill, forced tool choice or thinking settings that older models accept. A request built for one model can therefore fail validation at the fallback. The fallback path needs its own prompt variants and its own evaluation run; the Frontier Wire guide to comparing LLM evaluation frameworks covers the setup.

Streaming failures after output starts

Streaming breaks the simple retry model. Anthropic documents that an overloaded_error can arrive as an event inside a stream that has already returned HTTP 200. If text has been delivered by then, the user has seen part of an answer, and no gateway can splice a different model’s continuation into the same generation. Anthropic’s recovery advice is to capture the partial response and send a continuation request, and it notes that “tool use and extended thinking blocks cannot be partially recovered.” Bifrost’s documented handling reflects the same boundary: for Azure streams it buffers startup metadata so errors that arrive before any output can still trigger retries and fallbacks, but once text, reasoning or tool output arrives, later errors “remain stream errors.” Applications that stream need their own mid-stream error path.

Idempotency and side effects

Retrying a plain completion costs tokens but has no external effect. Agent requests are different: if a response triggered a tool call that wrote to a database or sent a message before the connection dropped, a retry can repeat the action. No model client can know what the application did with a partial response, so tool executions with side effects should carry their own idempotency keys at the tool layer.

Cost of falling back to pricier models

Fallback chains often point from a cheaper model to a more capable one, or to a provider without negotiated discounts, and an outage can move a large share of traffic to the expensive path for hours. Budgets should apply per provider as well as per team; in Bifrost, a governance budget-exceeded error on the primary cascades into the fallback chain like any other failure. The Frontier Wire analysis of why reasoning models cost more per query shows how quickly a switch to a thinking model changes the bill.

Silent failover

A fallback that works well can hide a primary that has been failing for days, so the retry and fallback rate belongs on a dashboard. Bifrost exports a bifrost_request_retries histogram and a bifrost_error_requests_total counter labeled by status code and a normalized error_type through Prometheus, and records each retry and fallback transition in a per-request routing log.

Testing LLM failover before an outage

Untested failover tends to break during the first real incident, because a fallback key was never provisioned, a fallback model was retired, or the backup account’s rate limits are far lower than the primary’s. Testing should inject each failure class from the error table on purpose.

A practical test plan covers:

  • Rate limits: synthetic 429s from one key should rotate to the next key before any cross-provider fallback.
  • Provider outage: 503 or 529 from the primary should end with the fallback serving the request.
  • Bad credentials: a revoked test key should rotate immediately, without backoff.
  • Non-retryable errors: invalid requests and spend-cap responses should not be retried.
  • Latency: client timeouts must exceed the total retry budget, or the client gives up while the gateway is still working.
  • Quality: the evaluation set should run against the fallback model, since a failover that produces wrong answers is still an incident.

Bifrost’s Mocker plugin returns configured success or error responses, such as a 429, with fixed or random latency and a per-rule activation probability, which makes these tests repeatable. Scheduled game days that disable the primary provider in staging catch configuration drift. The review of production-ready AI gateways summarizes how each product approaches failover.

Choosing where failover logic lives

One layer should own retries, fallbacks and load balancing, with other layers configured not to duplicate it. For one service calling one provider, SDK retries are often enough. Once several services call several providers, a gateway keeps key pools, fallback chains, budgets and failure metrics in one place. Buyers comparing products for that role can consult this enterprise AI gateway comparison and the Frontier Wire piece on how an LLM gateway differs from an API gateway.

Deployment constraints narrow the choice further. Regulated teams that cannot send prompts through a third-party service need a gateway they can host inside their own network, and the survey of AI gateways that run inside a private network covers which of them run without external databases or in air-gapped environments. The gateway must not become a new single point of failure; Bifrost addresses that in its enterprise tier with peer-to-peer clustering and zero-downtime deployments.

The routing layer is also where governance and security controls are applied. Beyond failover, Bifrost enforces governance controls such as virtual keys, budgets and rate limits at the gateway, along with security controls such as guardrails and audit logs in its enterprise tier.

Bifrost Edge, currently in alpha, extends those same gateway policies to AI traffic from desktop apps, browser AI and coding agents on employee machines, with guardrails enforced on each device through the same Bifrost instance.

Teams building a shortlist can weigh the self-hosted open-source LLM gateways against the top AI gateways for production workloads, which include managed services.

Reliable LLM failover comes down to a few habits: classify errors correctly, retry once in one place, balance across independent capacity, watch the fallback rate, and test the backup path before an incident does. Teams evaluating a gateway for that job can request a Bifrost demo, review the Bifrost source repository, or compare LLM gateways built for self-hosting before committing.

Sources

  1. Claude API errors (Anthropic)
  2. Claude API rate limits (Anthropic)
  3. Streaming messages: error events and error recovery (Anthropic)
  4. Anthropic status page incident history
  5. OpenAI API error codes
  6. OpenAI API rate limits
  7. OpenAI status page incident history
  8. OpenAI Python library README: retries and timeouts
  9. Site Reliability Engineering, chapter 22: Addressing Cascading Failures (Google)
  10. CircuitBreaker (Martin Fowler)
  11. Bifrost docs: retries and fallbacks
  12. Bifrost docs: keys management and weighted load balancing
  13. Bifrost docs: governance routing with virtual keys
  14. Bifrost docs: routing rules (CEL)
  15. Bifrost docs: provider routing
  16. Bifrost docs: adaptive load balancing (enterprise)
  17. Bifrost docs: circuit breaker (enterprise)
  18. Bifrost docs: session affinity
  19. Bifrost docs: Prometheus metrics
  20. Bifrost docs: Mocker plugin
  21. Bifrost docs: benchmarking
  22. Bifrost docs: Bifrost Edge overview (alpha)
  23. Bifrost source repository (GitHub, Apache 2.0)

Questions readers ask

What is the difference between LLM fallback and LLM failover?

The terms overlap. Fallback usually describes a single request moving to the next model or provider in a configured chain after the first one returns an error. Failover usually describes shifting a larger share of traffic away from a degraded provider, region or key for a period of time, often driven by a circuit breaker or health scores. Most production setups need both.

Which LLM API errors should be retried?

Retry connection errors, timeouts, 5xx responses, Anthropic's 529 overloaded_error and ordinary 429 rate-limit responses, using exponential backoff with jitter and honoring any retry-after header. Do not retry 400, 404 or 422 validation errors, and treat spend-cap or credit-exhausted 429s as non-retryable, because they keep failing until billing changes.

How do you handle LLM rate limits in production?

Spread traffic across several API keys or accounts with weighted load balancing, read the rate-limit response headers, back off with jitter on 429s, and cache repeated prompt prefixes where the provider excludes cached tokens from limits. Anthropic, for example, counts only uncached input tokens toward input-token limits on most models. When one key is exhausted, rotate to another before failing over to a different provider.

Does an AI gateway add latency to failover?

The gateway itself adds a small, measurable overhead; Bifrost reports 11 microseconds per request at 5,000 requests per second on a t3.xlarge in its own benchmark. The larger cost is the retry and backoff time spent on the failing provider before the fallback runs, which is set by retry counts and backoff caps, not by the proxy.

Can failover recover a streaming response that fails halfway?

Not transparently. Once tokens have reached the user, a different model cannot continue the same generation. Anthropic documents a capture-and-continue approach where the partial output is sent back with an instruction to continue, and notes that tool use and thinking blocks cannot be partially recovered. Gateways can only retry streams that fail before output begins.

How do you test LLM failover?

Inject failures deliberately in a staging environment. Return synthetic 429 and 5xx responses from a mock provider, revoke a test key, add latency, then confirm that the fallback served the request, that retry counts and fallback events appear in metrics, and that the fallback model's output still passes the application's evaluation suite.

More research