Independent reporting on artificial intelligence.


The Frontier Wire

Comparisons

Top 5 LLM gateways for Anthropic, OpenAI and Gemini in 2026

Most production teams now call at least two of the three big model providers. Five gateways that put OpenAI, Anthropic and Gemini behind one endpoint, measured on native SDK formats, cross-provider failover, spend control and where the proxy runs.

TL;DR

  • An LLM gateway puts OpenAI, Anthropic and Gemini behind one endpoint, so applications get one set of keys, budgets, retries, failover and logs instead of three integrations.
  • The deciding feature for a three-provider stack is native SDK support: whether the gateway accepts the OpenAI, Anthropic and Google GenAI request formats and translates between them, or only accepts the OpenAI format.
  • Bifrost ranks first: it serves all three SDK formats from one self-hosted Go binary, fails over across providers and enforces budgets in its Apache 2.0 build. LiteLLM, Kong AI Gateway, Cloudflare AI Gateway and Vercel AI Gateway follow.
  • Managed gateways remove operations work but route every prompt through a third party’s network; self-hosted gateways keep prompts and provider keys inside your own infrastructure.

An LLM gateway is the proxy that sits between applications and model providers, and for teams that use OpenAI, Anthropic and Google Gemini together it has become the simplest way to keep three providers manageable. Each provider has its own SDK, request format, error codes, rate limits and billing. Without a gateway, every application carries its own copy of that integration work, its own retry logic and its own copy of the provider keys. With one, applications call a single endpoint, and keys, budgets, failover and logging live in one place. This comparison ranks five gateways on how well they serve exactly that three-provider stack, with the criteria published before the ranking so a reader can reweight them.

What an LLM gateway does for a three-provider stack

An LLM gateway accepts a model request, decides which provider and key should serve it, converts it into that provider’s format, and returns the answer in the format the caller expects. For a stack built on OpenAI, Anthropic and Gemini, its most valuable job is translation, because the three providers do not share a request format.

OpenAI’s Chat Completions format is the closest thing the industry has to a common interface. Anthropic’s Messages API is shaped differently: the system prompt is a separate field, tool results are grouped into user turns, and extended thinking and prompt caching use their own parameters. Gemini’s generateContent API uses different role names and its own structure for tool calls and thinking. A gateway that only speaks the OpenAI format can still reach all three providers, but applications written against the Anthropic or Google SDKs have to be rewritten first.

Applications using the OpenAI, Anthropic and Google GenAI SDKs send requests to one LLM gateway, which translates each format and routes it to OpenAI, Anthropic or Gemini Figure 1: The gateway owns format translation, so an application can change provider without changing SDK.

The rest of the gateway’s work is the same as for any model traffic: authentication with gateway-issued keys rather than raw provider keys, per-team budgets and rate limits, retries and failover, and a request log with tokens and cost. The Frontier Wire’s explainer on LLM gateways versus API gateways covers why token-based limits and streaming make this a different job from conventional API management, and the ranking of the ten leading AI gateways covers the wider market.

For a longer treatment of the components, the overview of LLM gateway architecture, features and use cases published by the gateway’s maintainers walks through each layer of the request path.

Why the providers’ compatibility layers are not enough

Anthropic and Google both offer OpenAI-compatible endpoints, which raises an obvious question: why not point the OpenAI SDK at each provider and skip the gateway? Both providers answer it in their own documentation.

Anthropic describes its OpenAI SDK compatibility layer as “primarily intended to test and compare model capabilities”, and states that it “is not considered a long-term or production-ready solution for most use cases”. The same page lists what is lost: the strict flag for function calling is ignored, audio input is stripped, prompt caching is not supported, and most unsupported fields are silently ignored rather than rejected. Google’s Gemini OpenAI compatibility page is shorter but makes the same recommendation: teams not already using the OpenAI libraries should call the Gemini API directly.

The compatibility endpoints also solve only the format problem. Each still needs its own key, its own rate-limit handling and its own error handling. Anthropic, for example, returns a 529 overloaded_error when its API is under heavy load across all users, alongside the usual 429 and 5xx codes listed in its error reference. Deciding what to do with each of those codes, and where to send the request instead, is the gateway’s job. That is why the evaluation below weights native SDK support and cross-provider failover above raw provider count.

How the gateways were evaluated

Seven criteria were applied, weighted toward teams that run shared production traffic across all three providers. Each product was assessed from its own documentation, licence file and pricing page, read in October 2026.

CriterionWhat was checkedWhy it matters for OpenAI, Anthropic and Gemini
Native SDK formatsWhether the OpenAI, Anthropic Messages and Google GenAI formats are accepted, and whether each can route to any providerDecides whether existing code moves with a base-URL change
Translation fidelityHandling of system prompts, tool calls, thinking and prompt caching across providersSilent parameter loss changes model behaviour
Cross-provider failoverRetries, key rotation, fallback chains that mix providersKeeps the product up during one provider’s incident
Spend controlGateway keys, budgets and rate limits per team, project or keyThree providers means three bills to attribute
Deployment modelSelf-hosted, managed, or bothDecides where prompts and provider keys travel
OverheadPublished added latency, with load and hardwareMultiplies across agent loops
Licence and pricingLicence of the core, paid boundary, token markupTotal cost beyond provider tokens

LLM gateways compared at a glance

The table summarises how each gateway handles the three providers. “Paid boundary” lists notable capabilities that need a commercial licence or plan.

GatewayDeploymentSDK formats acceptedCross-provider fallbackBudgetsPublished overheadPaid boundaryCore licence or pricing
BifrostSelf-hosted binary, Docker, Kubernetes, Go SDKOpenAI, Anthropic, Google GenAI, Bedrock, LiteLLM, LangChainYes, ordered chain, each fallback with its own retriesVirtual keys with budgets and rate limits11 µs at 5,000 RPS (t3.xlarge)Clustering, adaptive load balancing, guardrails, SSO, RBAC, audit logsApache 2.0
LiteLLMSelf-hosted Python proxy or SDKOpenAI, Anthropic /v1/messages, provider pass-throughYes, between model groupsVirtual keys with budgets and rate limits8 ms p95 at 1,000 RPS (fake endpoint)SSO, SCIM, audit logs, admin roles, IP allowlistsMIT, commercial enterprise/ directory
Kong AI GatewaySelf-hosted data planes, KonnectOpenAI format; native Anthropic and Gemini formats only to the matching providerYes, in AI Proxy Advanced (3.10+ across mixed providers)Not assessedNot published in docs readAI Proxy Advanced, including load balancing and failoverApache 2.0 Kong Gateway
Cloudflare AI GatewayManaged, Cloudflare networkOpenAI-compatible endpoint, provider-native endpointsYes, via the Universal endpoint or dynamic routesBudget and rate-limit nodes in dynamic routesNot publishedWorkers Paid for higher log volume and LogpushCore free; 5% on unified billing credits
Vercel AI GatewayManagedOpenAI Chat Completions and Responses, Anthropic Messages, AI SDKYes, models fallback listTeam, project, key and user budgetsNot published in docs readCustom reporting and some controls are add-onsNo token markup

Bifrost leads because it is the only entry that accepts all three providers’ native SDK formats, routes any of them to any provider, and includes failover and budgets in an open-source build. A team already standardised on Kong, Cloudflare or Vercel could reasonably move that entry up, because consolidating on an existing platform can outweigh any single feature. A separate comparison of gateways for routing between OpenAI, Anthropic and Gemini covers the same question with a routing focus.

1. Bifrost

Bifrost is an open-source AI gateway written in Go by Maxim AI and published under Apache 2.0 in the Bifrost repository. It runs as a single binary, in Docker or on Kubernetes, or embeds into Go services as an SDK. It covers 25+ providers and 10,000+ models through one OpenAI-compatible API, including OpenAI, Anthropic, Gemini, Vertex AI, Bedrock and Azure; a walkthrough of configuring GPT, Gemini and Claude providers in one gateway shows the setup. It gets the longest write-up here because its documentation covers every criterion in the table, and because it is the entry most directly built around the three-SDK problem.

Native SDK formats. Bifrost exposes protocol adapters rather than a single OpenAI-shaped endpoint. The drop-in replacement model serves the OpenAI SDK at /openai, the Anthropic SDK at /anthropic and the Google GenAI SDK at /genai, alongside Bedrock, LiteLLM and LangChain integrations. Moving an existing application is a base-URL change plus a Bifrost virtual key in place of the provider key. Inside each SDK, models can be prefixed with a provider, so code written against the Anthropic SDK integration can call an OpenAI or Gemini model and receive an Anthropic-shaped response. The Google GenAI SDK integration works the same way in the other direction.

Translation fidelity. The provider pages document what each conversion does. For Anthropic, Bifrost extracts the system message into Anthropic’s separate field, merges consecutive tool messages into one user turn, maps OpenAI-style reasoning parameters to Anthropic’s thinking structure and renames fields such as max_completion_tokens to max_tokens. A top-level cache_control directive on /anthropic/v1/messages is forwarded unchanged, and an optional setting injects a cache breakpoint for agent clients that send none. For Gemini, it remaps roles, groups tool responses, preserves tool-call IDs and thought signatures, and maps thinking configuration. The provider guides for Anthropic, Gemini and OpenAI list the models, endpoints and field mappings for each.

Failover. Per the retries and fallbacks documentation, Bifrost separates two failure types. Per-key failures (401, 402, 403, 429) rotate to another key from the pool; rate-limit rotations still back off, because provider quotas are often shared across keys. Transient server failures (5xx, network errors) retry on the same key with exponential backoff and jitter. Validation errors such as 400 are not retried. When retries are exhausted, the request moves to the next entry in a fallbacks list of provider/model strings, each with its own retry budget, and the response records which provider served it. Retries default to zero, so they have to be configured per provider.

Spend control. Virtual keys are the unit of governance: each carries provider and model permissions, budgets and request and token rate limits, and keys roll up into teams and customers for hierarchical budgets. Budget resets run from one minute to one year, with calendar-aligned periods available. One virtual key can therefore cap a team’s combined spend across OpenAI, Anthropic and Gemini. The governance resource page summarises how keys, budgets and enterprise RBAC fit together.

Overhead. Bifrost’s published benchmark reports 11 µs of added overhead per request at 5,000 RPS on an AWS t3.xlarge, with a 100 percent success rate, in the t3.xlarge benchmark. A benchmark comparison against LiteLLM sets those figures beside a Python proxy under the same load.

Enterprise tier. The enterprise overview places high-availability clustering, adaptive load balancing, content guardrails, identity federation, RBAC, audit logs, log exports and in-VPC deployment in Bifrost Enterprise, which uses the same configuration schema as the open-source build.

Best for: teams that already have applications on two or three of the provider SDKs and want them behind one self-hosted endpoint, with failover and budgets, without rewriting client code.

2. LiteLLM

LiteLLM is a Python SDK and proxy server that calls more than 100 providers through the OpenAI format. It is the most common first gateway for Python teams and has the longest provider catalogue in this list.

SDK formats. Beyond OpenAI-format endpoints, LiteLLM’s /v1/messages endpoint accepts the Anthropic Messages format and routes it to any supported provider, including OpenAI, Bedrock, Vertex AI and Gemini, with cost tracking, logging, fallbacks and load balancing. For Gemini, its provider page distinguishes the gemini/ prefix, which uses a Google AI Studio API key, from Vertex AI, which needs full Google Cloud credentials, and documents a pass-through endpoint for the native Gemini API.

Failover. LiteLLM fallbacks move a request from one model group to another after the configured number of retries, so a failing model or provider fails over to a backup. Fallbacks are defined between named model groups in the router configuration rather than per request.

Spend control and paid boundary. The open-source proxy includes virtual keys, budgets and rate limits, spend tracking, an admin UI, load balancing, guardrails and an MCP gateway. The enterprise page adds SSO and SCIM, audit logs, organisation and team admins, secret managers, IP allowlists, per-team logging and multi-region support. The repository is MIT-licensed except an enterprise/ directory under a separate commercial licence, per its licence file.

Overhead. LiteLLM’s benchmark page reports 8 ms p95 latency at 1,000 requests per second against a fake OpenAI endpoint, and a separate high-throughput profile, available in nightly builds, reaching 3,000 requests per second with very large prompts.

Best for: Python-heavy teams that want an in-process SDK or a proxy with the widest provider catalogue, and that are comfortable operating a Python service.

3. Kong AI Gateway

Kong AI Gateway is the AI layer of Kong Gateway, whose core is Apache 2.0. It suits organisations that already route API traffic through Kong and want model traffic under the same control plane.

SDK formats. The open-source AI Proxy plugin, available from Kong Gateway 3.6, accepts OpenAI-format requests and translates them for one configured provider. From 3.10, setting llm_format to a native format such as anthropic or gemini passes requests upstream without transformation, but in that mode only the matching provider is supported. Cross-provider routing therefore uses the OpenAI format.

Failover and paid boundary. The AI Proxy Advanced plugin is part of Kong’s AI Gateway Enterprise offering and needs Kong Gateway 3.8 or later. It balances across targets with round-robin, consistent hashing, least connections, lowest latency, lowest usage by tokens or cost, semantic routing and priority-based failover. From 3.10, fallback works across targets with any supported format, so OpenAI, Anthropic and Gemini targets can sit in one group; earlier versions needed compatible formats between fallback targets. Client errors do not trigger failover. Its provider list includes OpenAI, Azure OpenAI, Bedrock, Anthropic, Gemini and Vertex AI.

Best for: enterprises already running Kong that want AI routing, rate limiting and failover in the same gateway as their other APIs, and that hold or plan an AI Gateway Enterprise licence.

4. Cloudflare AI Gateway

Cloudflare AI Gateway is a managed gateway that runs on Cloudflare’s network. Teams create a gateway in a Cloudflare account and change a base URL; there is no proxy to operate.

SDK formats. Cloudflare exposes provider-native endpoints for OpenAI, Anthropic, Google AI Studio, Vertex AI, Bedrock, Azure OpenAI and more than a dozen other providers, so each SDK can keep its own format when it calls its own provider. An OpenAI-compatible endpoint accepts {provider}/{model} names across providers; Cloudflare now marks it deprecated for single-model calls in favour of its REST API, while existing integrations keep working.

Failover. Fallbacks are defined on the Universal endpoint as an ordered array of provider requests, triggered by errors or request timeouts, and a cf-aig-step response header reports which step served the request. Dynamic routing adds versioned flows with conditional branches, percentage splits, model nodes and rate-limit and budget nodes that switch to a fallback when exceeded. Dynamic routes accept only the OpenAI chat completions shape; an Anthropic Messages request to a dynamic route returns a 400.

Pricing. Per the pricing page, dashboard analytics, caching and rate limiting are free on all plans, DLP scanning is free with two predefined profiles, guardrails are billed as Workers AI inference, and unified billing adds 5 percent to credit purchases while passing provider token prices through without markup.

Best for: teams already on Cloudflare that want a free, managed gateway with caching, analytics and basic failover, and that are comfortable with prompts passing through Cloudflare’s network.

5. Vercel AI Gateway

Vercel AI Gateway is a managed gateway that applications can call from any infrastructure, not only from Vercel deployments. It is the most direct option for teams already building with Vercel’s AI SDK.

SDK formats. The SDKs and APIs documentation lists the AI SDK, OpenAI Chat Completions, the OpenAI Responses API and the Anthropic Messages API. Google’s native GenAI format is not among the documented APIs, so Gemini models are reached through one of those formats rather than through the Google SDK.

Failover and spend control. A models array in providerOptions.gateway defines model fallbacks tried in order, and the same option works across every API format, so a Claude request can fall back to a Gemini model. Budgets apply at team, project, API key and user scope; every budget in scope is checked before each request.

Pricing. The pricing page states that the gateway charges no markup and no platform fee on tokens, including with bring-your-own keys, which require the paid tier. When a request with the customer’s own credentials fails, the gateway retries with Vercel’s credentials and charges that usage to the credit balance, a detail worth knowing for teams that need traffic to stay on their own provider contracts.

Best for: TypeScript teams on the AI SDK that want a managed gateway with model fallbacks and layered budgets and no token markup.

Failover across Anthropic, OpenAI and Gemini

Cross-provider failover is the main reliability reason to put all three providers behind one gateway. When Anthropic returns 529 or 5xx errors, or a key hits its rate limit, the gateway retries and then sends the request to an equivalent model at OpenAI or Google, and the application sees a normal response.

A request goes to Anthropic first; rate-limit, overload or server errors move it to OpenAI and then Gemini, while invalid-request errors return straight to the caller Figure 2: Only provider-side failures should trigger a fallback; a malformed request fails the same way everywhere.

Three design points separate a dependable setup from one that hides problems:

  • Classify errors before acting. Rate limits and server errors justify a retry or a fallback. Invalid requests do not, because they fail identically at every provider. Bifrost does not retry validation errors, and Kong’s documentation states that client errors do not trigger failover.
  • Match capability, not just availability. A fallback model must support the features the request uses: tool calling, the context length, structured output. Translation between formats keeps the request valid, but it cannot add a capability the fallback model lacks.
  • Record which provider answered. Bifrost returns the serving provider in the response, Cloudflare uses the cf-aig-step header, and Vercel records every routing attempt. Without that signal, quality changes caused by a silent fallback are hard to trace.

For the top-ranked gateway, the guide to enabling automatic fallback when a primary LLM provider fails and the article on routing, fallback and governance in the gateway walk through the configuration step by step.

The Frontier Wire’s guide to LLM failover and load balancing covers retry budgets and circuit breaking in more depth, and the companion comparison of gateways for failover across Bedrock, Vertex AI and Azure OpenAI covers the case where the same model is served by several clouds.

Self-hosted or managed

The second decision after SDK support is where the gateway runs. Bifrost, LiteLLM and Kong are self-hosted: the proxy runs in the team’s own cloud account or cluster, and prompts and provider keys stay inside that network until they reach the provider. Cloudflare and Vercel are managed: nothing to operate, but every request passes through the vendor’s network first.

Top lane shows an application calling a self-hosted LLM gateway inside its own network before the providers; bottom lane shows the application calling a managed gateway on the vendor's network Figure 3: With a managed gateway every prompt crosses a third party’s network; with a self-hosted one the team operates the proxy.

Neither model is better in general. The table maps common situations to the gateway that fits best.

SituationBest fitReason
Existing apps on the OpenAI, Anthropic and Google SDKsBifrostNative endpoints for all three formats, routed to any provider
Python services, long-tail providersLiteLLMWidest catalogue, in-process SDK option
API traffic already on KongKong AI GatewayOne control plane for APIs and models
Already on Cloudflare, want zero operationsCloudflare AI GatewayFree core features, provider-native endpoints
TypeScript apps on the AI SDKVercel AI GatewayModel fallbacks and budgets with no token markup
Prompts must not leave your networkBifrost, LiteLLM or KongSelf-hosted data plane

Teams with strict data-residency requirements should read the companion comparison of gateways for air-gapped and in-VPC deployments, and teams moving off LiteLLM should read the LiteLLM alternatives comparison alongside the project’s own page on replacing LiteLLM as an LLM gateway.

The bottom line

For a stack that runs on OpenAI, Anthropic and Gemini at once, the gateway that wins is the one that removes the most integration work without losing provider features in translation. On that test Bifrost comes first: three native SDK formats on one self-hosted endpoint, failover that mixes providers, virtual-key budgets and published microsecond-level overhead, all in the Apache 2.0 build. LiteLLM is the strongest alternative for Python teams, Kong for existing Kong estates, and Cloudflare and Vercel for teams that would rather not run a proxy at all. Whichever gateway a team chooses, it should test failover with real provider errors before trusting it in production, and compare latency at its own payload sizes. The full AI gateway rankings place these five among the wider market.

Sources

  1. Bifrost source repository (GitHub, Apache 2.0)
  2. Bifrost docs: drop-in replacement
  3. Bifrost docs: Anthropic SDK integration
  4. Bifrost docs: Google GenAI SDK integration
  5. Bifrost docs: retries and fallbacks
  6. Bifrost docs: virtual keys
  7. Bifrost docs: t3.xlarge benchmark
  8. Bifrost docs: enterprise overview
  9. LiteLLM docs: /v1/messages across providers
  10. LiteLLM docs: fallbacks
  11. LiteLLM docs: enterprise
  12. LiteLLM docs: benchmarks
  13. LiteLLM licence file
  14. Kong docs: AI Proxy plugin
  15. Kong docs: AI Proxy Advanced plugin
  16. Cloudflare AI Gateway: provider-native endpoints
  17. Cloudflare AI Gateway: fallbacks
  18. Cloudflare AI Gateway: dynamic routing
  19. Cloudflare AI Gateway: pricing
  20. Vercel AI Gateway: SDKs and APIs
  21. Vercel AI Gateway: model fallbacks
  22. Vercel AI Gateway: budgets
  23. Vercel AI Gateway: pricing
  24. Anthropic: OpenAI SDK compatibility
  25. Anthropic: Claude API errors
  26. Google: Gemini API OpenAI compatibility

Questions readers ask

What is an LLM gateway?

An LLM gateway is a proxy between applications and model providers. Applications send every request to one endpoint, and the gateway authenticates the caller, applies budgets and rate limits, translates the request into the provider's format, retries or fails over on errors, and logs tokens and cost. It replaces per-application integration code for each provider.

Which LLM gateway is the best for OpenAI, Anthropic and Gemini together?

On the criteria in this comparison, Bifrost ranks first because it accepts the OpenAI, Anthropic and Google GenAI SDK formats natively, translates between them, fails over across providers and enforces budgets in its open-source build. LiteLLM suits Python-heavy teams, Kong suits existing Kong users, and Cloudflare and Vercel suit teams that want a managed service.

Are LLM gateways free?

Several are. Bifrost, LiteLLM and Kong Gateway have open-source cores that cost nothing to run beyond infrastructure, with paid enterprise tiers for features such as SSO, audit logs and clustering. Cloudflare AI Gateway's core features are free, and Vercel AI Gateway charges no markup on tokens. Provider token costs apply in every case.

What is the difference between an MCP gateway and an LLM gateway?

An LLM gateway governs traffic from applications to models. An MCP gateway governs traffic from agents to tools exposed over the Model Context Protocol. Several products, including Bifrost, LiteLLM and Kong, now handle both, so one key can carry model permissions and tool permissions.

Can one gateway serve the OpenAI, Anthropic and Gemini SDKs at the same time?

Yes, if it exposes each provider's native API shape. Bifrost exposes OpenAI, Anthropic and Google GenAI endpoints on one server, and LiteLLM accepts OpenAI and Anthropic formats. Others accept the OpenAI format everywhere but pass native formats only to the matching provider, so cross-provider routing needs the OpenAI format.

More comparisons