Top 5 AI gateways for Claude Code, Codex CLI and Cursor in 2026
Claude Code, Codex CLI and Cursor each reach a model provider in a different way. Five gateways that put all three behind one set of developer keys, budgets and logs, compared on how each client connects, what governance applies and where the gateway runs.
TL;DR
- A Claude Code proxy is a gateway that holds the provider credential, issues each developer their own key, enforces budgets and logs every request; Claude Code reaches it through
ANTHROPIC_BASE_URL.- Claude Code, Codex CLI and Cursor connect in three different ways: the Anthropic Messages API, the OpenAI Responses API, and an OpenAI-compatible endpoint called from Cursor’s servers, so a shared gateway has to serve all three.
- Bifrost ranks first: one self-hosted Apache 2.0 binary with a native path for each client, virtual keys with budgets, and session pinning that keeps agent runs on a warm prompt cache.
- LiteLLM, Kong AI Gateway, Vercel AI Gateway and OpenRouter complete the list; the managed options are fastest to start, the self-hosted ones keep keys and prompts inside your network.
- Anthropic does not support routing Claude Code to non-Claude models through any gateway, and Cursor requires a gateway reachable from the internet.
Claude Code, OpenAI’s Codex CLI and Cursor each ship with their own login, their own billing relationship and their own way of reaching a model, so a team that uses all three ends up with three sets of credentials, three invoices and no single record of what was sent where. A gateway, often called a Claude Code proxy when it sits in front of Anthropic’s agent, fixes that by putting one control point between every developer machine and every provider. This guide explains how each client connects to a gateway, sets out the criteria that matter for coding traffic, and ranks five gateways against them.
What a Claude Code proxy does
A Claude Code proxy is a service your organization runs, or rents, between Claude Code and the model provider. Developers authenticate to the proxy with a credential it issues; the proxy authenticates to Anthropic, Amazon Bedrock or Google Cloud’s Agent Platform with a credential the organization holds. Budgets, rate limits and request logs are enforced in that one place.
Anthropic’s own gateway overview describes the same split: a developer credential that identifies the person, and a provider credential shared by all forwarded traffic. The benefits it lists are credentials kept server-side, usage attributed by developer or team, budgets and rate limits in one place, audit logging, and the ability to change provider in gateway configuration without touching developer machines. A longer explainer on Claude Code gateways, covering routing, governance and cost control walks through each of those in turn.
The same reasoning applies to Codex CLI and Cursor. Each agent can make dozens of model calls for a single task, so an unmanaged rollout turns into a cost and audit problem quickly. The Frontier Wire’s ranking of AI gateways for production traffic covers the general case; this guide narrows it to the three coding clients most teams run.
Figure 1: Cursor is the odd one out: its requests leave from Cursor’s servers, so the gateway must be reachable from the internet.
A gateway also changes how a subscription is used. Anthropic’s LLM gateway page explains that when Claude Code sends a gateway credential, requests are billed per token to whoever owns the provider key behind the gateway, and the developer’s claude.ai subscription limits no longer apply. Setting only the base URL, without a gateway credential, keeps the subscription as the active login.
How each coding agent connects to a gateway
Each of the three clients exposes a different configuration surface, and the gateway must speak the wire format that client sends. Claude Code speaks Anthropic’s Messages API, Codex CLI speaks OpenAI’s Responses API, and Cursor sends OpenAI-format requests from its own servers. The table summarises what each client documents.
| Client | How it is pointed at a gateway | Wire format | Constraint to know |
|---|---|---|---|
| Claude Code | ANTHROPIC_BASE_URL plus ANTHROPIC_AUTH_TOKEN; or ANTHROPIC_BEDROCK_BASE_URL / ANTHROPIC_VERTEX_BASE_URL with the matching CLAUDE_CODE_USE_* flag | Anthropic Messages, Bedrock InvokeModel or Agent Platform rawPredict | The gateway must forward anthropic-beta and anthropic-version and must not buffer streams |
| Codex CLI | A named [model_providers.<id>] block in config.toml with base_url, env_key and wire_api = "responses", or openai_base_url for a simple proxy | OpenAI Responses | Custom providers cannot reuse the IDs openai, ollama or lmstudio |
| Cursor | Your own provider key in Settings, Models; gateways document an OpenAI base URL override | OpenAI-compatible, sent from Cursor’s servers | Custom keys cover chat models only; Tab completion stays on Cursor’s models |
Claude Code. Anthropic’s gateway compatibility guide is unusually specific. A gateway must expose at least one of three formats, forward the anthropic-beta and anthropic-version headers unchanged, relay streaming events as they arrive, and keep passing the SSE ping events that stop Claude Code’s idle watchdog from aborting a long thinking pause. When the client speaks the Anthropic format but the gateway forwards to Bedrock or Agent Platform, bridging the difference in supported fields is the gateway’s job. For locked-down fleets, a managed setting of allowedProviders: ["customEndpoint"] makes the gateway the only destination a machine may use.
Codex CLI. OpenAI’s advanced configuration page defines custom model providers in config.toml, with a base URL, the environment variable holding the key, optional extra headers and an optional command that fetches short-lived bearer tokens. Codex also ships a built-in amazon-bedrock provider for teams that want Bedrock without a gateway.
Cursor. Cursor’s API key documentation lists OpenAI, Anthropic, Google, Azure OpenAI and AWS Bedrock keys, and states that every request is routed through Cursor’s servers for final prompt building. That has two consequences. A gateway behind a VPN or IP allowlist cannot receive the traffic, and Cursor’s zero data retention policy does not apply to requests made with your own keys. On Teams and Enterprise plans, Cursor also charges its token rate of $0.25 per million tokens on third-party model requests, including those made with your own key.
How the gateways were evaluated
Five gateways were assessed against eight criteria that matter specifically for coding-agent traffic, using each project’s own documentation read in October 2026. The weighting favours teams rolling agents out to many developers, where per-developer identity and spend matter more than a single engineer’s convenience.
| Criterion | What was checked | Why it matters for coding agents |
|---|---|---|
| Claude Code support | Anthropic-format endpoint, documented setup, cloud upstreams | Claude Code is the client most teams standardise first |
| Codex CLI support | Responses API endpoint, documented config.toml setup | Codex sends Responses requests, not Chat Completions |
| Cursor support | Documented Cursor endpoint and its network requirement | Cursor calls from its own servers |
| Per-developer keys and budgets | Keys per person or team, spend caps, rate limits | Agent loops spend quickly and need attribution |
| Prompt-cache continuity | Whether a session stays on the same provider and key | Long agent runs depend on cached context for cost |
| Model choice | Which providers each client can reach through the gateway | Claude on Bedrock or Agent Platform, GPT on Azure, open models |
| MCP tools | Whether the gateway also fronts MCP servers for the agents | Coding agents call tools as much as models |
| Deployment | Self-hosted, managed, or both | Decides where prompts, code and keys travel |
Overhead is not a separate criterion here. A coding-agent turn usually takes seconds of model time, so a gateway’s added latency matters less than whether it streams correctly. Bifrost’s published benchmark is noted in its entry for completeness.
The five gateways at a glance
The table compares the five gateways on the criteria above. “Not documented” means the vendor’s documentation read for this guide did not cover the point; it does not mean the feature is impossible.
| Gateway | Deployment | Claude Code | Codex CLI | Cursor | Per-developer budgets | MCP | Core licence |
|---|---|---|---|---|---|---|---|
| Bifrost | Self-hosted binary, Docker, Kubernetes | /anthropic; Claude on Anthropic, Bedrock, Vertex, Azure, or any configured model | /openai/v1, Responses API | /cursor, must be public | Virtual keys with budgets and rate limits | /mcp, tools filtered per key | Apache 2.0 |
| LiteLLM | Self-hosted Python proxy | Unified or /anthropic endpoint; daily compatibility matrix | /v1/responses, model catalog | /cursor, best effort, must be public | Virtual keys with budgets | MCP endpoint per server | MIT core, commercial enterprise/ |
| Kong AI Gateway | Self-hosted data planes, Konnect control plane | AI Proxy with llm_format: anthropic; guides for Anthropic, Bedrock, Vertex, Azure, Gemini, OpenAI | Konnect guide via OPENAI_BASE_URL | Not documented | AI Rate Limiting Advanced (Enterprise) | AI MCP Proxy (Enterprise) | Apache 2.0 Kong Gateway |
| Vercel AI Gateway | Managed | Dedicated /claude-code endpoint, model picker | Dedicated /codex/v1 endpoint | Dedicated /cursor/v1 endpoint | Team, project, API key and user budgets | Not documented as a gateway feature | Proprietary |
| OpenRouter | Managed | Anthropic-compatible endpoint; guaranteed only with Anthropic’s first-party provider | Custom provider in config.toml | /api/v1/cursor (beta) | Organizational budget controls | Not documented as a gateway feature | Proprietary |
1. Bifrost
Bifrost is an open-source AI gateway written in Go and published under Apache 2.0 by Maxim AI. It ranks first because its free, self-hosted build documents a native path for all three clients, applies one set of per-developer keys and budgets to all of them, and keeps each agent session pinned to the provider that served it. Its documentation covers each client in detail, which is why it gets the longest entry here.
Bifrost runs as a single binary, through Docker, or on Kubernetes from the open-source repository, and puts 25+ providers and 10,000+ models behind one OpenAI-compatible API; its Claude Code resource page summarises the coding-agent setup. Its published benchmark reports 11 µs of added overhead per request at 5,000 RPS on an AWS t3.xlarge.
Figure 2: Each client keeps its native wire format; the virtual key is the common identity across all of them.
Claude Code. The Claude Code guide sets ANTHROPIC_BASE_URL to the gateway’s /anthropic path and ANTHROPIC_AUTH_TOKEN to a virtual key, so the developer needs no Anthropic account login. Model slots can be pinned to Claude on Anthropic, Bedrock, Vertex or Azure (for example bedrock/global.anthropic.claude-sonnet-4-6), or rewritten at request time by routing rules that map an alias such as sonnet-model to any provider. A walkthrough for routing Claude Code through AWS Bedrock maps stable deployment names to Bedrock model IDs, and the Claude Code gateway guide covers cost tracking and multi-model routing. Developers can switch model mid-session with /model. The guide notes that any non-Claude model must handle Claude Code’s tool calls, and that Claude-specific server tools such as web search only exist on Claude models.
Codex CLI. The Codex guide defines a named bifrost provider pointing at /openai/v1 with wire_api = "responses". It recommends a named provider over openai_base_url because the built-in provider sends a web-search tool namespace that Bedrock rejects. Non-OpenAI models, such as anthropic/claude-sonnet-4-5 or gemini/gemini-2.5-pro, require Codex’s HTTPS mode rather than WebSockets, and a local catalog file makes them appear in Codex’s model picker with the right context window. The Codex CLI integration walkthrough shows the full configuration.
Cursor. The Cursor guide uses the OpenAI base URL override with a /cursor path on a publicly reachable deployment and a virtual key in the API key field. Models are named provider/model, so the agent mode can run on Claude while a cheaper model handles chat. The CLI agents overview lists the other terminals and editors that connect the same way.
Session pinning and failover. Claude Code sends x-claude-code-session-id and Codex sends session-id on every request. Bifrost’s session affinity uses those headers with no configuration, keeping a session on the provider and key that served it, so a long run keeps hitting the prompt cache it warmed. When a provider fails, retries and fallbacks rotate keys on 429 and authentication errors, back off on 5xx, and then move to the next provider in the chain.
Governance and tools. Virtual keys carry budgets, token and request rate limits, and provider and model allowlists, so each developer or team gets its own cap. The same gateway serves MCP tools at /mcp, filtered per virtual key, which pairs with the servers in The Frontier Wire’s guide to MCP servers for coding agents. How much context that saves is measured in an analysis of MCP gateway token costs in Claude Code and Codex CLI.
For laptops that are never configured to point at the gateway, Bifrost Edge, in early access, routes Claude Code through the same gateway at the machine level.
Enterprise tier. High-availability clustering, adaptive load balancing, OIDC identity federation, RBAC, audit logs, log exports and in-VPC deployment are part of Bifrost Enterprise, which uses the same configuration as the open-source build.
Best for: teams that run all three coding clients and want one self-hosted gateway, with per-developer budgets and prompt-cache-aware routing in the free build.
2. LiteLLM
LiteLLM is an open-source Python proxy that ranks second because it also documents all three clients and ships virtual keys and budgets in its core. Its Cursor support is explicitly best effort, and it publishes a daily compatibility matrix for Claude Code features by provider.
Claude Code. The Claude Code quickstart points ANTHROPIC_BASE_URL at the proxy’s unified endpoint or at an /anthropic pass-through, with a master key or a model-scoped virtual key as ANTHROPIC_AUTH_TOKEN. For rotating credentials it uses Claude Code’s apiKeyHelper to fetch a JWT. LiteLLM regenerates a Claude Code compatibility matrix daily across Haiku 4.5, Sonnet 4.6 and Opus 4.7 on Anthropic, Bedrock, Vertex AI and Azure, and recommends Bedrock’s Invoke route for Claude Code.
Codex CLI. The Codex setup defines a litellm provider in config.toml against /v1/responses, suggests raising stream idle timeouts for long agent turns, and can publish the gateway’s model catalog to Codex. The same file registers LiteLLM’s MCP endpoints.
Cursor. The Cursor integration uses a /cursor path that must be reachable from the internet. LiteLLM states that Cursor does not officially support AI gateways and that its support comes from reverse engineering Cursor’s API. Agent mode needs LiteLLM v1.97.0 or later, the Cursor CLI cannot target LiteLLM, and newer Cursor builds do not show the base URL override on every plan, in which case LiteLLM documents an Azure OpenAI workaround.
Licence. The core is MIT; code under the repository’s enterprise/ directory has a separate commercial licence.
Best for: Python-centric platform teams that already run LiteLLM for application traffic and want to extend it to coding agents.
3. Kong AI Gateway
Kong AI Gateway adds AI plugins to Kong Gateway, so it suits organizations that already run Kong for APIs. It ranks third because its Claude Code coverage is broad and its Codex guide is current, but it has no Cursor guide and its rate-limiting and MCP plugins are enterprise features.
Claude Code. Kong’s Claude Code guide configures the AI Proxy plugin with llm_format: anthropic, so the gateway accepts Claude’s native format, and raises the request body limit for long prompts. Companion guides route the same client to Bedrock, Vertex AI, Azure, Gemini, OpenAI and Hugging Face upstreams; the Bedrock version adds a policy that strips fields Claude Code sends that Bedrock rejects.
Codex CLI. Kong’s Codex guide creates a model with the agentic capability targeting the OpenAI Responses API, then points Codex’s OPENAI_BASE_URL at the gateway. It is written for Kong Konnect.
Governance and tools. The AI Proxy plugin itself has no enterprise label, but AI Proxy Advanced (load balancing), AI Rate Limiting Advanced and AI MCP Proxy are listed as available only in Kong’s AI Gateway Enterprise offering.
Best for: organizations with an existing Kong estate that want coding-agent traffic under the same control plane as their APIs.
4. Vercel AI Gateway
Vercel AI Gateway is a managed service with the most automated setup in this list: one CLI command configures Claude Code, Codex, Cursor and more than two dozen other agents. It ranks fourth because it runs only on Vercel’s infrastructure.
Setup. The coding agents guide documents vercel ai-gateway setup, which detects installed agents, shows a diff before writing configuration, and stores the key in the macOS Keychain rather than a plaintext file. Claude Code, Codex and Cursor each get a dedicated endpoint: /claude-code, /codex/v1 and /cursor/v1.
Claude Code. The Claude Code page enables gateway model discovery so the catalog appears in /model, and advises setting CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1 when routing to Bedrock or Vertex AI providers, because those reject some Anthropic beta headers.
Budgets. Budgets apply at four scopes: team, project, API key and user, and a request must pass every budget in scope. Requests made with your own provider keys are not counted against any budget.
Best for: teams already on Vercel that want coding agents connected in minutes and accept a managed data path.
5. OpenRouter
OpenRouter is a managed router with documented setups for Claude Code, Codex CLI and Cursor. It ranks fifth because its Claude Code support is guaranteed only when requests go to Anthropic’s own provider, and its Cursor integration is in beta.
For Claude Code, OpenRouter’s guide sets ANTHROPIC_BASE_URL to its Anthropic-compatible API, passes the OpenRouter key as ANTHROPIC_AUTH_TOKEN, blanks ANTHROPIC_API_KEY, and recommends making Anthropic’s first-party provider the top priority. It offers failover between providers that host Anthropic models and organizational budget controls. Codex connects through a custom provider in config.toml; Cursor uses a dedicated /api/v1/cursor base URL.
Best for: individual developers and small teams who want one managed key for many models and do not need to self-host.
Where Anthropic’s Claude apps gateway fits
Anthropic now ships its own self-hosted gateway inside the claude binary. The Claude apps gateway routes to Amazon Bedrock, Claude Platform on AWS, Google Cloud’s Agent Platform, Microsoft Foundry or the Anthropic API, signs developers in through the corporate identity provider, enforces model access by group, and emits OpenTelemetry usage metrics.
It is not ranked here because it serves only Claude clients. A team that runs Codex CLI or Cursor alongside Claude Code still needs a second gateway for those, and two gateways means two sets of budgets and logs. Anthropic’s own documentation also notes that its sign-in is a browser SSO step with no service-token flow, so CI pipelines must connect to the provider directly. For a Claude-only organization that wants no third-party component, it is a sound default.
Budgets, prompt caching and failover for agent traffic
Coding agents stress a gateway differently from chat applications. A single task can chain dozens of model calls that share a long, cached context, so the gateway’s routing decisions affect cost as much as uptime. Three behaviours matter most: spend caps per developer, keeping a session on one provider, and failing over without breaking a stream.
Figure 3: Session pinning keeps a long agent run on the provider whose prompt cache it has already warmed.
Spend caps. A per-developer key with a monthly cap is the simplest control that works. The Frontier Wire’s guide to LLM cost control at the gateway covers budget hierarchies, reset periods and the HTTP status codes a client sees when it hits one. A guide to governing Claude Code, Cursor and Codex at scale applies the same controls to a fleet of developers.
Prompt-cache continuity. Provider prompt caches are scoped to a provider and usually to a key. A load balancer that spreads one session across two keys pays full price for context it already cached. Gateways that read the session headers Claude Code and Codex send avoid that; round-robin balancing without session awareness does not.
Failover. Claude is available from Anthropic, Amazon Bedrock and Google Cloud’s Agent Platform, so a gateway can move Claude Code to another cloud when one returns errors, provided the fallback model accepts the same request fields. The companion piece on automatic failover across Bedrock, Vertex AI and Azure OpenAI covers that pattern, and the failover and load balancing explainer covers the mechanics.
Choosing by team and deployment
The right gateway depends more on where your code is allowed to travel and which clients your developers use than on any single feature. The table maps common situations to the entries above.
| Situation | Start with | Reason |
|---|---|---|
| All three clients, self-hosted, per-developer budgets | Bifrost | Native path for each client and budgets in the free build |
| Already running LiteLLM for applications | LiteLLM | Same keys and budgets extend to agents |
| Existing Kong platform, Claude Code and Codex only | Kong AI Gateway | One control plane for APIs and agents |
| Already on Vercel, managed is acceptable | Vercel AI Gateway | One command configures every installed agent |
| Individual developer wanting many models on one key | OpenRouter | Managed, no infrastructure |
| Claude Code only, SSO required, no third-party gateway | Claude apps gateway | Ships with Claude Code and tracks its releases |
Two checks apply to every option. First, confirm that the gateway forwards Claude Code’s beta headers and streams without buffering, because a gateway that strips them breaks features silently. Second, if Cursor is in scope, confirm that the gateway can be exposed to the internet behind authentication. A complete guide to choosing an AI gateway for Claude Code goes deeper on the Claude side. Teams standardising on one gateway for applications and agents alike should also read the sibling ranking of LLM gateways for OpenAI, Anthropic and Gemini traffic, and the broader AI gateway ranking for criteria beyond coding agents.
Limits of this comparison
This guide was assembled from each vendor’s own documentation, read in October 2026; no gateway was benchmarked for it. Coding-agent clients change quickly: Bifrost’s Claude Code guide, for example, records headers that Claude Code began enforcing in version 2.1.212, and LiteLLM notes that Cursor’s base URL override is not shown on every plan. Check each client’s current documentation before a rollout.
Routing Claude Code to non-Claude models works on several of these gateways, but Anthropic does not support it, and Claude-specific server tools are absent on other models. Treat it as an experiment, not a default; the enterprise guide to using Claude Code with non-Anthropic models covers the tool-calling checks to run first.
For most teams running more than one coding client, a self-hosted gateway that speaks each client’s native format, issues per-developer keys and keeps sessions on a warm cache is the strongest foundation, and on the criteria here Bifrost is the most complete of the five. Managed gateways remain the quickest start when code and prompts may leave your network.
Sources
- Claude Code docs: Gateway overview
- Claude Code docs: Other LLM gateways
- Claude Code docs: Gateway compatibility guide
- Claude Code docs: Claude apps gateway
- Claude Code docs: Amazon Bedrock
- OpenAI Codex docs: Advanced configuration, custom model providers
- Cursor docs: API keys
- Bifrost docs: CLI agents overview
- Bifrost docs: Claude Code
- Bifrost docs: Codex CLI
- Bifrost docs: Cursor
- Bifrost docs: Session affinity
- Bifrost docs: Benchmarking
- LiteLLM docs: Claude Code quickstart
- LiteLLM docs: Connect Codex CLI to LiteLLM
- LiteLLM docs: Cursor integration
- Kong docs: Use Claude Code with AI Gateway and Anthropic
- Kong docs: Use Codex with AI Gateway
- Vercel docs: Coding agents with AI Gateway
- Vercel docs: Claude Code with AI Gateway
- Vercel docs: AI Gateway budgets
Questions readers ask
What is a Claude Code proxy?
A Claude Code proxy is a gateway that Claude Code sends its model requests to instead of calling Anthropic directly. Claude Code is pointed at it with ANTHROPIC_BASE_URL and a gateway-issued credential. The gateway authenticates each developer, applies budgets and rate limits, forwards the request to the Anthropic API, Amazon Bedrock or Google Cloud's Agent Platform with the organization's key, and logs the exchange.
How do I make Claude Code use Amazon Bedrock?
Claude Code can call Bedrock directly by setting CLAUDE_CODE_USE_BEDROCK=1 with AWS credentials, or it can send Anthropic-format requests to a gateway that forwards them to Bedrock. The direct route needs no extra infrastructure. The gateway route adds per-developer keys, budgets and logs, and lets an administrator change the upstream without touching developer machines.
Can Claude Code use models other than Claude through a gateway?
Several gateways, including Bifrost, LiteLLM and Vercel AI Gateway, will route Claude Code requests to GPT, Gemini or open-weight models. Anthropic's documentation states that it does not support routing Claude Code to non-Claude models through any gateway, and the gateways themselves warn that the substitute model must handle Claude Code's tool calls. Treat it as unsupported by Anthropic.
Does Cursor work with a self-hosted AI gateway?
Only if the gateway is reachable from the internet. Cursor builds prompts on its own servers, so requests made with a custom key or base URL leave from Cursor's infrastructure rather than the developer's laptop. A gateway on a private network or behind a VPN cannot receive them. Custom keys also apply only to chat models; Tab completion stays on Cursor's own models.
Do developers on Claude subscriptions need a gateway?
Not for access, but a gateway changes billing. When Claude Code sends a gateway credential through ANTHROPIC_AUTH_TOKEN, requests are billed per token to the account behind the gateway and the developer's claude.ai subscription is not used. Setting only ANTHROPIC_BASE_URL routes traffic through the gateway while the subscription login stays the active credential.
