Independent reporting on artificial intelligence.


The Frontier Wire

Comparisons

Top 10 AI gateways in 2026, ranked on production criteria

An AI gateway now sits in front of most production model traffic. Ten options, from open-source binaries to cloud API platforms and managed routers, measured against ten published criteria before any ranking.

TL;DR

  • The ten AI gateways worth shortlisting in 2026 are Bifrost, Kong AI Gateway, LiteLLM, Cloudflare AI Gateway, Agent Router (formerly Envoy AI Gateway), agentgateway, Azure API Management, Vercel AI Gateway, OpenRouter and Apache APISIX.
  • They fall into three groups: standalone self-hosted gateways, AI features added to an existing API-gateway platform, and managed services that route traffic through the vendor’s network.
  • Bifrost ranks first on these criteria: an Apache 2.0 Go binary with the lowest vendor-reported overhead in the list and virtual keys, budgets, fallbacks and MCP tool filtering in its open-source build. Its clustering, adaptive load balancing, guardrails and SSO are enterprise-tier.
  • The right choice depends more on deployment model than on feature count. A team that already runs Kong, Envoy or Azure should check what its existing platform can do before adding another proxy.

An AI gateway is the proxy that sits between applications and model providers, and in 2026 it has become a standard part of production AI infrastructure. It gives every service one API, enforces who may spend what on which model, retries and fails over when a provider misbehaves, and records every request for cost and audit. The market has also split. Some products are standalone gateways you run yourself. Some are AI features inside an API platform a company already pays for. Some are managed services that you reach by changing a base URL. This guide ranks ten of them against ten production criteria, stated in full before the ranking so a reader can reweight them.

What an AI gateway does

The gateway’s job is easiest to see from the request path. An application sends a chat completion to one endpoint. The gateway authenticates the caller, checks the caller’s budget and rate limit, picks a provider and key, forwards the request, and on failure retries or moves to a fallback provider. On the way back it counts tokens, prices them, logs the exchange and optionally runs a guardrail over the response.

Each of those steps replaces code that would otherwise live in every application. That is why the category grew quickly and why it overlaps with conventional API management. The explainer on LLM gateways versus API gateways covers where the two differ: token-based limits, streaming, provider-specific request formats and per-request cost. The guide to running an AI gateway in front of every model call covers the operational case.

Two shifts define the 2026 market. First, agents. An agent may make dozens of model calls per task, so gateway overhead multiplies and spend control matters more. Second, the Model Context Protocol (MCP), the open protocol through which agents call tools. Several gateways now proxy MCP traffic as well as model traffic, so the same key that limits model spend can limit which tools an agent may call. The Frontier Wire’s comparison of MCP gateways covers that layer in depth.

This list covers the whole market, self-hosted and managed. Readers moving off LiteLLM specifically should read the LiteLLM alternatives comparison, which weights migration effort. Readers who need an open-source licence should read the sibling ranking of open-source AI gateways.

How the gateways were evaluated

The ten criteria below are weighted toward teams running shared, production model traffic. Every entry was assessed from its own documentation, repository, licence file, pricing page or release notes, read in late September and early October 2026.

CriterionWhat to checkWhy it matters
Added latencyVendor-reported overhead, the load and hardware it was measured onMultiplies across agent loops; vendor tests use mocked upstreams
Provider coverageNative providers and self-hosted model supportDecides whether every model you use fits behind one endpoint
OpenAI-compatible APIChat Completions format, other SDK formats acceptedDetermines whether adoption is a base-URL change
Failover and load balancingRetries, provider fallback chains, key and backend balancingKeeps the product up during provider incidents
Budgets, virtual keys, rate limitsPer-key and per-team spend caps, token limitsThe main reason finance and platform teams adopt a gateway
GuardrailsPII redaction, prompt-injection checks, content moderation, third-party hooksCentral policy instead of per-app filters
ObservabilityRequest logs, token and cost metrics, OpenTelemetry, PrometheusSpend attribution and incident debugging
MCP supportProxying MCP servers, per-consumer tool filteringAgents call tools as well as models
Licence and self-hostingOSI licence of the core, where the data plane can runDecides where prompts and keys live
Pricing modelFree core, enterprise tier, usage fee or markupTotal cost beyond provider tokens

A note on the first criterion. Overhead figures in this guide are the vendors’ own, measured on their chosen hardware with mocked providers. They are useful for order of magnitude only. The gap between a gateway that adds microseconds and one that adds milliseconds is real, but at a typical model response time of hundreds of milliseconds, governance and reliability features usually matter more. Benchmark at your own payload size and concurrency before deciding on latency.

AI gateways compared at a glance

The table summarises each gateway against the criteria. “Paid tier” lists notable controls that need a commercial licence or plan.

RankGatewayDeploymentCore licenceVendor-reported overheadMCPNotable paid-tier controlsPricing model
1BifrostSelf-hosted binary, Docker, Go SDKApache 2.011 µs at 5,000 RPS (t3.xlarge, mocked)Yes, with per-key tool filteringClustering, adaptive balancing, guardrails, SSO, RBAC, audit logsFree core; enterprise licence
2Kong AI GatewaySelf-hosted data planes; Konnect control planeApache 2.0 Kong GatewayNot published in docs readYes (Enterprise plugin)Advanced load balancing, MCP proxyEnterprise or Konnect subscription
3LiteLLMSelf-hosted Python proxy or SDKMIT, plus commercial enterprise/2 to 8 ms, p50 to p95 (4 instances, ~1,170 RPS)YesSSO, RBAC, audit logs, per-key guardrailsFree core; quoted enterprise
4Cloudflare AI GatewayManaged on Cloudflare’s networkProprietaryNot publishedSeparate MCP server portalsLog volume, Logpush, full DLPCore free; 5% on unified billing credits
5Agent RouterKubernetes on Envoy Gateway, or local CLIApache 2.0Not publishedYesGuardrails and cost attribution in Tetrate’s enterprise editionFree; commercial edition
6agentgatewayStandalone binary or KubernetesApache 2.0Not publishedYes, plus A2AEnterprise support from Solo.ioFree; commercial support
7Azure API ManagementAzure-managed; self-hosted gateway on Developer and PremiumProprietaryNot publishedYesVaries by service tierAzure tier pricing
8Vercel AI GatewayManagedProprietary“Sub-20ms” (vendor GA post)Not documented as a gateway featureNot tiered by featureNo token markup
9OpenRouterManaged router and billingProprietaryNo figure in docs readNo tool gatewayCustom BYOK allowance on enterprise5.5% card fee on credits; 5% BYOK fee above allowance
10Apache APISIXSelf-hosted API gatewayApache 2.0Not publishedNot confirmed in docs readNone in the core projectFree

The ranking rewards a standalone gateway that combines low overhead, a permissive licence, strong governance in the free build and MCP support. A team standardised on Kong, Envoy or Azure would reasonably move those entries up, because consolidating on an existing control plane can outweigh any single feature.

Leading AI gateways, ranks 1 to 3

The top three cover every criterion in some form. Two run as standalone services on your own infrastructure; one adds AI policy to a general API platform.

1. Bifrost

Bifrost is an open-source AI gateway written in Go by Maxim AI, published under Apache 2.0 in the Bifrost repository. It runs as one binary, through npx or Docker, or embeds into Go services as an SDK. It ranks first because it scores well on the most criteria at once without a paid licence: overhead, OpenAI compatibility, failover, budgets and MCP.

Latency. Bifrost’s benchmarking documentation reports 11 µs of added overhead per request at a sustained 5,000 requests per second on an AWS t3.xlarge (4 vCPUs, 16 GB), with a 100 percent success rate against mocked OpenAI calls. On a smaller t3.medium the same test reports 59 µs. These are vendor figures against a mocked upstream, but they are the lowest published in this list by roughly two orders of magnitude.

Coverage and compatibility. The provider documentation lists more than 30 providers, including OpenAI, Anthropic, Bedrock, Vertex AI, Azure, Gemini, Mistral, Groq, Cerebras, xAI, Ollama and vLLM, plus custom providers. The drop-in replacement model accepts the OpenAI, Anthropic and Google GenAI SDKs, LiteLLM and LangChain by changing only the base URL.

Reliability. Per the retries and fallbacks documentation, rate-limit and authentication failures (429, 401, 403) rotate to another key from the pool, transient 5xx errors retry on the same key with exponential backoff and jitter, and exhausted retries move the request to the next provider in a fallback chain, which gets its own retry budget. Keys are selected by weighted round robin.

Governance. Virtual keys carry their own budgets, with resets from one minute to one year, token and request rate limits, and provider and model allowlists that deny by default. Keys roll up into teams and customers for hierarchical budgets. The Bifrost governance overview describes the enforcement: a key over budget gets HTTP 402, a key over its rate limit gets 429, and both are enforced inline rather than reported afterwards.

Observability and MCP. Bifrost exports Prometheus metrics, OpenTelemetry traces using GenAI semantic conventions, and a request log with tokens, cost and latency, as described on its gateway observability page. Content logging can be switched off while keeping token and cost metrics. Semantic caching works with Redis or Valkey, Weaviate, Qdrant or Pinecone. Bifrost is also an MCP client and gateway: it aggregates upstream MCP servers behind /mcp, filters tools per virtual key, and by default treats model tool calls as suggestions that need a separate execution call.

Limits. The enterprise overview places clustering, adaptive load balancing, circuit breaking, guardrails, OIDC user provisioning, RBAC, audit logs, log exports, Virtual MCP bundles and in-VPC deployment on the enterprise tier. A team that needs high availability across nodes or central content guardrails is buying a licence, as with most entries here. The guardrails page lists integrations with AWS Bedrock Guardrails, Azure Content Safety and Google Model Armor, among others. Bifrost’s provider list is shorter than LiteLLM’s, so check long-tail providers before committing. Maxim AI, the company behind Bifrost, also ships Bifrost Edge, which extends the gateway to desktop apps and coding agents on employee machines; it is in alpha and was not assessed here.

Best fit. Teams that want a self-hosted gateway with low overhead and real governance in the free build. The Bifrost LLM gateway page summarises the open-source and enterprise split.

2. Kong AI Gateway

Kong AI Gateway is the AI layer of Kong Gateway, an API gateway built on Nginx and OpenResty with an Apache 2.0 core. Its documentation describes it as a connectivity and governance layer for LLM, MCP and agent-to-agent (A2A) traffic. It ranks second because it covers every criterion except published overhead and is the most complete option for organisations that already standardise API traffic on Kong.

Features. The documentation lists load balancing and failover across providers, semantic caching, prompt compression, token-cost-aware rate limiting, keyword and semantic prompt guards, PII redaction through its AI Sanitizer, third-party guardrails from AWS, Azure and GCP, and observability through OpenTelemetry, audit logs and Konnect dashboards.

Paid boundary. The AI Proxy Advanced plugin is part of the AI Gateway Enterprise offering and needs Kong Gateway 3.8 or later. It offers seven balancing strategies: weighted round robin, consistent hashing, least connections, lowest latency, lowest usage (by tokens or cost), semantic routing and priority-based failover, across more than 15 providers. The AI MCP Proxy plugin is also Enterprise-only (3.12 or later). It can proxy existing MCP servers or convert REST APIs into MCP tools, and from 3.13 it filters tool discovery per consumer with default and per-tool ACLs.

Recent release. Kong AI Gateway 2.2, announced on 30 September 2026, adds a passthrough mode for non-standard and self-hosted interfaces such as vLLM, Ollama and NVIDIA NIM, credential-based advanced rate limiting, Skills APIs and custom plugins for the AI control plane.

Limits. The features most teams want from an AI gateway are on the paid tier. Kong does not publish an AI-gateway overhead figure in the pages read for this guide, and commercial terms were not assessed.

Best fit. Enterprises already running Kong for API traffic that want AI and MCP policy in the same control plane.

3. LiteLLM

LiteLLM is a Python SDK and proxy server that exposes more than 100 providers through the OpenAI format. It has about 60,000 GitHub stars, the widest provider catalogue in this list, and the longest track record as a team’s first gateway. It ranks third, behind Kong, because its operational footprint and paid boundary weigh against it for large shared deployments.

Licence and pricing. The licence file places the repository under MIT, except an enterprise/ directory under a separate commercial licence. The enterprise page lists SSO, JWT authentication, audit logs, RBAC, IP allowlists, key rotation, secret-manager integration, tag-based budgets, team-based logging and per-key guardrails as enterprise features. Pricing “depends on your deployment size” and is quoted on request.

Latency. LiteLLM’s benchmark page reports gateway overhead of 2 to 8 ms (p50 to p95) with four instances of 4 CPU and 8 GB at about 1,170 requests per second against a fake OpenAI endpoint, with PostgreSQL attached. A high-throughput profile using a Rust token counter, PgBouncer and sidecars for spend tracking reports 3,000 requests per second.

Governance, reliability and MCP. Virtual keys, per-project spend tracking and budgets, load balancing across deployments and fallbacks are in the open-source proxy. Its MCP gateway exposes a fixed endpoint for tools over streamable HTTP, SSE and stdio, with key-level and team-level server permissions and OAuth 2.0 for upstream servers.

Limits. Production deployments need PostgreSQL and usually Redis, and the March 2026 PyPI compromise of versions 1.82.7 and 1.82.8 is documented in the LiteLLM alternatives piece. Many governance controls large organisations ask for first sit behind the commercial licence.

Best fit. Python-heavy teams that need long-tail providers, or that want an in-process SDK with no proxy at all.

Edge and Kubernetes gateways, ranks 4 to 6

These three trade breadth of governance for a particular deployment strength: a managed edge network, or a cloud-native proxy built for Kubernetes and agent traffic.

4. Cloudflare AI Gateway

Cloudflare AI Gateway proxies requests to Workers AI, OpenAI, Anthropic, Gemini and other providers, with analytics, logging, caching, rate limiting, retries and model fallbacks. An OpenAI-compatible endpoint accepts {provider}/{model} names across providers including Anthropic, OpenAI, Groq, Mistral, Vertex AI, xAI, DeepSeek and Cerebras, with Cloudflare-stored keys or your own. Dynamic routing adds conditional branches, percentage splits, and request-quota and budget nodes that switch to a fallback when exceeded. Guardrails moderate prompts and responses across providers.

Pricing. Per the pricing page, dashboard analytics, caching and rate limiting are free. Accounts created before 24 September 2026 keep 100,000 stored logs on Workers Free and 10 million per gateway on Workers Paid; newer accounts follow Workers Logs pricing. Guardrails are billed as Workers AI inference, full DLP needs a Zero Trust subscription, and unified billing adds 5 percent to credit purchases.

Limits. MCP is handled by a separate product, MCP server portals in Cloudflare One, rather than by the AI gateway itself. Budget controls are routing nodes, not a key hierarchy.

Best fit. Teams that want a free, managed gateway with caching and analytics, especially those already on Cloudflare.

5. Agent Router

Agent Router is the new name of Envoy AI Gateway. Per the rename announcement of 10 September 2026, the project left the CNCF Envoy umbrella for the Agentic AI Foundation. Its API group, its aigw CLI and its Apache 2.0 licence are unchanged, and existing manifests need no edits.

Features. The project site lists 17 providers out of the box, provider failover, model name virtualisation, token limits per team, application or model, provider credentials (including AWS, GCP and Azure identities) kept in the gateway, OpenTelemetry GenAI telemetry, and an MCP catalogue that aggregates many servers and filters tools by caller. Version 1.1 is current; 1.0 was the first stable release. It runs locally through aigw run, on Kubernetes, or as a hosted service from Tetrate.

Limits. Runtime guardrails and cost attribution are listed under Tetrate’s enterprise edition. Configuration is through Kubernetes custom resources, which suits platform teams and few others. No overhead figure is published on the pages read.

Best fit. Platform teams on Envoy Gateway that want AI traffic managed as Kubernetes resources.

6. agentgateway

agentgateway is an Apache 2.0 proxy written mainly in Rust, governed by the Linux Foundation after a donation from Solo.io. It handles LLM, MCP and A2A traffic in one data plane. The project site describes an OpenAI-compatible API in front of OpenAI, Anthropic, Bedrock, Gemini, Vertex and self-hosted models, token budgets and hard caps per key or team, semantic caching, failover with latency-aware and cost-aware routing, regex and model-based guardrails (including OpenAI moderation and Bedrock Guardrails), CEL-based RBAC, and OpenTelemetry metrics, logs and traces. It runs as a standalone binary or on Kubernetes with Gateway API support.

Limits. The provider list is shorter than the leaders’, budget features are newer, and the documentation for them is thinner than Bifrost’s or LiteLLM’s. No overhead figure is published on the pages read.

Best fit. Teams that want one open-source proxy for model, tool and agent-to-agent traffic, especially on Kubernetes.

Platform and managed gateways, ranks 7 to 10

The last four suit a narrower setting: an existing cloud API platform, two managed routers with nothing to operate, and a general API gateway with AI plugins.

7. Azure API Management

Azure API Management (APIM) adds an AI gateway to its existing API gateway; Microsoft notes it is “not a separate offering”. It manages APIs that follow the OpenAI Chat Completions or Responses schema, the Anthropic Messages API (in v2 tiers) and the Google Vertex AI API, and a unified model API in preview exposes several backends through one OpenAI-compatible endpoint.

Features. The llm-token-limit policy enforces tokens per minute and token quotas (hourly to yearly) on any counter key, returning 429 or 403. It runs on classic, v2, self-hosted and workspace gateways, though counts are tracked per gateway and not aggregated across an instance. Backends support round-robin, weighted, priority and session-aware balancing, with a circuit breaker that honours Retry-After. Semantic caching uses Azure Managed Redis, content moderation uses Azure AI Content Safety, and token metrics flow to Application Insights. APIM can also expose REST APIs as MCP servers and govern existing ones.

Limits. Capabilities vary by service tier, and the self-hosted gateway is available only on Developer and Premium, with a required connection back to Azure for configuration. Policies are written in APIM’s XML policy language.

Best fit. Organisations whose models run in Microsoft Foundry or Azure OpenAI and whose API estate is already on APIM.

8. Vercel AI Gateway

Vercel AI Gateway is a managed gateway reachable from any infrastructure, not only Vercel deployments. It accepts the AI SDK, OpenAI Chat Completions and Responses, and Anthropic Messages formats, covers text, image, video, speech, embeddings and reranking, and records provider attempts, latency, tokens and cost for each request. Ordered provider and model fallbacks are configurable. Vercel states it adds “zero markup to provider token prices, including with BYOK”, and its GA announcement claims sub-20 ms latency.

Budgets. Budgets apply at team, project, API key and user scope, stack, and return 402 when exceeded. They are soft caps: the request that crosses the limit completes. Spend on your own provider keys is metered separately and does not count against any budget.

Limits. No self-hosting, no documented MCP tool gateway, and guardrail options were not assessed here.

Best fit. Product teams that want managed routing and spend caps with no proxy to run.

9. OpenRouter

OpenRouter is a managed router and billing service with hundreds of models behind one OpenAI-compatible endpoint. Its provider routing balances by price by default, skips providers with outages in the last 30 seconds, falls back automatically, and lets callers pin an order, whitelist providers, refuse providers that store or train on data, or require zero-data-retention endpoints.

Pricing. The FAQ lists a 5.5 percent fee (minimum $0.80) on card credit purchases, 5 percent for crypto, and for bring-your-own-key use a 5 percent fee above an allowance of $25,000 per month on pay-as-you-go. Prompts and completions are not logged by default.

Limits. OpenRouter is a routing marketplace rather than a governance layer: there are no team hierarchies, guardrails or MCP tool gateway in the pages read, and all traffic and billing pass through it.

Best fit. Teams that want the widest model catalogue on one bill, or a fallback provider behind a self-hosted gateway.

10. Apache APISIX

Apache APISIX is an Apache 2.0 API gateway with a set of AI plugins. ai-proxy forwards to documented providers and OpenAI-compatible endpoints. ai-proxy-multi balances across instances of OpenAI, Anthropic, Azure OpenAI, Gemini, Vertex AI, Bedrock, DeepSeek, OpenRouter and others by weighted round robin, consistent hashing or semantic similarity, with fallback on rate-limit or error responses and active health checks. Token rate limiting uses provider-reported usage with local or Redis counters, and the project lists regex prompt guards, a Lakera Guard integration, exact and semantic caching, and metrics for latency, tokens and time to first token.

Limits. It has no virtual-key or hierarchical budget model comparable to the standalone gateways, and MCP proxying is not confirmed on the pages read. Everything is assembled from plugins, which is flexible but more work.

Best fit. Teams already running APISIX that want AI routing without a new component.

Choosing by deployment model

The ranking is a starting point. Most decisions follow from where the gateway must run.

  • Prompts and keys must stay in your network. Shortlist Bifrost, LiteLLM, agentgateway and Agent Router. Bifrost has the lightest runtime and the broadest free governance of the four; LiteLLM has the most providers.
  • You already run an API platform. Check Kong, Azure APIM or APISIX first. Adding AI policies to an existing control plane avoids a second proxy, at the cost of the platform’s licence boundaries.
  • You want nothing to operate. Cloudflare, Vercel and OpenRouter are base-URL changes. Compare where each draws the line on logging, retention and fees.
  • Agents call tools as well as models. Prefer gateways that govern MCP and model traffic with the same identity: Bifrost, LiteLLM, Kong Enterprise, Agent Router and agentgateway.

Whatever the choice, two practices carry over. Configure failover before you need it, as the guide to LLM failover and load balancing explains, and put budgets on every key from day one, as covered in the guide to LLM cost control with an AI gateway.

Limits of this ranking

No gateway was load-tested for this article. Every entry is described from its documentation, repository, licence file, pricing page or release notes as read between late September and 5 October 2026. All latency figures are vendor-reported, measured on the vendor’s hardware against mocked or fake upstreams, and were not reproduced; Kong, Agent Router, agentgateway, Azure APIM, APISIX and Cloudflare publish no overhead figure on the pages read, and OpenRouter’s documentation page on latency gives no number.

Several things were not covered. Enterprise pricing for Bifrost, Kong, LiteLLM and Tetrate is quoted privately and was not assessed. Google’s Apigee, which also offers AI gateway policies, was considered and left out of the ten in favour of products with clearer public documentation. Guardrail quality, which depends on the detection models behind each feature, was not tested. Feature boundaries between free and paid tiers move with each release, so check the linked page before a purchase.

For a narrower view focused on the leading five products, the production-ready comparison of LLM gateways weighs overhead and free-tier governance most heavily.

Sources

  1. Bifrost source repository (GitHub, Apache 2.0)
  2. Bifrost docs: benchmarking
  3. Bifrost docs: supported providers
  4. Bifrost docs: drop-in replacement
  5. Bifrost docs: retries and fallbacks
  6. Bifrost docs: virtual keys
  7. Bifrost docs: semantic caching
  8. Bifrost docs: MCP overview
  9. Bifrost docs: enterprise overview
  10. LiteLLM source repository (GitHub)
  11. LiteLLM license file (MIT, with enterprise directory carve-out)
  12. LiteLLM docs: benchmarks
  13. LiteLLM docs: enterprise features
  14. LiteLLM docs: MCP gateway
  15. Kong AI Gateway documentation
  16. Kong docs: AI Proxy Advanced plugin
  17. Kong docs: AI MCP Proxy plugin
  18. Kong blog: Introducing Kong AI Gateway 2.2 (30 September 2026)
  19. Kong Gateway source repository (GitHub, Apache 2.0)
  20. Cloudflare AI Gateway documentation
  21. Cloudflare AI Gateway: pricing
  22. Cloudflare AI Gateway: OpenAI-compatible endpoint
  23. Cloudflare AI Gateway: dynamic routing
  24. Cloudflare AI Gateway: guardrails
  25. Cloudflare One: MCP server portals
  26. Agent Router (formerly Envoy AI Gateway)
  27. Agent Router blog: Envoy AI Gateway is now Agent Router
  28. agentgateway source repository (GitHub, Apache 2.0)
  29. agentgateway project site
  30. Azure API Management: AI gateway capabilities
  31. Azure API Management: llm-token-limit policy
  32. Azure API Management: self-hosted gateway overview
  33. Vercel AI Gateway documentation
  34. Vercel AI Gateway: budgets and spend limits
  35. Vercel blog: AI Gateway is now generally available
  36. OpenRouter docs: quickstart
  37. OpenRouter docs: FAQ (fees, logging, BYOK)
  38. OpenRouter docs: provider routing
  39. Apache APISIX AI gateway
  40. Apache APISIX docs: ai-proxy-multi plugin

Questions readers ask

What is an AI gateway?

An AI gateway is a proxy between applications and model providers. Applications send requests to one endpoint, usually in the OpenAI format, and the gateway routes them to OpenAI, Anthropic, Bedrock, Vertex or a self-hosted model. On the way it applies keys, budgets, rate limits, retries, failover, caching, guardrails and logging, so each application does not have to implement them.

What is the difference between a self-hosted and a managed AI gateway?

A self-hosted gateway such as Bifrost, LiteLLM, Kong, Agent Router, agentgateway or APISIX runs on your infrastructure, so prompts and provider keys stay inside your network and you operate the service. A managed gateway such as Cloudflare AI Gateway, Vercel AI Gateway or OpenRouter runs on the vendor's network. You change a base URL and operate nothing, but every request passes through a third party.

Which AI gateway adds the least latency?

No independent benchmark covers all ten products under identical conditions. Among vendor-published figures, Bifrost reports 11 microseconds of added overhead at 5,000 requests per second on an AWS t3.xlarge against a mocked upstream, and LiteLLM reports 2 to 8 milliseconds (p50 to p95) with four instances at about 1,170 requests per second. Measure at your own payload size before relying on either.

Do I need an AI gateway if I only use one model provider?

Often not at first. The case appears when several teams share provider keys, when finance needs spend per team, when an outage at one provider should not take the product down, or when security needs a log of prompts and responses. Those needs usually arrive before a second provider does.

Are AI gateways and MCP gateways the same thing?

They overlap. An AI gateway governs traffic to models, and an MCP gateway governs traffic to tools exposed over the Model Context Protocol. Several products in this list, including Bifrost, LiteLLM, Kong, Agent Router and agentgateway, now handle both, so one key can carry model and tool permissions.

More comparisons