AI gateway vs API gateway: what changes when the traffic is tokens
API gateways were built for short request-response calls priced per request. Model traffic is priced per token, streamed for minutes and shaped differently by every provider. A mechanism-by-mechanism comparison, grounded in AWS, Kong, NGINX, OpenAI and Anthropic documentation, with Bifrost as the worked example.
TL;DR
- The core difference in AI gateway vs API gateway is the unit of account. API gateways meter requests; AWS API Gateway’s token bucket counts one token per request. Model providers meter tokens, and one request can consume thousands of times more than the next.
- An LLM gateway exists because model traffic breaks four assumptions of a request-level proxy: that requests cost about the same, that responses are short, that every upstream speaks one schema, and that the payload is opaque.
- Streaming changes failure handling. Anthropic’s API can send an
overloaded_errorinside a stream that already returned HTTP 200, so retries and token accounting have to work at the event level, not the status-code level.- Most production stacks run both layers: an API gateway on ingress for users, and an LLM gateway on egress to model providers. Kong AI Gateway extends an API gateway toward AI; purpose-built gateways such as Bifrost start from the model call.
An API gateway is a reverse proxy that authenticates, routes and rate-limits requests to backend services, while an LLM gateway (often called an AI gateway) sits between applications and model providers and applies policy in terms of tokens, models and prompts. That is the short answer to the AI gateway vs API gateway question. The differences are mechanical: they show up in how limits are counted, how long connections stay open, how errors arrive and what the proxy is allowed to read. Bifrost, an open-source AI gateway written in Go and released by Maxim AI under Apache 2.0, serves as the worked example throughout.
What an API gateway does
An API gateway is a single entry point in front of backend services. It terminates TLS, authenticates the caller, routes by host and path, applies rate limits and quotas, and logs the call. Its design assumes short request-response exchanges in which one request costs roughly as much as any other.
Rate limiting shows that assumption most clearly. Amazon API Gateway throttles with a token bucket “where a token counts for a request”, and its default account quota is 10,000 requests per second per Region with a burst bucket of 5,000. NGINX’s limit_req module uses a leaky bucket in requests per second or per minute, keyed on a variable such as the client address, and rejects excess requests with a 503 by default.
Kong’s Rate Limiting plugin counts requests per second through per year and returns a 429 when the count is exceeded. Envoy’s global rate limiting calls an external service for every new connection or HTTP request. In all four, the unit is the request or the connection; none opens the body to ask how much work the call will cause upstream, because for a typical REST call the answer barely varies.
Caching follows the same logic. Kong’s Proxy Cache plugin keys entries on a SHA-256 hash of the method, request, query parameters, headers and consumer groups. That suits GET /products/42, not two differently worded questions.
What an LLM gateway is
An LLM gateway is a proxy that model calls pass through on their way to OpenAI, Anthropic, Google, a cloud model service or a self-hosted inference server. It exposes one API for every provider, holds the provider credentials, and applies model-aware policy: token budgets, fallback chains, semantic caching, content guardrails and per-request cost logging.
“AI gateway” names the same layer. It exists because the upstream is different. A backend microservice is owned by the organization, speaks a schema the organization designed and bills nothing per call. A model provider is external, speaks its own schema, bills per token and enforces rate limits that are also counted in tokens.
Integration is usually a base-URL change. Bifrost accepts calls from existing OpenAI, Anthropic and Google GenAI SDKs at paths such as /openai and /anthropic, which makes it a drop-in replacement for direct provider calls, and translates each request for whichever of its more than 20 providers it is routed to. A production-grade LLM gateway comparison scores five products on overhead, failover, governance and MCP support, and the Frontier Wire’s overview of what an AI gateway does covers the general case for the layer.
AI gateway vs API gateway at a glance
The two layers overlap on transport and authentication and diverge on nearly everything the upstream bills for, per the vendor and provider documentation cited here.
| Dimension | Traditional API gateway | LLM gateway |
|---|---|---|
| Unit of rate limiting | Requests or connections per interval | Input and output tokens per interval, plus requests |
| Unit of cost | Roughly flat per request | Per token, varying by model, context length and cache status |
| Typical response | Buffered JSON, milliseconds to seconds | Server-sent event stream, seconds to minutes |
| Error signaling | HTTP status code before the body | Status code, or an error event inside a 200 stream |
| Upstream schema | The organization’s own API | A different schema per provider, normalized by the gateway |
| Failover target | Another instance of the same service | A different provider or model |
| Caching | Exact key on method, path, headers | Exact hash plus embedding similarity on the prompt |
| Payload inspection | Headers and schema validation | Prompt and completion content |
| Cost attribution | Request counts per consumer | Tokens and dollars per key, team and customer |
| Tool traffic | Not modeled | MCP tool discovery, filtering and execution |
| Telemetry | Latency, status, bytes | Tokens, cost, time to first token, model, provider |
The error-signaling row is the one most often missed: a gateway that decides retries from status codes alone will record a failed generation as a success. The payload-inspection row has the largest compliance consequence, because it brings prompt content into the proxy’s scope and logs. Teams that must keep that content inside their own network usually start with open-source LLM gateways for self-hosted deployments.
Rate limits measured in tokens
In the AI gateway vs API gateway comparison, rate limiting is the first mechanical split. Providers limit tokens as well as requests, so a gateway that counts only requests cannot predict when an upstream will start rejecting traffic. An LLM gateway estimates tokens before the call and reconciles them against the provider’s count afterward.
OpenAI measures rate limits in requests and tokens per minute and per day, applies whichever limit is hit first, and sets them per organization and project. Its accounting is conservative: a request’s rate-limit cost “is calculated as the maximum of max_tokens and the estimated number of tokens based on the character count of your request,” so a generous max_tokens spends capacity that may never be used.
Anthropic counts differently. Its rate limits are set per model class in requests, input tokens and output tokens per minute, enforced by a continuously refilling token bucket. For most Claude models, tokens read from the prompt cache do not count toward the input limit, and output limits are evaluated “in real time as output tokens are produced,” with max_tokens playing no part. A 429 names the exceeded limit and carries a retry-after header.
Three consequences follow for the gateway:
- Requests are not equal. A 50-token classification and a 200,000-token document summary each count as one request to AWS or NGINX, and very differently to the provider.
- The count is final only after the response. Output length is unknown until generation ends, so the gateway admits on an estimate and settles from the provider’s
usagefield. - Limits differ by provider and model. Spreading traffic across providers needs per-provider counters, not one global bucket.
Bifrost implements this with budgets and limits at the virtual key and provider-configuration levels. Each can carry a request limit and a token limit with reset windows from one minute to one year; when one provider configuration is exhausted, only that provider drops out of routing. A caller over its limit receives a 429 (token_limited or request_limited), and one over its spending cap receives a 402 (budget_exceeded), both enforced per virtual key. The distinction matters: a 429 invites a later retry, a 402 says money is the constraint. Most commercial and open-source AI gateways now offer token-denominated limits; they differ in where counters live and how many organizational levels they cover.
Streaming responses and long-lived connections
Streaming is the second split in the LLM gateway vs API gateway comparison. Most interactive model calls stream server-sent events, and an LLM gateway has to forward events as they arrive, hold connections open for minutes, and read the stream in transit to recover token counts and detect errors. Traditional gateways buffer the full response by default.
The WHATWG HTML standard defines server-sent events as a text/event-stream body in which a blank line dispatches each event. It also warns that “legacy proxy servers are known to, in certain cases, drop HTTP connections after a short timeout,” and suggests a comment line every 15 seconds or so.
AWS API Gateway shows what retrofitting involves. By default it “waits to receive the complete response before beginning transmission,” and REST integrations time out after 29 seconds. Its response streaming mode names generative AI chat as a use case and lifts the 29-second and 10 MB limits, with conditions: only HTTP_PROXY and AWS_PROXY integrations, streams of up to 15 minutes, a 5-minute idle timeout on Regional endpoints, and no endpoint caching, content encoding or VTL transformation, because each requires buffering.
The harder problem is what arrives inside the stream. Anthropic’s streaming documentation describes message_start, content block events, message_delta and message_stop, with ping events interleaved; token counts in message_delta are cumulative. During high usage the stream may carry an overloaded_error event, “which would normally correspond to an HTTP 529 in a non-streaming context.” The status line has already said 200. OpenAI’s streaming guide adds that streaming “makes it more difficult to moderate the content of the completions, as partial completions may be more difficult to evaluate.”
Bifrost handles accounting with a stream accumulator: each chunk is recorded as it passes, and when the provider’s final signal (such as a finish_reason) arrives, the accumulator rebuilds the complete message and computes token usage, cost and latency for the logging and governance plugins. Anyone comparing self-managed LLM gateway options should test this path directly: whether usage is recorded when a client disconnects mid-stream, and whether a mid-stream error reaches the caller or is logged as a success.
Provider schemas, fallbacks and routing
An LLM gateway normalizes different provider APIs into one request format, which is what lets it fail over from one provider or model to another without application changes. A traditional gateway fails over between identical instances and never translates.
OpenAI and Anthropic use different endpoints, message structures, streaming event names and usage fields. A fallback between them is not a retry against a second host; it is a translation of the request into another schema and of the response back. Kong’s AI Proxy plugin shows the pattern inside an API gateway, exposing route types such as llm/v1/chat and translating each request “to the configured target format.” Bifrost does the same at the core of its request pipeline, converting incoming bodies into an internal BifrostRequest structure so every plugin sees one shape.
Fallback semantics are where products differ most. A Bifrost request carries a fallbacks array of provider/model strings. With automatic fallbacks, the primary provider exhausts its own retry budget first, each fallback then receives a full retry budget, and each attempt is a fresh request, so caching, governance and logging plugins run again. The response’s extra_fields.provider records which provider served it, and a plugin can stop the chain so a policy rejection is not retried elsewhere.
Routing rules, written in the Common Expression Language (CEL), add a second layer. They can match on headers, team or customer, and on capacity signals such as budget_used or tokens_used as percentages. A rule like budget_used > 85 can move a team to a cheaper model before its budget runs out, which a gateway that counts only requests has no way to express. The companion piece on LLM failover and load balancing covers retry budgets in depth, and an evaluation of five LLM gateways sets out how five products handle failover.
Caching, guardrails and cost attribution
Three LLM-gateway features have no close API-gateway equivalent because each depends on reading the prompt: semantic caching, content guardrails and per-token cost attribution.
Semantic caching
Users phrase the same question differently, so exact-match caching rarely helps model traffic. Bifrost’s semantic caching runs two layers: a hash of the normalized request for exact replays, then an embedding-similarity search on misses, with a default cosine threshold of 0.8 and a default time-to-live of five minutes. Caching is opt-in per request through an x-bf-cache-key header, streamed responses replay chunk by chunk, and conversations longer than three messages are skipped by default. Kong’s AI Semantic Cache plugin is comparable and, per Kong, available only in its AI Gateway Enterprise offering. The threshold is a correctness setting: two close prompts that name different customers still need different answers.
Prompt and response guardrails
A traditional gateway validates headers and at most a request schema; an LLM gateway can inspect content. Bifrost’s guardrails, part of Bifrost Enterprise, run before the model or an MCP tool receives input, after generation, or both. CEL rules decide when to apply reusable profiles covering native secrets detection, custom regular expressions and LLM-based policy checks, plus services such as AWS Bedrock Guardrails, Azure Content Safety, Google Model Armor and Microsoft Presidio. A violation blocks the request with a 446 status or redacts the matching text. The streaming trade-off reappears here: detection-only rules observe a stream without delaying it, while rules that can block buffer output until evaluation completes.
Per-token cost attribution
In an API gateway, attribution means counting requests per consumer. In an LLM gateway it means multiplying tokens by a price that depends on model, context length and cache status. Bifrost’s model catalog syncs pricing every 24 hours when a config store is present, supports price tiers above 128,000 and 200,000 tokens of context, and prices a direct cache hit at zero and a semantic hit at the embedding cost only. Costs are deducted up a hierarchy of customer, team, virtual key and provider configuration, and a request proceeds only if every applicable budget has room. The Frontier Wire’s piece on LLM cost control with an AI gateway covers budget design, and hierarchical budgets are one of the clearer points of difference among open-source AI gateways that teams run themselves.
MCP and tool traffic
Agents add a second kind of upstream traffic: model-initiated tool calls, most often over the Model Context Protocol (MCP). An LLM gateway can govern which tools each caller may see and run, which an API gateway has no model for.
MCP’s transport specification encodes messages as JSON-RPC over stdio or Streamable HTTP, in which a single endpoint accepts POST and GET, may answer with an SSE stream, and can assign an Mcp-Session-Id that clients echo on later requests. An API gateway can proxy that traffic, but every tool call is a POST to the same path, so path routing and request counts reveal nothing about which tool ran.
Bifrost is both an MCP client, connecting to tool servers over stdio, HTTP or SSE, and an MCP gateway that exposes the aggregated tools at /mcp. By default it does not execute tool calls a model proposes; the application approves and executes them explicitly, with an optional agent mode for configured auto-approval. Tool filtering per virtual key is deny-by-default, and a caller’s header can narrow the allowed list but never widen it.
Tool definitions are also a token cost, since every schema in context is billed as input on every turn. Bifrost’s Code Mode replaces the tool list with four meta-tools and has the model write Python, run in a sandboxed Starlark interpreter, that calls tools on demand. In Bifrost’s own benchmark with 508 tools across 16 servers, input tokens fell 92.8 percent; the figure is vendor-reported and scales with tool count. The Frontier Wire survey of MCP gateways compares this layer across products, and a roundup of the best AI gateways for 2026 counts MCP support among its eight criteria.
Observability of prompts and completions
API gateway telemetry answers whether a request succeeded and how long it took. LLM gateway telemetry also records which model served it, how many tokens it used, what it cost and, where policy allows, what was asked and answered.
The OpenTelemetry semantic conventions for generative AI define spans, metrics and events for GenAI clients, MCP and specific providers, so model calls appear in the same tracing backend as the rest of a request. Bifrost’s OpenTelemetry integration emits spans in that format, with attributes such as gen_ai.request.model, gen_ai.provider.name, token counts and gen_ai.usage.cost, to Grafana Cloud, Datadog, New Relic, Honeycomb, Langfuse or any OTLP receiver.
Prompt logging carries a data-handling decision that request logs never did. Bifrost’s built-in logging stores input and output messages alongside tokens, cost and latency in SQLite or PostgreSQL, and a disable_content_logging setting drops the content while keeping usage metadata.
Inside an LLM gateway: a reference architecture
A purpose-built LLM gateway is a pipeline: normalize the request, authenticate and budget-check it, choose a provider, call it with retries and fallbacks, then account for the response. The Bifrost request flow is a concrete instance, in nine stages:
- Transport. A FastHTTP server parses the request into the unified
BifrostRequestschema. - Routing. A provider and key are chosen by weighted random selection, with health checks and circuit breaking.
- Pre-processing plugins. Authentication, governance, rate limiting and transformation run in sequence.
- MCP tool discovery. Tools are discovered, filtered and injected into the request.
- Memory pools. Objects come from pre-allocated pools to reduce garbage collection.
- Worker pools. Requests queue per provider and run in parallel.
- Provider call. Workers add credentials and call the provider with timeouts.
- Tool execution. Tool calls returned by the model run where configured.
- Post-processing plugins. Logging, caching and metrics run before the response is serialized.
Stages 4 and 8 exist only because models call tools, and stage 9 must run on streamed output, which is why the accumulator matters. The added work is small by the vendor’s measure: Bifrost’s benchmark methodology reports 11 microseconds of overhead per request at 5,000 requests per second on an AWS t3.xlarge and 59 microseconds on a t3.medium, against mocked OpenAI responses. Those are synthetic-upstream figures, and real deployments add network hops that exceed either number.
Bifrost Enterprise adds clustering for high availability, adaptive load balancing, the guardrails above, OIDC single sign-on, role-based access control, audit logs and in-VPC deployment. Teams planning to self-host an AI gateway should check which they need before choosing a tier; across the category, high availability and audit trails usually sit on the paid side.
Running both layers together
Most production stacks need both an API gateway and an LLM gateway because they govern different directions of traffic: the API gateway handles ingress from users to services, and the LLM gateway handles egress from services to model providers.
In a layered request, the API gateway terminates TLS, validates the user’s session, applies a per-user request limit and routes to an application service. That service builds a prompt and calls the LLM gateway with a virtual key for its team. The LLM gateway checks the token budget, applies guardrails, selects a provider, streams the response back and records tokens and cost. Each layer counts what it can see: users and requests at the edge, tokens and dollars at the model boundary.
| Concern | API gateway (ingress) | LLM gateway (egress to models) |
|---|---|---|
| End-user authentication | Owns it | Receives a service or team identity |
| Abuse and DDoS protection | Owns it | Not its role |
| Token budgets and spend caps | Cannot measure | Owns it |
| Provider keys | Should never see them | Holds and rotates them |
| Model fallback and routing | Not modeled | Owns it |
| Prompt and response guardrails | Not modeled | Owns it |
| Streaming to the end user | Must pass through unbuffered | Produces and meters the stream |
The streaming row is where layering most often breaks. If the edge gateway buffers responses or enforces a short integration timeout, as AWS does by default at 29 seconds, the stream never reaches the user intact. The fix is to enable streaming on the edge route or let model responses bypass it.
API gateways extended for AI
Some vendors collapse the layers by adding AI features to an API gateway. Kong AI Gateway is the clearest case. Its AI Rate Limiting Advanced plugin “uses the token data returned by the LLM provider” to limit by total, prompt or completion tokens or by cost, and other plugins cover schema normalization, semantic caching, prompt guards and PII sanitization. Kong lists that plugin as Enterprise-only.
For an organization standardized on Kong, extending the existing gateway avoids a second proxy. The trade-off is that model traffic inherits a general-purpose proxy’s configuration model, and features such as per-key MCP tool filtering or budget hierarchies have to be checked product by product. A side-by-side of the top LLM gateways alongside Kong AI Gateway sets the two approaches against each other, and a list of open-source gateway options for air-gapped environments covers the self-hosted end of the market.
Governance beyond the gateway
A gateway governs only the traffic configured to pass through it. Beyond routing, Bifrost applies governance and security controls (virtual keys, budgets, guardrails, audit logs) centrally, and Bifrost Edge, currently in alpha, extends that same governance and security to AI traffic on employee machines, including desktop chat apps, browser AI, coding agents and their MCP servers. Its endpoint security applies the gateway’s guardrails before a prompt leaves the device, covering traffic that neither an API gateway nor a server-side LLM gateway sees.
Choosing where to start
The LLM gateway vs API gateway decision is rarely either-or. A team that already runs an API gateway should keep it for ingress and add an LLM gateway once it uses a second provider, needs a fallback model, or must report model spending by team. The deciding test is whether anyone needs to know, per request, which model ran, how many tokens it used and who pays; that answer belongs in an LLM gateway. Teams evaluating one can request a Bifrost demo or review the open-source repository.
Sources
- Amazon API Gateway: throttle requests to your REST APIs
- Amazon API Gateway quotas
- Amazon API Gateway: stream the integration response for proxy integrations
- Kong: Rate Limiting plugin
- Kong: Proxy Cache plugin
- Kong: AI Rate Limiting Advanced plugin
- Kong: AI Proxy plugin
- Kong: AI Semantic Cache plugin
- Kong AI Gateway documentation
- NGINX: ngx_http_limit_req_module
- Envoy: global rate limiting architecture overview
- OpenAI API: rate limits
- OpenAI API: streaming responses
- Anthropic Claude API: rate limits
- Anthropic Claude API: streaming messages
- WHATWG HTML Living Standard: server-sent events
- Model Context Protocol specification: transports (2025-06-18)
- OpenTelemetry semantic conventions for generative AI
- Bifrost source repository (GitHub, Apache 2.0)
- Bifrost docs: request flow
- Bifrost docs: streaming framework
- Bifrost docs: fallbacks
- Bifrost docs: budget and limits
- Bifrost docs: virtual keys
- Bifrost docs: semantic caching
- Bifrost docs: guardrails
- Bifrost docs: model catalog
- Bifrost docs: MCP overview
- Bifrost docs: MCP Code Mode
- Bifrost docs: built-in observability
- Bifrost docs: OpenTelemetry integration
- Bifrost docs: MCP tool filtering per virtual key
- Bifrost docs: routing rules
- Bifrost docs: benchmarking
- Bifrost docs: enterprise overview
- Bifrost docs: Bifrost Edge overview (alpha)
- Bifrost docs: Bifrost Edge security and guardrails
Questions readers ask
What is an LLM gateway?
An LLM gateway is a proxy between applications and model providers such as OpenAI, Anthropic or AWS Bedrock. It exposes one API, holds the provider keys, and applies model-aware policy to every call, including token-based rate limits and budgets, fallback chains across providers, semantic caching, prompt and response guardrails, and per-request logging of tokens and cost.
What is the difference between an AI gateway and an API gateway?
An API gateway manages request traffic to backend services and measures it in requests per second. An AI gateway manages traffic to model providers and measures it in tokens and money. It understands provider schemas, streamed responses, model fallbacks and prompt content, which a request-level proxy does not inspect by default.
Do you need both an API gateway and an AI gateway?
Most production systems that serve external users do. The API gateway stays at the edge, authenticating users and protecting backend services. The AI gateway sits on the egress path between those services and the model providers, where it enforces token budgets, fallbacks and guardrails. A small internal tool calling one provider can often start without either.
How do AI gateways control costs and token usage?
They read the token counts that providers return with each response, price them against a model catalog, and deduct the result from budgets attached to keys, teams or customers. Requests that would exceed a token rate limit or a spending cap are rejected before they reach the provider, and cached answers avoid the provider call entirely.
How can AI gateways protect sensitive data?
An AI gateway can inspect prompt and response content in the request path. Guardrail rules detect secrets, personally identifiable information or unsafe content, then block the request or redact the matching text before it leaves the network. Content logging can also be switched off so that only token counts, cost and latency are stored.
How do you monitor and trace LLM activity through an AI gateway?
Route every model call through the gateway and export its telemetry. Gateways typically log each request's model, provider, tokens, cost, latency and status, publish Prometheus metrics, and emit OpenTelemetry spans that follow the GenAI semantic conventions, so traces land in the same backend as the rest of the application.
