LLM cost management: budgets, rate limits and virtual keys at the gateway
Output tokens cost five to six times more than input tokens, reasoning tokens bill as output, and shared API keys erase the link between a charge and its owner. This is how teams attribute and cap LLM spend, and what gateway-level governance in the open-source Bifrost gateway does about it.
TL;DR
- LLM cost management has three parts: metering tokens and cost per request, attributing each request to an owner, and enforcing budgets before the invoice rather than explaining it afterward.
- Output tokens cost five to six times more than input tokens at OpenAI, Anthropic and Google, and reasoning tokens are billed as output even when the API never shows them.
- Provider projects and workspaces can cap spend for one provider, but a gateway that issues its own virtual keys can apply budgets and rate limits across every provider, team and customer at once.
- Bifrost, an open-source gateway, checks budgets at the provider, virtual key, team and customer level on each request and blocks it with HTTP 402 when any one of them is spent.
- Gateway budgets only govern traffic that reaches the gateway; AI tools on employee laptops need an endpoint layer to bring that spend under the same limits.
LLM cost management is the practice of measuring, attributing and capping what an organization spends on language model APIs, and it gets harder with every provider, model and team added to the bill. A provider invoice reports tokens by model and by key, but not which feature, customer or engineer generated them, and it arrives after the money is spent. Many teams therefore move the controls into the request path, where a gateway can price each call, charge it to an owner and refuse it when a budget runs out. Bifrost, an open-source AI gateway written in Go and released by Maxim AI under the Apache 2.0 license, is one of the tools built around that model, and its governance features are examined below alongside the provider-native and application-level alternatives.
Why LLM spend is hard to attribute
LLM spend is hard to attribute because the unit of cost, the token, is generated inside shared infrastructure that rarely knows who asked for it. The bill is accurate in total and uninformative in detail.
Four mechanics produce most of the confusion:
- Shared keys. When one API key per provider is handed to every service, the provider’s usage report can only break spend down by that key. Everything behind it arrives as one number.
- Asymmetric pricing. Input and output tokens carry different prices, so two requests with the same total token count can differ in cost several times over depending on how much the model wrote.
- Invisible tokens. Reasoning models generate internal tokens before answering. Those tokens are billed but not returned, so logging the response text understates the cost.
- Tokenizer drift. Anthropic’s pricing page notes that Claude 4.7 and later models use a tokenizer that “produces approximately 30% more tokens for the same text,” so an unchanged prompt can cost more after a model upgrade at an identical per-token price.
The FinOps Foundation’s guide to building a generative AI cost and usage tracker, by James Barney of MetLife, ranks counting requests per key as least accurate and recording the provider’s actual input and output counts as most accurate. Unless a model’s tokenization is published, it warns, “any token count” produced outside the provider “will just be estimates.”
Token pricing: input, output and reasoning
Providers charge per million input tokens, per million output tokens, and at lower rates for cached input. Output is consistently the most expensive class, and reasoning or thinking tokens are billed at the output rate. The table uses list prices from each provider’s pricing page at the time of writing.
| Model | Input (per 1M tokens) | Cached input (per 1M) | Output (per 1M) | Output-to-input ratio |
|---|---|---|---|---|
| OpenAI gpt-6-sol | $2.00 | $0.20 | $10.00 | 5x |
| OpenAI gpt-6-luna | $0.10 | $0.01 | $0.50 | 5x |
| Anthropic Claude Sonnet 5.5 | $2.00 | $0.20 (cache hit) | $10.00 | 5x |
| Anthropic Claude Haiku 4.5 | $1.00 | $0.10 (cache hit) | $5.00 | 5x |
| Google Gemini 3.1 Pro Preview (prompts up to 200k) | $2.00 | not shown | $12.00 | 6x |
| Google Gemini 3.8 Flash (through Dec 31, 2026) | $0.75 | $0.075 | $3.75 | 5x |
Prices are from OpenAI, Anthropic and Google. Google lists the Gemini 3.8 Flash rates as doubling on January 1, 2027, a reminder that any cost model built on list prices needs a scheduled refresh.
A worked example shows why the ratio matters. A Claude Sonnet 5.5 request with 2,000 input tokens and 800 output tokens costs $0.004 for input and $0.008 for output, $0.012 in total: output is under 30 percent of the tokens and two thirds of the cost. If the model also spends 4,000 tokens thinking, those bill at the output rate and add $0.04, taking the request to $0.052, more than four times the original, with no change in the visible answer.
OpenAI’s reasoning guide states that reasoning tokens “are not visible via the API” but “are billed as output tokens,” and reports their count under output_tokens_details. Google labels its Gemini output column “Output price (including thinking tokens).” How effort settings change that hidden bill is covered in the reasoning-model era has a cost problem; metering has to read the provider’s usage fields, not count characters in the response.
Discounts pull the other way. Anthropic and OpenAI both price batch requests at 50 percent of standard rates, and Anthropic’s prompt cache reads cost 10 percent of base input on most models. A tracker that applies one flat rate per model overstates batched and cached traffic and understates reasoning-heavy traffic.
The FinOps view of AI spend
FinOps, the discipline of making engineering teams accountable for cloud spend, now treats AI as its own technology category. The FinOps Foundation names token-based billing, low forecast predictability and the difficulty of identifying who consumed a model’s output as the defining problems, and recommends tagging, showback, quotas and throttling in response.
The Foundation’s FinOps for AI overview, last updated in February 2026, singles out allocation: “More acute are the challenges of identifying the consumer of the model output, which is especially difficult when the consumers of the same model can be different interfaces/functional modules in the same user application.” The AI technology category of the FinOps Framework adds that “Predictability is generally lower, especially for the Crawl and Walk phases,” recommends shorter forecasting windows and more frequent anomaly reviews, and proposes cost per token as a metric that normalizes spend across vendors.
The tracker guide’s architectural advice is the most directly useful: a hub-and-spoke design with “API keys or authentication keys tied directly to a given use case,” so every request carries its owner. That is the design an AI gateway implements. Mapped onto FinOps practice:
- Allocation needs a credential per use case, team or customer, not per provider.
- Showback needs cost recorded per request with those identifiers attached.
- Quotas and throttling need a component in the request path that can refuse a call.
- Anomaly detection needs per-request data in a metrics system that can alert on it.
Approaches to LLM cost management
Teams use four broad approaches to LLM cost management: provider dashboards and limits, separate provider projects or keys per team, metering inside application code, and governance at a shared gateway. They differ in how many providers they cover, how fine the attribution is, and whether they can stop a request before it is billed.
| Approach | Attribution granularity | Enforcement before spend | Multi-provider | Main limitation |
|---|---|---|---|---|
| Provider dashboards and org limits | Per API key or project | Org-level hard limit only | No, one console per provider | Shared keys collapse all consumers into one line |
| Per-project provider keys (OpenAI projects, Anthropic workspaces) | Per project or workspace | Yes, per project or workspace | No, configured separately at each provider | Coarse units; key sprawl |
| App-level metering in code | As fine as the code records | Only where each team implements it | Yes, if every service does it | Inconsistent across services; no central kill switch |
| Gateway-level governance | Per virtual key, team, customer, user | Yes, on every request through the gateway | Yes | Governs only traffic routed through the gateway |
Provider controls have improved. OpenAI lets organization and project owners set hard spend limits that return a 429 with project_spend_limit_exceeded once reached, though it cautions that “Enforcement is not instantaneous” and recorded spend “can slightly exceed the configured amount.” Anthropic’s workspaces carry monthly spend limits and per-model rate limits, up to 100 workspaces per organization by default, and the automatically created Claude Code workspace is the only one with per-user monthly spend limits.
Each of these works inside one provider, so a team calling three providers maintains three sets of projects, keys and usage exports and joins them in a spreadsheet. Application-level metering solves the join but depends on every service implementing it identically, which the FinOps tracker guide flags as hard to enforce. A side-by-side review of LLM gateways scores governance depth alongside performance and deployment model.
Gateway-level governance for cost control
Gateway-level governance means the gateway, not the provider, issues the credentials applications use and attaches spending and access policy to each. Every request is authenticated, priced and checked against its budgets before it is forwarded, and provider keys never leave the gateway. The general architecture is described in what an AI gateway does.
- One policy language across providers. A budget on a virtual key applies whether the request goes to OpenAI, Anthropic or a self-hosted model.
- Enforcement at the request. A budget check before the call can reject it; a dashboard can only report it.
- Credentials separated from consumers. Capping one team’s key affects no one else, and provider keys rotate without redeploying applications.
The trade-off is that the gateway sits on the critical path, so deployment model and overhead matter. Teams that must keep prompts and cost data inside their own network usually look at on-premises AI gateway choices first.
Virtual keys, budgets and rate limits in Bifrost
In Bifrost, the virtual key is the primary governance entity. It authenticates a consumer, restricts which providers and models it can reach, and carries budgets and rate limits; keys attach to a team or a customer, and every applicable budget is checked on each request. The LLM gateways built for production on the market differ most in how deep this layer goes.
Virtual keys as the unit of attribution
Virtual keys are accepted in the header formats existing SDKs already send: the native x-bf-vk, Authorization: Bearer for OpenAI-compatible clients, x-api-key for Anthropic, x-goog-api-key for Gemini and api-key for Azure OpenAI. Each key attaches to exactly one team, exactly one customer, or neither, and can carry an expiry after which requests return 403.
One default matters in any rollout: governance is optional. Requests without a virtual key header pass through without checks until “Enforce Virtual Keys on Inference” is switched on, after which a request without a key is rejected with virtual_key_required (401).
Hierarchical budgets
Bifrost’s budget and limits model has four tiers: customer, team, virtual key, and provider configuration inside a virtual key. For a key under a team under a customer, checks run from provider config to virtual key to team to customer. “All applicable budgets must pass,” and a completed request’s cost is deducted from every level, so one exhausted budget anywhere in the chain blocks further requests with budget_exceeded (HTTP 402).
Reset periods are written as durations from 1m to 1Y, including 1M for a month and 1Q for a quarter. With calendar_aligned set, budgets reset on UTC calendar boundaries (midnight, Monday, the first of the month, or a fiscal quarter start set by quarter_start_month) instead of rolling windows, which lines gateway budgets up with the finance calendar.
The customer tier enables per-tenant cost control for companies that resell model access; in enterprise deployments a request can name the customer to charge with an x-bf-customer-id header.
Rate limits
Rate limits come in two parallel forms: request limits (for example, 100 requests per minute) and token limits (for example, 50,000 prompt and completion tokens per hour). They are set only at the virtual key and provider config levels, not on teams or customers, and exceeding one returns token_limited or request_limited with HTTP 429. When one provider inside a virtual key hits its budget or rate limit, routing excludes it while the key’s other providers stay available.
Token limits are the better guard for reasoning models. A cap of 100 requests per minute says nothing about a single request that generates 30,000 reasoning tokens; a token cap bounds the spend rate directly.
Model allow-lists, routing and semantic caching
Allow-lists, weighted routing and caching reduce cost by controlling which model answers and whether a paid call is needed at all. Bifrost applies allow-lists per provider inside each virtual key, splits traffic by weight, and offers a two-layer cache that can answer repeated requests without calling a provider.
Provider and model allow-lists
Provider access on a virtual key is deny-by-default unless allow_all_providers is set. Within each provider, allowed_models accepts a wildcard, an explicit list or regex: patterns, and blocked models take precedence. Virtual key routing follows the same rule: a request for one provider’s model is not sent to another provider unless that model is explicitly allowed there.
An allow-list is the simplest cost control available: a key for a summarization tool that can only reach a small model cannot run up a frontier-model bill, and the request returns model_blocked (403).
Weighted routing
When a virtual key lists several providers, requests are “distributed proportionally” by weight, so 0.8 and 0.2 send 80 percent of traffic for a shared model to the first. Shifting traffic toward a cheaper deployment of the same model needs no application change. Failover design is covered in LLM failover and load balancing.
Semantic caching
Semantic caching in Bifrost combines direct hash matching for exact repeats with embedding similarity for requests worded differently. Defaults are a five-minute TTL and a 0.8 similarity threshold. Caching activates only when a request carries a cache key (the x-bf-cache-key header or a configured default), and it needs a vector store such as Redis, Valkey, Weaviate, Qdrant or Pinecone.
The model catalog prices a direct cache hit at zero and a semantic match at the embedding call only. Caching pays off on repetitive traffic such as support questions; on unique agent calls it adds embedding cost with few hits.
Cost tracking and observability
Cost tracking is the part of LLM cost management that finance sees: every request priced from a maintained price list and recorded with its owner. Bifrost computes a USD cost per request, logs it with the virtual key, and exports cost and token counters to Prometheus with team and customer labels, giving finance and engineering one data source.
The model catalog resyncs its pricing sheet every 24 hours when a config store is present, and its cost calculation handles cache-read and cache-creation rates, long-context tiers, batch rates and media pricing, which addresses the flat-rate error described earlier.
Per-request records land in the built-in log store, SQLite by default and PostgreSQL for high-volume production, with token usage, USD cost, latency, provider, model and virtual key on each entry. Logging is asynchronous; Bifrost’s own measurement puts its overhead under 0.1 ms per request.
For dashboards and alerts, the Prometheus integration exposes bifrost_input_tokens_total, bifrost_output_tokens_total and bifrost_cost_total (USD), labeled with virtual_key_id, team_id, customer_id, provider and model. A Grafana panel of bifrost_cost_total grouped by team_id is a showback report, and an alert on its rate of change is the higher-frequency anomaly detection the FinOps Framework recommends. The 2026 roundup of top AI gateways lists observability among its eight criteria for readers comparing what other products ship.
Enterprise controls: RBAC, SSO and audit logs
Enterprise cost governance adds identity and accountability on top of budgets: who may raise a budget, how users receive keys, and a signed record of every change. Bifrost Enterprise is a strict superset of the open-source gateway with the same config.json schema, so keys and budgets carry over unchanged.
| Control | What it does for cost governance | Tier |
|---|---|---|
| Virtual keys, budgets, rate limits | Attribute and cap spend per key, team and customer | Open source |
| Model and provider allow-lists | Restrict which models each key can call | Open source |
| Semantic caching, cost logging, Prometheus metrics | Avoid repeat calls; record and export cost | Open source |
| RBAC | Limits who can create keys or change budgets | Enterprise |
| SSO and user provisioning (OIDC, SCIM) | Maps identity provider groups to teams and roles | Enterprise |
| Access profiles | Auto-issues per-user virtual keys with budgets and rate limits | Enterprise |
| Audit logs | Signed record of administrative changes, exportable to a SIEM | Enterprise |
Role-based access control ships with Admin, Developer and Viewer roles plus custom roles, assigned automatically from identity provider groups for Okta, Entra, Zitadel, Keycloak and Google Workspace. For cost control, RBAC answers the question budgets cannot: who is allowed to raise one.
Access profiles are templates covering allowed providers and models, budgets, rate limits and MCP tool access. Bifrost creates a per-user copy with isolated counters and issues a managed virtual key the user cannot weaken. When a user holds several profiles, access stacks but each request is charged to a single profile’s budget.
Audit logs record administrative operations with initiator, target and outcome, can be HMAC-signed for tamper evidence, and export as JSON or syslog. In a cost review, the audit log is where a budget increase or a new unrestricted key shows up with a name attached. Organizations that need those records on their own infrastructure often shortlist LLM gateways they can deploy themselves for that reason, and the Bifrost governance overview shows how the layers fit together.
Shadow AI spend on employee machines
Gateway budgets govern only traffic configured to reach the gateway. AI tools employees install on their own machines, such as desktop chat apps, browser AI and terminal coding agents, usually call providers directly, leaving that spend and its data flows outside every budget and audit trail. This ungoverned usage is commonly called shadow AI.
Coding agents make the cost side larger: they run long, tool-heavy sessions, and Anthropic’s pricing page notes that tool definitions and tool results bill as input tokens. On personal keys or expensed subscriptions, that spend appears only as reimbursements, with no project attribution and no cap. Provider-side options such as the per-user limits in Anthropic’s Claude Code workspace cover one tool on one provider.
Bifrost Edge, currently in alpha, extends the gateway’s governance and security controls to the endpoint. The Bifrost gateway remains the control plane and policy engine where virtual keys, budgets, rate limits, guardrails and audit logging are defined; Edge runs on macOS, Windows and Linux and routes each device’s AI traffic through that gateway. With Edge in place, existing virtual keys, budgets, audit logs and guardrails apply to the AI tools people actually use, not only to traffic someone pointed at the gateway.
- Identity-linked keys. Users sign in once through the organization’s SSO, which links the machine to the user and syncs their policies. No provider key is copied onto the laptop, and the menu bar or tray agent shows the active virtual key and its budget.
- App governance. With app governance, administrators decide which AI applications are allowed; blocked apps are stopped on the device before data leaves it, and newly detected apps and MCP servers enter an approval queue.
- Guardrails at the endpoint. Guardrails configured centrally, including secrets detection and PII patterns, evaluate prompts before they reach a provider and responses before they return.
- Fleet rollout. Edge deploys through MDM platforms including Jamf, Microsoft Intune, Kandji, Omnissa Workspace ONE and JumpCloud, with a managed configuration that carries only non-sensitive connection settings.
Supported applications today include Claude Desktop, the ChatGPT desktop app, Cursor, Codex, Claude Code, Codex CLI, OpenCode, and ChatGPT and Claude in the browser. As an alpha release, Edge is something to pilot on a subset of machines, with app coverage checked against the organization’s own tool inventory.
A rollout sequence for LLM cost management
A workable LLM cost management rollout moves from visibility to enforcement in stages, because budgets set before anyone has seen real usage are either too loose to matter or tight enough to break production.
- Route and observe. Put the gateway in front of existing traffic with governance optional, collect a few weeks of per-request cost, and reconcile the totals against provider invoices.
- Issue virtual keys per owner. Replace shared provider keys with one virtual key per service, team or customer, attached so the hierarchy matches the org chart and billing relationships.
- Restrict models. Set
allowed_modelson each key to what the workload needs, removing the most common source of surprise spend: a service quietly switched to a frontier model. - Set budgets from observed usage. Start at observed spend plus headroom, calendar-aligned to the finance month, with token rate limits on keys that call reasoning models.
- Enforce and alert. Make virtual keys mandatory and alert on cost by team before budgets are hit.
- Extend to people. Use access profiles for per-user keys and pilot endpoint coverage for employee AI tools.
Two limits apply. Gateway cost figures come from a price catalog, not the invoice, so negotiated discounts need separate reconciliation. And a gateway does not make prompts shorter; prompt design, model choice and reasoning effort still decide token counts. Teams choosing a product for this plan can compare open-source AI gateway projects on deployment dependencies or read an AI gateway buyer’s comparison that includes managed options; how the two gateway layers differ is covered in LLM gateway vs API gateway.
Next steps
LLM cost management works when every request carries an owner, a price and a limit before it reaches a provider. Provider workspaces cover part of that one provider at a time; a gateway covers it across providers, and an endpoint layer extends it to employee machines.
For teams that want to run the controls on their own infrastructure, a list of self-hostable open-source LLM gateways is a reasonable starting point, and a survey of the leading enterprise AI gateways covers the managed alternatives.
Teams evaluating Bifrost specifically can request a Bifrost demo, read the governance resource page, or review the source code on GitHub.
Sources
- FinOps Foundation: FinOps for AI overview
- FinOps Foundation: How to build a generative AI cost and usage tracker (Barney)
- FinOps Foundation: FinOps Framework, AI technology category
- OpenAI API pricing
- OpenAI reasoning models guide
- OpenAI spend limits guide
- Anthropic: Model pricing
- Anthropic: Workspaces
- Google Gemini API: Pricing
- Bifrost source repository (GitHub)
- Bifrost docs: governance overview
- Bifrost docs: virtual keys
- Bifrost docs: budget and limits
- Bifrost docs: virtual key routing
- Bifrost docs: semantic caching
- Bifrost docs: model catalog
- Bifrost docs: built-in observability
- Bifrost docs: Prometheus metrics
- Bifrost docs: enterprise overview
- Bifrost docs: role-based access control
- Bifrost docs: audit logs
- Bifrost docs: access profiles
- Bifrost docs: Bifrost Edge overview (alpha)
- Bifrost docs: Bifrost Edge app governance
- Bifrost docs: Bifrost Edge security and guardrails
- Bifrost docs: deploying Bifrost Edge with MDM
Questions readers ask
What is LLM cost management?
LLM cost management is the practice of measuring, attributing and limiting what an organization spends on language model APIs. It covers three jobs. Metering records tokens and cost per request. Attribution ties each request to a team, application, customer or user. Enforcement stops or reroutes traffic when a budget or rate limit is reached, before the provider invoice arrives rather than after it.
How do you track LLM costs per user or per team?
Give every consumer its own credential and record tokens and cost against that credential on every request. Provider projects and workspaces do this for one provider at a time. A gateway does it across providers by issuing virtual keys, pricing each request from a model catalog, and exporting cost with key, team and customer labels to logs and metrics that finance and engineering can both query.
Are reasoning tokens billed as output tokens?
Yes. OpenAI states that reasoning tokens are not visible through the API but are billed as output tokens, and Google's Gemini pricing page labels its output price as including thinking tokens. Because output is the most expensive token class, a reasoning model can cost several times more per request than its visible answer suggests, which is why per-request cost tracking matters more than per-request counts.
What is a virtual key in an AI gateway?
A virtual key is a credential issued by the gateway instead of a raw provider API key. Applications and users authenticate with it, and the gateway attaches policy to it, such as allowed providers and models, a spending budget with a reset period, and token or request rate limits. Provider keys stay inside the gateway, so revoking or capping one consumer does not touch anyone else.
Does semantic caching reduce LLM costs?
It can, for workloads where many requests ask for the same thing. In Bifrost, an exact-match cache hit is costed at zero and a semantic match costs only the embedding call used to find it. Savings depend entirely on how repetitive the traffic is, and the cache needs a vector store, a similarity threshold and a time-to-live chosen for the workload.
Are provider spend limits enough to control LLM costs?
They are a useful backstop but rarely sufficient on their own. OpenAI project limits and Anthropic workspace limits each cover one provider, enforce at a coarse unit, and OpenAI notes that enforcement is not instantaneous. Teams using several providers, or needing limits per customer or per user, usually add a layer that sees every request before it leaves the company.
