LiteLLM alternatives in 2026: gateways compared for production
LiteLLM is the default first gateway for many teams. Its production footprint, the March 2026 PyPI compromise and its license split send some of them looking elsewhere. Six alternatives, measured against published criteria, with a migration path.
TL;DR
- The LiteLLM alternatives worth shortlisting in 2026 are Bifrost, Kong AI Gateway, Agent Router (formerly Envoy AI Gateway), Apache APISIX, Cloudflare AI Gateway and OpenRouter. They split into self-hosted gateways and managed services.
- Teams leave LiteLLM for three documented reasons: the operational footprint of its Python proxy (PostgreSQL, Redis, worker recycling), the March 2026 PyPI compromise of versions 1.82.7 and 1.82.8, and a license split that puts SSO and audit logs behind an enterprise license.
- Bifrost ranks first here. It is an Apache 2.0 Go gateway with a
/litellmendpoint for existing LiteLLM SDK code, and its virtual keys, fallbacks, semantic caching and MCP gateway are in the open-source build.- Every self-hosted option, Bifrost included, keeps some enterprise controls on a paid tier. Compare where each vendor draws that line before comparing benchmarks.
LiteLLM is an open-source Python SDK and proxy that exposes more than 100 model providers through one OpenAI-format API, and for many engineering teams it was the first AI gateway they ran. Searches for LiteLLM alternatives tend to come later, once the proxy is carrying production traffic and its operating costs, its security exposure and its license boundary have started to matter. Bifrost, an open-source AI gateway written in Go by Maxim AI, is one of several gateways built for that stage, alongside API-gateway products with AI plugins and fully managed routing services. The comparison below states its criteria first, ranks six options against them, and ends with a migration path from LiteLLM.
Why teams move off LiteLLM
Teams replace LiteLLM for three reasons that trace to LiteLLM’s own documentation and the PyPI incident record: the proxy’s operational footprint at scale, two malicious releases on PyPI in March 2026, and a license boundary around the features large organizations ask for first.
None of this makes LiteLLM a poor product. The repository has close to 60,000 GitHub stars, the widest provider list in the category, and a description that now reads “Rust core with Python SDK”. The reasons below concern fit at scale, not quality.
Operational footprint
LiteLLM’s production best-practices page is candid about what a large deployment involves:
- One Uvicorn worker per pod, scaled horizontally, with 1 vCPU and 4 GiB of memory per pod. The page calls 4 GiB “a floor, not a target”, because the query engine of Prisma, the database client LiteLLM uses, grows its resident memory to fit the largest statement and does not return it.
- Worker recycling through
--max_requests_before_restart, to bound memory growth under sustained load. - PostgreSQL as the primary production database, with a per-worker connection-pool cap. The page’s own example, 100 replicas with four workers and ten connections each, comes to about 1,000 connections, more than a standard Postgres instance allows.
- Redis 7.0 or newer once there is more than one proxy instance, plus a Redis transaction buffer above 1,000 requests per second “to prevent spend-tracking deadlocks and connection exhaustion”.
These are reasonable instructions for a Python service that writes spend data to a relational database. The cost is a database tier and a cache tier to run and size alongside the gateway.
The March 2026 supply chain compromise
On 24 March 2026, attackers published LiteLLM versions 1.82.7 and 1.82.8 to PyPI with a credential stealer inside. The LiteLLM security update traces the compromise to a Trivy dependency in LiteLLM’s CI/CD scanning workflow. The payload harvested environment variables, SSH keys, cloud credentials, Kubernetes tokens and database passwords. The PyPI incident report puts total exposure at 2 hours 32 minutes and downloads of the bad versions at more than 119,000. PyPI estimated that 40 to 50 percent of LiteLLM installs were unpinned and fetching the latest release.
LiteLLM’s response was prompt. The official proxy Docker image was not affected because its dependencies are pinned, the team brought in Google’s Mandiant for forensics, and v1.83.0 shipped through a rebuilt pipeline. The lesson applies to every gateway: the gateway process holds every provider key a company has, so its supply chain deserves the scrutiny given to a secrets manager.
The license boundary
LiteLLM’s license file places the repository under MIT, except for an enterprise/ directory under a commercial license that allows production use only with a paid subscription. The enterprise features page lists single sign-on for the admin UI, audit logs, JWT authentication, IP allowlists, automated key rotation, secret-manager integrations, team-based logging and per-key guardrails. Pricing “depends on your deployment size” and is quoted on request.
Open-core licensing is normal in this category, and the question is where each vendor draws the line. A survey of open-source LLM gateways finds the same split across most self-hosted options. Buyers who need SSO and audit trails should expect to pay whichever gateway they pick.
Criteria for evaluating LiteLLM alternatives
A LiteLLM replacement should be judged on the problems that prompted the search, not on feature counts. The eight criteria below are weighted toward production operation; readers who weight them differently can re-rank the entries themselves.
| Criterion | What to check | Why it matters when leaving LiteLLM |
|---|---|---|
| Deployment model | Self-hosted binary, Kubernetes controller, API-gateway plugin or managed service | Decides where prompts and provider keys live |
| Runtime footprint | Required databases, caches and workers; memory at your request rate | The main operational complaint about the Python proxy |
| Gateway overhead | p99 latency added at your real payload size, measured by you | Vendor benchmarks use mocked upstreams and favorable settings |
| License and paid tier | OSI license of the core; which controls need a contract | SSO, RBAC and audit logs are usually paid everywhere |
| Governance | Virtual keys, budgets, rate limits, model allowlists | LiteLLM users depend on these today and need equivalents |
| Reliability | Retries, provider fallbacks, key-level load balancing | Replaces LiteLLM router configuration |
| Migration effort | LiteLLM SDK compatibility, OpenAI-compatible endpoint, model naming | Determines whether the switch touches application code |
| Agent and MCP support | MCP tool routing and per-key tool filtering | Increasingly part of the gateway’s job in 2026 |
This framework overlaps with the one used in a production-ready comparison of the top LLM gateways, which weighs performance and governance most heavily. The difference here is the emphasis on migration effort, since the reader here is assumed to run LiteLLM today. For the conceptual difference between these products and a conventional API gateway, see the explainer on LLM gateways versus API gateways.
LiteLLM alternatives compared
The six alternatives divide by deployment model. Four are self-hosted, although Kong can also be run with its managed control plane. Two are fully managed services that route requests through the vendor’s own infrastructure. The table summarizes each against the criteria, from each project’s own documentation as read in September 2026.
| Rank | Gateway | Deployment | Core license | Governance in free tier | LiteLLM migration path |
|---|---|---|---|---|---|
| 1 | Bifrost | Self-hosted binary, Docker, Go SDK; in-VPC on enterprise | Apache 2.0; paid enterprise tier | Virtual keys, budgets, rate limits, fallbacks, MCP tool filtering | /litellm endpoint; OpenAI and Anthropic SDK base-URL swap |
| 2 | Kong AI Gateway | Self-hosted data planes; Konnect managed control plane | Apache 2.0 Kong Gateway; AI Gateway Enterprise | Basic AI proxy and prompt guard; advanced balancing is Enterprise | OpenAI-format requests through the AI Proxy plugin |
| 3 | Agent Router | Kubernetes, on Envoy Gateway | Apache 2.0 | Token limits per team, app or model; provider fallback | OpenAI-compatible API; Kubernetes CRDs replace config |
| 4 | Apache APISIX | Self-hosted API gateway | Apache 2.0 | Token rate limiting, prompt guard, multi-provider balancing | OpenAI-compatible proxying via ai-proxy plugins |
| 5 | Cloudflare AI Gateway | Managed, on Cloudflare’s network | Proprietary service | Core features free; spend limits, rate limiting, caching | Base-URL change to a Cloudflare endpoint |
| 6 | OpenRouter | Managed routing and billing | Proprietary service | Provider routing and fallbacks; billing through OpenRouter | Base-URL change to openrouter.ai/api/v1 |
The ranking reflects this article’s weighting: a standalone gateway with low overhead, a permissive license, strong free-tier governance and the shortest migration from LiteLLM. A team already standardized on Kong or Envoy would reasonably weight consolidation higher and move entries two or three up.
1. Bifrost
Bifrost is the top pick among LiteLLM alternatives because it answers each of the three reasons teams leave: it runs as one Go binary that stores its configuration in embedded SQLite by default (PostgreSQL is optional), its core is Apache 2.0, and it ships an explicit compatibility path for LiteLLM code. The Bifrost gateway starts with npx -y @maximhq/bifrost or docker run -p 8080:8080 maximhq/bifrost, and the gateway setup guide shows a web console on port 8080 where providers, keys and fallbacks are configured.
Performance, as the vendor reports it
Bifrost’s benchmarking documentation reports 11 µs of added overhead per request at a sustained 5,000 requests per second on an AWS t3.xlarge, and 59 µs on a t3.medium, both with a 100 percent success rate against mocked OpenAI responses. A separate vendor benchmark against LiteLLM on a t3.medium at 500 requests per second for 60 seconds reports these results:
| Metric (vendor-run, t3.medium, 500 RPS) | Bifrost | LiteLLM |
|---|---|---|
| Throughput | 424 req/s | 44.84 req/s |
| p50 latency | 804 ms | 38.65 s |
| p99 latency | 1.68 s | 90.72 s |
| Peak memory | 120 MB | 372 MB |
| Success rate | 100% | 88.78% |
These are vendor figures, and LiteLLM’s configuration in that test was chosen by the vendor, not by LiteLLM. LiteLLM’s own benchmark page reports a median overhead of 2 ms with four instances (4 CPU, 8 GB each) at about 1,170 requests per second, which is a very different picture from the vendor test. The two tests are not comparable, but even on each vendor’s own numbers, Bifrost’s overhead is reported in microseconds and LiteLLM’s in milliseconds. That gap matters most for agent workloads that make dozens of model calls per task.
Open-source features that replace LiteLLM’s
The open-source build covers most of what LiteLLM proxy users configure today:
- Retries and fallbacks. Transient 5xx errors are retried with exponential backoff, and
429,401and403failures rotate to another key from the pool. When retries run out, the request moves to the next provider in the fallback chain, which gets its own retry budget. - Virtual keys. Each consumer gets a key with its own budget (resetting anywhere from one minute to one year), token and request rate limits, allowed providers and models, and an optional expiry. Keys roll up into teams and customers for hierarchical budgets.
- Semantic caching. Exact-match and similarity-based response caching, backed by Redis or Valkey, Weaviate, Qdrant or Pinecone.
- An MCP gateway. Bifrost connects to Model Context Protocol tool servers, filters tools per virtual key and can expose them to clients such as Claude Desktop. The MCP gateway roundup covers that side of the product in more depth.
Where the paid tier starts
The line is drawn in a familiar place. The Bifrost Enterprise overview lists clustering, adaptive load balancing, guardrails, OIDC identity federation (Okta, Microsoft Entra, Keycloak, Zitadel, Google Workspace), role-based access control, immutable audit logs, log exports and in-VPC deployment as enterprise features. There is no public price list, and a 14-day trial is offered.
Beyond routing, Bifrost applies governance and security controls centrally: virtual keys, budgets, guardrails and audit logs. Bifrost Edge, currently in alpha, extends that same governance to AI traffic on employee machines, with endpoint enforcement of the gateway’s guardrails for desktop apps, browser AI and coding agents.
Best for: engineering and platform teams that want a self-hosted gateway for production and regulated workloads. Bifrost suits teams that need microsecond-scale overhead, one control point for LLM and MCP traffic, and deployment inside their own network, and it is the shortest path off LiteLLM of any option here. It sits at or near the top of most lists of the best LLM gateways in 2026 for the same reasons.
Self-hosted alternatives
Three other self-hosted options suit teams that already run an API gateway or a Kubernetes-native traffic layer. Each adds AI routing to infrastructure that also carries ordinary API traffic, which means one less proxy to operate but ties the AI layer to a larger platform. A broader look at self-hostable AI gateways covers the same projects from a deployment angle.
2. Kong AI Gateway
Kong AI Gateway adds AI traffic policies to Kong Gateway, a Lua-based API gateway whose core repository is Apache 2.0. The open-source repository includes an ai-proxy plugin, which accepts OpenAI-format requests and translates them for each provider, alongside prompt-guard, prompt-template and request and response transformer plugins. Kong documents AI Proxy for providers including OpenAI, Anthropic, Azure OpenAI, Amazon Bedrock, Gemini, Vertex AI, Mistral, Cohere and Ollama. Kong’s documentation also lists a semantic cache, a PII sanitizer, semantic prompt and response guards, cost-based rate limiting, prompt compression, and MCP and Agent2Agent protocol support.
The boundary is significant. AI Proxy Advanced, the plugin that provides load balancing across providers (round-robin, lowest-latency, lowest-usage, semantic and priority-based failover), is documented as available only in Kong’s AI Gateway Enterprise offering. Production deployments usually run data planes in the customer’s environment connected to Kong’s Konnect control plane.
Best for: organizations that already run Kong for API management and want AI policy in the same platform, with the budget for its enterprise tier.
3. Agent Router (formerly Envoy AI Gateway)
Agent Router is the new name, since September 2026, of the Envoy AI Gateway project, which moved from the CNCF’s Envoy family to the Agentic AI Foundation under the Linux Foundation. The announcement says the code, maintainers, Apache 2.0 license and Kubernetes APIs are unchanged, and existing AIGatewayRoute manifests continue to apply. It is written in Go, runs on Envoy Proxy and Envoy Gateway, and offers one OpenAI-compatible API across 16 providers, with provider fallback, model-name virtualization, token limits per team, app or model, an MCP tool catalog filtered by caller, and OpenTelemetry GenAI tracing.
The trade-off is that Agent Router is configured through Kubernetes custom resources, which suits a platform team already running Envoy Gateway and is a heavy starting point for anyone else.
Best for: Kubernetes-first platform teams on Envoy Gateway that want an open governance model and no vendor tier.
4. Apache APISIX
Apache APISIX is an Apache 2.0 API gateway with more than 100 plugins, a set of them aimed at AI traffic. ai-proxy forwards requests to documented providers and OpenAI-compatible endpoints. ai-proxy-multi adds weighted round-robin or consistent-hashing balancing, semantic routing and retries. ai-rate-limiting enforces token quotas from provider-reported usage, and ai-prompt-guard allows or denies prompts by pattern. The AI gateway page also lists Redis-backed response caching and a Lakera Guard integration, but does not mention MCP.
Best for: teams already running APISIX that need token quotas and multi-provider balancing without adding a second proxy.
Managed alternatives
Managed gateways remove the operational footprint entirely. In exchange, prompts pass through a third party’s network, and governance is limited to what the service exposes. Teams with data-residency obligations cannot make that trade; for others, a managed service is the quickest exit from running a proxy.
5. Cloudflare AI Gateway
Cloudflare AI Gateway is a managed service on Cloudflare’s network that sits in front of providers including Workers AI, OpenAI, Anthropic, Google Gemini and Replicate. Integration is a base-URL change. Its feature list covers analytics, logging, caching, rate limiting, spend limits, dynamic routing with retries and model fallback, guardrails, data loss prevention and bring-your-own-keys. Cloudflare’s pricing page says the core features are free, and DLP scanning is free on all plans. Guardrails are billed as Workers AI inference, Logpush is a paid-plan feature, and Unified Billing adds a 5 percent fee on purchased credits.
Best for: teams already on Cloudflare that want observability, caching and rate limiting with nothing to deploy.
6. OpenRouter
OpenRouter is a hosted routing and billing service that exposes hundreds of models through https://openrouter.ai/api/v1, an OpenAI-compatible endpoint. It falls back to another provider automatically when one returns an error. Inference is billed at provider rates, with a fee on credit purchases (5.5 percent by card, 5 percent in crypto). Bring-your-own-key usage is free up to $25,000 a month on pay-as-you-go and carries a 5 percent fee above that. OpenRouter says prompts and completions are not logged by default, and it offers no self-hosted option.
Best for: prototypes and small teams that want many models behind one bill. It is weaker as a governance layer for a large organization, since budgets, access rules and audit trails live in a third-party account.
Performance claims and how to test them
Gateway overhead is the most-cited reason on LiteLLM alternatives pages and the claim least reliable without testing. Every vendor benchmark here, Bifrost’s and LiteLLM’s included, is run by the vendor on its own settings, and mocked upstreams hide the network and provider latency that dominate real requests.
A fair test for a team leaving LiteLLM takes about a week:
- Mirror production traffic. Send a copy of real requests, with real payload sizes and streaming, through the current LiteLLM proxy and each shortlisted gateway in the same region.
- Measure p99, not averages. Tail latency is where a Python worker under memory pressure and a compiled binary differ most, and it is what users notice.
- Record memory and restarts over the whole week.
- Break a provider on purpose. Revoke a key or point a provider at a failing endpoint, and check that fallbacks fire and the logs record the provider that served each request.
- Run the evaluation suite after any routing change. A fallback to a cheaper model can succeed on every request while degrading answers, the risk discussed in the guide to comparing LLM evaluation frameworks.
For a wider set of candidates and the results each vendor publishes, a comparison of leading AI gateways for production is a useful starting shortlist. Failover design, which the fourth step above tests, is covered in more detail in the Frontier Wire piece on LLM failover and load balancing.
Migrating from LiteLLM to Bifrost
Migrating from LiteLLM to Bifrost is mostly a base-URL change, because Bifrost accepts the request formats LiteLLM users already send. The LiteLLM SDK integration exposes a /litellm endpoint, so existing litellm.completion calls keep their provider-prefixed model names and only gain a base_url:
from litellm import completion
response = completion(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Hello!"}],
base_url="http://localhost:8080/litellm", # now routed through Bifrost
extra_headers={"x-bf-vk": "<BIFROST-VIRTUAL-KEY>"},
)
Services that call the LiteLLM proxy with the OpenAI or Anthropic SDK switch the same way. Under the drop-in replacement model, an OpenAI client points at http://localhost:8080/openai and presents a Bifrost virtual key instead of a LiteLLM key. The unified /v1/chat/completions endpoint also accepts the provider/model naming LiteLLM users are familiar with, such as openai/gpt-4o-mini.
Two details decide whether the switch is invisible to application code:
- Compatibility conversions. Bifrost’s LiteLLM compatibility plugin exists for code written against LiteLLM’s request handling. It converts text-completion requests to chat, converts chat to the Responses API for Responses-only models, and can drop or convert unsupported parameters. It is enabled globally or per request with an
x-bf-compatheader. - Provider coverage. The integration works only for providers both products support. A request to a provider only LiteLLM supports will fail, so compare provider lists before cutting over.
The configuration moves across by concept rather than by file:
| LiteLLM proxy concept | Bifrost equivalent |
|---|---|
| Virtual key with budget | Virtual key with budget and reset period, grouped into teams and customers |
| Router fallbacks and retries | Per-provider retries with key rotation, then a provider fallback chain |
| Deployment load balancing | Weighted keys per provider; adaptive balancing on the enterprise tier |
| Redis or semantic cache | Semantic caching plugin backed by Redis/Valkey, Weaviate, Qdrant or Pinecone |
| PostgreSQL spend logs | Built-in log store, Prometheus metrics and OpenTelemetry export |
| Enterprise SSO and audit logs | Enterprise OIDC, RBAC and audit logs |
The lowest-risk cutover routes one service through Bifrost first, compares a week of spend and latency, then moves the rest. Vendor migration notes and a feature matrix are collected on the Bifrost LiteLLM alternatives page. As with any vendor comparison, verify its LiteLLM column against LiteLLM’s current documentation. Teams planning production self-hosted LLM gateway deployments should also read the sizing notes in the guide to self-hosted LLM gateway deployments, and anyone building budgets into the new gateway can follow the guide to LLM cost control with an AI gateway.
Where LiteLLM still fits
LiteLLM remains a sensible choice in several situations. It has the broadest provider catalog here, more than 100 providers, so teams calling long-tail or regional providers may find no alternative covers them all. Its Python SDK runs inside an application with no proxy at all, which suits notebooks, research code and single-service products.
The case for moving is strongest for a team whose LiteLLM proxy has become shared production infrastructure, carrying many services and holding every provider key. At that point the database and cache tiers, the supply chain exposure and the paid boundary all matter at once. For teams that want to stay on open source while they decide, a roundup of open-source AI gateways you can run yourself is a sensible next read. The Frontier Wire explainer on running an AI gateway covers what the layer does before the product choice.
Limits and next steps
No gateway in this comparison was run under load for this article. Every product, LiteLLM included, is described from its documentation, pricing pages, repository and (for LiteLLM) the PyPI incident report, as read in September 2026. Benchmark figures are vendor-reported and were not reproduced, and feature boundaries change with each release, so re-check any figure on the linked page before a purchasing decision.
For a team leaving LiteLLM, Bifrost offers the shortest migration and the lightest runtime of the six options, with the governance most LiteLLM users rely on already in its open-source build. Kong, Agent Router and APISIX suit teams consolidating on an existing API gateway, and Cloudflare and OpenRouter suit teams that would rather run nothing. For a wider view of the market, the head-to-head review of enterprise LLM gateways scores the leading products on the same themes. Teams shortlisting LiteLLM alternatives can request a Bifrost demo or review the open-source repository.
Sources
- LiteLLM source repository and README (GitHub)
- LiteLLM license file (MIT, with enterprise directory carve-out)
- LiteLLM docs: production best practices
- LiteLLM docs: benchmarks
- LiteLLM docs: fallbacks (provider failover)
- LiteLLM docs: caching
- LiteLLM docs: enterprise features
- LiteLLM security update: suspected supply chain incident (March 2026)
- LiteLLM GitHub issue #24518: compromised PyPI packages, timeline and status
- PyPI blog: incident report, LiteLLM/Telnyx supply-chain attacks
- Bifrost documentation overview
- Bifrost source repository (GitHub, Apache 2.0)
- Bifrost docs: benchmarking
- Bifrost benchmark against LiteLLM (vendor page)
- Bifrost LiteLLM alternatives page (vendor page)
- Bifrost docs: LiteLLM SDK integration
- Bifrost docs: LiteLLM compatibility plugin
- Bifrost docs: drop-in replacement
- Bifrost docs: gateway setup
- Bifrost docs: retries and fallbacks
- Bifrost docs: virtual keys
- Bifrost docs: semantic caching
- Bifrost docs: enterprise overview
- Bifrost docs: Bifrost Edge overview (alpha)
- Bifrost docs: Edge security and guardrails
- Kong AI Gateway documentation
- Kong docs: AI Proxy Advanced plugin
- Kong Gateway source repository (GitHub, Apache 2.0)
- Agent Router (formerly Envoy AI Gateway)
- Agent Router blog: Envoy AI Gateway is now Agent Router
- Apache APISIX AI gateway
- Cloudflare AI Gateway documentation
- Cloudflare AI Gateway: pricing
- OpenRouter docs: quickstart
- OpenRouter docs: FAQ (fees, logging, BYOK)
Questions readers ask
What is LiteLLM used for?
LiteLLM gives applications one OpenAI-format API for more than 100 model providers. It ships as a Python SDK that runs inside an application and as a proxy server, the LiteLLM AI Gateway, that adds virtual keys, spend tracking, budgets, load balancing, fallbacks, semantic caching, guardrails and an admin dashboard. Teams use it to swap providers without rewriting code and to see model spend in one place.
Is LiteLLM free?
The core is free under the MIT license and can be self-hosted at no license cost, although production deployments need PostgreSQL, usually Redis, and compute. Code in the repository's enterprise directory is under a separate commercial license. Features such as single sign-on for the admin UI, audit logs, secret-manager integrations and IP allowlists need an enterprise license, priced by quote according to deployment size.
Is LiteLLM safe to use after the supply chain attack?
The malicious releases, versions 1.82.7 and 1.82.8 on PyPI, were quarantined on 24 March 2026, and LiteLLM shipped v1.83.0 through a rebuilt CI/CD pipeline. The official proxy Docker image was not affected because its dependencies are pinned. Anyone who installed either bad version should rotate every credential on the affected machine. Pinning versions with hashes is the durable fix, whichever gateway a team runs.
What is the difference between LiteLLM and OpenRouter?
LiteLLM is software you run yourself, so requests travel from your infrastructure straight to each provider using your own keys. OpenRouter is a hosted service. You call its endpoint, it routes to providers, and it bills you, taking a fee on credit purchases or, with your own keys, a fee above a monthly allowance. The choice is between operating a proxy and trusting a third party with the traffic.
What is the best LiteLLM alternative for self-hosting?
For a team that wants a standalone, self-hosted gateway, Bifrost is the strongest option in this comparison. It is an Apache 2.0 Go binary that needs no external database by default. It has a LiteLLM-compatible endpoint, and its virtual keys, fallbacks and MCP gateway are in the open-source build. Teams already running Kong, Envoy or APISIX should first check what their existing gateway can do.
Can Bifrost run existing LiteLLM code?
Yes, for providers both products support. Bifrost exposes a /litellm endpoint, so LiteLLM SDK code keeps working after its base_url points at the gateway. OpenAI and Anthropic SDK clients switch the same way, and a compatibility setting converts text-completion and chat requests for models that lack those endpoints. Requests to a provider only LiteLLM supports will fail, so check provider coverage first.
