Top 10 open-source AI gateways for self-hosting in 2026
An open-source AI gateway is only as open as its license file and its paid tier allow. Ten self-hosted LLM gateways, read at the level of the LICENSE file, the dependency list and the release history, with criteria published before the ranking.
TL;DR
- The ten open-source AI gateways ranked here are Bifrost, LiteLLM, agentgateway, Agent Router (formerly Envoy AI Gateway), Apache APISIX, Higress, Kong Gateway, Plano, New API and vLLM Semantic Router.
- Eight are Apache 2.0 at the repository level. LiteLLM is MIT with a commercially licensed enterprise directory, and New API is AGPLv3 with extra attribution terms.
- “Open source” rarely means “everything is free”. SSO, audit logs, guardrails, clustering and advanced load balancing sit in a paid tier for Bifrost, LiteLLM and Kong, and agentgateway has a commercial distribution from Solo.io.
- The footprint ranges from one process with embedded SQLite (Bifrost, New API, agentgateway in standalone mode) to a Kubernetes cluster running Envoy Gateway (Agent Router).
- Project health varies sharply: Kong’s open-source repository last shipped a release in June 2026, while TensorZero’s repository is now archived and is not ranked.
A self-hosted LLM gateway is a proxy that runs inside your own network, takes requests from applications in one format, and forwards them to model providers or to your own inference servers. It holds the provider keys, applies rate limits and budgets, retries failed calls, and logs what went where. Plenty of products now call themselves an open source AI gateway. This comparison is narrower than a general gateway list: it reads each project’s LICENSE file, marks where the free edition stops and the paid one starts, lists what each gateway needs to run, and records project health from GitHub as fetched on 5 October 2026. The criteria come first, so a reader who weights them differently can re-rank the list.
What “open source” means in this list
Three licenses appear in the ten repositories, and they impose different obligations.
- Apache 2.0 is a permissive license. You may use, modify and redistribute the code, including in commercial products, provided you keep the copyright and license notices and state significant changes. It also includes an explicit patent grant from contributors.
- MIT is shorter and also permissive. It asks only that the copyright notice and license text travel with the code.
- AGPLv3 (the GNU Affero General Public License) is a copyleft license. If you modify the software and let users interact with it over a network, you must offer them the modified source. That clause is what separates it from the ordinary GPL.
A license at the top of the repository does not settle the question. Many AI gateways are open core: the core is under an open license, and a set of features is sold separately, either as a commercial directory in the same repository or as a separate enterprise build. The practical question for a buyer is where that boundary falls relative to the controls they need. The same reading habit applies to model weights, as the Frontier Wire map of open-weight licensing shows: the headline label and the clauses often differ.
This list only includes projects whose gateway code is published and can be self-hosted. Managed services are excluded. For a broader list that includes hosted gateways, see the sibling piece on the best AI gateways in 2026.
Evaluation criteria
The criteria below are weighted toward teams that will operate the gateway themselves. License and the open-core boundary come first because they decide what you can legally and practically run without a contract.
| Criterion | What was checked | Why it matters for self-hosting |
|---|---|---|
| License | The LICENSE file and any NOTICE or directory carve-outs | Decides commercial use, redistribution and copyleft duties |
| Open-core boundary | Features the vendor documents as enterprise or paid | Governance controls are often the paid part |
| Language and runtime | Primary language, proxy engine (Envoy, OpenResty, native) | Shapes performance, extension model and who can debug it |
| Deployment footprint | Required databases, caches, control planes, Kubernetes | Every dependency is another system to patch and back up |
| LLM gateway features in the free edition | Unified API, fallbacks, rate limits, budgets, caching, MCP | The reason to deploy a gateway at all |
| Project health | Stars, contributors, latest release, releases since 5 April 2026 | Indicates whether fixes and provider updates will keep coming |
| Operational effort | Config model, upgrade path, components to run | The cost after the first install |
GitHub figures were read through the GitHub API on 5 October 2026. Contributor counts are the API’s count, which includes commit authors without GitHub accounts. Release counts include every published GitHub release, which inflates the figure for repositories that tag each sub-module separately.
The ten gateways at a glance
| Rank | Gateway | License | Language and engine | Minimum footprint | Paid tier |
|---|---|---|---|---|---|
| 1 | Bifrost | Apache 2.0 | Go, native binary | One process, embedded SQLite | Bifrost Enterprise |
| 2 | LiteLLM | MIT, enterprise directory commercial | Python SDK and proxy, “Rust core” | Proxy plus PostgreSQL; Redis when scaled out | LiteLLM Enterprise |
| 3 | agentgateway | Apache 2.0 | Rust, native proxy | One binary with a config file | Solo Enterprise for agentgateway |
| 4 | Agent Router | Apache 2.0 | Go control plane on Envoy | Kubernetes 1.32+ and Envoy Gateway 1.8.1+ | None documented |
| 5 | Apache APISIX | Apache 2.0 | Lua on OpenResty (NGINX) | APISIX plus etcd, or standalone file mode | None on the AI gateway page |
| 6 | Higress | Apache 2.0 | Go control plane, Envoy, Wasm plugins | One all-in-one container for development | Not covered |
| 7 | Kong Gateway | Apache 2.0 | Lua on OpenResty (NGINX) | DB-less or PostgreSQL | Kong AI Gateway Enterprise |
| 8 | Plano | Apache 2.0 | Rust and Envoy, plus routing models | Native binary or one container | Hosted routing models |
| 9 | New API | AGPLv3 plus Section 7 terms | Go backend, React console | One container with SQLite | Commercial license on request |
| 10 | vLLM Semantic Router | Apache 2.0 | Go, sits beside Envoy | Docker stack with dashboard | None documented |
The ranking rewards a clean license, governance features in the free edition, a small footprint and steady releases. A team that already runs Envoy, NGINX-based gateways or Istio would reasonably move Agent Router, APISIX, Kong or Higress up, because consolidating on an existing proxy saves more operational effort than any single feature.
Project health in October 2026
| Gateway | Repository | Stars | Contributors | Latest release | Releases since 5 Apr 2026 | Steward |
|---|---|---|---|---|---|---|
| Bifrost | maximhq/bifrost | 8,558 | 193 | transports v2.2.5, 2 Oct 2026 | 100+ (per-module tags) | Maxim AI |
| LiteLLM | BerriAI/litellm | 60,151 | 1,768 | v1.104.0, 3 Oct 2026 | 100+ | BerriAI |
| agentgateway | agentgateway/agentgateway | 5,174 | 263 | v1.6.0, 2 Oct 2026 | 22 | Linux Foundation |
| Agent Router | theagentrouter/agent-router | 2,184 | 153 | v1.1.0, 21 Aug 2026 | 6 | Agentic AI Foundation |
| Apache APISIX | apache/apisix | 17,192 | 519 | 3.19.0, 28 Sep 2026 | 4 | Apache Software Foundation |
| Higress | higress-group/higress | 9,490 | 211 | v2.2.5, 4 Oct 2026 | 5 | CNCF sandbox |
| Kong Gateway | Kong/kong | 44,238 | 453 | 3.9.3, 17 Jun 2026 | 2 | Kong Inc. |
| Plano | katanemo/plano | 7,075 | 43 | 0.4.37, 28 Sep 2026 | 19 | Katanemo |
| New API | QuantumNous/new-api | 49,272 | 331 | v1.0.0-rc.41, 30 Sep 2026 | 61 | QuantumNous |
| vLLM Semantic Router | vllm-project/semantic-router | 6,036 | 245 | v0.4.0, 27 Sep 2026 | 2 | vLLM project |
Stars measure attention, not quality. The more useful signals are the date of the last release and who controls the project. Four of the ten sit in a foundation (ASF, CNCF, Linux Foundation, Agentic AI Foundation), which makes a sudden license change harder. Five are run by a single company, which can move faster but can also change course, and Semantic Router sits under the vLLM project.
General-purpose LLM gateways
These two are built as standalone LLM gateways first, with a unified API, provider keys, budgets and fallbacks as the core product rather than plugins on a general proxy.
1. Bifrost
Bifrost is an AI gateway written in Go by Maxim AI, published on GitHub under Apache 2.0 with no directory carve-outs in the license file. It exposes one OpenAI-compatible API across more than 20 providers and runs as a single binary.
License and paid tier. The repository is Apache 2.0 throughout. The Bifrost LLM gateway page states that the open-source version is free and that adaptive load balancing, clustering, guardrails, RBAC, SSO, audit logs and in-VPC deployment need an enterprise license. The enterprise overview adds log exports to object storage and data warehouses, directory sync through OIDC providers, and custom Go plugins to that list.
Footprint. The setup guide starts the gateway with npx -y @maximhq/bifrost or docker run -p 8080:8080 maximhq/bifrost. Configuration and logs live in SQLite by default; PostgreSQL 16 or later is the option for production, and a file-only mode turns the config store off entirely. The Helm guide recommends PostgreSQL with three or more replicas and autoscaling for high availability.
Free-edition features. Virtual keys carry their own budgets (resetting from one minute to one year), token and request rate limits, and model allow and deny lists, and roll up into teams and customers. The README lists provider fallbacks, key-level load balancing, semantic caching, an MCP gateway, Prometheus metrics and tracing in the open-source build. Budgets and virtual keys are covered in more depth on the Bifrost AI governance page.
Performance, vendor-reported. Bifrost’s benchmark documentation reports 11 µs of added overhead per request at 5,000 requests per second on an AWS t3.xlarge, against mocked OpenAI responses. It was not reproduced for this article.
Limits: the provider list is shorter than LiteLLM’s 100-plus, the controls large organizations ask for first (SSO, RBAC, audit logs, guardrails, clustering) are paid, and the project is stewarded by one company rather than a foundation.
Best for: teams that want a standalone, self-hosted LLM gateway with budgets and fallbacks in the free edition and no mandatory database.
2. LiteLLM
LiteLLM is the most widely adopted project in this list by stars and contributors. It ships a Python SDK that runs inside an application and a proxy server that adds virtual keys, spend tracking, load balancing and an admin dashboard. The README now describes it as having a “Rust core with Python SDK”.
License and paid tier. The license file places everything under MIT except the enterprise/ directory, which has its own license. The enterprise page lists SSO for the admin UI beyond five users, JWT authentication, audit logs, RBAC with organizations, IP allowlists, key rotation, secret-manager integrations and per-team logging, with pricing that “depends on your deployment size”.
Footprint. The production guide asks for one Uvicorn worker per pod with 1 vCPU and 4 GiB of memory, PostgreSQL as the only production database, Redis 7.0 or later when running more than one instance, and worker recycling to bound memory growth. Docker images, Helm charts and Terraform modules are published.
Free-edition features. Unified API for more than 100 providers, virtual keys with spend tracking, load balancing, Redis-backed response caching, guardrails and MCP integration. The README cites “8ms P95 latency at 1k RPS”, a vendor figure.
Limits: the PostgreSQL and Redis tiers add operational weight, and the March 2026 supply chain incident, when versions 1.82.7 and 1.82.8 on PyPI carried a credential stealer, is a reminder to pin versions. The official proxy Docker image was not affected. The Frontier Wire review of LiteLLM alternatives covers both issues in detail.
Best for: teams that need the broadest provider coverage and are comfortable running a database-backed Python service.
Kubernetes-native AI gateways
The next two treat the Kubernetes Gateway API as the configuration surface. They suit platform teams that already manage ingress through custom resources.
3. agentgateway
agentgateway is a Rust proxy for agent traffic, covering model calls, MCP tool servers and agent-to-agent (A2A) messages. It is a Linux Foundation project under Apache 2.0.
License and paid tier. The repository is Apache 2.0. Solo.io sells a separate distribution, Solo Enterprise for agentgateway, announced in an October 2025 press release with MCP tool-server fingerprinting and versioning, token exchange for downstream services, and signed audit trails.
Footprint. The documentation describes two modes: a standalone binary or container driven by a local configuration file, and a Kubernetes deployment “using a built-in control plane and Gateway API”. No external database is listed for either.
Free-edition features. The README lists an OpenAI-compatible API for major providers, budget and spend controls, load balancing, failover, guardrails (regex, OpenAI moderation, Bedrock Guardrails, Model Armor), JWT, API key and OAuth authentication, CEL-based authorization, MCP with stdio, HTTP and SSE transports, and a built-in UI. It shipped 22 releases in six months, the most active cadence among the infrastructure-style gateways here.
Limits: policy is written as infrastructure configuration, which suits platform engineers more than application teams, and the documented LLM provider list is shorter than those of the standalone gateways.
Best for: teams that want one Rust data plane for LLM, MCP and A2A traffic, with or without Kubernetes.
A related note: kgateway, the CNCF sandbox project that began as Solo.io’s Gloo, was once a common answer for an Envoy-based AI gateway. Its README says that from version 2.3.0 its AI and agentic features moved to agentgateway, so it is no longer ranked as an AI gateway here.
4. Agent Router
Agent Router is the new name of Envoy AI Gateway. According to the project announcement, it left the CNCF Envoy project on 10 September 2026 to become a standalone project in the Agentic AI Foundation, with “same code, same maintainers”. CRDs, API groups, CLI names, container images and Helm charts keep their old names.
License and paid tier. Apache 2.0. The README does not mention a commercial edition.
Footprint. This is the heaviest install in the list. The prerequisites page requires Kubernetes 1.32 or later, Envoy Gateway 1.8.1 or later, kubectl and Helm, and recommends a fresh cluster. Rate limiting is installed as an optional add-on with its own values file. A standalone aigw run command exists for local use.
Free-edition features. The README describes an OpenAI-compatible API for hosted providers and self-hosted inference, token-based rate limiting, provider fallback and MCP integration, with credentials and quotas managed by the platform team.
Limits: six releases in six months and the lowest star count of the ten, the rename means older guides and search results use the Envoy name, and the Kubernetes dependency rules it out for small teams.
Best for: platform teams already running Envoy Gateway who want AI traffic expressed as Kubernetes resources.
API gateways with AI plugins
APISIX, Higress and Kong began as general API gateways and added AI features as plugins. Their strength is consolidation: one proxy for REST APIs and model traffic. The trade-offs between the two kinds of product are covered in the explainer on LLM gateways versus API gateways.
5. Apache APISIX
APISIX is an Apache Software Foundation project written in Lua on OpenResty, the NGINX distribution with an embedded Lua runtime. It is the most active foundation-governed API gateway here, with 519 contributors and a 3.19.0 release on 28 September 2026.
License and paid tier. Apache 2.0. The AI gateway page describes only the open-source project.
Footprint. In its traditional and decoupled modes, APISIX uses etcd as its configuration store. The installation guide also documents a standalone file-driven mode that reads YAML or JSON from disk and needs no etcd, and warns that the example etcd provisioned by Docker and Helm installs is not a production design.
Free-edition features. Seven AI plugins: ai-proxy, ai-proxy-multi for load balancing and semantic routing across model instances, ai-rate-limiting for token quotas, ai-prompt-guard, ai-cache with Redis-backed exact and semantic caching, ai-rag, and ai-lakera-guard.
Limits: the AI gateway page does not mention MCP, the etcd cluster is a real operational cost in traditional mode, and there are no LLM-specific virtual keys or hierarchical budgets.
Best for: teams that want a foundation-governed gateway for both APIs and models with no paid tier in sight.
6. Higress
Higress originated at Alibaba and is now a CNCF sandbox project. It is built on Istio and Envoy, with a Go control plane and plugins compiled to WebAssembly (Wasm) from Go, Rust or JavaScript.
License and paid tier. Apache 2.0. A commercial edition was not covered for this article.
Footprint. The README starts an all-in-one container with one docker run command that opens a UI console on port 8001 and gateway ports 8080 and 8443. Production installs use Helm on Kubernetes, and a standalone mode outside Kubernetes is documented.
Free-edition features. Multi-model load balancing, token-based rate limiting, response caching, observability and hosting of MCP servers through its plugin system. Configuration changes take effect without dropping long-lived connections, which matters for streaming responses.
Limits: AI capability depends on which Wasm plugins you enable and keep current, and the Istio heritage means more moving parts than a single-binary gateway.
Best for: teams on Istio or Envoy who want AI features as Wasm plugins and a console out of the box.
7. Kong Gateway
Kong Gateway is the most-starred API gateway in this list, with more than 44,000 GitHub stars, and is written in Lua on OpenResty. Its README now calls it an “API, LLM and MCP gateway”.
License and paid tier. The repository is Apache 2.0, and its plugin directory contains six AI plugins: ai-proxy, ai-prompt-decorator, ai-prompt-guard, ai-prompt-template, ai-request-transformer and ai-response-transformer. The AI Proxy plugin supports more than 15 providers. AI Proxy Advanced, which adds weighted, lowest-latency, lowest-usage and semantic load balancing, circuit breakers and priority failover, is documented as “only available as part of our AI Gateway Enterprise offering”.
Footprint. Kong runs in DB-less mode with declarative configuration or with PostgreSQL, and has an official Kubernetes ingress controller.
Project health. This is the main concern. The most recent GitHub releases are 3.9.2 and 3.9.3 from June 2026. A community discussion asking for an open-source 3.10 image has no official reply from Kong maintainers, while the enterprise documentation already refers to 3.10 features.
Limits: the useful AI routing features are enterprise-only, and the open-source release line appears to have stalled at 3.9.
Best for: existing Kong users who want basic AI proxying and prompt guards on a gateway they already operate.
Routing layers and key hubs
The last three are open-source and self-hostable, but each is a narrower tool than the gateways above.
8. Plano
Plano, from Katanemo, was previously called archgw (the old repository name redirects to it). It is an Envoy-based proxy for agentic applications, and its repository is mostly Rust: it routes between agents and between models, applies guardrail filters, and emits OpenTelemetry traces.
License and paid tier. Apache 2.0. The distinctive part is its use of small routing models, such as a 4-billion-parameter orchestrator. The README says these are “hosted free of charge in the US-central region” for a first run, and that production users should run the models locally or ask the company for API keys.
Footprint. The deployment docs describe a native mode (planoai up plano_config.yaml), a Docker Compose mode and a single-container Kubernetes deployment. Fully self-hosting also means serving the routing models yourself.
Limits: 43 contributors, the smallest base here, and a default that sends routing decisions to a hosted model until you change it.
Best for: teams building multi-agent applications who want routing by intent rather than by fixed model name.
9. New API
New API is a Go and React fork of One API, built as a model hub for aggregating providers and redistributing access through keys. It converts between OpenAI, Claude and Gemini request formats and ships a web console with a setup wizard.
License. AGPLv3, with additional terms under Section 7 in its NOTICE file: modified versions that present a user interface must keep the line “Frontend design and development by New API contributors” and a visible link to the original repository. The README offers a separate commercial license for organizations that cannot accept AGPL obligations.
Footprint. One container with SQLite to start. The README lists MySQL 5.7.8+ or PostgreSQL 9.6+ for production, optional Redis for shared rate limits across nodes, and optional ClickHouse for logs.
Limits: every release in the past six months is a v1.0.0 release candidate, the copyleft terms need legal review before modification, and the README devotes space to the regulatory duties of operators reselling API access, which signals the main audience.
Best for: teams that want a self-hosted key-distribution hub with a web console and are comfortable with AGPL.
10. vLLM Semantic Router
vLLM Semantic Router is a routing layer from the vLLM project that chooses a model per request based on signals such as intent, difficulty, modality, risk and user preference. It is Apache 2.0 and written mostly in Go.
What it is not. Its introduction is explicit: it “does not replace the gateway or the model servers. Envoy continues to carry traffic.” It ranks last for that reason. It complements a gateway rather than standing in for one.
Footprint. The installation guide needs Python 3.10+, Docker (or Podman on Linux), and a curl installer that starts the stack, with a dashboard on port 8700. Kubernetes support is through an operator.
Limits: two releases in six months, and no governance features of its own.
Best for: teams running several self-hosted models on vLLM who want quality and cost routing between them.
Operational effort compared
Footprint is where self-hosting costs show up. The table counts the systems you must run, not the features you might enable.
| Gateway | Smallest production shape | Config model | Extra systems at scale |
|---|---|---|---|
| Bifrost | Binary or container with SQLite | Web UI, API or config.json | PostgreSQL for multi-replica HA |
| LiteLLM | Proxy pods plus PostgreSQL | YAML plus database | Redis, connection pooling, worker recycling |
| agentgateway | Binary with config file | File or Gateway API resources | Kubernetes for dynamic config |
| Agent Router | Kubernetes with Envoy Gateway | CRDs | Rate-limit add-on |
| Apache APISIX | APISIX plus etcd, or file mode | Admin API or YAML | Production etcd cluster, Redis for AI cache |
| Higress | Helm on Kubernetes | Console, Ingress or Gateway API | Istio-style control plane |
| Kong Gateway | DB-less node | Declarative YAML or Admin API | PostgreSQL for DB mode |
| Plano | Native binary or container | YAML | GPU serving for routing models |
| New API | Container with SQLite | Web console | MySQL or PostgreSQL, Redis, ClickHouse |
| vLLM Semantic Router | Docker stack | Dashboard and config | An Envoy-based gateway in front |
Two patterns stand out. Gateways that embed SQLite or read a config file (Bifrost, agentgateway, New API, Kong in DB-less mode) are fastest to stand up and easiest to back up. Gateways that depend on Kubernetes controllers (Agent Router, Higress, agentgateway in its Kubernetes mode) cost more to install but fit teams whose change process already runs through GitOps. Whichever shape you pick, plan failover across providers from day one; the guide to LLM failover and load balancing covers the patterns.
Projects considered but not ranked
Several projects that appear in other open-source gateway lists did not make this one, each for a stated reason.
- TensorZero. Its GitHub repository is marked archived (read-only), and its last release, 2026.6.0, is dated 4 June 2026. An archived repository receives no fixes, which rules it out for new deployments.
- One API. The MIT-licensed project that New API forked from has had no release since v0.6.10 in February 2025, and its repository has no commits since then.
- kgateway. Its AI features moved to agentgateway, as noted above.
- LocalAI. An MIT-licensed engine for running models on local hardware. It serves models rather than governing access to many providers, so it belongs behind a gateway, not in place of one.
Choosing a self-hosted LLM gateway
The right choice depends more on what a team already runs than on feature counts.
- No existing gateway, want the least to operate: Bifrost or agentgateway in standalone mode. Both run as one process without a database.
- Need the longest provider list: LiteLLM, accepting PostgreSQL and Redis.
- Already on Envoy Gateway or Istio: Agent Router or Higress.
- Already on an NGINX-based gateway: APISIX, or Kong if you accept that the useful AI routing is paid.
- Need a key-distribution hub with a web console: New API, after a legal read of AGPL.
- Routing between self-hosted models by task: vLLM Semantic Router or Plano, behind one of the gateways above.
A good test before committing is to run the shortlisted gateways against your own traffic shape: your payload sizes, your streaming ratio, your provider mix. Vendor benchmarks use mocked upstreams and favorable settings. A survey of open-source LLM gateways for self-hosted deployments covers five of these projects with a deployment-sizing lens and is useful further reading, keeping in mind that it is published by the company behind Bifrost.
Limits of this comparison
No gateway was installed or load-tested for this article. Every statement comes from the project’s license file, README, documentation or release history as read on 5 October 2026, and benchmark figures are vendor-reported. GitHub statistics change daily and stars can be inflated, so treat them as rough signals. Paid-tier boundaries were taken from vendor documentation and can move with any release. The commercial offerings for Higress and APISIX were not reviewed, MCP gateway depth was not compared in detail (the Frontier Wire comparison of MCP gateways does that), and no legal advice is offered on license obligations. A change in Kong’s open-source release policy, a license change at any single-vendor project, or an un-archiving of TensorZero would change this ranking.
Sources
- Bifrost source repository (GitHub, Apache 2.0)
- Bifrost docs: gateway setup
- Bifrost docs: Helm deployment guide
- Bifrost docs: benchmarking
- Bifrost docs: enterprise overview
- Bifrost docs: virtual keys
- Bifrost LLM gateway product page
- Bifrost AI governance page
- LiteLLM source repository and README (GitHub)
- LiteLLM license file (MIT, with enterprise directory carve-out)
- LiteLLM docs: production best practices
- LiteLLM docs: enterprise features
- LiteLLM security update: supply chain incident (March 2026)
- agentgateway source repository (GitHub, Apache 2.0)
- agentgateway documentation
- Solo.io press release: Solo Enterprise for agentgateway (October 2025)
- Agent Router source repository (GitHub, Apache 2.0)
- Agent Router blog: Envoy AI Gateway is now Agent Router
- Agent Router docs: prerequisites
- Apache APISIX source repository (GitHub, Apache 2.0)
- Apache APISIX AI gateway
- Apache APISIX docs: installation and deployment modes
- Higress source repository and README (GitHub, Apache 2.0)
- Kong Gateway source repository and README (GitHub, Apache 2.0)
- Kong Gateway open-source AI plugins directory (GitHub)
- Kong docs: AI Proxy plugin
- Kong docs: AI Proxy Advanced plugin
- Kong GitHub discussion #14405: OSS image for 3.10.0
- Plano source repository and README (GitHub, Apache 2.0)
- Plano docs: deployment
- vLLM Semantic Router source repository (GitHub, Apache 2.0)
- vLLM Semantic Router docs: introduction
- vLLM Semantic Router docs: installation
- New API source repository and README (GitHub, AGPLv3)
- New API NOTICE file (AGPLv3 Section 7 additional terms)
- kgateway source repository and README (GitHub)
- TensorZero source repository (GitHub, archived)
- One API source repository (GitHub, MIT)
- Survey: five open-source LLM gateways for self-hosted deployments
Questions readers ask
What is an open-source AI gateway?
It is a proxy you run on your own infrastructure that sits between applications and model providers. It gives every application one API, usually OpenAI-compatible, and adds provider keys, routing, retries, fallbacks, rate limits, budgets and logging in one place. Open source means the code is published under a license such as Apache 2.0, MIT or AGPL that lets you run and modify it.
Which open-source AI gateway has the fewest dependencies?
Among the ten in this comparison, Bifrost and agentgateway can run as a single process with no external database. Bifrost stores its configuration in embedded SQLite by default, with PostgreSQL optional. New API also starts on SQLite. LiteLLM's production guide asks for PostgreSQL and, with more than one instance, Redis. APISIX in its traditional mode needs etcd, and Agent Router needs Kubernetes and Envoy Gateway.
Are open-source AI gateways free for commercial use?
The license usually allows it, but read the file. Apache 2.0 and MIT permit commercial use with attribution. LiteLLM is MIT except an enterprise directory under a commercial license. New API is AGPLv3 with extra attribution terms, which matters if you modify it and offer it to users over a network. Separately, most projects keep features such as SSO, audit logs or advanced load balancing in a paid edition.
Is Kong Gateway still open source?
The Kong/kong repository remains under Apache 2.0 and contains six AI plugins, including AI Proxy and AI Prompt Guard. Its most recent GitHub releases are 3.9.x patch releases from June 2026, and a community discussion about missing open-source 3.10 images has no official answer from Kong. Advanced AI features such as AI Proxy Advanced are documented as part of the AI Gateway Enterprise offering.
What is the best open-source AI gateway for Kubernetes?
It depends on what else runs in the cluster. agentgateway and Agent Router both use the Kubernetes Gateway API; agentgateway is a Rust proxy with its own control plane, while Agent Router configures Envoy Gateway and needs Kubernetes 1.32 or later. Higress suits teams already on Istio. Bifrost also ships a Helm chart for teams that want a standalone gateway inside the cluster.
