Top 5 AI gateways for multimodal LLM workloads in 2026
A multimodal LLM stack sends images, audio, video and files as well as text. Five gateways, measured endpoint by endpoint against the same published criteria, and which kind of team each one suits.
TL;DR
- A multimodal LLM workload sends images, audio, video and files as well as text, and each modality uses its own API endpoint, payload shape and pricing unit.
- A gateway that handles multimodal traffic well needs coverage per endpoint, not just per provider: vision chat, image generation, speech, transcription, realtime, video, files and batch.
- Bifrost ranks first on the published criteria: it covers every endpoint family in this guide in its Apache 2.0 build, caches speech, transcription and image responses, and prices each modality by its own unit.
- LiteLLM and Kong AI Gateway are the strongest self-hosted alternatives; Vercel and Cloudflare suit teams that want a managed service and nothing to operate.
Most production AI systems stopped being text-only some time ago. A support assistant reads screenshots, a sales tool transcribes calls, a marketing pipeline generates images and a voice agent streams audio in both directions. A multimodal LLM can handle several of those inputs in one model, and the providers behind them now span OpenAI, Anthropic, Google, AWS, ElevenLabs, Groq and many more. Each modality arrives on a different endpoint with a different request shape, file size and billing unit. This guide compares five gateways on how well they carry that traffic, with the criteria published before the ranking.
Multimodal LLM traffic at the gateway layer
A multimodal LLM is a model that accepts or produces more than one type of data, typically text with images, audio or video. The survey of multimodal large language models by Yin and colleagues describes the common design: a language model connected to encoders for other modalities. For the teams operating them, the important part is not the architecture but the interface.
That interface is not one endpoint. Vision is a chat request with an image attached. Image generation, speech synthesis and transcription each have their own routes in the OpenAI-compatible API format that most gateways speak. Realtime voice uses a long-lived WebSocket or WebRTC session; OpenAI’s Realtime API guide describes WebRTC for browsers and WebSocket for servers, and Google’s Gemini Live API streams audio, images and text over a stateful WebSocket. Bulk work goes through files and batch endpoints.
Figure 1: Multimodal support is a list of endpoints, not a single checkbox: each family has its own request shape, pricing unit and failure modes.
Payloads also grow. Anthropic’s vision documentation allows up to 600 images in one API request (100 for models with a 200,000-token context window) and caps standard requests at 32 MB. Gemini’s video understanding guide samples video at one frame per second by default. A gateway in the path has to forward those bodies without becoming the bottleneck, and has to log and price them in units that are not tokens.
The general case for running a gateway at all is covered in the Frontier Wire’s overview of the leading AI gateways and in the guide to running an AI gateway in front of every model call. This article narrows the question to one thing: which gateways carry multimodal AI traffic as completely as they carry chat.
How the gateways were evaluated
The criteria below weight endpoint coverage most heavily, because a gateway that routes chat but not transcription leaves the audio traffic outside its budgets, logs and failover. Every entry was assessed from its own documentation, read on 9 and 10 October 2026.
| Criterion | What to check | Why it matters for multimodal traffic |
|---|---|---|
| Vision input | Images and files inside chat requests, across providers | The most common multimodal call in production |
| Image generation and editing | /v1/images/generations and edits, and which providers back them | Image models change often; routing should not |
| Speech to text | /v1/audio/transcriptions, providers beyond OpenAI Whisper | Call and meeting transcription is high volume |
| Text to speech | /v1/audio/speech, specialist voice providers | Voice products need more than one voice vendor |
| Realtime | WebSocket or WebRTC sessions through the gateway | Voice agents bypass the gateway without it |
| Video | Video generation endpoints | The newest and most expensive modality |
| Files and batch | /v1/files, /v1/batches, cross-provider | Bulk media jobs run asynchronously |
| Caching beyond text | Which response types the cache stores | Repeated transcriptions and images are pure waste |
| Cost by modality | Pricing per image, per second, per character | Token-only pricing misstates media spend |
| Deployment | Self-hosted or managed, licence of the core | Decides where uploaded media and keys live |
Overhead matters less here than in pure chat workloads, because media requests run longer. Where a project publishes a figure it is listed, with its load and hardware.
The five gateways at a glance
“Not published” means the documentation read for this article did not state the capability; it does not mean the capability is absent.
| Gateway | Deployment | Vision chat | Image generation | Speech to text | Text to speech | Realtime | Video | Files and batch | Cache beyond text |
|---|---|---|---|---|---|---|---|---|---|
| Bifrost | Self-hosted, Apache 2.0 | Yes | Yes, plus edits and variations | Yes | Yes | WebSocket and WebRTC | Yes | Yes, cross-provider | Speech, transcription, images |
| LiteLLM | Self-hosted, MIT core | Yes | Yes, plus edits | Yes | Yes | Yes | Yes | Yes | Transcription |
| Kong AI Gateway | Self-hosted data plane | Yes | Yes, plus edits (3.11+) | Yes (3.11+) | Yes (3.11+) | Yes, AI Proxy Advanced | Yes (3.13+) | Yes (3.11+) | Semantic cache for chat, AI licence |
| Vercel AI Gateway | Managed | Yes | Yes, plus edits | Yes, beta | Yes, beta | Yes, beta | Yes, beta | Batch for text generation, beta | Provider prompt caching |
| Cloudflare AI Gateway | Managed | Yes, via providers | Yes, via providers | Yes, via providers | Yes, via providers | WebSocket, listed providers | Not published | Not published | Text and image, exact match |
The table makes the shape of the market clear. The two self-hosted standalone gateways cover every endpoint family. Kong reaches the same coverage on recent versions through its AI plugins. The managed services cover most modalities, with several in beta at Vercel and with Cloudflare acting mainly as a pass-through proxy to each provider’s own API.
1. Bifrost
Bifrost is an open-source AI gateway written in Go and published under Apache 2.0 in the Bifrost repository. It ranks first because its open-source build covers every endpoint family in the criteria, routes each one across several providers, and treats audio, images and video as first-class in caching, logging and cost, not only in routing. Its documentation is also the most explicit in this group about which provider supports which operation, which is why it receives the longest write-up here.
Endpoint coverage. The provider support matrix lists, provider by provider, support for chat, the Responses API, image generation, image edits and variations, embeddings, text to speech, speech to text, files, batch, rerank, OCR and video. Speech synthesis is available through OpenAI, Azure, ElevenLabs, Gemini, Groq, Hugging Face and Sarvam; transcription through OpenAI, Azure, ElevenLabs, Gemini, Groq, Mistral, Hugging Face, Sarvam and self-hosted vLLM; image generation through OpenAI, Azure, Bedrock, Gemini, Vertex AI, xAI, Replicate, Nebius, Runware and Hugging Face; and video through OpenAI, Azure, Gemini, Vertex AI, Replicate, Runway and Runware. The multimodal guide shows vision with multiple and base64 images, audio input and output in chat, speech in MP3, Opus, AAC and FLAC, and transcription with timestamps. Overall the gateway reaches 25+ providers and 10,000+ models through one OpenAI-compatible API.
Figure 2: Provider coverage differs by operation, so the useful question is which providers back each endpoint, not how many providers a gateway lists.
Per-provider setup is documented in guides such as ElevenLabs speech and transcription setup and Runway video generation setup, which list the request mappings and async polling for each. The project’s post on multimodal support in Bifrost walks through vision, speech and transcription requests end to end.
Realtime, files and batch. The API reference documents a realtime WebSocket session to realtime-capable providers, a WebRTC SDP exchange and ephemeral client secrets, so browser voice clients can be authorized without exposing provider keys. The files and batch integration lets the OpenAI SDK manage files and batch jobs across OpenAI, Anthropic, Bedrock and Gemini, with S3 locations configured for Bedrock.
Caching and cost. The semantic caching documentation lists what gets cached: chat and text completions, the Responses API, embeddings, transcriptions, speech and image generation, including streaming variants, with exact-match and embedding-based modes. The model catalog prices audio by characters, tokens or duration, images per image, per pixel or by tokens, and video by tokens or by seconds at banded resolutions, and the request log captures speech and transcription inputs and outputs and image URLs. Custom providers can restrict each provider instance to specific request types, for example allowing transcription but not speech, which gives platform teams a clean way to scope expensive modalities.
Performance. Bifrost’s published benchmark reports 11 µs of added overhead per request at 5,000 RPS on an AWS t3.xlarge, documented in its benchmarking guide.
Enterprise tier. Clustering, guardrails, OIDC user provisioning, audit logs and log exports to S3, GCS and BigQuery are part of Bifrost Enterprise, which runs the same configuration as the open-source build.
Best for: teams that run vision, voice and media generation side by side and want all of it under one self-hosted gateway with per-modality caching, cost and routing.
2. LiteLLM
LiteLLM is a Python SDK and proxy server with an MIT-licensed core. Its strength for multimodal work is breadth of endpoints and an unusually long list of specialist providers behind each.
Endpoint coverage. The image generation docs list OpenAI, Azure, Google AI Studio, Vertex AI, AWS Bedrock, Black Forest Labs, Recraft, OpenRouter, Xinference and Nscale. The transcription docs list OpenAI, Azure, Vertex AI, Gemini, Deepgram, Groq, Fireworks AI, OVHcloud and Mistral, and the text-to-speech docs add Gemini, Vertex AI and AWS Polly voices. A /videos endpoint follows OpenAI’s video generation specification for OpenAI, Azure, Gemini, Vertex AI and RunwayML, and an OCR endpoint follows Mistral’s format.
Realtime and batch. The realtime endpoint load-balances voice sessions across OpenAI, Azure and xAI, with Gemini and Vertex AI support listed. The batches API covers OpenAI, Azure, Vertex AI, Bedrock and Mistral batch jobs.
Caching. With caching enabled and no restriction, the proxy caching docs cover chat and text completions, embeddings, audio transcriptions, rerank, the Responses API and Anthropic-format messages.
Enterprise tier. SSO, audit logs and several governance controls sit in LiteLLM’s enterprise edition, and the docs mark batch cost tracking as enterprise-only.
Best for: Python-centric teams that want the widest list of specialist image and speech providers behind one proxy.
3. Kong AI Gateway
Kong AI Gateway adds AI plugins to Kong Gateway, so organizations that already route API traffic through Kong can extend the same data planes to multimodal LLM calls. Coverage depends on the gateway version, which is worth checking before planning around a modality.
Endpoint coverage. The AI Proxy plugin documentation maps route types to endpoints: embeddings, files, batches, the Responses API, speech, transcriptions, translations, image generation and image edits from Kong Gateway 3.11, and video generation from 3.13. The plugin translates OpenAI-format requests into each provider’s format, or passes a provider’s native format through while still recording analytics and cost.
Realtime. The AI Proxy Advanced plugin adds a realtime/v1/realtime route for bidirectional WebSocket streaming, alongside its load-balancing algorithms. That plugin requires an AI licence.
Caching. The AI Semantic Cache plugin stores chat requests in a vector database by meaning; it also requires an AI licence, and its documentation describes chat traffic rather than media responses.
Best for: organizations already standardized on Kong that want multimodal routes under their existing gateway policies.
4. Vercel AI Gateway
Vercel AI Gateway is a managed service reached through the AI SDK or OpenAI- and Anthropic-compatible APIs, with bring-your-own-key at no markup. Its multimodal surface has grown quickly, and several pieces are labeled beta.
Endpoint coverage. The image generation docs cover generating and editing images. Vision and file input work through chat, Messages and Responses APIs. Video generation, speech to text (with models such as openai/whisper-1 and openai/gpt-4o-transcribe) and text to speech from Google, Microsoft and OpenAI models are each in beta.
Realtime and batch. Realtime voice is exposed through the AI SDK’s experimental realtime interface. Batch processing, also beta, handles text generation jobs at half of standard token prices.
Caching. Automatic caching applies provider-side prompt caching markers rather than storing responses at the gateway.
Best for: teams already on Vercel that want image, video and voice models behind a managed endpoint with no infrastructure to run.
5. Cloudflare AI Gateway
Cloudflare AI Gateway is a managed proxy on Cloudflare’s network. Applications keep calling each provider’s own API through a gateway URL, and Cloudflare adds analytics, logging, caching, rate limits and fallbacks.
Endpoint coverage. The provider list includes media specialists alongside the large labs: Cartesia, Deepgram and ElevenLabs for speech, Fal AI and Ideogram for images, plus Replicate, OpenAI, Google AI Studio, Vertex AI, Bedrock and Azure OpenAI. Because requests use each provider’s native API, the media endpoints a team can reach are the ones each provider exposes.
Realtime. The realtime WebSockets API supports OpenAI, Google AI Studio, Cartesia, ElevenLabs, Fal AI and Deepgram on Workers AI, for text, audio and video interactions.
Caching. The caching documentation states that caching currently supports text and image responses, for identical requests only.
Best for: teams that want logging and caching across many media providers with a base-URL change and no gateway to operate.
Failover and governance across media providers
Failover is simpler for text than for media. Most chat models accept the same request and return an answer in the same shape, so a gateway can move a chat request from one provider to another and the application rarely notices; a walkthrough of automatic provider fallback shows the chat case. Media endpoints are less interchangeable, and a sensible fallback chain looks different for each modality.
| Modality | How interchangeable providers are | What a sensible fallback looks like |
|---|---|---|
| Speech to text | High: input is an audio file, output is text | A second transcription provider with a similar language list |
| Vision chat | High for description and extraction | Another vision-capable chat model; check image count and size limits |
| Image generation | Medium: same prompt, visibly different style | A second model approved by whoever owns the brand look |
| Text to speech | Low: voices are provider-specific | A pre-chosen equivalent voice on the second provider, or a clear failure |
| Realtime | Low within a session | Retry the session on another provider; an open session cannot move mid-call |
| Video | Low, and expensive per attempt | Usually fail fast and queue rather than retry on another model |
The practical consequence is that a multimodal stack needs per-modality routing rules rather than one global fallback list. OpenAI’s speech voices, for example, are named options such as alloy and nova that do not exist at ElevenLabs or Google, so a speech fallback has to map voices explicitly or it changes how the product sounds. Image fallbacks change the visual style of the output, which may be acceptable for internal drafts and unacceptable for customer-facing assets.
Governance follows the same logic. A key that may call a chat model should not automatically be able to generate video, which can cost far more per request than a chat completion. The gateways in this list handle that differently: some scope keys by model allowlist, some by route or endpoint, and Bifrost’s custom providers restrict each provider instance to named request types. Whichever mechanism a gateway uses, the test is simple: can a platform team give a team access to transcription without also granting image or video generation, and can it see the spend for each separately? Content screening raises the same question for images: a guide to screening text and image traffic with AWS Bedrock Guardrails at the gateway shows one way to apply a single policy to both.
Caching and cost for images and audio
Multimodal requests change two things a gateway does on every call: what it can cache and how it prices the result. The steps in the request path are the same as for chat, but the cache key is built from a file rather than a prompt, the payload can be megabytes, and the bill is counted in seconds, characters or images.
Figure 3: The steps match a chat request; what changes is the cache key, the payload size and the unit the gateway prices by.
Caching. Repeated media work is common: the same product photo described twice, the same help-center paragraph spoken aloud, the same recording transcribed by two services. Only some gateways cache those responses. The table below lists what each documents.
| Gateway | Media responses cached | Matching |
|---|---|---|
| Bifrost | Transcriptions, speech, image generation, plus chat and embeddings | Exact match and embedding similarity |
| LiteLLM | Transcriptions, plus chat, embeddings and rerank | Exact match by default, semantic options |
| Kong AI Gateway | Chat (semantic cache plugin) | Embedding similarity, AI licence |
| Vercel AI Gateway | Not published (provider prompt caching only) | Provider side |
| Cloudflare AI Gateway | Text and image responses | Exact match only |
Exact-match caching is safe for media because an identical file and identical parameters should produce an interchangeable result. Similarity matching needs more care; a technical deep dive on semantic caching explains the thresholds, and an analysis of the grey zone in semantic caches covers near-miss matches that look similar but should not be served.
Cost. Token counts misstate media spend. Transcription is billed by audio duration, speech by characters, image generation per image or by resolution, and video per second. A gateway that only counts tokens will show a voice product as nearly free. The Frontier Wire’s guide to LLM cost control with an AI gateway covers budgets and virtual keys; for multimodal stacks, check that the gateway’s price table includes the media units, and that budgets apply to media requests as well as chat. A breakdown of LLM cost optimization at the gateway covers caching, routing and budgets together.
Choosing a gateway for multimodal AI
The right choice depends mostly on where uploaded media may travel and on which modalities the product depends on today.
- Media must stay in your network. Shortlist the self-hosted options: Bifrost, LiteLLM and Kong. Of the three, Bifrost covers every endpoint family and media caching in its open-source build; LiteLLM has the longest list of specialist providers; Kong suits teams already running it.
- Voice is the core product. Check realtime support first, then speech providers. Bifrost, LiteLLM and Kong expose a realtime endpoint, Cloudflare proxies realtime WebSockets for six providers, and Vercel’s realtime voice is in beta.
- Image or video generation at volume. Prefer a gateway that routes generation across several providers and can cache identical requests, so a model change or an outage does not stop the pipeline. The Frontier Wire’s guide to LLM failover and load balancing applies to media endpoints too.
- Nothing to operate. Vercel and Cloudflare are base-URL changes. Compare which media features are still beta and what each logs about uploaded files.
Teams whose workloads are still mostly text should start with the broader ranking of AI gateways or the companion piece on LLM gateways for Anthropic, OpenAI and Gemini, and return to this list when media traffic grows. A longer look at bringing multimodal models to production with an AI gateway covers rollout patterns for mixed text, image and audio workloads.
Limits of this comparison
No gateway was load-tested with media payloads for this article. Every capability is described from documentation read on 9 and 10 October 2026, and beta labels are as the vendors published them on those dates. Upload size limits, streaming behavior for long audio and the quality of each provider’s models were not tested, and they vary by provider more than by gateway. Enterprise pricing for Bifrost, LiteLLM and Kong is quoted privately and was not assessed.
The ranking puts Bifrost first because it is the only gateway here whose open-source build covers every endpoint family in the criteria and caches speech, transcription and image responses. A team already committed to Kong or to Vercel’s platform would reasonably rank its own platform higher, and should check the version or beta status of each modality it needs before relying on it.
Sources
- A Survey on Multimodal Large Language Models (Yin et al., arXiv 2306.13549)
- Anthropic docs: vision
- OpenAI docs: Realtime API
- Google Gemini API: Live API
- Google Gemini API: video understanding
- Bifrost source repository (GitHub, Apache 2.0)
- Bifrost docs: supported providers and operation matrix
- Bifrost docs: multimodal support
- Bifrost docs: files and batch through the OpenAI SDK
- Bifrost docs: realtime API over WebSocket
- Bifrost docs: semantic caching
- Bifrost docs: model catalog and multimodal cost calculation
- Bifrost docs: custom providers and request-type control
- Bifrost docs: benchmarking
- Bifrost docs: enterprise overview
- LiteLLM docs: image generation
- LiteLLM docs: audio transcription
- LiteLLM docs: text to speech
- LiteLLM docs: realtime
- LiteLLM docs: video generation
- LiteLLM docs: batches
- LiteLLM docs: proxy caching
- Kong docs: AI Proxy plugin
- Kong docs: AI Proxy Advanced plugin
- Kong docs: AI Semantic Cache plugin
- Vercel docs: AI Gateway image generation
- Vercel docs: AI Gateway video generation
- Vercel docs: AI Gateway speech to text
- Vercel docs: AI Gateway realtime
- Vercel docs: AI Gateway batch processing
- Cloudflare docs: AI Gateway providers
- Cloudflare docs: AI Gateway caching
- Cloudflare docs: AI Gateway realtime WebSockets API
Questions readers ask
What is multimodal in LLM?
A multimodal LLM accepts or produces more than one kind of data, such as text, images, audio or video, in a single model. In practice that means sending a screenshot with a question, transcribing a call, generating speech or an image, or streaming live audio to a voice model, each through its own API endpoint.
What are examples of multimodal LLMs?
Widely used examples include OpenAI's GPT-4o family and its realtime and transcription models, Anthropic's Claude models with image input, and Google's Gemini models, which accept images, audio and video and power the Gemini Live API. Image models such as DALL-E and speech models such as Whisper usually sit alongside them in the same stack.
Which multimodal LLM is the best?
No single model leads every modality. Teams commonly mix a vision-capable chat model, a separate image model, a transcription model and a speech model from different providers. That mix is the main reason to put a gateway in front of them, so each modality can use the best available provider through one API and one key.
Can an AI gateway cache image and audio responses?
Some can. Bifrost's cache covers transcriptions, speech and image generation as well as chat and embeddings. Cloudflare AI Gateway caches identical text and image responses. LiteLLM's default cache covers chat, embeddings and transcriptions. Kong's semantic cache plugin targets chat requests and needs an AI licence.
Do AI gateways support the OpenAI Realtime API?
Several do. Bifrost, LiteLLM and Kong's AI Proxy Advanced plugin expose a realtime endpoint, Cloudflare AI Gateway proxies realtime WebSocket sessions for OpenAI, Google AI Studio and several voice providers, and Vercel AI Gateway offers realtime voice in beta through the AI SDK.
