Top 10 open-weight models in October 2026, ranked
Xiaomi, Z.ai, Moonshot, Alibaba and DeepSeek now publish weights within a few points of the closed frontier. Ten models, ranked on Artificial Analysis and Arena scores, with the licence clause, parameter count and GPU budget that decide whether a team can actually run each one.
TL;DR
- The ten best open-weight models in October 2026 are MiMo-V2.6-Pro, GLM-5.3, Kimi K3, Qwen3.8-2.4T-A95B, DeepSeek V4.1 Flash, Qwen3.8-27B, K2-Horizon-375B-A23B, MiniMax-M3, Inkling and Nemotron 3 Ultra.
- The top three sit within two points of each other on the Artificial Analysis Intelligence Index (46, 45, 44). On Arena’s crowd-voted text leaderboard, Kimi K3 is the highest-ranked open model, 16th overall.
- Licences split the list. Five models are MIT or Apache 2.0, one uses OpenMDW, and four add conditions that mostly bite companies selling inference: a security review at Z.ai, separate agreements at Moonshot and Qwen, and a notice-or-authorisation rule at MiniMax.
- Frontier open weights need a full GPU node at minimum. DeepSeek V4.1 Flash is the cheapest of the top group to host, and Qwen3.8-27B is the strongest model that fits on one GPU.
The best open-weight models of October 2026 come within a few points of the closed frontier on independent tests, but the gap between “downloadable” and “deployable” is wider than the leaderboards suggest. A model with a trillion parameters, a revenue clause aimed at hosting companies and a minimum footprint of eight Blackwell GPUs is open in a different sense from a 27 billion parameter Apache 2.0 model that runs on one card. This survey ranks ten models that were confirmed downloadable from Hugging Face on 5 October 2026, using independent benchmark scores first and then the licence text, parameter count, context window, hardware and provider coverage that decide whether each one can be used.
What changed since the summer
Four releases reshaped the open-weight field between June and September 2026. Moonshot published Kimi K3, a 2.8 trillion parameter model its card calls “the world’s first open 3T-class model”. Alibaba released Qwen3.8-2.4T-A95B, which the Qwen team describes as bringing “a Qwen-Max-class model to open release” for the first time. DeepSeek shipped V4.1 Flash, a model built to cut the memory cost of long contexts. And Xiaomi’s MiMo-V2.6-Pro, released on 21 September, took first place among open models on the Artificial Analysis index.
Two other facts frame the list. Meta has published no new Llama model on Hugging Face since the Llama 4 releases of April 2025; its current open release is the Apache 2.0 Muse Glimmer 30B. And the strongest closed models still lead. Artificial Analysis’s leaderboard puts the best proprietary score at 58 against the best open score of 46, and the top 24 places on its index are held by closed models.
The licences have moved too. Our survey of open-weight licensing found in September that most flagship releases ship under Apache 2.0 or MIT. That remains true for Xiaomi, DeepSeek, Thinking Machines and the Institute of Foundation Models. But Z.ai and Qwen have now joined Moonshot in writing custom licences for their largest models, each with a clause aimed at “Model as a Service” businesses, meaning companies that sell API access to the model.
How the models were chosen
Every model here had public weights on Hugging Face on 5 October 2026 and an independent score from Artificial Analysis. Each family appears once, represented by its strongest open checkpoint; siblings are named inside the entry. The six criteria below are weighted in the order listed, so capability decides the ranking and the rest decide whether a reader can act on it.
| Criterion | What was checked | Source |
|---|---|---|
| Capability | Artificial Analysis Intelligence Index v4.3.2 score; Arena text rank where listed; vendor benchmarks only as supporting detail | Artificial Analysis, Arena, model cards |
| Size | Total and active parameters for mixture-of-experts (MoE) models | Model cards, Hugging Face metadata |
| Context window | Native and maximum supported tokens | Model cards |
| Licence | Commercial use, user or revenue thresholds, attribution, prohibited uses | LICENSE file in each repository |
| Self-hosting cost | Minimum GPUs and precision in the vendor or vLLM recipe | Model cards, vLLM recipes |
| Availability | Number of API providers, listing on OpenRouter | Artificial Analysis, OpenRouter API |
Two terms recur below. A mixture-of-experts model stores many sub-networks (“experts”) but runs only a few per token, so its active parameter count sets the compute per token while its total count sets the memory needed to hold it. The Artificial Analysis Intelligence Index combines ten evaluations, including Terminal-Bench 4.0, Humanity’s Last Exam, SciCode, GDPval-AA and AA-LCR, into one score. It is run by a third party, which is why it leads the ranking here rather than vendor tables.
Top 10 open-weight models at a glance
| Rank | Model | Developer | Total / active params | Context | Licence | AA index | Minimum self-host footprint |
|---|---|---|---|---|---|---|---|
| 1 | MiMo-V2.6-Pro | Xiaomi | 1.02T / 42B | 1M | MIT | 46 | 8 GPUs (vLLM TP8); vendor example uses 2 nodes |
| 2 | GLM-5.3 | Z.ai | 753B / 40B | 1M | GLM-5.3 License | 45 | 8x H200 in FP8 |
| 3 | Kimi K3 | Moonshot AI | 2.8T / 104B | 1M | Kimi K3 License | 44 | 8x GB300 |
| 4 | Qwen3.8-2.4T-A95B | Alibaba | 2.4T / 95B | 262K, up to 1M | Qwen3.8-Max License | 40 | 8 Blackwell GPUs in NVFP4 |
| 5 | DeepSeek V4.1 Flash | DeepSeek | 552B / 16B | 1M | MIT | 39 | 4x GB200 or 8x H200 |
| 6 | Qwen3.8-27B | Alibaba | 27B dense | 262K, up to 1M | Apache 2.0 | 34 | 1 GPU |
| 7 | K2-Horizon-375B-A23B | Institute of Foundation Models | 375B / 23B | 512K | Apache 2.0 | 31 | 8x H200 in BF16 |
| 8 | MiniMax-M3 | MiniMax | ~428B / ~23B | 1M | MiniMax Community License | 29 | 8x H200 in BF16 |
| 9 | Inkling | Thinking Machines | 975B / 41B | 1M | Apache 2.0 | 25 | 4x GB200 in NVFP4 |
| 10 | Nemotron 3 Ultra | NVIDIA | 550B / 55B | Up to 1M | OpenMDW 1.1 | 23 | 8x B200 or 8x H200 |
Index scores are from Artificial Analysis as fetched on 5 October 2026 and use each model’s highest reasoning setting. Footprints are the smallest configuration named in the vendor card or vLLM recipe; production traffic usually needs more.
1. MiMo-V2.6-Pro
What it is. Xiaomi’s flagship, a sparse MoE with 1.02 trillion total and 42 billion active parameters, a 1 million token context and native input for text, image, video and audio, according to the MiMo-V2.6-Pro-RL model card. A second checkpoint, MiMo-V2.6-Pro-MOPD, followed on 27 September as a post-training upgrade that targets a failure mode Xiaomi calls tool-call repetition.
Licence. MIT, with no user or revenue conditions.
Strengths. It scores 46 on the Artificial Analysis index, first of 118 open-weight models in its class, and ranks 26th on Arena (1480, plus or minus 9). Xiaomi’s own table reports 89.9 on Terminal Bench 2.1 and 71.9 on DeepSWE v1.1, against 74.0 for Claude Opus 5 on the latter.
Limits. Artificial Analysis measured 41.2 output tokens per second, well below the 71.9 median for its class, and only four API providers. The Arena score rests on about 4,000 votes, fewer than its neighbours. The card’s SGLang example spans two nodes with 16-way tensor parallelism, and its vLLM example uses eight GPUs.
Best fit. Teams that want the strongest open model with the simplest licence and can run multi-GPU inference.
2. GLM-5.3
What it is. Z.ai’s August 2026 model, 753 billion total and 40 billion active parameters, with a 1,048,576 token context in its configuration. The GLM-5.3 card says it “uses the same base model as GLM-5.2” and that every gain comes from post-training. The main repository ships FP8 weights; a BF16 copy is separate.
Licence. The GLM-5.3 License is MIT-style with one added clause. A company that runs a Model as a Service business and has more than 10 billion US dollars in revenue over any 12 months “must pass Z.AI’s security review before using the Software or its derivative works for any commercial purpose.” The threshold excludes nearly every buyer. Note that Arena labels the model “MIT”; the repository’s licence file is the binding text.
Strengths. Its index score of 45 is second among open models, and Artificial Analysis counts 26 API providers, the most of any model here. Z.ai reports 88.2 on Terminal Bench 2.1, 84.5 on CyberGym and a GDPval-AA v2 rating of 1769, the last measured by Artificial Analysis. Arena places it 27th (1478, plus or minus 6).
Limits. It is verbose: 210 million output tokens during the index run against a median of 140 million, which raises the real cost per task. Z.ai also writes that cyber capability “developed faster than we expected”, which deserves a policy review before it is exposed to untrusted users. The vLLM recipe calls for 8x H200 in FP8, or 8x B200 for the full 1M context.
Best fit. Coding and agent workloads where broad provider choice matters. Teams wanting a pure MIT licence and lower cost can use the sibling GLM-5.3-Flash (320 billion total, 18 billion active, MIT), which scores 42.
3. Kimi K3
What it is. Moonshot AI’s 2.8 trillion parameter MoE with 104 billion active parameters, 896 experts of which 16 are selected per token, a 1M context and text and image input. Per the Kimi K3 card, it uses quantization-aware training “from the SFT stage onward” with MXFP4 weights and MXFP8 activations, so the published weights are already 4-bit.
Licence. The Kimi K3 License has two conditions. A Model as a Service business above 20 million US dollars in revenue over 12 months “must enter into a separate agreement with Moonshot AI”, and products above 100 million monthly users or 20 million dollars of monthly revenue must display “Kimi K3”. Section 4 exempts internal use and use “through Moonshot AI’s official products or certified inference partners”, a detail not covered in our September survey.
Strengths. It is the highest-ranked open model on Arena, 16th overall at 1488 (plus or minus 5) on more than 28,000 votes, and scores 44 on the Artificial Analysis index. Moonshot reports 91.2 on BrowseComp and 88.3 on Terminal-Bench 2.1.
Limits. It is the most expensive model here on hosted APIs, 3 dollars per million input tokens and 15 dollars per million output on Artificial Analysis’s reference pricing, and runs at 45 tokens per second. The vLLM recipe asks for “at least 8x GB300” and multi-node serving “for real production traffic”. Thinking is always on, and clients must pass reasoning_content back on every turn.
Best fit. Long-horizon agents and research tasks where output quality outweighs cost, run through one of 20 API providers rather than self-hosted.
4. Qwen3.8-2.4T-A95B
What it is. The open-weight form of Qwen3.8-Max: 2.4 trillion total and 95 billion active parameters, a hybrid of Gated DeltaNet linear attention and gated full attention, and a context of 262,144 tokens natively, extensible to about 1 million. The model card says the hosted Qwen3.8-Max adds vision input and other features, so the downloadable weights are text-only.
Licence. The Qwen3.8-Max License requires the model name on the interface above 100 million monthly users or 20 million dollars of monthly revenue. A licensee running a Model as a Service or “AI Work Assistant” business, defined as a coding or office-productivity product, with more than 50 million dollars in revenue over 12 months must obtain a separate licence. Internal use is exempt.
Strengths. Qwen reports 67.7 on SWE-bench Pro and 86.6 on Terminal Bench 2.1. Artificial Analysis scores the open weights at 40, sixth in its class.
Limits. The 45 that Artificial Analysis gives “Qwen3.8 Max (0902)” belongs to the hosted, proprietary API version, not to these weights. Output speed was 37.7 tokens per second, and only six providers serve it. At two bytes per parameter the BF16 weights come to roughly 4.9 TB; the vLLM recipe needs eight Blackwell GPUs in NVFP4 or sixteen in FP8.
Best fit. Organisations already standardised on Qwen that want Max-class quality on their own hardware and have the GPU budget for it.
5. DeepSeek V4.1 Flash
What it is. A multimodal MoE released on 10 September with 552 billion backbone parameters, 8 billion active during prefill and 16 billion during decode, and a separate 196 billion parameter “Engram” conditional-memory table. The DeepSeek-V4.1-Flash card says its key-value cache, the per-token memory a long conversation consumes, is 890 bytes per token, roughly a quarter of DeepSeek-V4-Flash.
Licence. MIT.
Strengths. It is the fastest model in the top five, measured at 213.2 tokens per second by Artificial Analysis, and among the cheapest at 0.30 dollars input and 1.20 dollars output per million tokens on DeepSeek’s API, with 24 providers. It scores 39 on the index and ranks 38th on Arena. DeepSeek reports 74.2 on DeepSWE v1.1 and 90.6 on Terminal-Bench 2.1 using its own harness.
Limits. It is very verbose (250 million tokens during the index run), and the release has no Jinja chat template; clients use DeepSeek’s encoding scripts or its deepseek-recipe library. The vLLM recipe puts the minimum at one GB200 NVL4 tray (768 GB) or one 8x H200 node.
Best fit. Input-heavy agent workloads where long context and cost per task matter more than the last few points of quality. The larger DeepSeek-V4-Pro-0813 (1.6 trillion total, 49 billion active, MIT) scores 36 on the index.
6. Qwen3.8-27B
What it is. A 27 billion parameter dense model with a vision encoder for image and video input, and a 262,144 token native context extensible to 1 million, per the Qwen3.8-27B card. It shares the hybrid attention design of the 2.4T model.
Licence. Apache 2.0.
Strengths. It scores 34 on the Artificial Analysis index, first of 142 models in its size class, where the median is 8. Qwen reports 61.7 on SWE-bench Pro, 73.0 on Terminal Bench 2.1 and 89.2 on GPQA Diamond. vLLM’s recipe says it fits on one Blackwell GPU in every precision, and its FP8 build runs on a single RTX 5090. Hugging Face’s counter showed more than 6.7 million downloads over the past month, the most of any model here.
Limits. Its Arena rank (99th) is far below its index score, which suggests the index’s agentic tasks flatter it relative to open-ended chat. It is also verbose, generating 200 million tokens during the index run.
Best fit. The default self-hosted model for teams with one GPU, edge deployments, and fine-tuning projects that need a permissive base.
7. K2-Horizon-375B-A23B
What it is. The flagship of the Institute of Foundation Models’ K2-Horizon family, released on 3 September: 375 billion total and 23 billion active parameters, a 524,288 token context and text-only input. The model card says the training data, recipe and training code are public, and intermediate checkpoints are published so changes can be studied across training.
Licence. Apache 2.0.
Strengths. It is the most open release in this list. Our licensing survey noted that the OSI’s Open Source AI Definition also asks for training data information and code; K2-Horizon publishes both, though we did not assess it against that bar. It scores 31 on Artificial Analysis, 15th among open models, and the developer reports 70.2 on Terminal-Bench 2.1 and 87.3 on GPQA Diamond.
Limits. Only one API provider serves it, it is not on Arena or OpenRouter, and Artificial Analysis measured a 20.29 second wait for the first token. The SGLang recipe was validated on 8x H200.
Best fit. Research groups, regulated buyers and anyone who needs to audit or reproduce training rather than just run the weights.
8. MiniMax-M3
What it is. A June 2026 model with about 428 billion total and 23 billion active parameters, a 1M context and native text, image and video input. Its MiniMax Sparse Attention design delivers, by the vendor’s account on the MiniMax-M3 card, “9x prefill and 15x decode speedups compared to M2 at 1M context”.
Licence. The most restrictive here. The MiniMax Community License grants rights “for non-commercial purposes”. Commercial use requires displaying “Built with MiniMax M3” and sending a one-time notice to MiniMax; above 20 million dollars in yearly revenue it requires “a separate, prior written authorization”. An appendix prohibits uses including “any military purpose”.
Strengths. It is efficient: Artificial Analysis recorded 120 million output tokens on the index, below the median, at 97.2 tokens per second and 0.30 dollars input per million tokens, across 15 providers. It scores 29.
Limits. Beyond the licence, it ranks 96th on Arena. The vLLM recipe calls 8x H200 or H20 “a tight single-node BF16 fit”.
Best fit. Long-context multimodal workloads, for teams that will file the notice and stay under the revenue line.
9. Inkling
What it is. Thinking Machines’ first open-weight model, released on 15 July: a 975 billion total, 41 billion active MoE that accepts text, image and audio input. The launch post says it was trained on 45 trillion tokens and supports up to 1M tokens of context, and the model card ships BF16 and NVFP4 weights.
Licence. Apache 2.0.
Strengths. It is fast, at 171.5 tokens per second on Artificial Analysis, and widely hosted: Together AI, Fireworks, Modal, Databricks and Baseten at launch, plus fine-tuning on Thinking Machines’ Tinker service. The vendor reports 77.6 percent on SWE-bench Verified and 74.1 percent on MCP Atlas, and it is one of few open models with audio input.
Limits. Its index score of 25 places it 29th among open models, below several smaller entries, and it sits 95th on Arena. The vLLM recipe pairs NVFP4 with 4x GB200 or 8x H200, and BF16 needs multiple nodes. A smaller Inkling-Small preview has 276 billion total and 12 billion active parameters.
Best fit. Teams that want an Apache 2.0 multimodal model with a managed fine-tuning path.
10. Nemotron 3 Ultra
What it is. NVIDIA’s 550 billion total, 55 billion active model, released on 4 June, built on a hybrid of Mamba-2 state-space layers, MoE layers and a few attention layers, and pre-trained in NVFP4. The Nemotron 3 Ultra card supports up to 1M tokens, though vLLM defaults to 256K.
Licence. OpenMDW 1.1, which grants rights without restriction and places no obligations on outputs, as our licensing survey covered.
Strengths. Transparency. The card says the model was pre-trained on about 20 trillion tokens, that “all datasets are disclosed”, that major portions of pre- and post-training data are released, and that the training recipe and RL environments are public. It states its hardware plainly: a minimum of 8x B200 or GB200, 16x H100 or 8x H200.
Limits. It scores 23 on the Artificial Analysis index and ranks 118th on Arena, the lowest capability in this list. NVIDIA reports 70.7 on SWE-Bench Verified, behind newer open models.
Best fit. Enterprises on NVIDIA infrastructure that want a documented, permissively licensed model and plan to fine-tune it.
Models that missed the list
Several strong releases were left out by the one-per-family rule or by a criterion.
- Gemma 4 (Google, Apache 2.0): the 31B dense and 26B A4B models score 15 to 17 on the index, but Gemma 4 31B ranks 76th on Arena, ahead of Qwen3.8-27B. The card lists five sizes down to E2B for phones.
- gpt-oss-120b (OpenAI, Apache 2.0): 117 billion total and 5.1 billion active parameters, it runs “on a single 80GB GPU” per its card, but at 10 to 12 on the index it has been overtaken since its August 2025 release.
- Muse Glimmer 30B (Meta, Apache 2.0): a 29.6 billion parameter agentic model distilled from Muse Spark, scoring 17, with a 131K context.
- Mistral Medium 3.5 (128B dense): scores 14, and its modified MIT licence withdraws all rights from companies with more than 20 million dollars in monthly revenue.
- Qwen3.8-Flash-Next (125 billion plus 51 billion n-gram embedding, 6 billion active): scores 40, but the Qwen Community License 1.0 requires a separate licence for any Model as a Service or AI Work Assistant business, with no revenue floor.
- MiMo-V2.6-Flash, GLM-5.3-Flash and DeepSeek V4-Pro-0813: strong siblings of ranked models, scoring 38, 42 and 36.
- Tencent Hy4-preview (Apache 2.0): published as a preview, so not ranked.
Licence terms side by side
The table below is drawn from the LICENSE file in each repository on 5 October 2026. “MaaS” means a business that gives third parties API access to the model.
| Model | Licence | Commercial use | Condition that matters |
|---|---|---|---|
| MiMo-V2.6-Pro | MIT | Yes | Notice only |
| GLM-5.3 | GLM-5.3 License | Yes | MaaS business above 10 billion dollars revenue in 12 months needs Z.ai security review |
| Kimi K3 | Kimi K3 License | Yes | MaaS above 20 million dollars in 12 months needs separate agreement; “Kimi K3” display above 100 million MAU or 20 million dollars monthly revenue; internal use exempt |
| Qwen3.8-2.4T-A95B | Qwen3.8-Max License | Yes | MaaS or AI Work Assistant above 50 million dollars in 12 months needs separate licence; name display above 100 million MAU |
| DeepSeek V4.1 Flash | MIT | Yes | Notice only |
| Qwen3.8-27B | Apache 2.0 | Yes | Notice only |
| K2-Horizon-375B-A23B | Apache 2.0 | Yes | Notice only |
| MiniMax-M3 | MiniMax Community License | With notice | “Built with MiniMax M3”; prior authorisation above 20 million dollars yearly revenue; prohibited-use appendix |
| Inkling | Apache 2.0 | Yes | Notice only |
| Nemotron 3 Ultra | OpenMDW 1.1 | Yes | None on use or outputs |
The pattern is consistent with our September findings. Permissive licences dominate, and the conditions that exist are aimed at inference resellers and very large consumer products rather than at companies building on the model. MiniMax is the exception: its grant starts from non-commercial use, so even a small company must file a notice.
Hardware for self-hosting
Memory is set by total parameters, not active ones. A trillion-parameter MoE that activates 42 billion parameters per token still has to hold all trillion in GPU memory, so it needs a full node even though each token is cheap to compute. That is why the cheapest frontier option is DeepSeek V4.1 Flash: its weights are smaller and its key-value cache is a fraction of its peers’, so one 8x H200 node or one GB200 tray serves it.
In practice the list falls into three hardware tiers:
- One GPU: Qwen3.8-27B; among the also-rans, Gemma 4 31B, Muse Glimmer and gpt-oss-120b.
- One eight-GPU node: GLM-5.3 (8x H200, FP8), DeepSeek V4.1 Flash, K2-Horizon, MiniMax-M3, Inkling in NVFP4, Nemotron 3 Ultra and MiMo-V2.6-Pro in its vLLM configuration.
- Multiple nodes or the newest Blackwell parts: Kimi K3 (at least 8x GB300) and Qwen3.8-2.4T-A95B (16 GPUs in FP8).
Self-hosting a node-class model makes economic sense when traffic is steady, data must stay in a controlled network, or a team wants to fine-tune. For bursty or experimental traffic, the same weights are available from inference providers: every model in the top ten except K2-Horizon is listed on OpenRouter’s model API, and several have free tiers there. Our analysis of the cost of reasoning models applies with force here, because the verbosity figures above multiply every per-token price.
Serving and routing open-weight models
Most teams will not choose one model. A common pattern is a self-hosted open model for high-volume or sensitive traffic, an inference provider for a larger open model, and a proprietary API for the hardest tasks. All three can speak the same protocol, because vLLM and SGLang expose OpenAI-compatible servers and nearly every provider does the same.
The piece that ties them together is a gateway: a proxy that gives applications one endpoint, holds the provider keys and decides where each request goes. Bifrost, an Apache 2.0 gateway written in Go by Maxim AI, is one example. Its vLLM provider documentation registers a self-hosted endpoint by URL, leaves the API key blank for local servers, and addresses models as vllm/<model_id>; SGLang and Ollama are supported the same way, alongside hosted providers such as Fireworks, Groq, Cerebras, Nebius and OpenRouter. A request then names its model in provider/model form and can carry a fallback chain, for example a self-hosted Qwen3.8-27B first, DeepSeek V4.1 Flash through a provider second and a closed model last. Each fallback gets its own retry budget, and the response records which provider answered. The vendor’s benchmark reports 11 µs of added overhead per request at 5,000 requests per second on an AWS t3.xlarge against a mocked upstream; real overhead should be measured at your own payload size.
Two operational points matter more with open weights than with closed APIs. First, the models here differ in how they handle reasoning: Kimi K3 needs its reasoning_content returned on every turn and DeepSeek V4.1 Flash ships no chat template, so test each one through the gateway before routing production traffic to it. Second, verbose models make cost tracking essential. Gateway-level observability for every model call that logs tokens, cost and latency per request shows quickly whether a cheaper model is really cheaper per task. Our guide to failover and load balancing covers the routing patterns, and a survey of open-source LLM gateways for self-hosted deployments compares the gateways that can front vLLM, SGLang and Ollama in air-gapped networks.
Choosing by use case
- Highest capability, simplest licence: MiMo-V2.6-Pro (MIT).
- Highest capability through hosted providers: GLM-5.3, with the widest provider coverage, or Kimi K3 for the best Arena showing at a higher price.
- Long context at low cost: DeepSeek V4.1 Flash.
- One GPU: Qwen3.8-27B, or Gemma 4 31B for chat-style work.
- Auditable training: K2-Horizon-375B-A23B or Nemotron 3 Ultra.
- Audio input and managed fine-tuning: Inkling.
Before committing, read the model card the way our guide to reading a model card like an auditor suggests: find which harness produced each vendor score and whether the comparison models ran at the same effort level. For how these models compare with closed ones, see our ranking of the best AI models.
Limits of this ranking
We did not run these models ourselves. Capability rests on Artificial Analysis and Arena, both third parties, and on vendor tables where neither covers a benchmark; vendor numbers are labelled as such. Artificial Analysis scores reflect each model’s maximum reasoning setting, which is the most expensive way to run it. Arena ranks carry confidence intervals that overlap among the top open models, so positions one to three should be read as a tie.
The one-entry-per-family rule hides strong siblings, and it favours the largest checkpoint in each family. Hardware footprints are minimums from vendor and vLLM recipes, not tested capacity plans. Licence summaries quote the files as they stood on 5 October 2026; vendors revise them in place, so re-read the file before each upgrade. Models available only through an API, including the hosted Qwen3.8-Max and every Anthropic, OpenAI and Google frontier model, were excluded by design. A new release from any of the five leading labs would likely change the top five within weeks.
Sources
- Artificial Analysis: LLM leaderboard (Intelligence Index v4.3.2)
- Artificial Analysis: MiMo-V2.6-Pro
- Artificial Analysis: GLM-5.3
- Artificial Analysis: Kimi K3
- Artificial Analysis: Qwen3.8 2.4T A95B
- Artificial Analysis: DeepSeek V4.1 Flash
- Artificial Analysis: Qwen3.8 27B
- Artificial Analysis: K2 Horizon 375B A23B
- Artificial Analysis: MiniMax-M3
- Artificial Analysis: Inkling
- Arena text leaderboard (updated 2 October 2026)
- MiMo-V2.6-Pro-RL model card (Xiaomi)
- GLM-5.3 model card (Z.ai)
- GLM-5.3 License
- Kimi K3 model card (Moonshot AI)
- Kimi K3 License
- Qwen3.8-2.4T-A95B model card (Alibaba Qwen)
- Qwen3.8-Max License
- Qwen Community License 1.0 (Qwen3.8-Flash-Next)
- DeepSeek-V4.1-Flash model card
- DeepSeek-V4-Pro-0813 model card
- Qwen3.8-27B model card
- K2-Horizon-375B-A23B model card (Institute of Foundation Models)
- MiniMax-M3 model card
- MiniMax Community License (MiniMax-M3)
- Inkling model card (Thinking Machines)
- Introducing Inkling (Thinking Machines, 15 July 2026)
- NVIDIA Nemotron 3 Ultra 550B-A55B model card
- Mistral Medium 3.5 licence (modified MIT)
- Gemma 4 31B model card (Google DeepMind)
- gpt-oss-120b model card (OpenAI)
- vLLM recipes: Kimi K3
- vLLM recipes: GLM-5.3
- vLLM recipes: DeepSeek-V4.1-Flash
- vLLM recipes: Qwen3.8-2.4T-A95B
- OpenRouter models API (listing and per-token prices)
- Bifrost docs: vLLM provider
- Bifrost docs: retries and fallbacks
- Bifrost docs: benchmarking
- Bifrost source repository (GitHub, Apache 2.0)
Questions readers ask
What is the best open-weight model in October 2026?
On the Artificial Analysis Intelligence Index v4.3.2, Xiaomi's MiMo-V2.6-Pro leads open-weight models with 46, ahead of GLM-5.3 at 45 and Kimi K3 at 44. On the Arena text leaderboard of 2 October 2026, Kimi K3 is the highest-ranked open-weight model at 16th overall. The top three are close enough that licence and hardware should decide between them.
Which open-weight models are fully permissive for commercial use?
Among the ten ranked here, MiMo-V2.6-Pro and DeepSeek V4.1 Flash are MIT, and Qwen3.8-27B, K2-Horizon-375B-A23B and Inkling are Apache 2.0. Nemotron 3 Ultra uses OpenMDW 1.1, which grants rights without restriction. GLM-5.3, Kimi K3, Qwen3.8-2.4T-A95B and MiniMax-M3 add revenue or scale conditions, mostly aimed at companies that resell inference.
What is the best open-weight model that runs on a single GPU?
Qwen3.8-27B, a 27 billion parameter dense model under Apache 2.0, scores 34 on the Artificial Analysis index, the highest of any model its size. vLLM's recipe says it fits on one Blackwell GPU in every precision. Gemma 4 31B and gpt-oss-120b also run on one data-centre GPU, at lower benchmark scores.
How much hardware does a frontier open-weight model need?
The top models need a full server node or more. vLLM's recipes call for at least eight GB300 GPUs for Kimi K3, eight H200s in FP8 for GLM-5.3, and eight to sixteen Blackwell or Hopper GPUs for Qwen3.8-2.4T-A95B depending on precision. DeepSeek V4.1 Flash is the cheapest of the frontier group, fitting on one four-GPU GB200 tray or one eight-GPU H200 node.
Can open-weight and proprietary models be used behind one API?
Yes. vLLM and SGLang expose OpenAI-compatible endpoints, and most inference providers do the same. A gateway such as Bifrost can register self-hosted endpoints alongside hosted providers and route between them through one OpenAI-compatible interface, with fallbacks from one to the other.
