Best Open-Weight Models for Coding in 2026
Qwen3.8-Max, DeepSeek V4-Flash-0731, Kimi K3, GLM-5.2, Llama 4 Maverick — official benchmark and pricing data, plus how to pick for local or API work.
The open-weight coding field in August 2026 is a five-model race: DeepSeek V4-Flash-0731, Qwen3.8-Max, Kimi K3, GLM-5.2, and Llama 4 Maverick. This is the hub page of our open-weight cluster — the numbers below come from the official model cards, not from third-party trackers, because the trackers disagree with each other (more on that in our V4-Flash vs Qwen3.8-Max reconciliation). Two newer pages extend the standard: Qwen3.8-Max vs Kimi K3: The Numbers Beat — the first shared official DeepSWE board rows for both models — and the Open-Weight LLM API Pricing 2026 reconciliation with all five rate cards.
The field, in one table
| Model | Active params | Context | License | API price (in/out, per M) | Best for |
|---|---|---|---|---|---|
| DeepSeek V4-Flash-0731 | 13B of 284B | 1M | MIT | $0.14 / $0.28 | Cost-driven agentic coding |
| Qwen3.8-Max | 95B of 2.4T | 1M (4M TokenBox) | Apache-2.0 | ~$2.00 / $6.00 | Long-horizon, high-capability work |
| Kimi K3 | 104B of 2.8T, hybrid linear attention | 1M | Kimi K3 license (custom) | $3.00 / $15.00 | Largest-scale open serving |
| GLM-5.2 | 753B total (8 experts per token) | 1M | MIT | $1.40 / $4.40 (Mistral-hosted) | European-hosted agentic coding |
| Llama 4 Maverick | 404B total | 1M | Llama 4 Community | Free weights | Ecosystem default, local-first |
Why the leaderboards disagree — and what to trust
Open any two aggregator pages and you get two different winners for the same models. Artificial Analysis currently indexes Qwen3.8-Max at 58 against V4-Flash at 42 (~$1.18 vs $0.07); orcarouter puts the same match at 58 vs 52 ($2.00 vs ~$0.15); eesel’s table lists Qwen3.8-Max at $2/$6. All three have updated pages this week, and all three disagree — on both price and index.
Ground truth is the official model card and the official API pricing page, cross-checked against the release notes. Benchmark numbers in this article are cited to the cards themselves; we treat trackers as directional only.
DeepSeek V4-Flash-0731 — the price-quality play
The July 31 retrain of the flat V4-Flash line is the cheapest serious coding model at frontier-adjacent scores. The official card reports Terminal-Bench 2.1 at 82.7, NL2Repo at 54.2, DeepSWE at 54.4, and Cybergym at 76.7 — measured on the release build, not a preview. At $0.14 in / $0.28 out it is roughly 15x cheaper than Qwen3.8-Max’s API and about 30x cheaper than Kimi K3 on input.
The catch: it is a 284B model, so “cheap API” and “cheap to self-host” are different questions — expect ~160GB in INT4 (official estimate), which is why our local VRAM guide is a required read before buying GPUs. Our full coverage: DeepSeek V4-Flash: 100x Cut in Frontier Model Cost — and when coverage says “V4-Flash Max,” our 0731 vs Max disambiguation shows which model is actually meant.
Qwen3.8-Max — the long-context workhorse
Alibaba’s sparse MoE activates 95B of 2.4T parameters per token and ships Apache-2.0 with a 1M context (4M via TokenBox) — the only model in this list with true long-horizon serving economics at the frontier tier. Launch-reported numbers: SWE-bench Pro 65.6, DeepSWE 1.1 at 65.93, and a GS-AGI 82.3 the team says clears Opus 4.8 — all vendor-reported, and the official DeepSWE board measured the model at 57.5 ± 3 (xhigh) on launch week, so treat the launch table as claims until the card is public (details in The Numbers Beat). On Terminal-Bench 2.1 the launch table cites 79.7 — comparing against a pre-0731 DeepSeek checkpoint, so treat any direct Qwen-vs-DeepSeek TB2.1 delta with the checkpoint dates in hand (we do exactly that in the reconciled comparison).
Self-hosting needs 418GB at BF16, or ~218GB quantized — a cluster proposition, not a workstation one. API pricing lands around $2/$6. Our full Qwen3.8-Max analysis.
Kimi K3 — the largest open model ever
Moonshot’s 2.8T-parameter hybrid-linear-attention release (weights live since July 27) is the biggest open checkpoint ever published and the only one with native vision in this list. It led the Artificial Analysis Elo at 1,547 on release and API-priced at $3/$15 undercuts every closed premium tier at this capability level. The hybrid attention makes 1M-context serving feasible at this scale — the architecture question most trackers ignore. See our Kimi K3 coverage before committing a cluster to it — for the head-to-head no tracker will call, our Kimi K3 vs DeepSeek V4-Flash comparison, and for the shared official board rows, Qwen3.8-Max vs Kimi K3: The Numbers Beat.
GLM-5.2 — the European-hosted option
Z.ai’s GLM-5.2 (MIT, 753B total, 8 of 128 experts activated per token, 1M context) is the agentic-coding release that just became geopolitically interesting: Mistral hosts it at $1.40 in / $0.26 cached in / $4.40 out per M tokens on its global and EU regional endpoints — the first third-party open model on Mistral’s sovereign compute (announced alongside Regional Endpoints GA — see our Mistral sovereign AI report and the pricing reconciliation). The US regional endpoint is not yet listed for GLM-5.2 (official docs). For EU teams under the active EU AI Act obligations, that’s the strongest compliance story of the five.
Llama 4 Maverick — the ecosystem default
No longer the strongest model in the list, but still the one Ollama, Open WebUI, and most local stacks ship by default: 404B total parameters, 1M context, free weights under the Llama 4 Community license. If your constraint is “runs today on the hardware we already own,” Maverick beats the newer models on integration and community tooling alone — the 2026 open-weight landscape page covers the decision framework in depth.
How to choose
- API-first and cost-driven: V4-Flash-0731, unless your workload needs Qwen3.8-Max’s long-context economics at scale.
- Capability ceiling, budget secondary: Qwen3.8-Max or Kimi K3 — pick Kimi for vision and scale, Qwen for licensing simplicity and SWE benchmarks.
- EU compliance in the picture: GLM-5.2 on Mistral’s regional endpoints.
- Self-hosting today: Llama 4 Maverick, then work up the VRAM reality guide.
What to watch
- DeepSeek V4-Pro is still officially “coming soon” — the 0731 retrain’s TB2.1 sweep makes the Pro release the most anticipated open-weight event this quarter.
- Mistral’s next hosted models — GLM-5.2 was explicitly the first of third-party open models on regional endpoints.
- Qwen’s cadence — a generational jump every few months makes the August-September window the likely next drop.
First published August 11, 2026 — checkpoint dates for every benchmark are cited at source.