Qwen3.8-Max vs Kimi K3: The Numbers Beat
Moonshot's card reports GPQA 93.5 and Terminal-Bench 2.1 88.3; Qwen's launch numbers are vendor-reported. One official board measures both — and disagrees.
Two of the summer’s biggest open-weight releases — Qwen3.8-Max and Kimi K3 — have never been comparable on a shared official basis. This article, part of the open-weight coding hub, changes that with one finding: the official DeepSWE leaderboard now lists both models on the same harness. The rows contradict one side’s launch materials — the numbers, not the commentary, are the story.
The two models
| Qwen3.8-Max | Kimi K3 | |
|---|---|---|
| Total / active params | 2.4T / 95B (official) | 2.8T / 104B (official) |
| Context | 1M (text-only) | 1M (native vision) |
| License | Apache-2.0 (unverified — official card gated at publish) | Kimi K3 license, tag “other” (official) |
| Release | August 3, 2026 | Weights live July 27, 2026 |
| Official price (per 1M) | CNY 14.988 in / 44.965 out (Intl); CNY 12 / 36 (Beijing) | $0.30 cache-hit / $3.00 in / $15.00 out |
Qwen3.8-Max steps 95B of a 2.4T sparse MoE per token — text-only, 1M context — per Qwen’s launch pages. Kimi K3 activates 104B of 2.8T through hybrid linear attention (Kimi Delta Attention over gated-MLA blocks, 896 experts), with native vision and MXFP4/MXFP8 quantized training, per Moonshot’s official model card. Everything else below is benchmark, and benchmark is where these two stop being comparable.
The shared instrument
The official DeepSWE leaderboard — the harness Moonshot’s own card cites as deepswe.datacurve.ai — now carries rows for both models on the same mini-swe-agent harness (fetched August 12):
| Model (board id) | Effort config | Official score |
|---|---|---|
| kimi-k3 | max | 68.5% |
| qwen3-8-max | xhigh | 57.5% ± 3 (CI 54.8–60.1) |
| deepseek-v4-flash | max | 53.3% |
| glm-5-2 | max | 43.8% |
Context: the board also lists Claude Opus 5 at 73.6, GPT-5.6 Sol at 72.7, and GPT-5.6 Terra at 69.6 — all mini-swe-agent. Kimi K3 holds the top open-weight row on this official instrument. The caveat that matters: effort configurations differ (max vs xhigh), so the row pair is a shared-basis data point, not a controlled experiment.
The other number in the headline: Qwen’s launch materials reported DeepSWE 1.1 at 65.93. The official board measured 57.5 (CI 54.8–60.1) — the launch claim sits outside the confidence interval. Harness-version differences are possible, but the official instrument did not reproduce the vendor figure at publish time. We are documenting the discrepancy, not adjudicating it.
What Moonshot’s card says officially
Kimi K3’s official card publishes a full agent-coding sweep on the Kimi Code harness (max effort, temp 1.0, topp 0.95): GPQA Diamond 93.5, DeepSWE 67.5, Terminal-Bench 2.1 88.3, ProgramBench 77.8, FrontierSWE 81.2, SWE-Marathon 42.0, Kimi Code Bench 72.9. On Terminal-Bench, the only official cross-model chain available: Kimi K3 at 88.3 vs GLM-5.2 at 82.7 (GLM’s own official launch figure, footnoted on Moonshot’s card). Qwen’s Terminal-Bench 2.1 79.7 is vendor-reported — it has no official board row at publish time — so no Qwen-to-anyone TB2.1 delta exists on official ground.
What other coverage missed: every “Qwen vs Kimi beats” headline this month was derived from non-comparable vendor claims. The official board gives the first shared, independently measured pair — and it says the gap is real: ~11 points on DeepSWE, with Kimi carrying the higher-effort row at a lower effort config to Qwen’s score still above it. The “numbers beat” is official; the pre-launch-claim narrative was not.
Pricing: official rows, one currency problem
Kimi K3’s rate card is unambiguous: $0.30 cache-hit, $3.00 cache-miss input, $15.00 output per 1M tokens (official platform docs). Qwen3.8-Max’s official rate card is CNY-only — CNY 14.988 / 44.965 per 1M International, CNY 12 / 36 Beijing — with no official USD list price at publish; the widely-cited $2/$6 is a conversion ($2.11/$6.33 at prevailing rates, labeled an estimate under our rules). Both models sit far above DeepSeek V4-Flash-0731 at $0.14/$0.28 — with the Flash pair and the “Max” label sorted in our Kimi K3 vs DeepSeek V4-Flash and 0731 vs Max pieces — and the full table, now including GLM-5.2’s official Mistral-hosted price, is on our pricing reconciliation page.
Which one should you run
- Pick Kimi K3 when the workload wants vision in the 1M window, or when official benchmark rows matter to your vendor risk review — it owns the top open-weight DeepSWE board row and the official card’s terminal sweep.
- Pick Qwen3.8-Max when Apache-2.0 licensing is a hard requirement (pending the card gate opening), CNY pricing fits your treasury, or your workload is text-only long-context where the 4M TokenBox option applies (vendor-reported).
- Split the fleet, per the hub’s choice rules: Kimi for the vision-heavy planning layer, Qwen for the long-horizon text layer, and DeepSeek’s Flash for the cost floor.
What to watch
- The card gate. Qwen’s official Hugging Face card was login-gated at publish; opening it resolves the Apache-2.0 question and the launch-claim discrepancies.
- Effort-level rows. Kimi’s board row is “max” against Qwen’s “xhigh” — an official xhigh row for Kimi K3 (or a max row for Qwen) would make the pair directly controlled.
- Moonshot’s Kimi Code harness publication — the official card’s numbers remain uncrossed by independent instruments until then.
Official source
- Moonshot AI model card: moonshotai/Kimi-K3 (fetched 2026-08-12) · official license tag
other, 2.8T params - Official DeepSWE leaderboard: deepswe.datacurve.ai (fetched 2026-08-12)
- Qwen3.8-Max announcement: qwen.ai/blog?id=qwen3.8 · Alibaba Cloud model info and model-pricing (fetched 2026-08-12)
First published August 12, 2026. Rows marked vendor-reported could not be traced to an official page at publish time; official tracking of disputed figures is logged in ops/CORRECTIONS.md.