Skip to content
Newsroom
Models 2h ago by Rajat Jain

Kimi K3 vs DeepSeek V4-Flash: No Shared Basis, One Split

The honest answer is non-comparability: trackers won't share a weighted basis and 'V4-Flash' means three different models. Official numbers, disambiguated.

Kimi K3 vs DeepSeek V4-Flash: No Shared Basis, One Split

The leaderboard hosting this matchup says it best: “No shared weighted benchmark basis supports a winner.” That quote — from the tracker that ranks Kimi K3 fifth and still refuses to call the pair — is not a data gap, it is the thesis. Kimi K3 and DeepSeek V4-Flash are not two answers to the same question; they are answers to one shared question (agentic coding) measured on different instruments, and “V4-Flash” quietly means three different releases. This article disambiguates the variants and reconciles the pair on official numbers only. It is part of our open-weight coding hub and the sibling of our DeepSeek V4-Flash vs Qwen3.8-Max reconciliation.

What each model actually is

Kimi K3DeepSeek V4-Flash-0731
Total / active params2.8T / hybrid linear attention284B / 13B
Context1M, native vision1M
LicenseKimi K3 license (custom, official tag “other”)MIT
ReleaseWeights live July 27, 2026July 31, 2026 (retrain)
Official API price$3.00 / $15.00$0.14 / $0.28

Two intentional differences explain most of the surface score gap, before any benchmark is consulted: Kimi K3 activates the entire 2.8T sparse structure through a hybrid linear-attention design (Kimi Delta Attention). V4-Flash-0731 activates just 13B of 284B per token. Both are “1M context, permissive license” models; they are not the same size of bet.

The disambiguation nobody else does: which “V4-Flash”?

Almost every thread, tracker, and tweet comparing Kimi K3 to “V4-Flash” is comparing against a different object:

  • V4-Flash-0731 — the July 31 retrain of the flat line. 284B/13B, $0.14/$0.28. This is the model whose official card reports the Terminal-Bench 2.1 sweep (82.7).
  • The “Max” label — a max-effort evaluation setting of the single V4-Flash-0731 model, not a variant: no separate Flash-Max exists (see the disambiguation). Early “Kimi K3 vs V4-Flash” writeups citing a “Flash-Max” benchmarked the max-effort configuration of earlier preview checkpoints.
  • V4-Flash (flat) — the original July release the 0731 retrains, still the API default for some providers at $0.14/$0.28, and the one most cost-trackers index.

A “Kimi K3 vs V4-Flash” comparison that doesn’t state which checkpoint it means is meaningless because the numbers it reports depend on the build named. Everything below compares Kimi K3 against the current release build, V4-Flash-0731 — the only Flash variant with a current official price sheet. For the official word on what “Max” is (an effort label, not a model), see our 0731 vs Max disambiguation.

Benchmarks: official numbers only

DeepSeek’s official card for the 0731 build is unambiguous: Terminal-Bench 2.1 82.7, NL2Repo 54.2, DeepSWE 54.4, Cybergym 76.7. Kimi K3’s official position was once thinner — Moonshot published Elo and a technical report, not per-task numbers — but the official model card now publishes a full sweep on the Kimi Code harness: GPQA Diamond 93.5, DeepSWE 67.5, Terminal-Bench 2.1 88.3, ProgramBench 77.8, FrontierSWE 81.2, SWE-Marathon 42.0, Kimi Code Bench 72.9 — and the official DeepSWE leaderboard (mini-swe-agent harness) lists kimi-k3 at 68.5 (max effort), 15 points above deepseek-v4-flash at 53.3.

That asymmetry has narrowed since first publication: Moonshot now publishes an official per-task sweep, and the one shared official instrument — the DeepSWE board — places kimi-k3 (68.5) well above deepseek-v4-flash (53.3), while Kimi K3 still has no official Terminal-Bench row outside its own harness. The legitimate conclusions are (a) Kimi K3 competes at the frontier with vision in the context window and now carries official per-task numbers, and (b) on the shared official board the gap to DeepSeek’s terminal results is no longer abstract — and conflating independent instruments (Kimi Code harness vs DeepSeek’s release card vs the DeepSWE board) is still how “Kimi beats DeepSeek” and “DeepSeek sweeps everything” takes both form in the same week. See Qwen3.8-Max vs Kimi K3: The Numbers Beat for the shared-basis rows, and the Open-Weight LLM API Pricing 2026 table for the rate cards.

Pricing: the variant trap in the wild

SourceKimi K3”V4-Flash” (checkpoint-dependent)
Official cards (current)$3.00 / $15.00$0.14 / $0.28 (0731 only)
Community threads (r/LocalLLaMA)“weighing the 2.8T license risk""extremely cheap to run” — the flat line at $0.14 input

The raw list-price split is 21x on input against the 0731 — the only verifiable delta at the official list price. The reconciliation rule from our Qwen comparison applies verbatim: budget on the official card, and state the checkpoint before quoting a delta.

Which team should pick which

  • Pick Kimi K3 if your workload wants the 2.8T frontier position: scale-adjacent quality, native vision in a 1M-token window, and the architectural bet on hybrid attention (sub-quadratic serving of long context). Budget: $3/$15, and hardware that can hold a 2.8T footprint if you self-host.
  • Pick V4-Flash-0731 if the load is terminal-shaped and cost-linear: high-volume agentic loops, sweep-and-fix coding, per-call economics at $0.14 input. The 13B-active serving profile also makes VPC hosting realistic.
  • Split the fleet, as our coding hub decision rule notes: Kimi K3 for the planning/vision-heavy layer, V4-Flash-0731 for the grind.

What to watch

  • Kimi K3 terminal numbers: the fastest way this comparison changes is Moonshot publishing DeepSWE/Terminal-Bench results under official methodology.
  • V4-Pro: DeepSeek’s “coming soon” build inherits the 0731 trajectory; the Pro card is the next event for this whole cluster.
  • Tracker convergence: when the shared-basis language disappears from the hosting leaderboard, the comparison will finally have numbers everyone can cite.

First published August 12, 2026. Official numbers cited to the DeepSeek V4-Flash release card and Moonshot’s Kimi K3 report; see the open-weight hub for the full field.

#Models #Kimi #DeepSeek #Open Source