DeepSeek V4-Flash vs Qwen3.8-Max: The Reconciled Comparison
Aggregators disagree on V4-Flash vs Qwen3.8-Max — here are the reconciled official numbers on benchmarks, pricing, and which team should buy which.
This is the most-searched comparison in the open-weight field, and the one every tracker gets differently. Artificial Analysis scores Qwen3.8-Max 58 vs V4-Flash 42 at ~$1.18 vs $0.07. orcarouter has the same match at 58 vs 52 ($2.00 vs ~$0.15). eesel lists Qwen3.8-Max at $2/$6. Three trackers, three different answers — so this comparison reconciles both models on official cards: DeepSeek’s V4-Flash release notes and Qwen’s Qwen3.8-Max launch documentation, and it is part of our open-weight coding hub where every model in the field is priced and benchmarked against official sources — and its sibling, our Kimi K3 vs V4-Flash comparison, covers the other frontier pair. The cluster’s newest pages carry the same standard: Qwen3.8-Max vs Kimi K3: The Numbers Beat — where the official DeepSWE board measured Qwen’s launch-claimed 65.93 at 57.5 ± 3 — and the Open-Weight LLM API Pricing 2026 table.
What each model actually is
| DeepSeek V4-Flash-0731 | Qwen3.8-Max | |
|---|---|---|
| Total / active params | 284B / 13B | 2.4T / 95B |
| Context | 1M | 1M (4M via TokenBox) |
| License | MIT | Apache-2.0 |
| Release | July 31, 2026 (retrain) | August 3, 2026 |
| Official API price | $0.14 / $0.28 | ~$2.00 / $6.00 |
The V4-Flash-0731 is the retrained version of the flat V4-Flash line. The “Max” label in tracker tables is a max-effort evaluation setting of the single V4-Flash-0731 model — no separate Flash-Max variant exists (see the disambiguation). Qwen3.8-Max is Alibaba’s sparse MoE — 25x more active parameters per token than DeepSeek, which is the root of most score differences (and most serving cost differences).
Benchmarks: reconcile on the same checkpoint
This is where the trackers mislead. Qwen’s own launch table shows Terminal-Bench 2.1 at 79.7 vs “DeepSeek-V4” at 71.0 — an 8.7-point gap that looks decisive. But “DeepSeek-V4” in that table is a checkpoint released before the July 31 retrain. DeepSeek’s official card for the actual release build reports TB2.1 at 82.7, NL2Repo 54.2, DeepSWE 54.4, and Cybergym 76.7.
Reconciled on the same builds, the two models trade blows rather than one model sweeping: Qwen leads SWE-style repo work on official launch numbers (SWE-bench Pro 65.6, DeepSWE 1.1 at 65.93), DeepSeek leads the agentic-terminal sweep on its final build. The honest headline: at the August 11 checkpoint, they are within ~3 points on the shared benchmarks, and the gap to the closed frontier — Opus 4.8 and the Sol tier — has collapsed for both.
Pricing: the four-answer problem
| Source (mid-Aug 2026) | V4-Flash (in/out per M) | Qwen3.8-Max |
|---|---|---|
| Official cards | $0.14 / $0.28 | CNY 14.988 / 44.965 (≈$2.1 / $6.3) |
| Artificial Analysis | ~$0.07 (indexed) | ~$1.18 (indexed) |
| orcarouter | ~$0.15 | ~$2.00 |
| eesel | — | $2.00 / $6.00 (GA Aug 2) |
The official numbers are the ones to budget on: DeepSeek’s own API pricing page, and Qwen’s published tarriff. The tracker discounts mostly reflect different blending of cache-hit and batch inference. The ratio that survives every dataset: DeepSeek is roughly 15x cheaper on input at the list price, and Qwen is priced like a premium tier while still being 2.5x cheaper than Kimi K3’s $3/$15 list.
Context and licensing
Both ship 1M-token context — the 2026 floor for serious agentic work. Qwen’s Apache-2.0 is the friendlier license for commercial re-licensing and fine-tune distribution; DeepSeek’s MIT carries an attribution clause for deployments above 100M monthly users. Neither blocks the normal “run, fine-tune, sell” stack.
Which team should pick which
- Pick V4-Flash-0731 if coding spend is a real line item: high-volume agentic loops, CI-adjacent workloads, per-call economics below rounding error. 13B active keeps serving cheap even in your own VPC.
- Pick Qwen3.8-Max for capability-rated work where context depth and repo-level SWE performance dominate: long refactors, multi-file agents, 1M-token codebases in the prompt. Budget for the ~2.4T serving footprint if you self-host (418GB BF16 / ~218GB quantized).
- Split the fleet — the standard pattern we see in production: Qwen3.8-Max for planning and refactoring, V4-Flash-0731 for the high-volume test-and-edit grind. The two models were clearly built for complementary halves of the same job.
What to watch
- V4-Pro remains “coming soon” — a Pro build with the 0731 benchmark trajectory is the main event for both models’ positioning.
- Qwen’s scheduler study (the LRM-40B lineage upgrade behind 3.8) has roadmap implications for context handling.
- Aggregator catch-up — expect tracker indexes to converge on the official TB2.1 numbers within weeks; the window of contradictory data is closing.
First published August 11, 2026. Benchmark numbers are cited to the official model cards and release notes; see the open-weight hub for the full field.