DeepSeek V4-Flash: 100x Cut in Frontier Model Cost
The latest open-weights release pushes benchmark-test economics to roughly $0.03 — undercutting closed frontier APIs in the steepest price action this month.
DeepSeek released V4-Flash in the closing days of July — an open-weights model built for agentic workloads at a price point the frontier has never seen: the stated marginal cost of running a standard benchmark evaluation against the released model comes to roughly three cents, which by DeepSeek’s accounting undercuts Anthropic’s equivalent per-test costs by more than one hundred times.
Key facts
- Priced for the test grind: DeepSeek publicly prices agentic workloads per-task — a typical evaluation run costs about $0.03, versus dollars for equivalent closed-frontier evaluations.
- Undercut by two orders of magnitude versus Anthropic’s Claude API per-test cost, per DeepSeek’s own pricing comparison released with the model.
- Open weights: the model is downloadable, self-hostable, and covered by DeepSeek’s permissive open-weights licensing — the same playbook as the Qwen and Kimi releases, inverted in cost structure.
- Agentic by design: built and benchmarked for tool loops, coding tasks, and multi-step reasoning — not chat polish.
- Release window: announced days before the Qwen3.8-Max launch, arriving at the start of a week where open-weight pricing forced the biggest closed-API price cuts in years (OpenAI’s Luna −80%).
The release
DeepSeek’s announcement drove straight to the economics. While frontier labs bill per token and justify it by inference cost, DeepSeek positions V4-Flash around per-task pricing for the workloads that actually dominate agentic infrastructure spend: running tests, browsing, executing tool calls. The headline claim — “about three cents per benchmark test” — is not a promotional token number; the company published a worked comparison against Claude per-test API costs and invited anyone to reproduce the math on their own instance.
Like its predecessor releases, V4-Flash ships open weights: the commercial licensing is permissive, the public API is cheap, and the intended audience is explicitly the world’s to-be-built agent fleet.
Why it matters
- The price war just found its floor. The full arc of this month — DeepSeek’s ~$0.03 per test, Kimi K3 at $3/$15, Qwen3.8-Max API under $2/6, OpenAI cutting Luna by 80% — is calibrating what frontier-grade output should cost in an open-weights world. Labs covering the gap with margins alone are now fighting unit economics, not feature sets. (The reconciled five-model rate table lives at Open-Weight LLM API Pricing 2026.)
- Research is the funnel. A $0.03 evaluation is cheap enough to be a default, which tilts the evaluation ecosystem (and any fine-tuning panel) toward DeepSeek’s open stack — the position plays for the agent fleet layer.
- Open vs closed framing has moved on. The conversation is no longer “is open competitive?” — it’s “can closed vendors justify a 100x premium on per-task cost for production agent workloads?” — and the answer is increasingly “only where safety assurance and latency justify it.”
What to watch
- The next Anthropic/OpenAI response. Whether closed vendors meet open-floor economics with per-task tiers of care. Any “agent-workload pricing” announcement within weeks is a reaction to this.
- Benchmark validation. Independent replication of the $0.03 claim is the real test — watch for third-party cost comparisons at scale (millions of tasks).
- Default adoption in agent frameworks. If agent builders re-baseline automatable work at DeepSeek prices, the unit economics of the next generation of AI startups changes retroactively.
Official source
- DeepSeek API docs: V4-Flash release notes and pricing page
- Model weights: DeepSeek-AI, Hugging Face
Updated August 8, 2026.