Skip to content
Newsroom
Models 2h ago by Rajat Jain

Open-Weight LLM API Pricing 2026: Reconciled

Official rate cards for DeepSeek V4-Flash-0731 and V4-Pro, Kimi K3, Qwen3.8-Max, GLM-5.2, Llama 4 — cache terms stated, tracker spread as evidence, not fact.

Open-Weight LLM API Pricing 2026: Reconciled

Every open-weight pricing roundup this month quotes different numbers for the same models. Artificial Analysis indexes DeepSeek V4-Flash at ~$0.07; the official rate card says $0.14. eesel lists Qwen3.8-Max at $2/$6 in USD; the official first-party card is quoted in CNY. BenchLM’s “Flash (Max)” profile is the July 31 release under an effort label (see the disambiguation). This page is the cluster’s single source of truth, reconciled on official rate cards fetched at publish time — cache terms stated in the same row, derived deltas only where both inputs are official (hub, with the Qwen-vs-Kimi match, the V4-Flash vs Qwen3.8-Max reconciliation, and the Kimi-vs-DeepSeek comparisons alongside).

The official rate card (per 1M tokens)

ModelInput (cache miss)Cache hitOutputCurrencySource (fetched 2026-08-12)
DeepSeek V4-Flash-0731$0.14$0.0028$0.28USDapi-docs.deepseek.com/quick_start/pricing
DeepSeek V4-Pro$0.435$0.003625$0.87USDapi-docs.deepseek.com/quick_start/pricing
Kimi K3$3.00$0.30$15.00USDplatform.kimi.ai/docs/pricing
Qwen3.8-MaxCNY 14.988CNY 44.965CNY (Intl)help.aliyun.com model-pricing
Qwen3.8-Max (Beijing)CNY 12CNY 1.5 implicitCNY 36CNYhelp.aliyun.com/qwen3-8-max model info
GLM-5.2 (Mistral-hosted)$1.40 (€1.19)$0.26 (€0.22)$4.40 (€3.74)USD / EURdocs.mistral.ai/models/zai-glm-5-2
Llama 4 Maverickno first-party API tierweights: 404B, Llama 4 Community License

Notes: Qwen has no official USD list price at publish — the common “$2/$6” is a conversion (≈ $2.11/$6.33 at prevailing rates, labeled an estimate). Qwen’s International cache-hit rate is not published on a fetched official page; Beijing implicit-cache CNY 1.5 is official. GLM-5.2’s Mistral price covers the global + EU regional endpoints; the US regional endpoint is not live yet (official docs). Llama 4 Maverick’s official HF metadata: 404B total params, 1M context, Llama 4 Community License.

Why aggregators disagree (evidence, not facts)

  • Cache blending. Official spread between cache-miss input and cache-hit is 50x at DeepSeek ($0.14 vs $0.0028) and 10x at Kimi ($3.00 vs $0.30). Trackers blend these differently, so “price per token” differs by an order of magnitude across sites using the same official card.
  • Effort labels. “Flash-Max” and “Max Effort” are evaluation configurations, not models — the source of most duplicate rows (full mechanics in the disambiguation).
  • Currency. Qwen’s card is CNY-only and GLM’s Mistral card quotes EUR alongside USD; roundups that convert at different rates and round differently produce different “identical” tables.
  • Checkpoint drift. DeepSeek’s price history moved the flat line to $0.14/$0.28, and the company has announced a “significant” price increase expected in the near future (official pricing page, fetched today) — any table not dated against the changelog quotes a ghost.

Reconciled deltas (derived from official rows above)

  • Flash vs Pro: 3.1x input / 3.1x output — two tiers of one family, identical cache structure.
  • Kimi K3 vs Flash-0731: 21.4x input / 53.6x output at list — the largest official gap in the table.
  • GLM-5.2 (Mistral) vs Flash-0731: 10x input / 15.7x output.
  • Kimi K3 vs GLM-5.2 (Mistral): 2.14x input, 3.4x output, both official USD rows.
  • Qwen3.8-Max vs Flash-0731: ≈15x input after converting Qwen’s CNY row at prevailing rates — the conversion step is labeled an estimate.
  • Qwen3.8-Max vs GLM-5.2 (Mistral): ≈1.5x input, same conversion caveat.

Every delta above derives from two official rows; none is quoted from an aggregator. That is the rule this page exists to enforce — and the same standard behind the corrections log when a figure fails it.

Terms that will change these numbers

  • DeepSeek’s announced price increase — the current card is the low-water mark.
  • Qwen publishing a USD rate card — until then, CNY is the only official statement.
  • Mistral’s US regional endpoint for GLM-5.2 — Euro-priced tier may diverge.
  • Llama 4 successor — Meta’s cadence makes the “no API tier” row time-limited.

Official source

First published August 12, 2026. Every row traceable to the listed official page fetched this session; figures without an official page at publish time are omitted or labeled provisional per ops/CORRECTIONS.md.

#Models #Pricing #DeepSeek #Qwen #Kimi #Open Source