Skip to content
Newsroom
Industry Jul 30, 2026 by Rajat Jain

OpenAI Slashes GPT-5.6 API Prices: Luna Down 80%

The cheapest frontier-tier price OpenAI has ever offered, plus a Sol Fast tier that doubles throughput for a 2x premium — announced by Sam Altman on X.

OpenAI Slashes GPT-5.6 API Prices: Luna Down 80%

OpenAI announced major price cuts across the GPT-5.6 family on July 30, announced directly on X by Sam Altman and detailed on the OpenAI pricing page: the Luna small model drops 80% to $0.20 per million input tokens, the strongest price the company has ever quoted for a frontier-tier model, and Terra drops 20% to $2.00.

Key facts

  • Luna (small): −80%, from $1.00 to $0.20/M input ($1.20/M output). Designed for high-volume agentic workloads.
  • Terra (mid): −20% to $2.00/M input, $12.00/M output (previously $2.50 / $15).
  • Sol (frontier): unchanged at $5.00/M input, $30.00/M output.
  • New Sol Fast mode: 2x the price of Sol Standard ($10 / $60) for roughly 2.5x end-to-end throughput — aimed at latency-sensitive applications.
  • Announcement channel: Sam Altman posted the cuts directly (“major price cuts today”) as pricing page updated the same day. Right to an official pricing page, no press release needed.

What changed and why

The cuts are marketed as a cost-engineering story: OpenAI says improvements to its inference stack have brought per-token serving costs down faster than model capability is growing, so the savings flow downstream. The Luna cut — an 80% drop — is the loudest signal.

The strategic context is the quarter’s running price war:

  • From below: open-weight competitors — Moonshot’s Kimi K3, Alibaba’s Qwen3.8-Max, and DeepSeek’s V4-Flash — are pricing open-weights models aggressively, pushing the effective cost floor of frontier-grade output down in weeks, not quarters.
  • From the side: Anthropic and Google have been cutting rates on mid-tier Claude and Gemini models, squeezing the same mid-market quadrant that Terra occupies.

Luna at $0.20/M is the direct answer to the developer math: agent loops that read a lot of tokens are where the volume lives, so making the small model cheaper is a bigger total-cost lever than any mid-tier discount.

Why it matters for developers

  • Agent economics improve immediately. For agent frameworks that push tens of millions of input tokens daily, Luna at −80% is a step-change in gross margin — expect a new wave of “agent at scale” applications.
  • The new Sol Fast option is effectively a premium for latency: for synchronous, user-facing tools (chat, copilots, coding assistants), the 2.5x throughput at 2x price is priced for product teams, not batch work.
  • Pressure on the rest of the market. With Sol unchanged, a developer comparing against a Tier-2 frontier model has to justify performance via output quality, not price. Expect Anthropic and Google to respond with their own cuts within weeks (they already did — see our Qwen and Kimi briefings).
  • Open-weight parity check: a 2.4T open model (Qwen3.8-Max) at ~$2/M input, vs Luna at $0.20/M input — the spread defines the current arbitrage for serious builders.

What to watch (as of August 8)

  • Follow-on cuts: whether Anthropic and Google match or beat Terra/Luna pricing in the next earnings-adjacent window.
  • Sol Fast adoption: throughput premium is only worth it if users actually notice the difference — check developer chatter for which workloads actually buy the 2x tier.
  • Margin year: how the price war shows up in cloud-inference margins at the next earnings call.

Official source

Updated August 8, 2026.

#Industry #Pricing #APIs