The AI price war went nuclear this week: DeepSeek V4 Flash at $0.14/M tokens and GPT-5.6 Luna down 80% — in the same 24 hours
If you blinked this week you missed the cheapest frontier-class AI has ever been. On July 30, OpenAI cut GPT-5.6 Luna prices by 80%. Less than a day later, DeepSeek's V4 Flash graduated from preview to an official public beta — the '0731' build — at prices starting from $0.14 per million input tokens. We maintain a structured catalog of 45 AI providers with real prices, free tiers, and rate limits, and weeks like this are exactly why. This article is not a rewrite of press releases. It is what changed in numbers, where each number came from, how we re-verified it against our catalog, and what it means if you actually pay for tokens. Prices below are per million tokens (input / output) unless noted.
What OpenAI announced on July 30 — with provenance
GPT-5.6 Luna dropped from $1.00 to $0.20 per million input tokens and from $6.00 to $1.20 per million output tokens — an 80% cut on both sides, effective immediately. GPT-5.6 Terra got a smaller 20% cut, and the flagship Sol kept its pricing but gained a new Fast mode. That is the OpenAI announcement itself (July 30, 2026). Per reporting around the announcement, Sol autonomously rewrote parts of its own production inference stack — GPU kernels and speculative-decoding paths — and the efficiency gains funded the price drop. We flag that last sentence as reporting, not as a provider price fact, because it is a narrative about how the cut became possible rather than a number you can invoice against. What is invoiceable is the table in our provider directory: the Luna entry was updated July 30 to $0.20 / $1.20, with the long-context surcharge preserved. If you push long contexts, prompts over ~272K input tokens bill at 2x input and 1.5x output, and cache writes run 1.25x the uncached rate. Third-party trackers listed the same new Luna price within hours, which is our cross-check that we are not looking at a cached page.
What DeepSeek shipped on July 31 — and what 'from $0.14' means
DeepSeek V4 Flash is now an official public-beta release. The architecture is unchanged from the preview — this is DeepSeek's efficiency model, not a new flagship — but the pricing is the point. DeepSeek's own docs for the 0731 build list the entry price as $0.14 input / $0.28 output per million tokens, and third-party trackers we cross-checked July 31 show the same floor. That undercuts even the newly-cut Luna on the input side by $0.06/M. One thing to watch, because it will bite you if you miss it: DeepSeek has announced a peak-hours pricing policy that is not yet active as of July 31. The docs describe higher peak pricing in a future window, but no active surcharge is in effect today. If you build on V4 Flash, build with that in mind — pin a price alert or keep a fallback model in your router. We've added V4 Flash to our DeepSeek provider entry today, and the free-tier hub notes where DeepSeek's free lane sits, because at $0.14/M the free tier versus paid math is closer than ever.
| Date | Provider | Move | New price |
|---|---|---|---|
| Jul 9–10 | Meta | Muse Spark 1.1 — Meta's first paid model API, then the cheapest paid frontier | $1.25 / $4.25 |
| Jul 12 | OpenAI | GPT-5.6 Sol usage limits temporarily relaxed | — |
| Jul 30 | OpenAI | GPT-5.6 Luna cut 80%; Terra cut 20%; Sol gains Fast mode | $0.20 / $1.20 (Luna) |
| Jul 31 | DeepSeek | V4 Flash '0731' official public beta | from $0.14 / $0.28 |
The math: what 80% actually saves you
Percentages sound dramatic; token bills are arithmetic. Take a real workload. A small app that processes 5 million input tokens and 2 million output tokens per day paid $17/day for Luna at the old $1/$6. At the new $0.20/$1.20 it pays $3.40/day. That is $13.60/day left on the table, or about $408/month. Even a light personal workload of 1M input / 0.5M output per month goes from $4/month to $0.80/month. The table below is not a forecast — it is today's catalog price times your volume. If you pinned Luna at $1/$6 in a budget spreadsheet, your projections are now wrong by 5x in your favor. Re-run them, and re-run any per-request cost you baked into pricing for your own users.
| Monthly volume (in / out) | Luna at old $1 / $6 | Luna at new $0.20 / $1.20 | DeepSeek V4 Flash at $0.14 / $0.28 |
|---|---|---|---|
| 1M / 0.5M (heavy personal chat) | $4.00 | $0.80 | $0.28 |
| 5M / 2M (small app, ~170K in / 67K out per day) | $17.00 | $3.40 | $1.26 |
| 50M / 20M (growing product) | $170.00 | $34.00 | $12.60 |
| 500M / 200M (scale) | $1,700.00 | $340.00 | $126.00 |
How we track this so you don't have to trust a blog post
We keep the provider directory as data, not as a blog post someone wrote once. Each of the 45 entries carries a verified_at timestamp, source_url, and structured fields for free-tier terms, rate limits, and subscription plans. For this story we touched two entries. OpenAI: source_url is the provider pricing page, verified_at set to 2026-07-30, payg_rates updated to Luna $0.20/$1.20 and Terra's 20% cut, rate-limit note preserved for the 272K-token long-context surcharge. DeepSeek: added V4 Flash to models_available, verified_at 2026-07-31, free_tier note unchanged, rate_limits note updated to mention the not-yet-active peak policy. Third-party trackers are our second pair of eyes, not our source of truth — when Artificial Analysis and our catalog agree within hours, we publish. When they disagree, we re-fetch the provider page and hold the post. That discipline is the whole point of the free-tier hub and the plans table: the article is a snapshot, the catalog is live. Methodology limits: provider pages can cache, exchange rates can confuse, and 'from $0.14' in DeepSeek's docs is a floor tied to cache-hit and context specifics — your effective price will be a little higher whenever you miss the cache. We note that explicitly here so you do not budget on the floor alone.
What to do this week if you pay for tokens
- Re-run your unit economics. Any price you derived from Luna $1/$6 is stale. Update the per-request cost in your pricing page, your pricing model docs, and your customer-facing estimates before you undercharge yourself for another month.
- Don't re-platform on one day's floor. V4 Flash at $0.14 is the cheapest input price we have ever recorded, but 'from $0.14' plus a dormant peak policy means the number has an asterisk. Use Luna at $0.20 as your planning anchor and treat V4 Flash as upside.
- Keep two keys warm. This is the exact week where bring-your-own-key pays off: hold both OpenAI and DeepSeek keys, route chat and summarization to V4 Flash, keep reasoning on Luna or Sol, and switch in config rather than in code. Our own /chat runs entirely in your browser with your keys for exactly this test — put DeepSeek in one column and Luna in the other and compare them on your actual prompts today, not on someone's benchmark.
- Rate limits are the new price. At $0.14/M, a free tier with 1,000 requests/day is worth more than a 10% cheaper token. Check the free-tier comparison before you chase another $0.02 of price.
- Watch the floor, not the flagship. Sol's new Fast mode is a speed story more than a price story; the market-moving numbers this week were on the efficiency models, not the frontier one.
If you'd rather not re-platform every time the market sneezes, this is the case for the tooling shape we run ourselves: keep accounts at multiple providers, and switch models per task rather than per quarter. Our plans table shows what a paid commitment looks like once you outgrow free tiers, and the provider directory carries the live price for every entry on this page. When prices change again — and they will — the directory reflects it first. That is not a slogan; it is how the catalog is built: capture, cross-check, timestamp, publish.