DeepSeek's open-weights shockwave, one year on
DeepSeek's R1 landed in January 2025 and permanently reset what frontier AI costs to build and buy. A year-plus on, the price war, the efficiency pivot, and the open-weights wave it unleashed are still the industry's defining forces.
On January 20, 2025, a relatively unknown lab in Hangzhou released a reasoning model under an MIT license, with weights anyone could download, and pricing roughly 27 times cheaper than the closest competitor. Within a week, Nvidia had suffered the largest single-day market-cap loss in stock-market history — roughly $590 billion on January 27 — and the DeepSeek app had displaced ChatGPT at the top of the App Store's free charts. A year-plus on, the dust has long settled, and what's left behind is not a single upset but a permanently rewired industry. Here's what actually changed, and what was always overstated.
The release that broke the pricing umbrella#
DeepSeek-R1 was a 671-billion-parameter mixture-of-experts model that activated only 37 billion parameters per token. DeepSeek's technical report (arXiv 2501.12948) showed it matching OpenAI's o1 on the reasoning benchmarks that mattered:
| Benchmark | DeepSeek-R1 | OpenAI o1 |
|---|---|---|
| AIME 2024 (pass@1) | 79.8% | 79.2% |
| MATH-500 | 97.3% | 96.4% |
| Codeforces (percentile) | 96.3 | 96.6 |
| SWE-bench Verified | 49.2% | 48.9% |
| GPQA Diamond | 71.5% | 78.0% |
The numbers were essentially a tie — and that was the whole point. R1 wasn't dramatically better than o1. It was as good while costing $0.55 per million input tokens and $2.19 per million output tokens, versus $15 and $60 for o1. Benchmark parity at 27x the discount is what detonated the industry's pricing structure.
The training-cost claim deserves its caveat. DeepSeek reported roughly $5.58 million for the final training run, and critics — including analysts at Janus Henderson — rightly noted that figure covers one run only, excluding prior experiments, the cost of the teacher models, and the data pipelines beneath it. The headline number was apples-to-oranges next to Western lab budgets. But even multiplying it several times over, R1 was orders of magnitude cheaper than anything comparable, and the market reacted to what that implied: that the compute moat protecting incumbents was far thinner than their capex narratives suggested.
The price war never really stopped#
R1 didn't just undercut rivals; it gave every buyer in the market a credible anchor price. Within roughly 90 days, OpenAI, Google, and Anthropic had all cut flagship API pricing, with Google's Gemini cuts reported at up to 85%. Inference economics — the single biggest barrier to building AI-native products — were rewritten in a quarter.
The dynamic kept compounding through 2025 and 2026:
- September 2025: DeepSeek slashed V3.2 API pricing more than 50% overnight, pulling input costs down to fractions of a cent per million tokens and dragging Chinese rivals like ByteDance and Tencent down with it.
- Mid-2026: DeepSeek shipped the V4 generation and made a 75% price cut permanent, arguing — plausibly — that it was passing through genuine efficiency gains in long-context inference, not subsidizing losses.
- July 2026: OpenAI responded by cutting GPT-5.6 Luna pricing 80% to $0.20/$1.20 per million tokens, explicitly undercutting DeepSeek's V4 Pro on inputs. OpenAI attributed the cut to infrastructure efficiency gains — speculative decoding, better GPU utilization — which is itself a telling admission: the efficiency playbook DeepSeek forced on the industry was now everyone's playbook.
By late 2026, per-token prices had fallen 60–80% across the board versus the GPT-4 era. The most honest signal of how permanent the reset was: when DeepSeek introduced surge pricing on its API in August 2026 — doubling rates during Beijing business hours — nobody questioned the base rate anymore. The price war's victor was allowed to act like a utility.
Open weights went from heresy to strategy#
The deepest structural change was philosophical. Before R1, the industry's working assumption was that the best models were closed because openness surrendered the moat. R1 proved the inverse: openness was the moat. The MIT license let anyone run, fine-tune, and commercialize the weights, and the ecosystem did the rest.
DeepSeek simultaneously released distilled variants built on Qwen 2.5 and Llama 3.1 bases, in sizes from 1.5B to 70B, that put frontier-adjacent reasoning onto commodity hardware and laptops. Enterprise adoption exploded precisely because self-hosting solved the two objections that kept regulated industries off the API — data sovereignty and vendor lock-in. Startups integrated R1 distills into agents and copilots at near-zero marginal model cost, shifting competition upward from "who has the smartest model" to "who builds the most useful integrated application."
The copycats validated the thesis. Meta pushed the open-weights frontier with Llama releases, Alibaba's Qwen family iterated at a furious cadence, and Mistral kept the European flag flying — all now operating in a world where a capable open model is the expected baseline rather than a surprise. Yann LeCun's framing of R1 as a victory for open-source innovation holds up: DeepSeek itself built on Llama and Qwen bases, making the shock a demonstration of what the open ecosystem could compound.
The backlash was real, too#
R1's aftereffects weren't all disruption-positive, and a clear-eyed retrospective has to count them:
- Geopolitics and scrutiny. A Chinese lab matching Western frontier models under US chip-export restrictions drew exactly the response you'd expect: harder scrutiny of weight proliferation, debates over export-control efficacy, and enterprise policies splitting between "self-host it and it's fine" and "keep it away from sensitive data."
- Distillation ethics and disputes. DeepSeek's training relied on distillation techniques — including, reportedly, from models whose terms of service prohibit exactly that — sparking disputes with OpenAI over terms-of-use violations. The industry still hasn't settled what fair-use distillation looks like.
- The $5.6 million myth. The training-cost headline was weaponized far beyond what the paper claimed, and it warped investment discourse for months. Efficiency is real; "frontier models cost lunch money" was never the claim.
- Market myopia. The Nvidia crash recovered quickly — Jevons paradox in action: cheaper AI meant vastly more AI usage, which meant more GPUs, not fewer. The panic priced in a demand collapse that never came.
Takeaway: the era of "bigger is better" is over#
What changed permanently: (1) frontier-quality inference is a commodity, with per-token prices 60–80% below the GPT-4 era and no floor in sight; (2) open weights are a competitive strategy, not a concession, with the ecosystem's compounding power now an accepted industry force; (3) efficiency — MoE sparsity, RL-based reasoning, distillation, speculative decoding — is the primary axis of competition, displacing raw scale; and (4) labs no longer get to price on mystique; every buyer has a cheap, credible reference point.
What didn't change: capital still matters. The labs that survived the price war are the ones with the infrastructure to serve tokens at those margins. DeepSeek didn't end the AI spending spree — it redirected it from training budgets into serving infrastructure and applications. Twenty months on, the shockwave looks less like a revolution and more like a market correction that arrived all at once: intelligence turned out to be a function of strategy as much as scale, and there's no going back to pricing that pretends otherwise.