447×
smaller than the teacher
1.5B student vs 671B R1
27×
cheaper output tokens
R1 $2.19 vs o1 $60.00 / M (early-2025)
$589B
wiped off Nvidia in one day
27 Jan 2025 · −17% · biggest ever
The story, in order
Click any point on the timeline for a plain-language recap and its source.
Tiny students, real scores
Model size (billions of parameters, log scale) vs AIME 2024 score
Selected student
Other students
o1-mini reference
R1-Distill-Qwen-1.5B
Distilled from DeepSeek-R1 (671B) · open weights
The price shock that started the panic
API price per million tokens, DeepSeek R1 vs OpenAI o1. Both reasoning models, early-2025 launch prices
OpenAI o1
DeepSeek R1
The same reasoning job cost about 27× less on R1, roughly 96% cheaper. That, more than any single benchmark, is what rattled the markets.
What distillation actually is (and isn’t)
Distillation is compression, not magic. A big, expensive “teacher” model trains a small, cheap “student” to imitate its answers. The student ends up nearly as good at a specific job at a fraction of the size and cost. The idea comes from a 2015 Google paper by Hinton, Vinyals & Dean. Think of a master chef writing down a recipe so a line cook reproduces the dish without 20 years of training.
The bit the headlines blurred
- Classic distillation needs the teacher’s internal “soft” outputs: you need access to the model’s guts.
- The DeepSeek row was mostly about the cheaper cousin: training on another model’s API text outputs (“synthetic data”, or less kindly, “copying homework”). This is what allegedly breached OpenAI’s Terms of Service.
- OpenAI and Microsoft said they found evidence and opened a probe; investor David Sacks cited “substantial evidence.” No public, independently verifiable proof was released, so this stays an allegation.
The $5.6 million myth
The famous $5.6M was real, but it was the electricity-and-rental bill for the final training run only (2.788M H800 GPU-hours × $2/hour = $5.576M). DeepSeek’s paper explicitly excludes prior research, failed experiments, data, salaries, and the chips themselves. Analysts (SemiAnalysis) put DeepSeek’s total server capital expenditure near $1.3 billion, a different thing entirely. It was never the price of building the company.