DeepSeek-R1 is an AI model developed by DeepSeek (China), first published in January 2025. It works in the language domain, on tasks such as language modeling/generation, code generation, quantitative reasoning and question answering.
Training it took an estimated 3.5×10²⁴ FLOP of compute (estimation method: operation counting). The model has 671,000,000,000 parameters. It was trained on roughly 14.8T datapoints. The compute alone is estimated at $7M in 2023 dollars.
Access: Open weights (unrestricted). Its weights are openly available. It is built on top of DeepSeek-V3. Epoch AI rates the confidence of this record as confident.
| 01 | Epoch Capabilities Index | ECI Score | 139.7 |
| 02 | LiveBench | Global average | 71.6 |
| 03 | ForecastBench | Overall score | 60 |
| 04 | Aider Polyglot | Percent correct | 56.9 |
| 05 | Lech Mazur Writing | Mean score | 8.3 |
| 06 | Algotune | Score | 1.7 |
| 07 | MATH Level 5 | mean_score | 0.93 |
| 08 | GPQA Diamond | mean_score | 0.69 |
| 09 | OTIS Mock AIME 2024-2025 | mean_score | 0.53 |
| 10 | METR Time Horizons | average_score | 0.52 |
| 11 | WeirdML | Accuracy | 0.36 |
| 12 | SciCode | Score | 0.36 |
| 13 | Balrog | Average progress | 0.35 |
| 14 | Fictionlivebench | 120k token score | 0.33 |
| 15 | SimpleBench | Score (AVG@5) | 0.31 |
| 16 | ARC-AGI | Score | 0.16 |
| 17 | ARC-AGI-2 | Score | 0.01 |
| 18 | Critpt | Accuracy | 0.01 |