En vivo
Benchmarks

MATH Level 5

Models scored
108
Best score
0.98
Metric
mean_score

MATH Level 5 is an AI evaluation tracked by Epoch AI, with 108 scored model versions on record. Epoch AI runs this evaluation directly, so the scores are reproducible against their published logs.

Scored models span June 2023 to October 2025. The best score recorded is 0.98, measured as mean_score. Models come from OpenAI, Anthropic, Alibaba, DeepSeek, Google DeepMind, Mistral AI, xAI, Meta AI and others.

Full record
Score metric
mean_score
Models scored
108
Best score recorded
0.98
Earliest model scored
Jun 13, 2023
Most recent model scored
Oct 15, 2025
Third-party reported
No
Leaderboard
01gpt-5-2025-08-07_highOpenAI0.98
02gpt-5-2025-08-07_mediumOpenAI0.98
03gpt-5-mini-2025-08-07_highOpenAI0.98
04o4-mini-2025-04-16_highOpenAI0.98
05o3-2025-04-16_highOpenAI0.98
06claude-sonnet-4-5-20250929_32KAnthropic0.98
07qwen3-max-2025-09-23Alibaba0.97
08gpt-5-mini-2025-08-07_mediumOpenAI0.97
09DeepSeek-R1-0528DeepSeek0.97
10o3-mini-2025-01-31_highOpenAI0.96
11claude-haiku-4-5-20251001_32KAnthropic0.96
12gemini-2.5-pro-preview-05-06Google DeepMind0.96
13gemini-2.5-pro-preview-03-25Google DeepMind0.96
14gpt-5-nano-2025-08-07_mediumOpenAI0.95
15o3-mini-2025-01-31_mediumOpenAI0.95
16gpt-5-nano-2025-08-07_highOpenAI0.95
17o1-2024-12-17_highOpenAI0.95
18o1-2024-12-17_mediumOpenAI0.94
19DeepSeek-R1DeepSeek0.93
20claude-3-7-sonnet-20250219_64KAnthropic0.91
21grok-3-mini-beta_lowxAI0.91
22claude-3-7-sonnet-20250219_32KAnthropic0.9
23DeepSeek-R1-Distill-Llama-70BDeepSeek0.9
24o1-mini-2024-09-12_highOpenAI0.89
25grok-3-betaxAI0.89
26grok-3-mini-beta_highxAI0.88
27gpt-4.1-mini-2025-04-14OpenAI0.87
28DeepSeek-R1-Distill-Qwen-14BDeepSeek0.87
29claude-haiku-4-5-20251001Anthropic0.87
30claude-3-7-sonnet-20250219_16KAnthropic0.86
31claude-opus-4-20250514Anthropic0.85
32claude-sonnet-4-20250514Anthropic0.84
33o1-mini-2024-09-12_mediumOpenAI0.84
34gemini-2.0-pro-exp-02-05Google DeepMind0.83
35gpt-4.1-2025-04-14OpenAI0.83
36gemini-2.0-flash-001Google DeepMind,Google0.82
37o1-preview-2024-09-12OpenAI0.82
38mistral-medium-2505Mistral AI0.82
39gpt-4.5-preview-2025-02-27OpenAI0.79
40DeepSeek-V3-0324DeepSeek0.76
41gemma-3-27b-itGoogle DeepMind0.74
42Llama-4-Maverick-17B-128E-Instruct-FP8Meta AI0.73
43gemini-1.5-pro-002Google DeepMind0.7
44gpt-4.1-nano-2025-04-14OpenAI0.7
45qwen3-235b-a22bAlibaba0.69
46claude-3-7-sonnet-20250219Anthropic0.68
47qwen-max-2025-01-25Alibaba0.67
48qwen-plus-2025-01-25Alibaba0.65
49phi-4Microsoft Research0.65
50DeepSeek-V3DeepSeek0.65
51grok-2-1212xAI0.64
52qwen2.5-72b-instructAlibaba0.63
53Llama-4-Scout-17B-16E-InstructMeta AI0.62
54gemini-1.5-flash-002Google DeepMind0.62
55claude-3-5-sonnet-20241022Anthropic0.57
56qwen-turbo-2024-11-01Alibaba0.56
57qwen2.5-32b-instructAlibaba0.56
58gpt-4o-2024-08-06OpenAI0.53
59gpt-4o-mini-2024-07-18OpenAI0.53
60claude-3-5-sonnet-20240620Anthropic0.52
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks