Live
Benchmarks

FrontierMath Tier 4

Models scored
72
Best score
0.48
Metric
mean_score

FrontierMath Tier 4 is an AI evaluation tracked by Epoch AI, with 72 scored model versions on record. Epoch AI runs this evaluation directly, so the scores are reproducible against their published logs.

Scored models span June 2024 to May 2026. The best score recorded is 0.48, measured as mean_score. Models come from Anthropic, Alibaba, Google, Z.ai (Zhipu AI), Google DeepMind, Moonshot, OpenAI, Meta AI and others.

Full record
Score metric
mean_score
Models scored
72
Best score recorded
0.48
Earliest model scored
Jun 20, 2024
Most recent model scored
May 28, 2026
Third-party reported
No
Leaderboard
01gdm-ai-co-mathematicianGoogle DeepMind0.48
02gpt-5.5-pro-pre-release_xhighOpenAI0.4
03gpt-5.5-pro-pre-release_highOpenAI0.4
04gpt-5.4-pro-2026-03-05-web-appOpenAI0.38
05gpt-5.5-pre-release_xhighOpenAI0.35
06gpt-5.2-pro-2025-12-11-webappOpenAI0.31
07claude-opus-4-8_maxAnthropic0.31
08gpt-5.4-2026-03-05_xhighOpenAI0.27
09claude-opus-4-7_xhighAnthropic0.23
10claude-opus-4-6_maxAnthropic0.23
11claude-opus-4-6_32KAnthropic0.21
12claude-opus-4-6_64KAnthropic0.21
13gpt-5.2-2025-12-11_xhighOpenAI0.19
14gpt-5.2-2025-12-11_highOpenAI0.19
15gemini-3-pro-previewGoogle DeepMind0.19
16gemini-3.1-pro-previewGoogle DeepMind0.17
17gpt-5.2-2025-12-11_mediumOpenAI0.17
18muse-sparkMeta AI0.15
19gpt-5-pro-2025-10-06_highOpenAI0.15
20gemini-3.5-flash_highGoogle0.15
21claude-opus-4-6Anthropic0.15
22kimi-k2.6Moonshot0.15
23glm-5.1Z.ai (Zhipu AI)0.13
24gpt-5-2025-08-07_highOpenAI0.13
25gpt-5.1-2025-11-13_highOpenAI0.13
26gemini-2.5-deep-think-2025-08-01-webappGoogle,Google DeepMind0.1
27qwen3.6-plusAlibaba0.08
28claude-sonnet-4-6_16KAnthropic0.08
29gpt-5-mini-2025-08-07_highOpenAI0.06
30gpt-5.2-2025-12-11_lowOpenAI0.06
31o4-mini-2025-04-16_highOpenAI0.06
32gpt-5-2025-08-07_mediumOpenAI0.06
33gpt-5.4-nano-2026-03-17_highOpenAI0.06
34fireworks/kimi-k2p5Moonshot0.04
35claude-opus-4-5-20251101_32KAnthropic0.04
36o3-mini-2025-01-31_highOpenAI0.04
37claude-opus-4-20250514_27KAnthropic0.04
38gemini-2.5-flashGoogle DeepMind0.04
39gemini-2.5-proGoogle DeepMind0.04
40claude-opus-4-1-20250805_27KAnthropic0.04
41gpt-5-mini-2025-08-07_mediumOpenAI0.04
42claude-sonnet-4-5-20250929_32KAnthropic0.04
43gpt-5.1-2025-11-13_mediumOpenAI0.04
44claude-opus-4-5-20251101Anthropic0.04
45gemini-3-flash-previewGoogle DeepMind0.04
46qwen3.6-max-previewAlibaba0.04
47zai-org/GLM-4.6Z.ai (Zhipu AI),Tsinghua University0.02
48glm-5Z.ai (Zhipu AI)0.02
49fireworks/deepseek-v3p2DeepSeek0.02
50claude-haiku-4-5-20251001_32KAnthropic0.02
51claude-opus-4-5-20251101_16KAnthropic0.02
52o3-2025-04-16_highOpenAI0.02
53o4-mini-2025-04-16_mediumOpenAI0.02
54claude-sonnet-4-5-20250929Anthropic0.02
55gpt-5-nano-2025-08-07_mediumOpenAI0.02
56qwen3.5-plusAlibaba0.02
57grok-4-0709xAI0.02
58gemini-2.5-pro-preview-06-05Google DeepMind0.02
59gpt-5.4-mini-2026-03-17_highOpenAI0.02
60grok-4-heavy-web-appxAI0.02
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks