En direct
Benchmarks

OTIS Mock AIME 2024-2025

Models scored
173
Best score
1
Metric
mean_score

OTIS Mock AIME 2024-2025 is an AI evaluation tracked by Epoch AI, with 173 scored model versions on record. Epoch AI runs this evaluation directly, so the scores are reproducible against their published logs.

Scored models span March 2023 to July 2026. The best score recorded is 1, measured as mean_score. Models come from Anthropic, OpenAI, Moonshot, DeepSeek, Google, xAI, Z.ai (Zhipu AI), Alibaba and others.

Full record
Score metric
mean_score
Models scored
173
Best score recorded
1
Earliest model scored
Mar 14, 2023
Most recent model scored
Jul 24, 2026
Third-party reported
No
Leaderboard
01gpt-5.5-pro-pre-release_xhighOpenAI1
02gpt-5.6-sol_maxOpenAI1
03gpt-5.5-pre-release_xhighOpenAI1
04gpt-5.6-terra_maxOpenAI1
05claude-fable-5_maxAnthropic1
06claude-opus-5_maxAnthropic0.99
07claude-opus-4-8_maxAnthropic0.98
08gpt-5.6-luna_maxOpenAI0.98
09claude-opus-4-7_xhighAnthropic0.98
10grok-4.5_highxAI0.98
11gpt-5.4-2026-03-05_highOpenAI0.98
12kimi-k3_maxMoonshot0.97
13deepseek-v4-pro_maxDeepSeek0.97
14kimi-k2.7-codeMoonshot0.96
15gpt-5.2-2025-12-11_highOpenAI0.96
16kimi-k2.6Moonshot0.96
17gpt-5.2-2025-12-11_xhighOpenAI0.96
18gemini-3.1-pro-previewGoogle DeepMind0.96
19gemini-3.5-flash_highGoogle0.96
20gpt-5.4-2026-03-05_mediumOpenAI0.96
21gpt-5.4-2026-03-05_xhighOpenAI0.95
22qwen3.7-maxAlibaba0.95
23claude-sonnet-5_xhighAnthropic0.95
24claude-opus-4-6_64KAnthropic0.94
25gpt-5.2-2025-12-11_mediumOpenAI0.94
26grok-4.3_highxAI0.93
27claude-opus-4-6_32KAnthropic0.93
28gemini-3-flash-previewGoogle DeepMind0.93
29glm-5.1Z.ai (Zhipu AI)0.92
30grok-4.20-0309-reasoningxAI0.92
31fireworks/kimi-k2p5Moonshot0.92
32gemini-3-pro-previewGoogle DeepMind0.91
33gpt-5-2025-08-07_highOpenAI0.91
34qwen3.6-max-previewAlibaba0.91
35qwen3.6-plusAlibaba0.91
36muse-sparkMeta AI0.89
37openai/gpt-oss-120b_highOpenAI0.89
38gpt-5.1-2025-11-13_highOpenAI0.89
39deepseek-reasonerDeepSeek0.88
40gpt-5.4-nano-2026-03-17_highOpenAI0.88
41gpt-5-2025-08-07_mediumOpenAI0.87
42gpt-5.4-mini-2026-03-17_highOpenAI0.87
43Qwen3-235B-A22B-Thinking-2507Alibaba0.87
44gpt-5-mini-2025-08-07_highOpenAI0.87
45glm-5.2_maxZ.ai (Zhipu AI)0.86
46claude-opus-4-5-20251101_32KAnthropic0.86
47qwen3.6-flashAlibaba0.86
48claude-sonnet-4-6_32KAnthropic0.86
49qwen3.5-flashAlibaba0.86
50gpt-5.1-2025-11-13_mediumOpenAI0.86
51qwen3.5-plusAlibaba0.85
52gpt-5.4-2026-03-05_lowOpenAI0.84
53gemini-2.5-proGoogle DeepMind0.84
54grok-4-0709xAI0.84
55o3-2025-04-16_highOpenAI0.84
56glm-4.7Z.ai (Zhipu AI)0.83
57kimi-k2-thinking-turboMoonshot0.83
58claude-sonnet-4-6_mediumAnthropic0.82
59claude-opus-4-5-20251101_16KAnthropic0.82
60o4-mini-2025-04-16_highOpenAI0.82
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks