OTIS Mock AIME 2024-2025 is an AI evaluation tracked by Epoch AI, with 173 scored model versions on record. Epoch AI runs this evaluation directly, so the scores are reproducible against their published logs.
Scored models span March 2023 to July 2026. The best score recorded is 1, measured as mean_score. Models come from Anthropic, OpenAI, Moonshot, DeepSeek, Google, xAI, Z.ai (Zhipu AI), Alibaba and others.
| 01 | gpt-5.5-pro-pre-release_xhigh | OpenAI | 1 |
| 02 | gpt-5.6-sol_max | OpenAI | 1 |
| 03 | gpt-5.5-pre-release_xhigh | OpenAI | 1 |
| 04 | gpt-5.6-terra_max | OpenAI | 1 |
| 05 | claude-fable-5_max | Anthropic | 1 |
| 06 | claude-opus-5_max | Anthropic | 0.99 |
| 07 | claude-opus-4-8_max | Anthropic | 0.98 |
| 08 | gpt-5.6-luna_max | OpenAI | 0.98 |
| 09 | claude-opus-4-7_xhigh | Anthropic | 0.98 |
| 10 | grok-4.5_high | xAI | 0.98 |
| 11 | gpt-5.4-2026-03-05_high | OpenAI | 0.98 |
| 12 | kimi-k3_max | Moonshot | 0.97 |
| 13 | deepseek-v4-pro_max | DeepSeek | 0.97 |
| 14 | kimi-k2.7-code | Moonshot | 0.96 |
| 15 | gpt-5.2-2025-12-11_high | OpenAI | 0.96 |
| 16 | kimi-k2.6 | Moonshot | 0.96 |
| 17 | gpt-5.2-2025-12-11_xhigh | OpenAI | 0.96 |
| 18 | gemini-3.1-pro-preview | Google DeepMind | 0.96 |
| 19 | gemini-3.5-flash_high | 0.96 | |
| 20 | gpt-5.4-2026-03-05_medium | OpenAI | 0.96 |
| 21 | gpt-5.4-2026-03-05_xhigh | OpenAI | 0.95 |
| 22 | qwen3.7-max | Alibaba | 0.95 |
| 23 | claude-sonnet-5_xhigh | Anthropic | 0.95 |
| 24 | claude-opus-4-6_64K | Anthropic | 0.94 |
| 25 | gpt-5.2-2025-12-11_medium | OpenAI | 0.94 |
| 26 | grok-4.3_high | xAI | 0.93 |
| 27 | claude-opus-4-6_32K | Anthropic | 0.93 |
| 28 | gemini-3-flash-preview | Google DeepMind | 0.93 |
| 29 | glm-5.1 | Z.ai (Zhipu AI) | 0.92 |
| 30 | grok-4.20-0309-reasoning | xAI | 0.92 |
| 31 | fireworks/kimi-k2p5 | Moonshot | 0.92 |
| 32 | gemini-3-pro-preview | Google DeepMind | 0.91 |
| 33 | gpt-5-2025-08-07_high | OpenAI | 0.91 |
| 34 | qwen3.6-max-preview | Alibaba | 0.91 |
| 35 | qwen3.6-plus | Alibaba | 0.91 |
| 36 | muse-spark | Meta AI | 0.89 |
| 37 | openai/gpt-oss-120b_high | OpenAI | 0.89 |
| 38 | gpt-5.1-2025-11-13_high | OpenAI | 0.89 |
| 39 | deepseek-reasoner | DeepSeek | 0.88 |
| 40 | gpt-5.4-nano-2026-03-17_high | OpenAI | 0.88 |
| 41 | gpt-5-2025-08-07_medium | OpenAI | 0.87 |
| 42 | gpt-5.4-mini-2026-03-17_high | OpenAI | 0.87 |
| 43 | Qwen3-235B-A22B-Thinking-2507 | Alibaba | 0.87 |
| 44 | gpt-5-mini-2025-08-07_high | OpenAI | 0.87 |
| 45 | glm-5.2_max | Z.ai (Zhipu AI) | 0.86 |
| 46 | claude-opus-4-5-20251101_32K | Anthropic | 0.86 |
| 47 | qwen3.6-flash | Alibaba | 0.86 |
| 48 | claude-sonnet-4-6_32K | Anthropic | 0.86 |
| 49 | qwen3.5-flash | Alibaba | 0.86 |
| 50 | gpt-5.1-2025-11-13_medium | OpenAI | 0.86 |
| 51 | qwen3.5-plus | Alibaba | 0.85 |
| 52 | gpt-5.4-2026-03-05_low | OpenAI | 0.84 |
| 53 | gemini-2.5-pro | Google DeepMind | 0.84 |
| 54 | grok-4-0709 | xAI | 0.84 |
| 55 | o3-2025-04-16_high | OpenAI | 0.84 |
| 56 | glm-4.7 | Z.ai (Zhipu AI) | 0.83 |
| 57 | kimi-k2-thinking-turbo | Moonshot | 0.83 |
| 58 | claude-sonnet-4-6_medium | Anthropic | 0.82 |
| 59 | claude-opus-4-5-20251101_16K | Anthropic | 0.82 |
| 60 | o4-mini-2025-04-16_high | OpenAI | 0.82 |