FrontierMath Tier 4 is an AI evaluation tracked by Epoch AI, with 72 scored model versions on record. Epoch AI runs this evaluation directly, so the scores are reproducible against their published logs.
Scored models span June 2024 to May 2026. The best score recorded is 0.48, measured as mean_score. Models come from Anthropic, Alibaba, Google, Z.ai (Zhipu AI), Google DeepMind, Moonshot, OpenAI, Meta AI and others.
| 01 | gdm-ai-co-mathematician | Google DeepMind | 0.48 |
| 02 | gpt-5.5-pro-pre-release_xhigh | OpenAI | 0.4 |
| 03 | gpt-5.5-pro-pre-release_high | OpenAI | 0.4 |
| 04 | gpt-5.4-pro-2026-03-05-web-app | OpenAI | 0.38 |
| 05 | gpt-5.5-pre-release_xhigh | OpenAI | 0.35 |
| 06 | gpt-5.2-pro-2025-12-11-webapp | OpenAI | 0.31 |
| 07 | claude-opus-4-8_max | Anthropic | 0.31 |
| 08 | gpt-5.4-2026-03-05_xhigh | OpenAI | 0.27 |
| 09 | claude-opus-4-7_xhigh | Anthropic | 0.23 |
| 10 | claude-opus-4-6_max | Anthropic | 0.23 |
| 11 | claude-opus-4-6_32K | Anthropic | 0.21 |
| 12 | claude-opus-4-6_64K | Anthropic | 0.21 |
| 13 | gpt-5.2-2025-12-11_xhigh | OpenAI | 0.19 |
| 14 | gpt-5.2-2025-12-11_high | OpenAI | 0.19 |
| 15 | gemini-3-pro-preview | Google DeepMind | 0.19 |
| 16 | gemini-3.1-pro-preview | Google DeepMind | 0.17 |
| 17 | gpt-5.2-2025-12-11_medium | OpenAI | 0.17 |
| 18 | muse-spark | Meta AI | 0.15 |
| 19 | gpt-5-pro-2025-10-06_high | OpenAI | 0.15 |
| 20 | gemini-3.5-flash_high | 0.15 | |
| 21 | claude-opus-4-6 | Anthropic | 0.15 |
| 22 | kimi-k2.6 | Moonshot | 0.15 |
| 23 | glm-5.1 | Z.ai (Zhipu AI) | 0.13 |
| 24 | gpt-5-2025-08-07_high | OpenAI | 0.13 |
| 25 | gpt-5.1-2025-11-13_high | OpenAI | 0.13 |
| 26 | gemini-2.5-deep-think-2025-08-01-webapp | Google,Google DeepMind | 0.1 |
| 27 | qwen3.6-plus | Alibaba | 0.08 |
| 28 | claude-sonnet-4-6_16K | Anthropic | 0.08 |
| 29 | gpt-5-mini-2025-08-07_high | OpenAI | 0.06 |
| 30 | gpt-5.2-2025-12-11_low | OpenAI | 0.06 |
| 31 | o4-mini-2025-04-16_high | OpenAI | 0.06 |
| 32 | gpt-5-2025-08-07_medium | OpenAI | 0.06 |
| 33 | gpt-5.4-nano-2026-03-17_high | OpenAI | 0.06 |
| 34 | fireworks/kimi-k2p5 | Moonshot | 0.04 |
| 35 | claude-opus-4-5-20251101_32K | Anthropic | 0.04 |
| 36 | o3-mini-2025-01-31_high | OpenAI | 0.04 |
| 37 | claude-opus-4-20250514_27K | Anthropic | 0.04 |
| 38 | gemini-2.5-flash | Google DeepMind | 0.04 |
| 39 | gemini-2.5-pro | Google DeepMind | 0.04 |
| 40 | claude-opus-4-1-20250805_27K | Anthropic | 0.04 |
| 41 | gpt-5-mini-2025-08-07_medium | OpenAI | 0.04 |
| 42 | claude-sonnet-4-5-20250929_32K | Anthropic | 0.04 |
| 43 | gpt-5.1-2025-11-13_medium | OpenAI | 0.04 |
| 44 | claude-opus-4-5-20251101 | Anthropic | 0.04 |
| 45 | gemini-3-flash-preview | Google DeepMind | 0.04 |
| 46 | qwen3.6-max-preview | Alibaba | 0.04 |
| 47 | zai-org/GLM-4.6 | Z.ai (Zhipu AI),Tsinghua University | 0.02 |
| 48 | glm-5 | Z.ai (Zhipu AI) | 0.02 |
| 49 | fireworks/deepseek-v3p2 | DeepSeek | 0.02 |
| 50 | claude-haiku-4-5-20251001_32K | Anthropic | 0.02 |
| 51 | claude-opus-4-5-20251101_16K | Anthropic | 0.02 |
| 52 | o3-2025-04-16_high | OpenAI | 0.02 |
| 53 | o4-mini-2025-04-16_medium | OpenAI | 0.02 |
| 54 | claude-sonnet-4-5-20250929 | Anthropic | 0.02 |
| 55 | gpt-5-nano-2025-08-07_medium | OpenAI | 0.02 |
| 56 | qwen3.5-plus | Alibaba | 0.02 |
| 57 | grok-4-0709 | xAI | 0.02 |
| 58 | gemini-2.5-pro-preview-06-05 | Google DeepMind | 0.02 |
| 59 | gpt-5.4-mini-2026-03-17_high | OpenAI | 0.02 |
| 60 | grok-4-heavy-web-app | xAI | 0.02 |