Proofbench is an AI evaluation tracked by Epoch AI, with 44 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span August 2025 to July 2026. The best score recorded is 0.78, measured as Accuracy. Models come from OpenAI, Anthropic, Meta AI, Z.ai (Zhipu AI), xAI, Google, Alibaba, Google DeepMind and others.
| 01 | claude-opus-5_max | Anthropic | 0.78 |
| 02 | claude-fable-5_max | Anthropic | 0.77 |
| 03 | gpt-5.6-sol_max | OpenAI | 0.77 |
| 04 | gpt-5.6-terra_xhigh | OpenAI | 0.71 |
| 05 | claude-opus-4-8_max | Anthropic | 0.69 |
| 06 | claude-sonnet-5_max | Anthropic | 0.66 |
| 07 | gpt-5.4-2026-03-05_xhigh | OpenAI | 0.56 |
| 08 | gpt-5.6-luna_max | OpenAI | 0.54 |
| 09 | claude-opus-4-7_max | Anthropic | 0.54 |
| 10 | claude-opus-4-6_max | Anthropic | 0.5 |
| 11 | gpt-5.5_xhigh | OpenAI | 0.5 |
| 12 | claude-sonnet-4-6_max | Anthropic | 0.45 |
| 13 | muse-spark-1.1 | Meta AI | 0.39 |
| 14 | claude-opus-4-5-20251101_unknown | Anthropic | 0.36 |
| 15 | glm-5.2_max | Z.ai (Zhipu AI) | 0.35 |
| 16 | grok-4.5_high | xAI | 0.3 |
| 17 | gemini-3.5-flash_high | 0.29 | |
| 18 | gemini-3.1-pro-preview | Google DeepMind | 0.26 |
| 19 | qwen3.7-max | Alibaba | 0.26 |
| 20 | mimo-v2.5-pro | Xiaomi Corp | 0.24 |
| 21 | glm-5.1 | Z.ai (Zhipu AI) | 0.22 |
| 22 | gpt-5.4-mini-2026-03-17_xhigh | OpenAI | 0.21 |
| 23 | gemini-3-pro-preview | Google DeepMind | 0.2 |
| 24 | claude-sonnet-4-5-20250929_unknown | Anthropic | 0.19 |
| 25 | MiniMax-M3 | MiniMax | 0.19 |
| 26 | gpt-5-2025-08-07_high | OpenAI | 0.18 |
| 27 | muse-spark | Meta AI | 0.17 |
| 28 | mimo-v2.5 | Xiaomi Corp | 0.16 |
| 29 | kimi-k2.6 | Moonshot | 0.16 |
| 30 | gemini-3-flash-preview | Google DeepMind | 0.15 |
| 31 | gpt-5.2-2025-12-11_xhigh | OpenAI | 0.15 |
| 32 | grok-4.20-0309-reasoning | xAI | 0.14 |
| 33 | gpt-5-nano-2025-08-07_high | OpenAI | 0.12 |
| 34 | grok-4.3_high | xAI | 0.11 |
| 35 | deepseek-v4-pro_max | DeepSeek | 0.1 |
| 36 | gpt-5.1-codex-max | OpenAI | 0.09 |
| 37 | gpt-5-mini-2025-08-07_high | OpenAI | 0.09 |
| 38 | fireworks/deepseek-v3p2 | DeepSeek | 0.08 |
| 39 | glm-4.7 | Z.ai (Zhipu AI) | 0.06 |
| 40 | gpt-5.4-nano-2026-03-17_high | OpenAI | 0.05 |
| 41 | grok-4-1-fast-reasoning | xAI | 0.04 |
| 42 | MiniMax-M2.5 | MiniMax | 0.04 |
| 43 | MiniMax-M2.7 | MiniMax | 0.03 |
| 44 | nemotron-3-ultra | NVIDIA | 0.02 |