ALE Bench is an AI evaluation tracked by Epoch AI, with 87 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span April 2025 to June 2026. The best score recorded is 2041.3, measured as Performance. Models come from Anthropic, OpenAI, Google DeepMind, Alibaba, xAI, Moonshot, Z.ai (Zhipu AI), DeepSeek and others.
| 01 | claude-fable-5_high | Anthropic | 2041.3 |
| 02 | gpt-5.5_xhigh | OpenAI | 1943 |
| 03 | gpt-5.3-codex_xhigh | OpenAI | 1655.2 |
| 04 | gpt-5.4-2026-03-05_high | OpenAI | 1607 |
| 05 | gpt-5.5_medium | OpenAI | 1589.4 |
| 06 | claude-opus-4-8_high | Anthropic | 1563.8 |
| 07 | gpt-5.4-2026-03-05_medium | OpenAI | 1520.7 |
| 08 | claude-opus-4-8_none | Anthropic | 1411.8 |
| 09 | gemini-3-flash-preview | Google DeepMind | 1367.2 |
| 10 | claude-sonnet-4-6_medium | Anthropic | 1327.3 |
| 11 | claude-opus-4-7 | Anthropic | 1323 |
| 12 | gpt-5.2-codex | OpenAI | 1299.9 |
| 13 | gpt-5.2-2025-12-11_high | OpenAI | 1293.5 |
| 14 | gpt-5.2-2025-12-11_medium | OpenAI | 1249.8 |
| 15 | gpt-5.1-codex | OpenAI | 1244.9 |
| 16 | gpt-5.1-codex-max | OpenAI | 1208.8 |
| 17 | gpt-5.1-2025-11-13_high | OpenAI | 1192.2 |
| 18 | qwen3.7-max | Alibaba | 1189.4 |
| 19 | gpt-5.4-mini-2026-03-17_high | OpenAI | 1188.6 |
| 20 | gemini-3-pro-preview | Google DeepMind | 1176.8 |
| 21 | gpt-5-2025-08-07_high | OpenAI | 1162.5 |
| 22 | gemini-3.1-pro-preview | Google DeepMind | 1160.6 |
| 23 | grok-4-20 | xAI | 1150.3 |
| 24 | gpt-5.5_none | OpenAI | 1127.6 |
| 25 | kimi-k2.6 | Moonshot | 1092.7 |
| 26 | gpt-5.4-2026-03-05_none | OpenAI | 1086 |
| 27 | glm-5.2_high | Z.ai (Zhipu AI) | 1047 |
| 28 | claude-opus-4-5-20251101_16K | Anthropic | 1025.4 |
| 29 | glm-5.2_max | Z.ai (Zhipu AI) | 1010.2 |
| 30 | deepseek-v4-pro_high | DeepSeek | 1006.1 |
| 31 | gpt-5.4-nano-2026-03-17_high | OpenAI | 1004.5 |
| 32 | claude-opus-4-6 | Anthropic | 996.5 |
| 33 | grok-4-3 | xAI | 944.2 |
| 34 | o3-2025-04-16_high | OpenAI | 933.5 |
| 35 | gemma-4-26b-a4b | Google DeepMind | 927.2 |
| 36 | gemma-4-31b-it | Google DeepMind | 925.5 |
| 37 | gemini-3.5-flash_high | 911 | |
| 38 | mimo-v2.5-pro | Xiaomi Corp | 899.8 |
| 39 | glm-5.1 | Z.ai (Zhipu AI) | 887.1 |
| 40 | kimi-k2.7-code | Moonshot | 886.2 |
| 41 | o4-mini-2025-04-16_high | OpenAI | 826.2 |
| 42 | kimi-k2.5 | Moonshot | 821.6 |
| 43 | gpt-5-2025-08-07_minimal | OpenAI | 807.6 |
| 44 | DeepSeek-R1-0528 | DeepSeek | 804.1 |
| 45 | gpt-5-mini-2025-08-07_high | OpenAI | 799.8 |
| 46 | gemini-3.1-flash-lite | 797.7 | |
| 47 | claude-sonnet-4-5-20250929_32K | Anthropic | 796.1 |
| 48 | mercury-2 | Inception Labs | 785.6 |
| 49 | gemini-2.5-pro_32K | Google DeepMind | 785.5 |
| 50 | mimo-v2-pro | Xiaomi Corp | 785.2 |
| 51 | glm-5 | Z.ai (Zhipu AI) | 765.6 |
| 52 | mimo-v2-flash | Xiaomi Corp | 738 |
| 53 | gpt-5-nano-2025-08-07_high | OpenAI | 718.7 |
| 54 | deepseek-v4-flash_high | DeepSeek | 678.2 |
| 55 | claude-opus-4-1-20250805_16K | Anthropic | 674.8 |
| 56 | qwen3.6-plus | Alibaba | 670.1 |
| 57 | gemini-2.5-flash | Google DeepMind | 661.9 |
| 58 | claude-sonnet-4-20250514_32K | Anthropic | 655.4 |
| 59 | claude-haiku-4-5-20251001_32K | Anthropic | 653.5 |
| 60 | MiniMax-M3 | MiniMax | 640 |