Live
Benchmarks

ALE Bench

Models scored
87
Best score
2041.3
Metric
Performance

ALE Bench is an AI evaluation tracked by Epoch AI, with 87 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span April 2025 to June 2026. The best score recorded is 2041.3, measured as Performance. Models come from Anthropic, OpenAI, Google DeepMind, Alibaba, xAI, Moonshot, Z.ai (Zhipu AI), DeepSeek and others.

Full record
Score metric
Performance
Models scored
87
Best score recorded
2041.3
Earliest model scored
Apr 5, 2025
Most recent model scored
Jun 16, 2026
Third-party reported
Yes
Leaderboard
01claude-fable-5_highAnthropic2041.3
02gpt-5.5_xhighOpenAI1943
03gpt-5.3-codex_xhighOpenAI1655.2
04gpt-5.4-2026-03-05_highOpenAI1607
05gpt-5.5_mediumOpenAI1589.4
06claude-opus-4-8_highAnthropic1563.8
07gpt-5.4-2026-03-05_mediumOpenAI1520.7
08claude-opus-4-8_noneAnthropic1411.8
09gemini-3-flash-previewGoogle DeepMind1367.2
10claude-sonnet-4-6_mediumAnthropic1327.3
11claude-opus-4-7Anthropic1323
12gpt-5.2-codexOpenAI1299.9
13gpt-5.2-2025-12-11_highOpenAI1293.5
14gpt-5.2-2025-12-11_mediumOpenAI1249.8
15gpt-5.1-codexOpenAI1244.9
16gpt-5.1-codex-maxOpenAI1208.8
17gpt-5.1-2025-11-13_highOpenAI1192.2
18qwen3.7-maxAlibaba1189.4
19gpt-5.4-mini-2026-03-17_highOpenAI1188.6
20gemini-3-pro-previewGoogle DeepMind1176.8
21gpt-5-2025-08-07_highOpenAI1162.5
22gemini-3.1-pro-previewGoogle DeepMind1160.6
23grok-4-20xAI1150.3
24gpt-5.5_noneOpenAI1127.6
25kimi-k2.6Moonshot1092.7
26gpt-5.4-2026-03-05_noneOpenAI1086
27glm-5.2_highZ.ai (Zhipu AI)1047
28claude-opus-4-5-20251101_16KAnthropic1025.4
29glm-5.2_maxZ.ai (Zhipu AI)1010.2
30deepseek-v4-pro_highDeepSeek1006.1
31gpt-5.4-nano-2026-03-17_highOpenAI1004.5
32claude-opus-4-6Anthropic996.5
33grok-4-3xAI944.2
34o3-2025-04-16_highOpenAI933.5
35gemma-4-26b-a4bGoogle DeepMind927.2
36gemma-4-31b-itGoogle DeepMind925.5
37gemini-3.5-flash_highGoogle911
38mimo-v2.5-proXiaomi Corp899.8
39glm-5.1Z.ai (Zhipu AI)887.1
40kimi-k2.7-codeMoonshot886.2
41o4-mini-2025-04-16_highOpenAI826.2
42kimi-k2.5Moonshot821.6
43gpt-5-2025-08-07_minimalOpenAI807.6
44DeepSeek-R1-0528DeepSeek804.1
45gpt-5-mini-2025-08-07_highOpenAI799.8
46gemini-3.1-flash-liteGoogle797.7
47claude-sonnet-4-5-20250929_32KAnthropic796.1
48mercury-2Inception Labs785.6
49gemini-2.5-pro_32KGoogle DeepMind785.5
50mimo-v2-proXiaomi Corp785.2
51glm-5Z.ai (Zhipu AI)765.6
52mimo-v2-flashXiaomi Corp738
53gpt-5-nano-2025-08-07_highOpenAI718.7
54deepseek-v4-flash_highDeepSeek678.2
55claude-opus-4-1-20250805_16KAnthropic674.8
56qwen3.6-plusAlibaba670.1
57gemini-2.5-flashGoogle DeepMind661.9
58claude-sonnet-4-20250514_32KAnthropic655.4
59claude-haiku-4-5-20251001_32KAnthropic653.5
60MiniMax-M3MiniMax640
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks