Humanity's Last Exam is an AI evaluation tracked by Epoch AI, with 51 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span September 2024 to April 2026. The best score recorded is 0.46, measured as Accuracy. Models come from Google DeepMind, OpenAI, Meta AI, Anthropic, Moonshot, Google, Z.ai (Zhipu AI),Tsinghua University, Google DeepMind,Google and others.
| 01 | gemini-3.1-pro-preview | Google DeepMind | 0.46 |
| 02 | gpt-5.4-pro-2026-03-05_unknown | OpenAI | 0.44 |
| 03 | muse-spark | Meta AI | 0.41 |
| 04 | gemini-3-pro-preview | Google DeepMind | 0.38 |
| 05 | gpt-5.4-2026-03-05_xhigh | OpenAI | 0.36 |
| 06 | claude-opus-4-7_unknown | Anthropic | 0.36 |
| 07 | claude-opus-4-6_max | Anthropic | 0.34 |
| 08 | gpt-5-pro-2025-10-06_unknown | OpenAI | 0.32 |
| 09 | gpt-5.2-2025-12-11_unknown | OpenAI | 0.28 |
| 10 | gpt-5-2025-08-07_high | OpenAI | 0.25 |
| 11 | gpt-5-2025-08-07_unknown | OpenAI | 0.25 |
| 12 | claude-opus-4-5-20251101_unknown | Anthropic | 0.25 |
| 13 | kimi-k2.5 | Moonshot | 0.24 |
| 14 | gpt-5.1-2025-11-13_unknown | OpenAI | 0.24 |
| 15 | gemini-2.5-pro-preview-06-05 | Google DeepMind | 0.22 |
| 16 | o3-2025-04-16_high | OpenAI | 0.2 |
| 17 | gpt-5-mini-2025-08-07_unknown | OpenAI | 0.19 |
| 18 | o3-2025-04-16_medium | OpenAI | 0.19 |
| 19 | claude-opus-4-6 | Anthropic | 0.19 |
| 20 | gemini-2.5-pro-exp-03-25 | Google DeepMind | 0.18 |
| 21 | o4-mini-2025-04-16_high | OpenAI | 0.18 |
| 22 | gemini-2.5-pro-preview-05-06 | Google DeepMind | 0.18 |
| 23 | o4-mini-2025-04-16_medium | OpenAI | 0.14 |
| 24 | claude-sonnet-4-5-20250929_unknown | Anthropic | 0.14 |
| 25 | gemini-2.5-flash-preview-04-17 | Google DeepMind | 0.12 |
| 26 | claude-opus-4-1-20250805_unknown | Anthropic | 0.12 |
| 27 | gemini-2.5-flash-preview-05-20 | Google DeepMind | 0.11 |
| 28 | claude-opus-4-20250514_unknown | Anthropic | 0.11 |
| 29 | gemini-3.1-flash-lite | 0.09 | |
| 30 | glm-4.5 | Z.ai (Zhipu AI),Tsinghua University | 0.08 |
| 31 | GLM-4.5-Air | Z.ai (Zhipu AI),Tsinghua University | 0.08 |
| 32 | o1-pro-2025-03-19 | OpenAI | 0.08 |
| 33 | claude-3-7-sonnet-20250219_unknown | Anthropic | 0.08 |
| 34 | o1-2024-12-17_unknown | OpenAI | 0.08 |
| 35 | claude-sonnet-4-20250514_unknown | Anthropic | 0.08 |
| 36 | gpt-5.1-2025-11-13_none | OpenAI | 0.07 |
| 37 | gemini-2.0-flash-thinking-exp-01-21 | Google DeepMind,Google | 0.07 |
| 38 | Llama-4-Maverick-17B-128E-Instruct | Meta AI | 0.06 |
| 39 | gpt-4.5-preview-2025-02-27 | OpenAI | 0.05 |
| 40 | gpt-4.1-2025-04-14 | OpenAI | 0.05 |
| 41 | gemini-1.5-pro-002 | Google DeepMind | 0.05 |
| 42 | mistral-medium-2505 | Mistral AI | 0.05 |
| 43 | amazon.nova-pro-v1:0 | Amazon | 0.04 |
| 44 | claude-3-5-sonnet-20241022 | Anthropic | 0.04 |
| 45 | amazon.nova-lite-v1:0 | Amazon | 0.04 |
| 46 | gpt-4o-2024-11-20 | OpenAI | 0.03 |