Live
Benchmarks

SimpleQA Verified

Models scored
68
Best score
0.77
Metric
mean_score

SimpleQA Verified is an AI evaluation tracked by Epoch AI, with 68 scored model versions on record. Epoch AI runs this evaluation directly, so the scores are reproducible against their published logs.

Scored models span October 2024 to July 2026. The best score recorded is 0.77, measured as mean_score. Models come from Anthropic, Moonshot, xAI, OpenAI, Z.ai (Zhipu AI), DeepSeek, Alibaba, Google and others.

Full record
Score metric
mean_score
Models scored
68
Best score recorded
0.77
Earliest model scored
Oct 22, 2024
Most recent model scored
Jul 24, 2026
Third-party reported
No
Leaderboard
01gemini-3.1-pro-previewGoogle DeepMind0.77
02gemini-3-pro-previewGoogle DeepMind0.73
03gpt-5.6-sol_maxOpenAI0.72
04gemini-3.5-flash_highGoogle0.68
05claude-fable-5_xhighAnthropic0.68
06qwen3-max-2025-09-23Alibaba0.67
07gemini-3-flash-previewGoogle DeepMind0.67
08muse-sparkMeta AI0.66
09gpt-5.5-pro-pre-release_xhighOpenAI0.65
10gpt-5.5-pre-release_xhighOpenAI0.63
11qwen3.7-maxAlibaba0.59
12deepseek-v4-pro_maxDeepSeek0.57
13qwen3.6-max-previewAlibaba0.57
14claude-opus-5_maxAnthropic0.57
15gemini-2.5-proGoogle DeepMind0.56
16grok-4.5_highxAI0.54
17o3-2025-04-16_highOpenAI0.53
18claude-opus-4-7_xhighAnthropic0.51
19gpt-5-2025-08-07_highOpenAI0.51
20Qwen3-235B-A22B-Thinking-2507Alibaba0.5
21qwen3.6-plusAlibaba0.49
22gpt-5.1-2025-11-13_highOpenAI0.49
23grok-4-0709xAI0.48
24gpt-5.4-pro-2026-03-05_xhighOpenAI0.48
25claude-opus-4-6_32KAnthropic0.46
26gpt-5.4-2026-03-05_xhighOpenAI0.45
27claude-opus-4-6Anthropic0.43
28gpt-5.6-terra_maxOpenAI0.43
29kimi-k3_maxMoonshot0.43
30claude-opus-4-5-20251101_32KAnthropic0.42
31gpt-5.6-luna_maxOpenAI0.42
32claude-opus-4-6_maxAnthropic0.41
33claude-opus-4-8_maxAnthropic0.4
34kimi-k2.7-codeMoonshot0.39
35gpt-5.2-2025-12-11_xhighOpenAI0.39
36kimi-k2.6Moonshot0.39
37gpt-5.2-2025-12-11_highOpenAI0.38
38glm-5.2_maxZ.ai (Zhipu AI)0.38
39grok-4.3_highxAI0.38
40glm-5.1Z.ai (Zhipu AI)0.37
41gpt-5.2-2025-12-11_mediumOpenAI0.35
42grok-4.20-0309-reasoningxAI0.35
43claude-opus-4-1-20250805_27KAnthropic0.35
44gpt-5.2-2025-12-11_lowOpenAI0.35
45fireworks/kimi-k2p5Moonshot0.34
46kimi-k2-thinking-turboMoonshot0.32
47glm-4.7Z.ai (Zhipu AI)0.32
48claude-sonnet-4-6_32KAnthropic0.29
49gpt-5.4-mini-2026-03-17_highOpenAI0.29
50deepseek-reasonerDeepSeek0.28
51DeepSeek-R1-0528DeepSeek0.27
52qwen3.5-plusAlibaba0.26
53claude-sonnet-5_xhighAnthropic0.25
54o4-mini-2025-04-16_highOpenAI0.24
55claude-sonnet-4-5-20250929_59KAnthropic0.24
56claude-sonnet-5_maxAnthropic0.22
57o4-mini-2025-04-16_lowOpenAI0.22
58qwen3.6-flashAlibaba0.21
59grok-3-mini-beta_highxAI0.21
60gpt-5-mini-2025-08-07_highOpenAI0.21
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks