En vivo
Benchmarks

Critpt

Models scored
74
Best score
0.32
Metric
Accuracy

Critpt is an AI evaluation tracked by Epoch AI, with 74 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span July 2024 to July 2026. The best score recorded is 0.32, measured as Accuracy. Models come from OpenAI, Anthropic, Moonshot, Z.ai (Zhipu AI), Google DeepMind, Meta AI, Alibaba, Google and others.

Full record
Score metric
Accuracy
Models scored
74
Best score recorded
0.32
Earliest model scored
Jul 23, 2024
Most recent model scored
Jul 24, 2026
Third-party reported
Yes
Leaderboard
01gpt-5.6-sol_maxOpenAI0.32
02gpt-5.5-pro_xhighOpenAI0.31
03gpt-5.6-terra_maxOpenAI0.3
04gpt-5.4-pro-2026-03-05_xhighOpenAI0.3
05claude-opus-5_maxAnthropic0.29
06gpt-5.6-sol_xhighOpenAI0.29
07claude-fable-5_maxAnthropic0.29
08gpt-5.5_xhighOpenAI0.27
09gemini-3-deep-think-preview0.26
10gpt-5.5_highOpenAI0.25
11gpt-5.4-2026-03-05_xhighOpenAI0.23
12kimi-k3_maxMoonshot0.23
13glm-5.2_maxZ.ai (Zhipu AI)0.21
14claude-opus-4-8_maxAnthropic0.21
15gpt-5.6-luna_maxOpenAI0.21
16gpt-5.5_mediumOpenAI0.19
17gemini-3.1-pro-previewGoogle DeepMind0.18
18claude-sonnet-5_maxAnthropic0.17
19muse-spark-1.1Meta AI0.15
20qwen3.7-maxAlibaba0.13
21gemini-3.5-flash_highGoogle0.13
22deepseek-v4-pro_maxDeepSeek0.13
23gpt-5-2025-08-07_highOpenAI0.13
24claude-opus-4-7_maxAnthropic0.12
25muse-sparkMeta AI0.11
26gpt-5.4-mini-2026-03-17_xhighOpenAI0.1
27kimi-k2.7-codeMoonshot0.1
28gpt-5.4-nano-2026-03-17_xhighOpenAI0.09
29gpt-5.5_lowOpenAI0.08
30kimi-k2.6Moonshot0.08
31grok-4.3_highxAI0.08
32deepseek-v4-flash_maxDeepSeek0.07
33gemini-3-pro-previewGoogle DeepMind0.07
34gpt-5.1-2025-11-13_unknownOpenAI0.05
35glm-5.1Z.ai (Zhipu AI)0.05
36mimo-v2.5-proXiaomi Corp0.04
37MiniMax-M3MiniMax0.04
38nemotron-3-ultraNVIDIA0.03
39claude-sonnet-4-6_maxAnthropic0.03
40qwen3.6-plusAlibaba0.03
41gemini-2.5-proGoogle DeepMind0.02
42gpt-5.5_noneOpenAI0.01
43gemma-4-31b-itGoogle DeepMind0.01
44gpt-oss-20b_highOpenAI0.01
45o3-2025-04-16_highOpenAI0.01
46gemini-3.1-flash-liteGoogle0.01
47openai/gpt-oss-120b_highOpenAI0.01
48claude-sonnet-4-5-20250929_unknownAnthropic0.01
49gemini-2.5-flashGoogle DeepMind0.01
50DeepSeek-R1DeepSeek0.01
51o4-mini-2025-04-16_highOpenAI0.01
52MiniMax-M2.7MiniMax0.01
53claude-opus-4-20250514_unknownAnthropic0
54qwen3.5-9BAlibaba0
55qwen3.6-35b-a3bAlibaba0
56claude-sonnet-4-20250514_unknownAnthropic0
57Qwen3-32BAlibaba0
58Llama-3.1-8B-InstructMeta AI0
59gpt-4.1-mini-2025-04-14OpenAI0
60gpt-5-2025-08-07_minimalOpenAI0
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks