Live
Benchmarks

WeirdML

Models scored
146
Best score
0.89
Metric
Accuracy

WeirdML is an AI evaluation tracked by Epoch AI, with 146 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span June 2023 to July 2026. The best score recorded is 0.89, measured as Accuracy. Models come from OpenAI, Anthropic, Google DeepMind, Z.ai (Zhipu AI), Google, Moonshot, xAI, DeepSeek and others.

Full record
Score metric
Accuracy
Models scored
146
Best score recorded
0.89
Earliest model scored
Jun 13, 2023
Most recent model scored
Jul 9, 2026
Third-party reported
Yes
Leaderboard
01gpt-5.6-sol_highOpenAI0.89
02claude-fable-5_highAnthropic0.88
03gpt-5.5_xhighOpenAI0.85
04gpt-5.5_highOpenAI0.84
05claude-opus-4-8_xhighAnthropic0.83
06gpt-5.3-codexOpenAI0.79
07gpt-5.3-codex_xhighOpenAI0.79
08claude-opus-4-6_highAnthropic0.78
09claude-opus-4-6_unknownAnthropic0.78
10gpt-5.4-2026-03-05_xhighOpenAI0.78
11claude-opus-4-7Anthropic0.76
12claude-opus-4-7_highAnthropic0.76
13claude-opus-4-7_unknownAnthropic0.76
14claude-opus-4-8_mediumAnthropic0.76
15claude-opus-4-7_maxAnthropic0.75
16gpt-5.2-2025-12-11_xhighOpenAI0.72
17gemini-3.1-pro-previewGoogle DeepMind0.72
18claude-opus-4-8_noneAnthropic0.7
19glm-5.2_maxZ.ai (Zhipu AI)0.7
20gemini-3-pro-previewGoogle DeepMind0.7
21claude-sonnet-5_highAnthropic0.69
22glm-5.2_highZ.ai (Zhipu AI)0.67
23gpt-5.5_noneOpenAI0.67
24claude-sonnet-4-6_mediumAnthropic0.66
25claude-opus-4-6Anthropic0.66
26claude-opus-4-5-20251101_16KAnthropic0.64
27gpt-5.2-2025-12-11_mediumOpenAI0.63
28gemini-3.5-flash_highGoogle0.63
29gemini-3-flash-previewGoogle DeepMind0.62
30gpt-5.1-2025-11-13_highOpenAI0.61
31gpt-5-2025-08-07_highOpenAI0.61
32gpt-5-pro-2025-10-06_highOpenAI0.6
33gpt-5.4-mini-2026-03-17_highOpenAI0.6
34o3-pro-2025-06-10_highOpenAI0.58
35gpt-5.4-2026-03-05_noneOpenAI0.57
36gpt-5.4-pro-2026-03-05_noneOpenAI0.57
37glm-5.1Z.ai (Zhipu AI)0.57
38kimi-k2.6Moonshot0.56
39gpt-5-codexOpenAI0.55
40gpt-5-codex_highOpenAI0.55
41kimi-k2.7-codeMoonshot0.54
42gemini-2.5-pro_16KGoogle DeepMind0.54
43gpt-5-mini-2025-08-07_highOpenAI0.53
44o4-mini-2025-04-16_highOpenAI0.53
45o3-2025-04-16_highOpenAI0.52
46grok-4-20xAI0.52
47gemma-4-31b-itGoogle DeepMind0.52
48gemini-3.1-flash-liteGoogle0.52
49grok-4-3xAI0.5
50gpt-5.2-2025-12-11_noneOpenAI0.5
51gpt-5.2-2025-12-11_lowOpenAI0.5
52gpt-5.4-nano-2026-03-17_highOpenAI0.49
53deepseek-v4-pro_maxDeepSeek0.49
54gpt-oss-120b_highOpenAI0.48
55glm-5Z.ai (Zhipu AI)0.48
56openai/gpt-oss-120b_highOpenAI0.48
57claude-sonnet-4-5-20250929_16KAnthropic0.48
58o1-preview-2024-09-12OpenAI0.48
59DeepSeek-V3.2-SpecialeDeepSeek0.47
60claude-sonnet-4-5-20250929Anthropic0.47
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks