WeirdML is an AI evaluation tracked by Epoch AI, with 146 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span June 2023 to July 2026. The best score recorded is 0.89, measured as Accuracy. Models come from OpenAI, Anthropic, Google DeepMind, Z.ai (Zhipu AI), Google, Moonshot, xAI, DeepSeek and others.
| 01 | gpt-5.6-sol_high | OpenAI | 0.89 |
| 02 | claude-fable-5_high | Anthropic | 0.88 |
| 03 | gpt-5.5_xhigh | OpenAI | 0.85 |
| 04 | gpt-5.5_high | OpenAI | 0.84 |
| 05 | claude-opus-4-8_xhigh | Anthropic | 0.83 |
| 06 | gpt-5.3-codex | OpenAI | 0.79 |
| 07 | gpt-5.3-codex_xhigh | OpenAI | 0.79 |
| 08 | claude-opus-4-6_high | Anthropic | 0.78 |
| 09 | claude-opus-4-6_unknown | Anthropic | 0.78 |
| 10 | gpt-5.4-2026-03-05_xhigh | OpenAI | 0.78 |
| 11 | claude-opus-4-7 | Anthropic | 0.76 |
| 12 | claude-opus-4-7_high | Anthropic | 0.76 |
| 13 | claude-opus-4-7_unknown | Anthropic | 0.76 |
| 14 | claude-opus-4-8_medium | Anthropic | 0.76 |
| 15 | claude-opus-4-7_max | Anthropic | 0.75 |
| 16 | gpt-5.2-2025-12-11_xhigh | OpenAI | 0.72 |
| 17 | gemini-3.1-pro-preview | Google DeepMind | 0.72 |
| 18 | claude-opus-4-8_none | Anthropic | 0.7 |
| 19 | glm-5.2_max | Z.ai (Zhipu AI) | 0.7 |
| 20 | gemini-3-pro-preview | Google DeepMind | 0.7 |
| 21 | claude-sonnet-5_high | Anthropic | 0.69 |
| 22 | glm-5.2_high | Z.ai (Zhipu AI) | 0.67 |
| 23 | gpt-5.5_none | OpenAI | 0.67 |
| 24 | claude-sonnet-4-6_medium | Anthropic | 0.66 |
| 25 | claude-opus-4-6 | Anthropic | 0.66 |
| 26 | claude-opus-4-5-20251101_16K | Anthropic | 0.64 |
| 27 | gpt-5.2-2025-12-11_medium | OpenAI | 0.63 |
| 28 | gemini-3.5-flash_high | 0.63 | |
| 29 | gemini-3-flash-preview | Google DeepMind | 0.62 |
| 30 | gpt-5.1-2025-11-13_high | OpenAI | 0.61 |
| 31 | gpt-5-2025-08-07_high | OpenAI | 0.61 |
| 32 | gpt-5-pro-2025-10-06_high | OpenAI | 0.6 |
| 33 | gpt-5.4-mini-2026-03-17_high | OpenAI | 0.6 |
| 34 | o3-pro-2025-06-10_high | OpenAI | 0.58 |
| 35 | gpt-5.4-2026-03-05_none | OpenAI | 0.57 |
| 36 | gpt-5.4-pro-2026-03-05_none | OpenAI | 0.57 |
| 37 | glm-5.1 | Z.ai (Zhipu AI) | 0.57 |
| 38 | kimi-k2.6 | Moonshot | 0.56 |
| 39 | gpt-5-codex | OpenAI | 0.55 |
| 40 | gpt-5-codex_high | OpenAI | 0.55 |
| 41 | kimi-k2.7-code | Moonshot | 0.54 |
| 42 | gemini-2.5-pro_16K | Google DeepMind | 0.54 |
| 43 | gpt-5-mini-2025-08-07_high | OpenAI | 0.53 |
| 44 | o4-mini-2025-04-16_high | OpenAI | 0.53 |
| 45 | o3-2025-04-16_high | OpenAI | 0.52 |
| 46 | grok-4-20 | xAI | 0.52 |
| 47 | gemma-4-31b-it | Google DeepMind | 0.52 |
| 48 | gemini-3.1-flash-lite | 0.52 | |
| 49 | grok-4-3 | xAI | 0.5 |
| 50 | gpt-5.2-2025-12-11_none | OpenAI | 0.5 |
| 51 | gpt-5.2-2025-12-11_low | OpenAI | 0.5 |
| 52 | gpt-5.4-nano-2026-03-17_high | OpenAI | 0.49 |
| 53 | deepseek-v4-pro_max | DeepSeek | 0.49 |
| 54 | gpt-oss-120b_high | OpenAI | 0.48 |
| 55 | glm-5 | Z.ai (Zhipu AI) | 0.48 |
| 56 | openai/gpt-oss-120b_high | OpenAI | 0.48 |
| 57 | claude-sonnet-4-5-20250929_16K | Anthropic | 0.48 |
| 58 | o1-preview-2024-09-12 | OpenAI | 0.48 |
| 59 | DeepSeek-V3.2-Speciale | DeepSeek | 0.47 |
| 60 | claude-sonnet-4-5-20250929 | Anthropic | 0.47 |