Critpt is an AI evaluation tracked by Epoch AI, with 74 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span July 2024 to July 2026. The best score recorded is 0.32, measured as Accuracy. Models come from OpenAI, Anthropic, Moonshot, Z.ai (Zhipu AI), Google DeepMind, Meta AI, Alibaba, Google and others.
| 01 | gpt-5.6-sol_max | OpenAI | 0.32 |
| 02 | gpt-5.5-pro_xhigh | OpenAI | 0.31 |
| 03 | gpt-5.6-terra_max | OpenAI | 0.3 |
| 04 | gpt-5.4-pro-2026-03-05_xhigh | OpenAI | 0.3 |
| 05 | claude-opus-5_max | Anthropic | 0.29 |
| 06 | gpt-5.6-sol_xhigh | OpenAI | 0.29 |
| 07 | claude-fable-5_max | Anthropic | 0.29 |
| 08 | gpt-5.5_xhigh | OpenAI | 0.27 |
| 09 | gemini-3-deep-think-preview | 0.26 | |
| 10 | gpt-5.5_high | OpenAI | 0.25 |
| 11 | gpt-5.4-2026-03-05_xhigh | OpenAI | 0.23 |
| 12 | kimi-k3_max | Moonshot | 0.23 |
| 13 | glm-5.2_max | Z.ai (Zhipu AI) | 0.21 |
| 14 | claude-opus-4-8_max | Anthropic | 0.21 |
| 15 | gpt-5.6-luna_max | OpenAI | 0.21 |
| 16 | gpt-5.5_medium | OpenAI | 0.19 |
| 17 | gemini-3.1-pro-preview | Google DeepMind | 0.18 |
| 18 | claude-sonnet-5_max | Anthropic | 0.17 |
| 19 | muse-spark-1.1 | Meta AI | 0.15 |
| 20 | qwen3.7-max | Alibaba | 0.13 |
| 21 | gemini-3.5-flash_high | 0.13 | |
| 22 | deepseek-v4-pro_max | DeepSeek | 0.13 |
| 23 | gpt-5-2025-08-07_high | OpenAI | 0.13 |
| 24 | claude-opus-4-7_max | Anthropic | 0.12 |
| 25 | muse-spark | Meta AI | 0.11 |
| 26 | gpt-5.4-mini-2026-03-17_xhigh | OpenAI | 0.1 |
| 27 | kimi-k2.7-code | Moonshot | 0.1 |
| 28 | gpt-5.4-nano-2026-03-17_xhigh | OpenAI | 0.09 |
| 29 | gpt-5.5_low | OpenAI | 0.08 |
| 30 | kimi-k2.6 | Moonshot | 0.08 |
| 31 | grok-4.3_high | xAI | 0.08 |
| 32 | deepseek-v4-flash_max | DeepSeek | 0.07 |
| 33 | gemini-3-pro-preview | Google DeepMind | 0.07 |
| 34 | gpt-5.1-2025-11-13_unknown | OpenAI | 0.05 |
| 35 | glm-5.1 | Z.ai (Zhipu AI) | 0.05 |
| 36 | mimo-v2.5-pro | Xiaomi Corp | 0.04 |
| 37 | MiniMax-M3 | MiniMax | 0.04 |
| 38 | nemotron-3-ultra | NVIDIA | 0.03 |
| 39 | claude-sonnet-4-6_max | Anthropic | 0.03 |
| 40 | qwen3.6-plus | Alibaba | 0.03 |
| 41 | gemini-2.5-pro | Google DeepMind | 0.02 |
| 42 | gpt-5.5_none | OpenAI | 0.01 |
| 43 | gemma-4-31b-it | Google DeepMind | 0.01 |
| 44 | gpt-oss-20b_high | OpenAI | 0.01 |
| 45 | o3-2025-04-16_high | OpenAI | 0.01 |
| 46 | gemini-3.1-flash-lite | 0.01 | |
| 47 | openai/gpt-oss-120b_high | OpenAI | 0.01 |
| 48 | claude-sonnet-4-5-20250929_unknown | Anthropic | 0.01 |
| 49 | gemini-2.5-flash | Google DeepMind | 0.01 |
| 50 | DeepSeek-R1 | DeepSeek | 0.01 |
| 51 | o4-mini-2025-04-16_high | OpenAI | 0.01 |
| 52 | MiniMax-M2.7 | MiniMax | 0.01 |
| 53 | claude-opus-4-20250514_unknown | Anthropic | 0 |
| 54 | qwen3.5-9B | Alibaba | 0 |
| 55 | qwen3.6-35b-a3b | Alibaba | 0 |
| 56 | claude-sonnet-4-20250514_unknown | Anthropic | 0 |
| 57 | Qwen3-32B | Alibaba | 0 |
| 58 | Llama-3.1-8B-Instruct | Meta AI | 0 |
| 59 | gpt-4.1-mini-2025-04-14 | OpenAI | 0 |
| 60 | gpt-5-2025-08-07_minimal | OpenAI | 0 |