CL Bench is an AI evaluation tracked by Epoch AI, with 23 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span April 2025 to March 2026. The best score recorded is 0.28, measured as Overall. Models come from OpenAI, xAI, Anthropic, Google DeepMind, Alibaba, Moonshot, Z.ai (Zhipu AI), Xiaomi Corp and others.
| 01 | gpt-5.4-2026-03-05_xhigh | OpenAI | 0.28 |
| 02 | gpt-5.1-2025-11-13_high | OpenAI | 0.24 |
| 03 | grok-4-20 | xAI | 0.22 |
| 04 | gpt-5.1-2025-11-13_unknown | OpenAI | 0.21 |
| 05 | claude-opus-4-5-20251101_unknown | Anthropic | 0.21 |
| 06 | gemini-3.1-pro-preview | Google DeepMind | 0.21 |
| 07 | claude-opus-4-6_unknown | Anthropic | 0.21 |
| 08 | qwen3.6-plus | Alibaba | 0.2 |
| 09 | qwen3.5-plus | Alibaba | 0.2 |
| 10 | kimi-k2.5 | Moonshot | 0.19 |
| 11 | glm-5 | Z.ai (Zhipu AI) | 0.19 |
| 12 | gpt-5.2-2025-12-11_unknown | OpenAI | 0.18 |
| 13 | gpt-5.2-2025-12-11_high | OpenAI | 0.18 |
| 14 | o3-2025-04-16_high | OpenAI | 0.18 |
| 15 | moonshotai/Kimi-K2-Thinking | Moonshot | 0.18 |
| 16 | zai-org/glm-4.7 | Z.ai (Zhipu AI) | 0.16 |
| 17 | gemini-3-pro-preview | Google DeepMind | 0.16 |
| 18 | mimo-v2-pro | Xiaomi Corp | 0.16 |
| 19 | qwen3-max-2025-09-23 | Alibaba | 0.14 |
| 20 | DeepSeek-V3.2-Exp_thinking | DeepSeek | 0.13 |
| 21 | deepseek/deepseek-v3.2 | DeepSeek | 0.12 |
| 22 | kimi-k2-thinking | Moonshot | 0.12 |
| 23 | MiniMax-M2.5 | MiniMax | 0.11 |