Live
Benchmarks

CL Bench

Models scored
23
Best score
0.28
Metric
Overall

CL Bench is an AI evaluation tracked by Epoch AI, with 23 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span April 2025 to March 2026. The best score recorded is 0.28, measured as Overall. Models come from OpenAI, xAI, Anthropic, Google DeepMind, Alibaba, Moonshot, Z.ai (Zhipu AI), Xiaomi Corp and others.

Full record
Score metric
Overall
Models scored
23
Best score recorded
0.28
Earliest model scored
Apr 16, 2025
Most recent model scored
Mar 31, 2026
Third-party reported
Yes
Leaderboard
01gpt-5.4-2026-03-05_xhighOpenAI0.28
02gpt-5.1-2025-11-13_highOpenAI0.24
03grok-4-20xAI0.22
04gpt-5.1-2025-11-13_unknownOpenAI0.21
05claude-opus-4-5-20251101_unknownAnthropic0.21
06gemini-3.1-pro-previewGoogle DeepMind0.21
07claude-opus-4-6_unknownAnthropic0.21
08qwen3.6-plusAlibaba0.2
09qwen3.5-plusAlibaba0.2
10kimi-k2.5Moonshot0.19
11glm-5Z.ai (Zhipu AI)0.19
12gpt-5.2-2025-12-11_unknownOpenAI0.18
13gpt-5.2-2025-12-11_highOpenAI0.18
14o3-2025-04-16_highOpenAI0.18
15moonshotai/Kimi-K2-ThinkingMoonshot0.18
16zai-org/glm-4.7Z.ai (Zhipu AI)0.16
17gemini-3-pro-previewGoogle DeepMind0.16
18mimo-v2-proXiaomi Corp0.16
19qwen3-max-2025-09-23Alibaba0.14
20DeepSeek-V3.2-Exp_thinkingDeepSeek0.13
21deepseek/deepseek-v3.2DeepSeek0.12
22kimi-k2-thinkingMoonshot0.12
23MiniMax-M2.5MiniMax0.11
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks