Live
Benchmarks

Surface Evolver Bench

Models scored
17
Best score
0.95
Metric
Mean score

Surface Evolver Bench is an AI evaluation tracked by Epoch AI, with 17 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span August 2025 to July 2026. The best score recorded is 0.95, measured as Mean score. Models come from Anthropic, OpenAI, Moonshot, xAI, Z.ai (Zhipu AI), DeepSeek, MiniMax, Alibaba and others.

Full record
Score metric
Mean score
Models scored
17
Best score recorded
0.95
Earliest model scored
Aug 5, 2025
Most recent model scored
Jul 16, 2026
Third-party reported
Yes
Leaderboard
01claude-fable-5_highAnthropic0.95
02gpt-5.6-sol_xhighOpenAI0.93
03kimi-k3_maxMoonshot0.93
04gpt-5.5_highOpenAI0.88
05claude-opus-4-8_highAnthropic0.88
06gpt-5.6-terra_xhighOpenAI0.84
07gpt-5.5_mediumOpenAI0.81
08grok-4.5_highxAI0.74
09claude-opus-4-8_noneAnthropic0.66
10gpt-5.6-luna_mediumOpenAI0.6
11glm-5.2_highZ.ai (Zhipu AI)0.56
12MiniMax-M3MiniMax0.53
13kimi-k2.7-codeMoonshot0.49
14qwen3.6-35b-a3bAlibaba0.44
15deepseek-v4-pro_highDeepSeek0.4
16gemma-4-31b-itGoogle DeepMind0.31
17gpt-oss-120bOpenAI0.25
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks