Live
Benchmarks

Frontiercode

Models scored
14
Best score
0.54
Metric
Main score

Frontiercode is an AI evaluation tracked by Epoch AI, with 14 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span April 2026 to July 2026. The best score recorded is 0.54, measured as Main score. Models come from Anthropic, OpenAI, xAI, Moonshot, Z.ai (Zhipu AI), DeepSeek, MiniMax.

Full record
Score metric
Main score
Models scored
14
Best score recorded
0.54
Earliest model scored
Apr 16, 2026
Most recent model scored
Jul 24, 2026
Third-party reported
Yes
Leaderboard
01claude-fable-5_unknownAnthropic0.54
02claude-opus-5_maxAnthropic0.53
03gpt-5.6-sol_unknownOpenAI0.47
04claude-opus-4-8_unknownAnthropic0.47
05gpt-5.5_unknownOpenAI0.43
06claude-sonnet-5_unknownAnthropic0.43
07grok-4.5_unknownxAI0.42
08gpt-5.6-terra_unknownOpenAI0.41
09gpt-5.6-luna_unknownOpenAI0.4
10claude-opus-4-7_unknownAnthropic0.39
11kimi-k2.7-codeMoonshot0.3
12glm-5.2_noneZ.ai (Zhipu AI)0.24
13deepseek-v4-pro_noneDeepSeek0.18
14MiniMax-M3MiniMax0.15
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks