Live
Benchmarks

Gbaeval

Models scored
16
Best score
0.74
Metric
Overall score

Gbaeval is an AI evaluation tracked by Epoch AI, with 16 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span February 2026 to June 2026. The best score recorded is 0.74, measured as Overall score. Models come from Anthropic, OpenAI, Google, MiniMax, Moonshot, Google DeepMind, Alibaba, Z.ai (Zhipu AI).

Full record
Score metric
Overall score
Models scored
16
Best score recorded
0.74
Earliest model scored
Feb 5, 2026
Most recent model scored
Jun 16, 2026
Third-party reported
Yes
Leaderboard
01claude-fable-5_unknownAnthropic0.74
02claude-opus-4-8_unknownAnthropic0.71
03gpt-5.5_unknownOpenAI0.53
04claude-sonnet-4-6_unknownAnthropic0.49
05claude-opus-4-6_unknownAnthropic0.44
06claude-opus-4-7_unknownAnthropic0.44
07gpt-5.4-2026-03-05_unknownOpenAI0.32
08gemini-3.5-flash_unknownGoogle0.07
09MiniMax-M3MiniMax0.01
10kimi-k2.6Moonshot0.01
11kimi-k2.7-codeMoonshot0.01
12gemini-3.1-pro-previewGoogle DeepMind0.01
13qwen3.7-maxAlibaba0
14glm-5.2_unknownZ.ai (Zhipu AI)0
15MiniMax-M2.7MiniMax0
16glm-5.1Z.ai (Zhipu AI)0
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks