Live
Benchmarks

Posttrainbench

Models scored
30
Best score
0.34
Metric
Average (%)

Posttrainbench is an AI evaluation tracked by Epoch AI, with 30 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span September 2025 to June 2026. The best score recorded is 0.34, measured as Average (%). Models come from Z.ai (Zhipu AI), Anthropic, OpenAI, Google DeepMind, Moonshot, MiniMax, Alibaba.

Full record
Score metric
Average (%)
Models scored
30
Best score recorded
0.34
Earliest model scored
Sep 24, 2025
Most recent model scored
Jun 16, 2026
Third-party reported
Yes
Leaderboard
01glm-5.2_maxZ.ai (Zhipu AI)0.34
02claude-opus-4-8_maxAnthropic0.34
03claude-opus-4-8_highAnthropic0.34
04claude-opus-4-7_xhighAnthropic0.29
05gpt-5.5_xhighOpenAI0.25
06claude-opus-4-6_unknownAnthropic0.25
07claude-opus-4-6Anthropic0.23
08claude-opus-4-6_highAnthropic0.23
09gemini-3.1-pro-previewGoogle DeepMind0.22
10gpt-5.2-2025-12-11_mediumOpenAI0.21
11gpt-5.2-2025-12-11_noneOpenAI0.21
12gpt-5.4-2026-03-05_highOpenAI0.2
13gpt-5.1-codex-maxOpenAI0.2
14gemini-3-pro-previewGoogle DeepMind0.18
15gpt-5.3-codex_highOpenAI0.18
16claude-opus-4-5-20251101Anthropic0.17
17gpt-5.2-codexOpenAI0.17
18claude-sonnet-4-6Anthropic0.16
19glm-5Z.ai (Zhipu AI)0.14
20gpt-5.3-codexOpenAI0.14
21kimi-k2.5Moonshot0.1
22claude-sonnet-4-5-20250929Anthropic0.1
23MiniMax-M2.5MiniMax0.1
24MiniMax-M2.1MiniMax0.09
25glm-4.7Z.ai (Zhipu AI)0.07
26qwen3-max-2025-09-23Alibaba0.07
27kimi-k2-thinkingMoonshot0.07
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks