라이브
Benchmarks

GSO

Models scored
36
Best score
0.44
Metric
Score OPT@1

GSO is an AI evaluation tracked by Epoch AI, with 36 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span October 2024 to April 2026. The best score recorded is 0.44, measured as Score OPT@1. Models come from Anthropic, OpenAI, Google DeepMind, Moonshot, Alibaba, Z.ai (Zhipu AI),Tsinghua University.

Full record
Score metric
Score OPT@1
Models scored
36
Best score recorded
0.44
Earliest model scored
Oct 22, 2024
Most recent model scored
Apr 23, 2026
Third-party reported
Yes
Leaderboard
01claude-opus-4-7Anthropic0.44
02claude-opus-4-7_highAnthropic0.44
03claude-opus-4-6_highAnthropic0.41
04gpt-5.5_xhighOpenAI0.4
05claude-opus-4-6Anthropic0.33
06claude-opus-4-6_unknownAnthropic0.33
07gpt-5.4-2026-03-05_xhighOpenAI0.31
08gpt-5.2-2025-12-11_highOpenAI0.27
09claude-opus-4-5-20251101Anthropic0.27
10claude-opus-4-5-20251101_unknownAnthropic0.26
11gpt-5.4-2026-03-05_highOpenAI0.25
12gemini-3.1-pro-previewGoogle DeepMind0.23
13gemini-3-pro-previewGoogle DeepMind0.19
14claude-sonnet-4-5-20250929_unknownAnthropic0.15
15claude-sonnet-4-5-20250929Anthropic0.15
16gpt-5.1-2025-11-13_unknownOpenAI0.14
17gpt-5.1-2025-11-13_highOpenAI0.14
18gemini-3-flash-previewGoogle DeepMind0.1
19o3-2025-04-16_highOpenAI0.09
20claude-opus-4-20250514Anthropic0.07
21gpt-5-2025-08-07_highOpenAI0.07
22claude-opus-4-20250514_unknownAnthropic0.07
23Kimi-K2-InstructMoonshot0.05
24Qwen3-Coder-480B-A35B-InstructAlibaba0.05
25claude-sonnet-4-20250514_unknownAnthropic0.05
26claude-sonnet-4-20250514Anthropic0.05
27claude-3-5-sonnet-20241022Anthropic0.05
28gemini-2.5-pro-preview-06-05Google DeepMind0.04
29gemini-2.5-proGoogle DeepMind0.04
30claude-3-7-sonnet-20250219_unknownAnthropic0.04
31claude-3-7-sonnet-20250219Anthropic0.04
32o4-mini-2025-04-16_highOpenAI0.04
33GLM-4.5-AirZ.ai (Zhipu AI),Tsinghua University0.03
34o3-mini-2025-01-31_lowOpenAI0.01
35o3-mini-2025-01-31_highOpenAI0.01
36gpt-4o-2024-11-20OpenAI0
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks