Live
Benchmarks

SWE-bench Verified

Models scored
35
Best score
0.83
Metric
mean_score

SWE-bench Verified is an AI evaluation tracked by Epoch AI, with 35 scored model versions on record. Epoch AI runs this evaluation directly, so the scores are reproducible against their published logs.

Scored models span November 2024 to June 2026. The best score recorded is 0.83, measured as mean_score. Models come from Z.ai (Zhipu AI), Alibaba, DeepSeek, Google, Moonshot, OpenAI, Anthropic, Google DeepMind.

Full record
Score metric
mean_score
Models scored
35
Best score recorded
0.83
Earliest model scored
Nov 20, 2024
Most recent model scored
Jun 16, 2026
Third-party reported
No
Leaderboard
01claude-opus-4-7_maxAnthropic0.83
02gpt-5.5-pre-release_xhighOpenAI0.81
03gemini-3.5-flash_highGoogle0.79
04claude-opus-4-6Anthropic0.79
05glm-5.2_maxZ.ai (Zhipu AI)0.79
06deepseek-v4-pro_maxDeepSeek0.78
07qwen3.7-maxAlibaba0.77
08gpt-5.4-2026-03-05_highOpenAI0.77
09claude-opus-4-5-20251101Anthropic0.77
10qwen3.6-max-previewAlibaba0.77
11kimi-k2.6Moonshot0.77
12gemini-3.1-pro-preview-customtoolsGoogle DeepMind0.76
13gemini-3-flash-previewGoogle DeepMind0.75
14claude-sonnet-4-6Anthropic0.75
15gpt-5.3-codex_highOpenAI0.75
16glm-5.1Z.ai (Zhipu AI)0.74
17kimi-k2.5Moonshot0.74
18gpt-5.2-2025-12-11_highOpenAI0.74
19gpt-5-2025-08-07_highOpenAI0.74
20claude-opus-4-1-20250805Anthropic0.73
21gemini-3-pro-previewGoogle DeepMind0.73
22glm-5Z.ai (Zhipu AI)0.72
23gpt-5-2025-08-07_mediumOpenAI0.71
24claude-sonnet-4-5-20250929Anthropic0.71
25claude-opus-4-20250514Anthropic0.71
26gpt-5.1-2025-11-13_highOpenAI0.68
27gpt-5-mini-2025-08-07_mediumOpenAI0.65
28o3-2025-04-16_mediumOpenAI0.62
29claude-3-7-sonnet-20250219Anthropic0.61
30qwen3.6-plusAlibaba0.58
31gemini-2.5-proGoogle DeepMind0.58
32gpt-4.1-2025-04-14OpenAI0.49
33gpt-4o-2024-11-20OpenAI0.31
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks