Live
Benchmarks

Deepswe

Models scored
37
Best score
0.73
Metric
Pass@1

Deepswe is an AI evaluation tracked by Epoch AI, with 37 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span February 2026 to July 2026. The best score recorded is 0.73, measured as Pass@1. Models come from OpenAI, Anthropic, Meta AI, Z.ai (Zhipu AI), Moonshot, Google DeepMind.

Full record
Score metric
Pass@1
Models scored
37
Best score recorded
0.73
Earliest model scored
Feb 17, 2026
Most recent model scored
Jul 9, 2026
Third-party reported
Yes
Leaderboard
01gpt-5.6-sol_maxOpenAI0.73
02gpt-5.6-sol_xhighOpenAI0.71
03claude-fable-5_xhighAnthropic0.7
04claude-fable-5_maxAnthropic0.7
05gpt-5.6-terra_maxOpenAI0.7
06gpt-5.6-sol_highOpenAI0.69
07claude-fable-5_highAnthropic0.69
08gpt-5.6-luna_maxOpenAI0.67
09gpt-5.5_xhighOpenAI0.67
10gpt-5.5_highOpenAI0.64
11gpt-5.6-sol_mediumOpenAI0.61
12gpt-5.6-terra_xhighOpenAI0.6
13claude-opus-4-8_maxAnthropic0.59
14gpt-5.6-luna_xhighOpenAI0.57
15claude-opus-4-8_xhighAnthropic0.54
16gpt-5.5_mediumOpenAI0.54
17claude-sonnet-5_maxAnthropic0.54
18gpt-5.6-terra_highOpenAI0.54
19muse-spark-1.1Meta AI0.53
20claude-opus-4-8_highAnthropic0.52
21gpt-5.4-2026-03-05_xhighOpenAI0.52
22claude-sonnet-5_xhighAnthropic0.5
23claude-opus-4-8_mediumAnthropic0.49
24claude-sonnet-5_highAnthropic0.48
25gpt-5.6-sol_lowOpenAI0.45
26gpt-5.6-luna_highOpenAI0.44
27glm-5.2_maxZ.ai (Zhipu AI)0.44
28claude-opus-4-8_lowAnthropic0.41
29glm-5.2_highZ.ai (Zhipu AI)0.36
30gpt-5.6-terra_mediumOpenAI0.35
31kimi-k2.7-codeMoonshot0.31
32claude-sonnet-4-6_highAnthropic0.3
33gpt-5.5_lowOpenAI0.27
34gpt-5.6-terra_lowOpenAI0.24
35gemini-3.1-pro-previewGoogle DeepMind0.12
36gpt-5.6-luna_mediumOpenAI0.11
37gpt-5.6-luna_lowOpenAI0.02
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks