Live
Benchmarks

Vending-Bench 2

Models scored
54
Best score
10936.8
Metric
Score

Vending-Bench 2 is an AI evaluation tracked by Epoch AI, with 54 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span June 2025 to July 2026. The best score recorded is 10936.76, measured as Score. Models come from Anthropic, OpenAI, Z.ai (Zhipu AI), Moonshot, Google DeepMind, Google, Alibaba, xAI and others.

Full record
Score metric
Score
Models scored
54
Best score recorded
10936.8
Earliest model scored
Jun 17, 2025
Most recent model scored
Jul 9, 2026
Third-party reported
Yes
Leaderboard
01claude-opus-4-7_unknownAnthropic10936.8
02gpt-5.6-sol_unknownOpenAI9619.4
03glm-5.2_unknownZ.ai (Zhipu AI)8313.8
04claude-opus-4-6_unknownAnthropic8017.6
05gpt-5.5_unknownOpenAI7523.8
06gpt-5.6-terra_unknownOpenAI7343.2
07claude-sonnet-4-6_unknownAnthropic7204.1
08claude-sonnet-5_unknownAnthropic6377.7
09kimi-k2.6Moonshot6204.6
10gpt-5.4-2026-03-05_unknownOpenAI6144.2
11gpt-5.3-codexOpenAI5940.1
12claude-opus-4-8_unknownAnthropic5787.4
13claude-fable-5_highAnthropic5680.3
14glm-5.1Z.ai (Zhipu AI)5634.4
15gemini-3-pro-previewGoogle DeepMind5478.2
16gemini-3.5-flash_unknownGoogle5396.4
17qwen3.6-plusAlibaba5114.9
18kimi-k2.7-codeMoonshot5082.9
19claude-fable-5_lowAnthropic5018.5
20claude-opus-4-5-20251101_unknownAnthropic4967.1
21claude-fable-5_maxAnthropic4966.6
22grok-4-20xAI4662.8
23claude-fable-5Anthropic4529.9
24glm-5Z.ai (Zhipu AI)4432.1
25claude-fable-5_mediumAnthropic4339.8
26qwen3.6-max-previewAlibaba4254.2
27gpt-5.6-luna_unknownOpenAI4094.7
28grok-4.5_unknownxAI3887.4
29claude-sonnet-4-5-20250929_unknownAnthropic3838.7
30gemini-3.1-pro-preview-customtoolsGoogle DeepMind3774.3
31gemini-3-flash-previewGoogle DeepMind3634.7
32gpt-5.2-2025-12-11_unknownOpenAI3591.3
33deepseek-v4-pro_unknownDeepSeek3284.5
34claude-opus-4-8_maxAnthropic2992.3
35glm-4.7Z.ai (Zhipu AI)2376.8
36MiniMax-M3MiniMax2157.8
37gpt-5.1-2025-11-13_unknownOpenAI1473.4
38kimi-k2.5Moonshot1198.5
39grok-4-1-fast-reasoningxAI1106.6
40DeepSeek-V3.2-Exp_unknownDeepSeek1034
41gemini-3.1-pro-previewGoogle DeepMind911.2
42gemini-2.5-proGoogle DeepMind573.6
43gemini-2.5-flashGoogle DeepMind548.8
44qwen3.5-flashAlibaba462.7
45claude-haiku-4-5-20251001_unknownAnthropic458.9
46qwen3.5-27BAlibaba202
47MiniMax-M2MiniMax160.6
48qwen3-max-2025-09-23Alibaba71.6
49grok-4-3xAI35.3
50qwen3.5-plusAlibaba0.54
51Qwen3-235B-A22B-Thinking-2507Alibaba-11.34
52gpt-oss-120bOpenAI-21.53
53MiniMax-M2.5MiniMax-23.16
54gpt-5-mini-2025-08-07_unknownOpenAI-31.18
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks