Vending-Bench 2 is an AI evaluation tracked by Epoch AI, with 54 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span June 2025 to July 2026. The best score recorded is 10936.76, measured as Score. Models come from Anthropic, OpenAI, Z.ai (Zhipu AI), Moonshot, Google DeepMind, Google, Alibaba, xAI and others.
| 01 | claude-opus-4-7_unknown | Anthropic | 10936.8 |
| 02 | gpt-5.6-sol_unknown | OpenAI | 9619.4 |
| 03 | glm-5.2_unknown | Z.ai (Zhipu AI) | 8313.8 |
| 04 | claude-opus-4-6_unknown | Anthropic | 8017.6 |
| 05 | gpt-5.5_unknown | OpenAI | 7523.8 |
| 06 | gpt-5.6-terra_unknown | OpenAI | 7343.2 |
| 07 | claude-sonnet-4-6_unknown | Anthropic | 7204.1 |
| 08 | claude-sonnet-5_unknown | Anthropic | 6377.7 |
| 09 | kimi-k2.6 | Moonshot | 6204.6 |
| 10 | gpt-5.4-2026-03-05_unknown | OpenAI | 6144.2 |
| 11 | gpt-5.3-codex | OpenAI | 5940.1 |
| 12 | claude-opus-4-8_unknown | Anthropic | 5787.4 |
| 13 | claude-fable-5_high | Anthropic | 5680.3 |
| 14 | glm-5.1 | Z.ai (Zhipu AI) | 5634.4 |
| 15 | gemini-3-pro-preview | Google DeepMind | 5478.2 |
| 16 | gemini-3.5-flash_unknown | 5396.4 | |
| 17 | qwen3.6-plus | Alibaba | 5114.9 |
| 18 | kimi-k2.7-code | Moonshot | 5082.9 |
| 19 | claude-fable-5_low | Anthropic | 5018.5 |
| 20 | claude-opus-4-5-20251101_unknown | Anthropic | 4967.1 |
| 21 | claude-fable-5_max | Anthropic | 4966.6 |
| 22 | grok-4-20 | xAI | 4662.8 |
| 23 | claude-fable-5 | Anthropic | 4529.9 |
| 24 | glm-5 | Z.ai (Zhipu AI) | 4432.1 |
| 25 | claude-fable-5_medium | Anthropic | 4339.8 |
| 26 | qwen3.6-max-preview | Alibaba | 4254.2 |
| 27 | gpt-5.6-luna_unknown | OpenAI | 4094.7 |
| 28 | grok-4.5_unknown | xAI | 3887.4 |
| 29 | claude-sonnet-4-5-20250929_unknown | Anthropic | 3838.7 |
| 30 | gemini-3.1-pro-preview-customtools | Google DeepMind | 3774.3 |
| 31 | gemini-3-flash-preview | Google DeepMind | 3634.7 |
| 32 | gpt-5.2-2025-12-11_unknown | OpenAI | 3591.3 |
| 33 | deepseek-v4-pro_unknown | DeepSeek | 3284.5 |
| 34 | claude-opus-4-8_max | Anthropic | 2992.3 |
| 35 | glm-4.7 | Z.ai (Zhipu AI) | 2376.8 |
| 36 | MiniMax-M3 | MiniMax | 2157.8 |
| 37 | gpt-5.1-2025-11-13_unknown | OpenAI | 1473.4 |
| 38 | kimi-k2.5 | Moonshot | 1198.5 |
| 39 | grok-4-1-fast-reasoning | xAI | 1106.6 |
| 40 | DeepSeek-V3.2-Exp_unknown | DeepSeek | 1034 |
| 41 | gemini-3.1-pro-preview | Google DeepMind | 911.2 |
| 42 | gemini-2.5-pro | Google DeepMind | 573.6 |
| 43 | gemini-2.5-flash | Google DeepMind | 548.8 |
| 44 | qwen3.5-flash | Alibaba | 462.7 |
| 45 | claude-haiku-4-5-20251001_unknown | Anthropic | 458.9 |
| 46 | qwen3.5-27B | Alibaba | 202 |
| 47 | MiniMax-M2 | MiniMax | 160.6 |
| 48 | qwen3-max-2025-09-23 | Alibaba | 71.6 |
| 49 | grok-4-3 | xAI | 35.3 |
| 50 | qwen3.5-plus | Alibaba | 0.54 |
| 51 | Qwen3-235B-A22B-Thinking-2507 | Alibaba | -11.34 |
| 52 | gpt-oss-120b | OpenAI | -21.53 |
| 53 | MiniMax-M2.5 | MiniMax | -23.16 |
| 54 | gpt-5-mini-2025-08-07_unknown | OpenAI | -31.18 |