Cybench is an AI evaluation tracked by Epoch AI, with 22 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span February 2024 to February 2026. The best score recorded is 0.93, measured as Unguided % Solved. Models come from Anthropic, xAI, OpenAI, Mistral AI, Meta AI, Google DeepMind.
| 01 | claude-opus-4-6_unknown | Anthropic | 0.93 |
| 02 | claude-opus-4-5-20251101_unknown | Anthropic | 0.82 |
| 03 | claude-sonnet-4-5-20250929_unknown | Anthropic | 0.6 |
| 04 | claude-sonnet-4-5-20250929 | Anthropic | 0.55 |
| 05 | grok-4-0709 | xAI | 0.43 |
| 06 | claude-opus-4-1-20250805 | Anthropic | 0.42 |
| 07 | grok-4-1 | xAI | 0.39 |
| 08 | claude-opus-4-20250514 | Anthropic | 0.38 |
| 09 | claude-sonnet-4-20250514 | Anthropic | 0.35 |
| 10 | grok-4-fast | xAI | 0.3 |
| 11 | o3-mini-2025-01-31_medium | OpenAI | 0.23 |
| 12 | claude-3-7-sonnet-20250219 | Anthropic | 0.2 |
| 13 | gpt-4.5-preview-2025-02-27 | OpenAI | 0.17 |
| 14 | claude-3-5-sonnet-20240620 | Anthropic | 0.17 |
| 15 | gpt-4o-2024-11-20 | OpenAI | 0.13 |
| 16 | claude-3-opus-20240229 | Anthropic | 0.1 |
| 17 | o1-preview-2024-09-12 | OpenAI | 0.1 |
| 18 | o1-mini-2024-09-12_medium | OpenAI | 0.1 |
| 19 | Llama-3.1-405B-Instruct | Meta AI | 0.07 |
| 20 | gemini-1.5-pro-001-feb24 | Google DeepMind | 0.07 |
| 21 | Mixtral-8x22B-Instruct-v0.1 | Mistral AI | 0.07 |
| 22 | Meta-Llama-3-70B-Instruct | Meta AI | 0.05 |