LiveBench is an AI evaluation tracked by Epoch AI, with 54 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span February 2024 to November 2025. The best score recorded is 82.35, measured as Global average. Models come from Google DeepMind, Anthropic, OpenAI, Alibaba, DeepSeek, Google DeepMind,Google, xAI, Meta AI and others.
| 01 | gemini-2.5-pro-exp-03-25 | Google DeepMind | 82.3 |
| 02 | gpt-5.1-2025-11-13_high | OpenAI | 78.8 |
| 03 | claude-3-7-sonnet-20250219 | Anthropic | 76.1 |
| 04 | o3-mini-2025-01-31_high | OpenAI | 75.9 |
| 05 | o1-2024-12-17_high | OpenAI | 75.7 |
| 06 | QwQ-32B | Alibaba | 72 |
| 07 | DeepSeek-R1 | DeepSeek | 71.6 |
| 08 | o3-mini-2025-01-31_medium | OpenAI | 70 |
| 09 | gpt-4.5-preview-2025-02-27 | OpenAI | 69 |
| 10 | gemini-2.0-flash-thinking-exp-01-21 | Google DeepMind,Google | 66.9 |
| 11 | DeepSeek-V3-0324 | DeepSeek | 66.9 |
| 12 | gemini-2.0-pro-exp-02-05 | Google DeepMind | 65.1 |
| 13 | gemini-exp-1206 | Google DeepMind,Google | 64.1 |
| 14 | o3-mini-2025-01-31_low | OpenAI | 62.5 |
| 15 | qwen2.5-max | Alibaba | 62.3 |
| 16 | gemini-2.0-flash-001 | Google DeepMind,Google | 61.5 |
| 17 | DeepSeek-V3 | DeepSeek | 60.5 |
| 18 | gemini-2.0-flash-exp | Google DeepMind,Google | 59.3 |
| 19 | claude-3-5-sonnet-20241022 | Anthropic | 59 |
| 20 | o1-mini-2024-09-12_medium | OpenAI | 57.8 |
| 21 | gpt-4o-2024-08-06 | OpenAI | 55.3 |
| 22 | DeepSeek-R1-Distill-Llama-70B | DeepSeek | 54.5 |
| 23 | grok-2-1212 | xAI | 54.3 |
| 24 | gemini-2.0-flash-lite | Google DeepMind | 54.3 |
| 25 | gemini-2.0-flash-lite-preview-02-05 | Google DeepMind | 53.2 |
| 26 | Dracarys2-72B-Instruct | 52.6 | |
| 27 | gpt-4o-2024-11-20 | OpenAI | 52.2 |
| 28 | learnlm-1.5-pro-experimental | 52.2 | |
| 29 | Llama-3.3-70B-Instruct | Meta AI | 50.2 |
| 30 | gemma-3-27b-it | Google DeepMind | 50 |
| 31 | claude-3-opus-20240229 | Anthropic | 49.2 |
| 32 | mistral-large-2411 | Mistral AI | 48.4 |
| 33 | sonar | 46.9 | |
| 34 | Qwen2.5-Coder-32B-Instruct | Alibaba | 46.2 |
| 35 | Dracarys2-Llama-3.1-70B-Instruct | 46.2 | |
| 36 | DeepSeek-R1-Distill-Qwen-32B | DeepSeek | 45.5 |
| 37 | mistral-small-2503 | Mistral AI | 44 |
| 38 | amazon.nova-pro-v1:0 | Amazon | 43.5 |
| 39 | claude-3-5-haiku-20241022 | Anthropic | 43.5 |
| 40 | mistral-small-2501 | Mistral AI | 42.5 |
| 41 | phi-4 | Microsoft Research | 41.6 |
| 42 | gpt-4o-mini-2024-07-18 | OpenAI | 41.3 |
| 43 | QwQ-32B-Preview | Alibaba | 40.3 |
| 44 | gemma-2-27b-it | Google DeepMind | 38.2 |
| 45 | amazon.nova-lite-v1:0 | Amazon | 36.4 |
| 46 | c4ai-command-r-plus-08-2024 | Cohere,Cohere for AI | 31.8 |
| 47 | amazon.nova-micro-v1:0 | Amazon | 29.6 |
| 48 | gemma-2-9b-it | Google DeepMind | 28.7 |
| 49 | c4ai-command-r-08-2024 | 27.5 | |
| 50 | Phi-3-small-8k-instruct | Microsoft | 24 |
| 51 | Phi-3-mini-4k-instruct | Microsoft | 22.4 |
| 52 | OLMo-2-1124-13B-Instruct | Allen Institute for AI,University of Washington,New York University (NYU) | 22.1 |