Live
Benchmarks

LiveBench

Models scored
54
Best score
82.3
Metric
Global average

LiveBench is an AI evaluation tracked by Epoch AI, with 54 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span February 2024 to November 2025. The best score recorded is 82.35, measured as Global average. Models come from Google DeepMind, Anthropic, OpenAI, Alibaba, DeepSeek, Google DeepMind,Google, xAI, Meta AI and others.

Full record
Score metric
Global average
Models scored
54
Best score recorded
82.3
Earliest model scored
Feb 29, 2024
Most recent model scored
Nov 13, 2025
Third-party reported
Yes
Leaderboard
01gemini-2.5-pro-exp-03-25Google DeepMind82.3
02gpt-5.1-2025-11-13_highOpenAI78.8
03claude-3-7-sonnet-20250219Anthropic76.1
04o3-mini-2025-01-31_highOpenAI75.9
05o1-2024-12-17_highOpenAI75.7
06QwQ-32BAlibaba72
07DeepSeek-R1DeepSeek71.6
08o3-mini-2025-01-31_mediumOpenAI70
09gpt-4.5-preview-2025-02-27OpenAI69
10gemini-2.0-flash-thinking-exp-01-21Google DeepMind,Google66.9
11DeepSeek-V3-0324DeepSeek66.9
12gemini-2.0-pro-exp-02-05Google DeepMind65.1
13gemini-exp-1206Google DeepMind,Google64.1
14o3-mini-2025-01-31_lowOpenAI62.5
15qwen2.5-maxAlibaba62.3
16gemini-2.0-flash-001Google DeepMind,Google61.5
17DeepSeek-V3DeepSeek60.5
18gemini-2.0-flash-expGoogle DeepMind,Google59.3
19claude-3-5-sonnet-20241022Anthropic59
20o1-mini-2024-09-12_mediumOpenAI57.8
21gpt-4o-2024-08-06OpenAI55.3
22DeepSeek-R1-Distill-Llama-70BDeepSeek54.5
23grok-2-1212xAI54.3
24gemini-2.0-flash-liteGoogle DeepMind54.3
25gemini-2.0-flash-lite-preview-02-05Google DeepMind53.2
26Dracarys2-72B-Instruct52.6
27gpt-4o-2024-11-20OpenAI52.2
28learnlm-1.5-pro-experimental52.2
29Llama-3.3-70B-InstructMeta AI50.2
30gemma-3-27b-itGoogle DeepMind50
31claude-3-opus-20240229Anthropic49.2
32mistral-large-2411Mistral AI48.4
33sonar46.9
34Qwen2.5-Coder-32B-InstructAlibaba46.2
35Dracarys2-Llama-3.1-70B-Instruct46.2
36DeepSeek-R1-Distill-Qwen-32BDeepSeek45.5
37mistral-small-2503Mistral AI44
38amazon.nova-pro-v1:0Amazon43.5
39claude-3-5-haiku-20241022Anthropic43.5
40mistral-small-2501Mistral AI42.5
41phi-4Microsoft Research41.6
42gpt-4o-mini-2024-07-18OpenAI41.3
43QwQ-32B-PreviewAlibaba40.3
44gemma-2-27b-itGoogle DeepMind38.2
45amazon.nova-lite-v1:0Amazon36.4
46c4ai-command-r-plus-08-2024Cohere,Cohere for AI31.8
47amazon.nova-micro-v1:0Amazon29.6
48gemma-2-9b-itGoogle DeepMind28.7
49c4ai-command-r-08-202427.5
50Phi-3-small-8k-instructMicrosoft24
51Phi-3-mini-4k-instructMicrosoft22.4
52OLMo-2-1124-13B-InstructAllen Institute for AI,University of Washington,New York University (NYU)22.1
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks