实时
Benchmarks

MMLU

Models scored
217
Best score
0.88
Metric
EM

MMLU is an AI evaluation tracked by Epoch AI, with 217 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span August 2021 to February 2025. The best score recorded is 0.88, measured as EM. Models come from Amazon, Baichuan, Cerebras Systems, DeepMind, Anthropic, Cohere,Cohere for AI, DeepSeek, Databricks and others.

Full record
Score metric
EM
Models scored
217
Best score recorded
0.88
Earliest model scored
Aug 5, 2021
Most recent model scored
Feb 26, 2025
Third-party reported
Yes
Leaderboard
01gpt-4o-2024-11-20OpenAI0.88
02claude-3-5-sonnet-20241022Anthropic0.87
03DeepSeek-V3DeepSeek0.87
04gemini-1.5-pro-002Google DeepMind0.87
05claude-3-5-sonnet-20240620Anthropic0.86
06gpt-4-0314OpenAI0.86
07Llama-3.3-70B-InstructMeta AI0.86
08gemini-1.5-pro-001Google DeepMind0.86
09qwen2.5-72b-instructAlibaba0.85
10Qwen2.5-72BAlibaba0.85
11phi-4Microsoft Research0.85
12claude-3-opus-20240229Anthropic0.85
13Llama-3.1-405B-InstructMeta AI0.84
14Llama-3.1-405BMeta AI0.84
15gpt-4o-2024-08-06OpenAI0.84
16gpt-4o-2024-05-13OpenAI0.84
17gemini-1.5-pro-001-feb24Google DeepMind0.83
18gpt-4-0613OpenAI0.82
19qwen2-72b-instructAlibaba0.82
20amazon.nova-pro-v1:0Amazon0.82
21gpt-4o-mini-2024-07-18OpenAI0.82
22gpt-4-turbo-2024-04-09OpenAI0.81
23Llama-3.2-90B-Vision-InstructMeta AI0.8
24Llama-3.1-70B-InstructMeta AI0.8
25mistral-large-2407Mistral AI0.8
26qwen2.5-14b-instructAlibaba0.8
27gemini-2.0-flash-expGoogle DeepMind,Google0.8
28gpt-4-turboOpenAI0.8
29Yi-large01.AI0.79
30Meta-Llama-3-70B-InstructMeta AI0.79
31Qwen2.5-Coder-32BAlibaba0.79
32claude-2.0Anthropic0.79
33DeepSeek-V2DeepSeek0.78
34Phi-3-medium-128k-instructMicrosoft0.78
35gemini-1.5-flash-001Google DeepMind0.78
36Mixtral-8x22B-v0.1Mistral AI0.78
37gemini-1.5-flash-0514Google DeepMind0.78
38amazon.nova-lite-v1:0Amazon0.77
39claude-1.3Anthropic0.77
40Yi-34B01.AI0.76
41claude-3-sonnet-20240229Anthropic0.76
42gemma-2-27b-itGoogle DeepMind0.76
43Phi-3-small-8k-instructMicrosoft0.76
44Qwen2.5-Coder-14B0.75
45qwen1.5-32BAlibaba0.74
46claude-3-5-haiku-20241022Anthropic0.74
47gemini-1.5-flash-002Google DeepMind0.74
48claude-3-haiku-20240307Anthropic0.74
49Yi-34B-Chat01.AI0.73
50claude-2.1Anthropic0.73
51claude-instant-1.1Anthropic0.73
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks