MMLU is an AI evaluation tracked by Epoch AI, with 217 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span August 2021 to February 2025. The best score recorded is 0.88, measured as EM. Models come from Amazon, Baichuan, Cerebras Systems, DeepMind, Anthropic, Cohere,Cohere for AI, DeepSeek, Databricks and others.
| 01 | gpt-4o-2024-11-20 | OpenAI | 0.88 |
| 02 | claude-3-5-sonnet-20241022 | Anthropic | 0.87 |
| 03 | DeepSeek-V3 | DeepSeek | 0.87 |
| 04 | gemini-1.5-pro-002 | Google DeepMind | 0.87 |
| 05 | claude-3-5-sonnet-20240620 | Anthropic | 0.86 |
| 06 | gpt-4-0314 | OpenAI | 0.86 |
| 07 | Llama-3.3-70B-Instruct | Meta AI | 0.86 |
| 08 | gemini-1.5-pro-001 | Google DeepMind | 0.86 |
| 09 | qwen2.5-72b-instruct | Alibaba | 0.85 |
| 10 | Qwen2.5-72B | Alibaba | 0.85 |
| 11 | phi-4 | Microsoft Research | 0.85 |
| 12 | claude-3-opus-20240229 | Anthropic | 0.85 |
| 13 | Llama-3.1-405B-Instruct | Meta AI | 0.84 |
| 14 | Llama-3.1-405B | Meta AI | 0.84 |
| 15 | gpt-4o-2024-08-06 | OpenAI | 0.84 |
| 16 | gpt-4o-2024-05-13 | OpenAI | 0.84 |
| 17 | gemini-1.5-pro-001-feb24 | Google DeepMind | 0.83 |
| 18 | gpt-4-0613 | OpenAI | 0.82 |
| 19 | qwen2-72b-instruct | Alibaba | 0.82 |
| 20 | amazon.nova-pro-v1:0 | Amazon | 0.82 |
| 21 | gpt-4o-mini-2024-07-18 | OpenAI | 0.82 |
| 22 | gpt-4-turbo-2024-04-09 | OpenAI | 0.81 |
| 23 | Llama-3.2-90B-Vision-Instruct | Meta AI | 0.8 |
| 24 | Llama-3.1-70B-Instruct | Meta AI | 0.8 |
| 25 | mistral-large-2407 | Mistral AI | 0.8 |
| 26 | qwen2.5-14b-instruct | Alibaba | 0.8 |
| 27 | gemini-2.0-flash-exp | Google DeepMind,Google | 0.8 |
| 28 | gpt-4-turbo | OpenAI | 0.8 |
| 29 | Yi-large | 01.AI | 0.79 |
| 30 | Meta-Llama-3-70B-Instruct | Meta AI | 0.79 |
| 31 | Qwen2.5-Coder-32B | Alibaba | 0.79 |
| 32 | claude-2.0 | Anthropic | 0.79 |
| 33 | DeepSeek-V2 | DeepSeek | 0.78 |
| 34 | Phi-3-medium-128k-instruct | Microsoft | 0.78 |
| 35 | gemini-1.5-flash-001 | Google DeepMind | 0.78 |
| 36 | Mixtral-8x22B-v0.1 | Mistral AI | 0.78 |
| 37 | gemini-1.5-flash-0514 | Google DeepMind | 0.78 |
| 38 | amazon.nova-lite-v1:0 | Amazon | 0.77 |
| 39 | claude-1.3 | Anthropic | 0.77 |
| 40 | Yi-34B | 01.AI | 0.76 |
| 41 | claude-3-sonnet-20240229 | Anthropic | 0.76 |
| 42 | gemma-2-27b-it | Google DeepMind | 0.76 |
| 43 | Phi-3-small-8k-instruct | Microsoft | 0.76 |
| 44 | Qwen2.5-Coder-14B | 0.75 | |
| 45 | qwen1.5-32B | Alibaba | 0.74 |
| 46 | claude-3-5-haiku-20241022 | Anthropic | 0.74 |
| 47 | gemini-1.5-flash-002 | Google DeepMind | 0.74 |
| 48 | claude-3-haiku-20240307 | Anthropic | 0.74 |
| 49 | Yi-34B-Chat | 01.AI | 0.73 |
| 50 | claude-2.1 | Anthropic | 0.73 |
| 51 | claude-instant-1.1 | Anthropic | 0.73 |