Live
Benchmarks

GSM8K

Models scored
170
Best score
0.94
Metric
EM

GSM8K is an AI evaluation tracked by Epoch AI, with 170 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span June 2020 to November 2024. The best score recorded is 0.94, measured as EM. Models come from OpenAI, Meta AI, Mistral AI, MosaicML, Technology Innovation Institute, Large Model Systems Organization,University of California (UC) Berkeley, Hugging Face,BigScience, Google DeepMind and others.

Full record
Score metric
EM
Models scored
170
Best score recorded
0.94
Earliest model scored
Jun 22, 2020
Most recent model scored
Nov 29, 2024
Third-party reported
Yes
Leaderboard
01DeepSeek-Coder-V2-InstructDeepSeek0.94
02Qwen2.5-Coder-14B-Instruct0.94
03Qwen2.5-Coder-32B-InstructAlibaba0.93
04gpt-4-0314OpenAI0.92
05gpt-4o-mini-2024-07-18OpenAI0.91
06Qwen2.5-Coder-32BAlibaba0.91
07gpt-4-0613OpenAI0.9
08Qwen2.5-Coder-14B0.89
09Phi-3.5-MoE-instructMicrosoft0.89
10DeepSeek-Coder-V2-Lite-Instruct0.88
11claude-instant-1.2Anthropic0.87
12Qwen2.5-Coder-7B-InstructAlibaba0.87
13Phi-3.5-mini-instructMicrosoft0.86
14DeepSeek-Coder-V2-BaseDeepSeek0.86
15gemma-2-9bGoogle DeepMind0.85
16Mistral-Nemo-Base-2407Mistral AI0.84
17Qwen2.5-Coder-7BAlibaba0.84
18Llama-3.1-8B-InstructMeta AI0.82
19gemini-1.5-flash-001Google DeepMind0.82
20claude-instant-1.1Anthropic0.81
21Qwen2.5-Coder-3B-Instruct0.81
22text-davinci-0030.78
23Yi-34B-Chat01.AI0.76
24Qwen2.5-Coder-3B0.76
25Mixtral-8x7B-v0.1Mistral AI0.74
26StableBeluga2Stability AI0.7
27Llama-2-70b-hfMeta AI0.7
28Yi-34B01.AI0.67
29DeepSeek-Coder-V2-Lite-Base0.67
30Qwen2.5-Coder-1.5BAlibaba0.66
31internlm-20b0.63
32Qwen-14BAlibaba0.61
33Qwen-14B-ChatAlibaba0.61
34Llama-2-70b-chatMeta AI0.59
35gpt-3.5-turbo-0613OpenAI0.58
36starcoder2-15bHugging Face,ServiceNow,NVIDIA,BigCode0.58
37code-davinci-002OpenAI0.57
38PaLM 540BGoogle Research0.56
39LLaMA-65BMeta AI0.54
40Mistral-7B-v0.1Mistral AI0.54
41falcon-180BTechnology Innovation Institute0.54
42falcon-11bTechnology Innovation Institute0.54
43Baichuan-2-13B-BaseBaichuan0.53
44Qwen-7BAlibaba0.52
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks