GSM8K is an AI evaluation tracked by Epoch AI, with 170 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span June 2020 to November 2024. The best score recorded is 0.94, measured as EM. Models come from OpenAI, Meta AI, Mistral AI, MosaicML, Technology Innovation Institute, Large Model Systems Organization,University of California (UC) Berkeley, Hugging Face,BigScience, Google DeepMind and others.
| 01 | DeepSeek-Coder-V2-Instruct | DeepSeek | 0.94 |
| 02 | Qwen2.5-Coder-14B-Instruct | 0.94 | |
| 03 | Qwen2.5-Coder-32B-Instruct | Alibaba | 0.93 |
| 04 | gpt-4-0314 | OpenAI | 0.92 |
| 05 | gpt-4o-mini-2024-07-18 | OpenAI | 0.91 |
| 06 | Qwen2.5-Coder-32B | Alibaba | 0.91 |
| 07 | gpt-4-0613 | OpenAI | 0.9 |
| 08 | Qwen2.5-Coder-14B | 0.89 | |
| 09 | Phi-3.5-MoE-instruct | Microsoft | 0.89 |
| 10 | DeepSeek-Coder-V2-Lite-Instruct | 0.88 | |
| 11 | claude-instant-1.2 | Anthropic | 0.87 |
| 12 | Qwen2.5-Coder-7B-Instruct | Alibaba | 0.87 |
| 13 | Phi-3.5-mini-instruct | Microsoft | 0.86 |
| 14 | DeepSeek-Coder-V2-Base | DeepSeek | 0.86 |
| 15 | gemma-2-9b | Google DeepMind | 0.85 |
| 16 | Mistral-Nemo-Base-2407 | Mistral AI | 0.84 |
| 17 | Qwen2.5-Coder-7B | Alibaba | 0.84 |
| 18 | Llama-3.1-8B-Instruct | Meta AI | 0.82 |
| 19 | gemini-1.5-flash-001 | Google DeepMind | 0.82 |
| 20 | claude-instant-1.1 | Anthropic | 0.81 |
| 21 | Qwen2.5-Coder-3B-Instruct | 0.81 | |
| 22 | text-davinci-003 | 0.78 | |
| 23 | Yi-34B-Chat | 01.AI | 0.76 |
| 24 | Qwen2.5-Coder-3B | 0.76 | |
| 25 | Mixtral-8x7B-v0.1 | Mistral AI | 0.74 |
| 26 | StableBeluga2 | Stability AI | 0.7 |
| 27 | Llama-2-70b-hf | Meta AI | 0.7 |
| 28 | Yi-34B | 01.AI | 0.67 |
| 29 | DeepSeek-Coder-V2-Lite-Base | 0.67 | |
| 30 | Qwen2.5-Coder-1.5B | Alibaba | 0.66 |
| 31 | internlm-20b | 0.63 | |
| 32 | Qwen-14B | Alibaba | 0.61 |
| 33 | Qwen-14B-Chat | Alibaba | 0.61 |
| 34 | Llama-2-70b-chat | Meta AI | 0.59 |
| 35 | gpt-3.5-turbo-0613 | OpenAI | 0.58 |
| 36 | starcoder2-15b | Hugging Face,ServiceNow,NVIDIA,BigCode | 0.58 |
| 37 | code-davinci-002 | OpenAI | 0.57 |
| 38 | PaLM 540B | Google Research | 0.56 |
| 39 | LLaMA-65B | Meta AI | 0.54 |
| 40 | Mistral-7B-v0.1 | Mistral AI | 0.54 |
| 41 | falcon-180B | Technology Innovation Institute | 0.54 |
| 42 | falcon-11b | Technology Innovation Institute | 0.54 |
| 43 | Baichuan-2-13B-Base | Baichuan | 0.53 |
| 44 | Qwen-7B | Alibaba | 0.52 |