HellaSwag is an AI evaluation tracked by Epoch AI, with 113 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span November 2019 to December 2024. The best score recorded is 0.95, measured as Overall accuracy. Models come from OpenAI, Baichuan, Hugging Face,BigScience, Cerebras Systems, DeepSeek, Databricks, Technology Innovation Institute, Google DeepMind and others.
| 01 | gpt-4-0314 | OpenAI | 0.95 |
| 02 | gpt-4-32k-0314 | OpenAI | 0.95 |
| 03 | Llama-3.1-405B | Meta AI | 0.89 |
| 04 | falcon-180B | Technology Innovation Institute | 0.89 |
| 05 | DeepSeek-V3 | DeepSeek | 0.89 |
| 06 | DeepSeek-V2 | DeepSeek | 0.87 |
| 07 | PaLM 2-L | 0.87 | |
| 08 | Mixtral-8x7B-v0.1 | Mistral AI | 0.87 |
| 09 | text-davinci-003 | 0.85 | |
| 10 | Llama-2-70b-hf | Meta AI | 0.85 |
| 11 | falcon-40b | Technology Innovation Institute | 0.85 |
| 12 | Qwen2.5-72B | Alibaba | 0.85 |
| 13 | LLaMA-65B | Meta AI | 0.84 |
| 14 | StableBeluga2 | Stability AI | 0.84 |
| 15 | PaLM 2-M | 0.84 | |
| 16 | PaLM 540B | Google Research | 0.84 |
| 17 | Qwen2.5-Coder-32B | Alibaba | 0.83 |
| 18 | falcon-11b | Technology Innovation Institute | 0.83 |
| 19 | LLaMA-33B | Meta AI | 0.83 |
| 20 | Phi-3-medium-128k-instruct | Microsoft | 0.82 |
| 21 | Megatron-Turing NLG 530B | Microsoft,NVIDIA | 0.82 |
| 22 | Nemotron-4 15B | NVIDIA | 0.82 |
| 23 | gemma-7b | Google DeepMind | 0.82 |
| 24 | PaLM 2-S | 0.82 | |
| 25 | text-davinci-002 | OpenAI | 0.81 |
| 26 | Mistral-7B-v0.1 | Mistral AI | 0.81 |
| 27 | Llama-2-13b | Meta AI | 0.81 |
| 28 | Qwen2.5-Coder-14B | 0.8 | |
| 29 | text-davinci-001 | OpenAI | 0.79 |
| 30 | LLaMA-13B | Meta AI | 0.79 |
| 31 | Gopher (280B) | DeepMind | 0.79 |
| 32 | opt-175b | Meta AI | 0.79 |
| 33 | falcon-7b | Technology Innovation Institute | 0.78 |
| 34 | internlm-20b | 0.78 | |
| 35 | davinci | OpenAI | 0.78 |