BIG-Bench Hard is an AI evaluation tracked by Epoch AI, with 88 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span February 2023 to December 2024. The best score recorded is 0.89, measured as Average. Models come from Alibaba, Google DeepMind, Mistral AI, Meta AI, DeepSeek, 01.AI, Baichuan, NVIDIA and others.
| 01 | gemini-1.5-pro-001 | Google DeepMind | 0.89 |
| 02 | DeepSeek-V3 | DeepSeek | 0.88 |
| 03 | gemini-1.5-pro-001-feb24 | Google DeepMind | 0.84 |
| 04 | Llama-3.1-405B | Meta AI | 0.83 |
| 05 | Phi-3-medium-128k-instruct | Microsoft | 0.81 |
| 06 | Qwen2.5-72B | Alibaba | 0.8 |
| 07 | Phi-3-small-8k-instruct | Microsoft | 0.79 |
| 08 | DeepSeek-V2 | DeepSeek | 0.79 |
| 09 | gpt-4-0613 | OpenAI | 0.75 |
| 10 | Yi-34B-Chat | 01.AI | 0.72 |
| 11 | Phi-3-mini-4k-instruct | Microsoft | 0.72 |
| 12 | StableBeluga2 | Stability AI | 0.69 |
| 13 | Llama-2-70b-hf | Meta AI | 0.65 |
| 14 | gpt-3.5-turbo-0613 | OpenAI | 0.62 |
| 15 | phi-2 | Microsoft | 0.59 |
| 16 | Nemotron-4 15B | NVIDIA | 0.59 |
| 17 | Llama-2-70b-chat | Meta AI | 0.58 |
| 18 | LLaMA-65B | Meta AI | 0.58 |
| 19 | Llama-2-13b-chat | Meta AI | 0.58 |
| 20 | Mistral-7B-v0.1 | Mistral AI | 0.56 |
| 21 | gemma-7b | Google DeepMind | 0.55 |
| 22 | Qwen-14B-Chat | Alibaba | 0.55 |
| 23 | Yi-34B | 01.AI | 0.54 |
| 24 | Qwen-14B | Alibaba | 0.53 |
| 25 | internlm-20b | 0.53 | |
| 26 | LLaMA-33B | Meta AI | 0.5 |
| 27 | Baichuan-2-13B-Base | Baichuan | 0.49 |
| 28 | Baichuan2-13B-Chat | 0.47 | |
| 29 | Yi-6B-Chat | 01.AI | 0.47 |
| 30 | Llama-2-13b | Meta AI | 0.47 |
| 31 | Qwen-7B | Alibaba | 0.45 |
| 32 | Llama-2-34b | Meta AI | 0.44 |
| 33 | vicuna-13b-v1.1 | 0.43 | |
| 34 | Baichuan-13B-Base | Baichuan | 0.43 |
| 35 | Yi-6B | 01.AI | 0.43 |
| 36 | internlm-chat-20b | 0.42 | |
| 37 | Baichuan-2-7B-Base | Baichuan | 0.42 |