Live
Benchmarks

BIG-Bench Hard

Models scored
88
Best score
0.89
Metric
Average

BIG-Bench Hard is an AI evaluation tracked by Epoch AI, with 88 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span February 2023 to December 2024. The best score recorded is 0.89, measured as Average. Models come from Alibaba, Google DeepMind, Mistral AI, Meta AI, DeepSeek, 01.AI, Baichuan, NVIDIA and others.

Full record
Score metric
Average
Models scored
88
Best score recorded
0.89
Earliest model scored
Feb 24, 2023
Most recent model scored
Dec 26, 2024
Third-party reported
Yes
Leaderboard
01gemini-1.5-pro-001Google DeepMind0.89
02DeepSeek-V3DeepSeek0.88
03gemini-1.5-pro-001-feb24Google DeepMind0.84
04Llama-3.1-405BMeta AI0.83
05Phi-3-medium-128k-instructMicrosoft0.81
06Qwen2.5-72BAlibaba0.8
07Phi-3-small-8k-instructMicrosoft0.79
08DeepSeek-V2DeepSeek0.79
09gpt-4-0613OpenAI0.75
10Yi-34B-Chat01.AI0.72
11Phi-3-mini-4k-instructMicrosoft0.72
12StableBeluga2Stability AI0.69
13Llama-2-70b-hfMeta AI0.65
14gpt-3.5-turbo-0613OpenAI0.62
15phi-2Microsoft0.59
16Nemotron-4 15BNVIDIA0.59
17Llama-2-70b-chatMeta AI0.58
18LLaMA-65BMeta AI0.58
19Llama-2-13b-chatMeta AI0.58
20Mistral-7B-v0.1Mistral AI0.56
21gemma-7bGoogle DeepMind0.55
22Qwen-14B-ChatAlibaba0.55
23Yi-34B01.AI0.54
24Qwen-14BAlibaba0.53
25internlm-20b0.53
26LLaMA-33BMeta AI0.5
27Baichuan-2-13B-BaseBaichuan0.49
28Baichuan2-13B-Chat0.47
29Yi-6B-Chat01.AI0.47
30Llama-2-13bMeta AI0.47
31Qwen-7BAlibaba0.45
32Llama-2-34bMeta AI0.44
33vicuna-13b-v1.10.43
34Baichuan-13B-BaseBaichuan0.43
35Yi-6B01.AI0.43
36internlm-chat-20b0.42
37Baichuan-2-7B-BaseBaichuan0.42
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks