Live
Benchmarks

BoolQ

Models scored
136
Best score
0.91
Metric
Score

BoolQ is an AI evaluation tracked by Epoch AI, with 136 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span November 2019 to August 2024. The best score recorded is 0.91, measured as Score. Models come from MosaicML, Technology Innovation Institute, Baichuan, Meta AI, Stability AI, Alibaba, Mistral AI, Google DeepMind and others.

Full record
Score metric
Score
Models scored
136
Best score recorded
0.91
Earliest model scored
Nov 5, 2019
Most recent model scored
Aug 17, 2024
Third-party reported
Yes
Leaderboard
01T5-11BGoogle0.91
02PaLM 2-L0.91
03T5-3BGoogle0.9
04Inflection-1Inflection AI0.9
05StableBeluga2Stability AI0.89
06falcon-180BTechnology Innovation Institute0.89
07gpt-4o-mini-2024-07-18OpenAI0.89
08PaLM 540BGoogle Research0.89
09PaLM 2-M0.89
10Llama-2-70b-hfMeta AI0.89
11text-davinci-0030.88
12PaLM 2-S0.88
13text-davinci-002OpenAI0.88
14internlm-20b0.88
15Mistral-7B-v0.1Mistral AI0.87
16LLaMA-65BMeta AI0.87
17gpt-3.5-turbo-0613OpenAI0.87
18Qwen-14BAlibaba0.86
19LLaMA-33BMeta AI0.86
20gemini-1.5-flash-001Google DeepMind0.86
21gemma-2-9bGoogle DeepMind0.86
22T5-Large0.85
23mpt-30b-instructMosaicML0.85
24Megatron-Turing NLG 530BMicrosoft,NVIDIA0.85
25PaLM 62B0.85
26Phi-3.5-MoE-instructMicrosoft0.85
27Llama-2-34bMeta AI0.84
28Chinchilla (70B)DeepMind0.84
29vicuna-13b-v1.10.83
30Mistral-7B-Instruct-v0.2Mistral AI0.83
31gemma-7bGoogle DeepMind0.83
32falcon-40bTechnology Innovation Institute0.83
33Llama-3.1-8B-InstructMeta AI0.83
34Mistral-Nemo-Base-2407Mistral AI0.82
35Llama-2-13bMeta AI0.82
36T5-Base0.81
37vicuna-13b-v1.3Large Model Systems Organization,University of California (UC) Berkeley0.81
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks