BoolQ is an AI evaluation tracked by Epoch AI, with 136 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span November 2019 to August 2024. The best score recorded is 0.91, measured as Score. Models come from MosaicML, Technology Innovation Institute, Baichuan, Meta AI, Stability AI, Alibaba, Mistral AI, Google DeepMind and others.
| 01 | T5-11B | 0.91 | |
| 02 | PaLM 2-L | 0.91 | |
| 03 | T5-3B | 0.9 | |
| 04 | Inflection-1 | Inflection AI | 0.9 |
| 05 | StableBeluga2 | Stability AI | 0.89 |
| 06 | falcon-180B | Technology Innovation Institute | 0.89 |
| 07 | gpt-4o-mini-2024-07-18 | OpenAI | 0.89 |
| 08 | PaLM 540B | Google Research | 0.89 |
| 09 | PaLM 2-M | 0.89 | |
| 10 | Llama-2-70b-hf | Meta AI | 0.89 |
| 11 | text-davinci-003 | 0.88 | |
| 12 | PaLM 2-S | 0.88 | |
| 13 | text-davinci-002 | OpenAI | 0.88 |
| 14 | internlm-20b | 0.88 | |
| 15 | Mistral-7B-v0.1 | Mistral AI | 0.87 |
| 16 | LLaMA-65B | Meta AI | 0.87 |
| 17 | gpt-3.5-turbo-0613 | OpenAI | 0.87 |
| 18 | Qwen-14B | Alibaba | 0.86 |
| 19 | LLaMA-33B | Meta AI | 0.86 |
| 20 | gemini-1.5-flash-001 | Google DeepMind | 0.86 |
| 21 | gemma-2-9b | Google DeepMind | 0.86 |
| 22 | T5-Large | 0.85 | |
| 23 | mpt-30b-instruct | MosaicML | 0.85 |
| 24 | Megatron-Turing NLG 530B | Microsoft,NVIDIA | 0.85 |
| 25 | PaLM 62B | 0.85 | |
| 26 | Phi-3.5-MoE-instruct | Microsoft | 0.85 |
| 27 | Llama-2-34b | Meta AI | 0.84 |
| 28 | Chinchilla (70B) | DeepMind | 0.84 |
| 29 | vicuna-13b-v1.1 | 0.83 | |
| 30 | Mistral-7B-Instruct-v0.2 | Mistral AI | 0.83 |
| 31 | gemma-7b | Google DeepMind | 0.83 |
| 32 | falcon-40b | Technology Innovation Institute | 0.83 |
| 33 | Llama-3.1-8B-Instruct | Meta AI | 0.83 |
| 34 | Mistral-Nemo-Base-2407 | Mistral AI | 0.82 |
| 35 | Llama-2-13b | Meta AI | 0.82 |
| 36 | T5-Base | 0.81 | |
| 37 | vicuna-13b-v1.3 | Large Model Systems Organization,University of California (UC) Berkeley | 0.81 |