OpenBookQA is an AI evaluation tracked by Epoch AI, with 70 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span November 2019 to April 2024. The best score recorded is 0.88, measured as Accuracy. Models come from Cerebras Systems, Databricks, Technology Innovation Institute, Google DeepMind, Google, OpenAI, EleutherAI,LAION, EleutherAI and others.
| 01 | Phi-3-small-8k-instruct | Microsoft | 0.88 |
| 02 | Phi-3-mini-4k-instruct | Microsoft | 0.88 |
| 03 | Phi-3-medium-128k-instruct | Microsoft | 0.87 |
| 04 | gpt-3.5-turbo-1106 | OpenAI | 0.86 |
| 05 | Mixtral-8x7B-v0.1 | Mistral AI | 0.86 |
| 06 | Meta-Llama-3-8B-Instruct | Meta AI | 0.83 |
| 07 | Mistral-7B-v0.1 | Mistral AI | 0.8 |
| 08 | gemma-7b | Google DeepMind | 0.79 |
| 09 | phi-2 | Microsoft | 0.74 |
| 10 | PaLM 540B | Google Research | 0.68 |
| 11 | text-davinci-001 | OpenAI | 0.65 |
| 12 | falcon-180B | Technology Innovation Institute | 0.64 |
| 13 | GLaM (MoE) | 0.63 | |
| 14 | LLaMA-65B | Meta AI | 0.6 |
| 15 | Llama-2-70b-hf | Meta AI | 0.6 |
| 16 | LLaMA-33B | Meta AI | 0.59 |
| 17 | Llama-2-7b | Meta AI | 0.59 |
| 18 | Llama-2-34b | Meta AI | 0.58 |
| 19 | PaLM 2-M | 0.57 | |
| 20 | LLaMA-7B | Meta AI | 0.57 |
| 21 | Llama-2-13b | Meta AI | 0.57 |
| 22 | falcon-40b | Technology Innovation Institute | 0.57 |
| 23 | LLaMA-13B | Meta AI | 0.56 |
| 24 | PaLM 2-S | 0.56 | |
| 25 | mpt-30b | MosaicML | 0.52 |
| 26 | falcon-7b | Technology Innovation Institute | 0.52 |
| 27 | mpt-7b | MosaicML | 0.51 |
| 28 | PaLM 62B | 0.5 | |
| 29 | xgen-7b-8k-base | Salesforce | 0.4 |
| 30 | RedPajama-INCITE-7B-Base | 0.4 | |
| 31 | opt-13b | 0.4 | |
| 32 | dolly-v2-12b | Databricks | 0.39 |
| 33 | open_llama_7b | 0.39 | |
| 34 | gpt-neox-20b | EleutherAI | 0.39 |
| 35 | gpt-j-6b | EleutherAI,LAION | 0.38 |
| 36 | phi-1_5 | Microsoft | 0.37 |