Live
Benchmarks

OpenBookQA

Models scored
70
Best score
0.88
Metric
Accuracy

OpenBookQA is an AI evaluation tracked by Epoch AI, with 70 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span November 2019 to April 2024. The best score recorded is 0.88, measured as Accuracy. Models come from Cerebras Systems, Databricks, Technology Innovation Institute, Google DeepMind, Google, OpenAI, EleutherAI,LAION, EleutherAI and others.

Full record
Score metric
Accuracy
Models scored
70
Best score recorded
0.88
Earliest model scored
Nov 5, 2019
Most recent model scored
Apr 23, 2024
Third-party reported
Yes
Leaderboard
01Phi-3-small-8k-instructMicrosoft0.88
02Phi-3-mini-4k-instructMicrosoft0.88
03Phi-3-medium-128k-instructMicrosoft0.87
04gpt-3.5-turbo-1106OpenAI0.86
05Mixtral-8x7B-v0.1Mistral AI0.86
06Meta-Llama-3-8B-InstructMeta AI0.83
07Mistral-7B-v0.1Mistral AI0.8
08gemma-7bGoogle DeepMind0.79
09phi-2Microsoft0.74
10PaLM 540BGoogle Research0.68
11text-davinci-001OpenAI0.65
12falcon-180BTechnology Innovation Institute0.64
13GLaM (MoE)Google0.63
14LLaMA-65BMeta AI0.6
15Llama-2-70b-hfMeta AI0.6
16LLaMA-33BMeta AI0.59
17Llama-2-7bMeta AI0.59
18Llama-2-34bMeta AI0.58
19PaLM 2-M0.57
20LLaMA-7BMeta AI0.57
21Llama-2-13bMeta AI0.57
22falcon-40bTechnology Innovation Institute0.57
23LLaMA-13BMeta AI0.56
24PaLM 2-S0.56
25mpt-30bMosaicML0.52
26falcon-7bTechnology Innovation Institute0.52
27mpt-7bMosaicML0.51
28PaLM 62B0.5
29xgen-7b-8k-baseSalesforce0.4
30RedPajama-INCITE-7B-Base0.4
31opt-13b0.4
32dolly-v2-12bDatabricks0.39
33open_llama_7b0.39
34gpt-neox-20bEleutherAI0.39
35gpt-j-6bEleutherAI,LAION0.38
36phi-1_5Microsoft0.37
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks