TriviaQA is an AI evaluation tracked by Epoch AI, with 115 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span December 2021 to December 2024. The best score recorded is 0.88, measured as EM. Models come from Anthropic, OpenAI, Meta AI, Google Research, Google DeepMind, Mistral AI, DeepSeek, Alibaba and others.
| 01 | Llama-2-70b-hf | Meta AI | 0.88 |
| 02 | claude-2.0 | Anthropic | 0.88 |
| 03 | claude-1.3 | Anthropic | 0.87 |
| 04 | PaLM 2-L | 0.86 | |
| 05 | LLaMA-65B | Meta AI | 0.86 |
| 06 | gpt-3.5-turbo-1106 | OpenAI | 0.86 |
| 07 | gpt-4-0613 | OpenAI | 0.85 |
| 08 | Llama-2-34b | Meta AI | 0.85 |
| 09 | LLaMA-33B | Meta AI | 0.84 |
| 10 | DeepSeek-V3 | DeepSeek | 0.83 |
| 11 | Llama-3.1-405B | Meta AI | 0.83 |
| 12 | Mixtral-8x7B-v0.1 | Mistral AI | 0.82 |
| 13 | PaLM 2-M | 0.82 | |
| 14 | PaLM 540B | Google Research | 0.81 |
| 15 | DeepSeek-V2 | DeepSeek | 0.8 |
| 16 | falcon-40b | Technology Innovation Institute | 0.8 |
| 17 | Llama-2-13b | Meta AI | 0.8 |
| 18 | claude-instant-1.1 | Anthropic | 0.79 |
| 19 | claude-instant-1.2 | Anthropic | 0.79 |
| 20 | LLaMA-13B | Meta AI | 0.78 |
| 21 | GLaM (MoE) | 0.76 | |
| 22 | PaLM 2-S | 0.75 | |
| 23 | Mistral-7B-v0.1 | Mistral AI | 0.75 |
| 24 | Phi-3-medium-128k-instruct | Microsoft | 0.74 |
| 25 | Llama-2-7b | Meta AI | 0.74 |
| 26 | mpt-30b | MosaicML | 0.74 |
| 27 | gemma-7b | Google DeepMind | 0.72 |
| 28 | Qwen2.5-72B | Alibaba | 0.72 |