实时
Benchmarks

TriviaQA

Models scored
115
Best score
0.88
Metric
EM

TriviaQA is an AI evaluation tracked by Epoch AI, with 115 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span December 2021 to December 2024. The best score recorded is 0.88, measured as EM. Models come from Anthropic, OpenAI, Meta AI, Google Research, Google DeepMind, Mistral AI, DeepSeek, Alibaba and others.

Full record
Score metric
EM
Models scored
115
Best score recorded
0.88
Earliest model scored
Dec 8, 2021
Most recent model scored
Dec 26, 2024
Third-party reported
Yes
Leaderboard
01Llama-2-70b-hfMeta AI0.88
02claude-2.0Anthropic0.88
03claude-1.3Anthropic0.87
04PaLM 2-L0.86
05LLaMA-65BMeta AI0.86
06gpt-3.5-turbo-1106OpenAI0.86
07gpt-4-0613OpenAI0.85
08Llama-2-34bMeta AI0.85
09LLaMA-33BMeta AI0.84
10DeepSeek-V3DeepSeek0.83
11Llama-3.1-405BMeta AI0.83
12Mixtral-8x7B-v0.1Mistral AI0.82
13PaLM 2-M0.82
14PaLM 540BGoogle Research0.81
15DeepSeek-V2DeepSeek0.8
16falcon-40bTechnology Innovation Institute0.8
17Llama-2-13bMeta AI0.8
18claude-instant-1.1Anthropic0.79
19claude-instant-1.2Anthropic0.79
20LLaMA-13BMeta AI0.78
21GLaM (MoE)Google0.76
22PaLM 2-S0.75
23Mistral-7B-v0.1Mistral AI0.75
24Phi-3-medium-128k-instructMicrosoft0.74
25Llama-2-7bMeta AI0.74
26mpt-30bMosaicML0.74
27gemma-7bGoogle DeepMind0.72
28Qwen2.5-72BAlibaba0.72
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks