Live
Benchmarks

ARC AI2

Models scored
134
Best score
0.95
Metric
Challenge score

ARC AI2 is an AI evaluation tracked by Epoch AI, with 134 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span November 2019 to December 2024. The best score recorded is 0.95, measured as Challenge score. Models come from 01.AI, Google Research, Technology Innovation Institute, OpenAI, Meta AI, MosaicML, Baichuan, Stability AI and others.

Full record
Score metric
Challenge score
Models scored
134
Best score recorded
0.95
Earliest model scored
Nov 5, 2019
Most recent model scored
Dec 26, 2024
Third-party reported
Yes
Leaderboard
01Llama-3.1-405BMeta AI0.95
02DeepSeek-V3DeepSeek0.95
03Qwen2.5-72BAlibaba0.94
04DeepSeek-V2DeepSeek0.92
05Phi-3-medium-128k-instructMicrosoft0.92
06Phi-3-small-8k-instructMicrosoft0.91
07gpt-3.5-turbo-1106OpenAI0.87
08Mixtral-8x7B-v0.1Mistral AI0.87
09claude-instant-1.2Anthropic0.86
10StableBeluga2Stability AI0.86
11claude-instant-1.1Anthropic0.86
12PaLM 540BGoogle Research0.85
13text-davinci-002OpenAI0.85
14Phi-3-mini-4k-instructMicrosoft0.85
15Qwen-14BAlibaba0.84
16Meta-Llama-3-8B-InstructMeta AI0.83
17internlm-20b0.82
18Mistral-7B-v0.1Mistral AI0.79
19gemma-7bGoogle DeepMind0.78
20Llama-2-70b-hfMeta AI0.78
21phi-2Microsoft0.76
22Qwen-7BAlibaba0.75
23Qwen2.5-Coder-32BAlibaba0.7
24LLaMA-65BMeta AI0.69
25internlm-7b0.69
26PaLM 2-L0.69
27falcon-180BTechnology Innovation Institute0.68
28LLaMA-33BMeta AI0.68
29Qwen2.5-Coder-14B0.66
30PaLM 2-M0.65
31DeepSeek-Coder-V2-BaseDeepSeek0.64
32falcon-40bTechnology Innovation Institute0.62
33chatglm2-6b0.61
34Qwen2.5-Coder-7BAlibaba0.61
35Llama-2-13bMeta AI0.6
36PaLM 2-S0.6
37DeepSeek-Coder-V2-Lite-Base0.57
38Yi-9B0.56
39Nemotron-4 15BNVIDIA0.56
40INTELLECT-1-InstructPrime Intellect,Hugging Face,Arcee AI0.55
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks