ARC AI2 is an AI evaluation tracked by Epoch AI, with 134 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span November 2019 to December 2024. The best score recorded is 0.95, measured as Challenge score. Models come from 01.AI, Google Research, Technology Innovation Institute, OpenAI, Meta AI, MosaicML, Baichuan, Stability AI and others.
| 01 | Llama-3.1-405B | Meta AI | 0.95 |
| 02 | DeepSeek-V3 | DeepSeek | 0.95 |
| 03 | Qwen2.5-72B | Alibaba | 0.94 |
| 04 | DeepSeek-V2 | DeepSeek | 0.92 |
| 05 | Phi-3-medium-128k-instruct | Microsoft | 0.92 |
| 06 | Phi-3-small-8k-instruct | Microsoft | 0.91 |
| 07 | gpt-3.5-turbo-1106 | OpenAI | 0.87 |
| 08 | Mixtral-8x7B-v0.1 | Mistral AI | 0.87 |
| 09 | claude-instant-1.2 | Anthropic | 0.86 |
| 10 | StableBeluga2 | Stability AI | 0.86 |
| 11 | claude-instant-1.1 | Anthropic | 0.86 |
| 12 | PaLM 540B | Google Research | 0.85 |
| 13 | text-davinci-002 | OpenAI | 0.85 |
| 14 | Phi-3-mini-4k-instruct | Microsoft | 0.85 |
| 15 | Qwen-14B | Alibaba | 0.84 |
| 16 | Meta-Llama-3-8B-Instruct | Meta AI | 0.83 |
| 17 | internlm-20b | 0.82 | |
| 18 | Mistral-7B-v0.1 | Mistral AI | 0.79 |
| 19 | gemma-7b | Google DeepMind | 0.78 |
| 20 | Llama-2-70b-hf | Meta AI | 0.78 |
| 21 | phi-2 | Microsoft | 0.76 |
| 22 | Qwen-7B | Alibaba | 0.75 |
| 23 | Qwen2.5-Coder-32B | Alibaba | 0.7 |
| 24 | LLaMA-65B | Meta AI | 0.69 |
| 25 | internlm-7b | 0.69 | |
| 26 | PaLM 2-L | 0.69 | |
| 27 | falcon-180B | Technology Innovation Institute | 0.68 |
| 28 | LLaMA-33B | Meta AI | 0.68 |
| 29 | Qwen2.5-Coder-14B | 0.66 | |
| 30 | PaLM 2-M | 0.65 | |
| 31 | DeepSeek-Coder-V2-Base | DeepSeek | 0.64 |
| 32 | falcon-40b | Technology Innovation Institute | 0.62 |
| 33 | chatglm2-6b | 0.61 | |
| 34 | Qwen2.5-Coder-7B | Alibaba | 0.61 |
| 35 | Llama-2-13b | Meta AI | 0.6 |
| 36 | PaLM 2-S | 0.6 | |
| 37 | DeepSeek-Coder-V2-Lite-Base | 0.57 | |
| 38 | Yi-9B | 0.56 | |
| 39 | Nemotron-4 15B | NVIDIA | 0.56 |
| 40 | INTELLECT-1-Instruct | Prime Intellect,Hugging Face,Arcee AI | 0.55 |