PIQA is an AI evaluation tracked by Epoch AI, with 113 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span November 2019 to December 2024. The best score recorded is 0.89, measured as Score. Models come from Meta AI, Mistral AI, Google DeepMind, MosaicML, Technology Innovation Institute, Baichuan, Stability AI, Alibaba and others.
| 01 | gpt-4o-mini-2024-07-18 | OpenAI | 0.89 |
| 02 | Phi-3.5-MoE-instruct | Microsoft | 0.89 |
| 03 | gemini-1.5-flash-002 | Google DeepMind | 0.88 |
| 04 | Llama-3.1-405B | Meta AI | 0.86 |
| 05 | falcon-180B | Technology Innovation Institute | 0.85 |
| 06 | DeepSeek-V3-Base | DeepSeek | 0.85 |
| 07 | Inflection-1 | Inflection AI | 0.84 |
| 08 | DeepSeek-V2 | DeepSeek | 0.84 |
| 09 | gemma-2-9b | Google DeepMind | 0.84 |
| 10 | Mixtral-8x7B-v0.1 | Mistral AI | 0.84 |
| 11 | Mistral-Nemo-Base-2407 | Mistral AI | 0.83 |
| 12 | StableBeluga2 | Stability AI | 0.83 |
| 13 | Megatron-Turing NLG 530B | Microsoft,NVIDIA | 0.83 |
| 14 | falcon-40b | Technology Innovation Institute | 0.83 |
| 15 | Mistral-7B-v0.1 | Mistral AI | 0.83 |
| 16 | Llama-2-70b-hf | Meta AI | 0.83 |
| 17 | LLaMA-65B | Meta AI | 0.83 |
| 18 | Qwen2.5-72B | Alibaba | 0.83 |
| 19 | Nemotron-4 15B | NVIDIA | 0.82 |
| 20 | LLaMA-33B | Meta AI | 0.82 |
| 21 | text-davinci-001 | OpenAI | 0.82 |
| 22 | PaLM 540B | Google Research | 0.82 |
| 23 | Mistral-7B-Instruct-v0.2 | Mistral AI | 0.82 |
| 24 | Mistral-7B-Instruct-v0.1 | Mistral AI | 0.82 |
| 25 | Llama-2-34b | Meta AI | 0.82 |
| 26 | mpt-30b | MosaicML | 0.82 |
| 27 | Chinchilla (70B) | DeepMind | 0.82 |
| 28 | Gopher (280B) | DeepMind | 0.82 |
| 29 | gemma-7b | Google DeepMind | 0.81 |
| 30 | Llama-3.1-8B-Instruct | Meta AI | 0.81 |
| 31 | Phi-3.5-mini-instruct | Microsoft | 0.81 |
| 32 | Llama-2-13b | Meta AI | 0.81 |
| 33 | mpt-7b | MosaicML | 0.81 |