ScienceQA is an AI evaluation tracked by Epoch AI, with 26 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span January 2022 to August 2024. The best score recorded is 0.91, measured as Score. Models come from Salesforce Research, OpenAI, Meta AI, Anthropic, Google DeepMind.
| 01 | Phi-3.5-vision-instruct | 0.91 | |
| 02 | gpt-4o-2024-05-13 | OpenAI | 0.89 |
| 03 | gemini-1.0-pro-vision | Google DeepMind | 0.8 |
| 04 | falcon-11B-vlm | 0.75 | |
| 05 | blip2-opt-2.7b | Salesforce Research | 0.74 |
| 06 | text-davinci-001 | OpenAI | 0.74 |
| 07 | llama3-llava-next-8b | 0.74 | |
| 08 | llava-v1.6-vicuna-13b | 0.74 | |
| 09 | llava-v1.6-mistral-7b | 0.73 | |
| 10 | MM1-7B-Chat | 0.73 | |
| 11 | claude-3-haiku-20240307 | Anthropic | 0.72 |
| 12 | llava-v1.6-vicuna-7b | 0.71 | |
| 13 | InternVL-Chat-ViT-6B-Vicuna-13B | 0.7 | |
| 14 | MM1-3B-Chat | 0.69 | |
| 15 | Qwen-VL-Chat | 0.68 | |
| 16 | llava-v1.5-7b | 0.67 | |
| 17 | InternVL-Chat-ViT-6B-Vicuna-7B | 0.66 | |
| 18 | instructblip-vicuna-13b | 0.63 | |
| 19 | instructblip-vicuna-7b | 0.6 | |
| 20 | Llama-2-13b | Meta AI | 0.56 |
| 21 | LLaMA-13B | Meta AI | 0.43 |
| 22 | Llama-2-7b | Meta AI | 0.43 |
| 23 | LLaMA-7B | Meta AI | 0.36 |