लाइव
Benchmarks

ScienceQA

Models scored
26
Best score
0.91
Metric
Score

ScienceQA is an AI evaluation tracked by Epoch AI, with 26 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span January 2022 to August 2024. The best score recorded is 0.91, measured as Score. Models come from Salesforce Research, OpenAI, Meta AI, Anthropic, Google DeepMind.

Full record
Score metric
Score
Models scored
26
Best score recorded
0.91
Earliest model scored
Jan 27, 2022
Most recent model scored
Aug 16, 2024
Third-party reported
Yes
Leaderboard
01Phi-3.5-vision-instruct0.91
02gpt-4o-2024-05-13OpenAI0.89
03gemini-1.0-pro-visionGoogle DeepMind0.8
04falcon-11B-vlm0.75
05blip2-opt-2.7bSalesforce Research0.74
06text-davinci-001OpenAI0.74
07llama3-llava-next-8b0.74
08llava-v1.6-vicuna-13b0.74
09llava-v1.6-mistral-7b0.73
10MM1-7B-Chat0.73
11claude-3-haiku-20240307Anthropic0.72
12llava-v1.6-vicuna-7b0.71
13InternVL-Chat-ViT-6B-Vicuna-13B0.7
14MM1-3B-Chat0.69
15Qwen-VL-Chat0.68
16llava-v1.5-7b0.67
17InternVL-Chat-ViT-6B-Vicuna-7B0.66
18instructblip-vicuna-13b0.63
19instructblip-vicuna-7b0.6
20Llama-2-13bMeta AI0.56
21LLaMA-13BMeta AI0.43
22Llama-2-7bMeta AI0.43
23LLaMA-7BMeta AI0.36
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks