Live
Benchmarks

Vpct

Models scored
38
Best score
0.91
Metric
Correct

Vpct is an AI evaluation tracked by Epoch AI, with 38 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span June 2024 to December 2025. The best score recorded is 0.91, measured as Correct. Models come from Google DeepMind, OpenAI, Anthropic.

Full record
Score metric
Correct
Models scored
38
Best score recorded
0.91
Earliest model scored
Jun 20, 2024
Most recent model scored
Dec 17, 2025
Third-party reported
Yes
Leaderboard
01gemini-3-pro-previewGoogle DeepMind0.91
02gpt-5.2-2025-12-11_xhighOpenAI0.84
03gemini-3-flash-previewGoogle DeepMind0.73
04gpt-5.2-2025-12-11_highOpenAI0.67
05gpt-5-2025-08-07_highOpenAI0.66
06gpt-5-2025-08-07_mediumOpenAI0.63
07gpt-5.1-2025-11-13_highOpenAI0.59
08o4-mini-2025-04-16_mediumOpenAI0.57
09gpt-5.1-2025-11-13_mediumOpenAI0.53
10o3-2025-04-16_mediumOpenAI0.52
11gemini-2.5-pro-preview-03-25Google DeepMind0.48
12gemini-2.5-pro-preview-06-05Google DeepMind0.46
13gemini-2.5-flash-preview-09-2025Google DeepMind0.46
14gpt-4.5-preview-2025-02-27OpenAI0.45
15gemini-robotics-er-1.5-preview0.41
16gemini-2.5-pro-preview-05-06Google DeepMind0.41
17gpt-5-mini-2025-08-07_mediumOpenAI0.4
18claude-opus-4-5-20251101_32KAnthropic0.4
19gpt-4o-2024-11-20OpenAI0.4
20claude-sonnet-4-5-20250929_32KAnthropic0.4
21claude-3-7-sonnet-20250219Anthropic0.39
22gpt-5-mini-2025-08-07_highOpenAI0.39
23gemini-2.5-flash-preview-04-17Google DeepMind0.38
24claude-opus-4-20250514_16KAnthropic0.38
25claude-sonnet-4-5-20250929Anthropic0.38
26gpt-5-nano-2025-08-07_highOpenAI0.37
27o1-2024-12-17_mediumOpenAI0.37
28gpt-5-nano-2025-08-07_mediumOpenAI0.35
29claude-3-7-sonnet-20250219_64KAnthropic0.35
30claude-opus-4-1-20250805Anthropic0.35
31claude-sonnet-4-20250514_32KAnthropic0.34
32gpt-4o-mini-2024-07-18OpenAI0.34
33claude-opus-4-1-20250805_16KAnthropic0.33
34claude-3-5-sonnet-20241022Anthropic0.33
35claude-opus-4-20250514Anthropic0.33
36claude-3-5-sonnet-20240620Anthropic0.33
37gemini-2.5-flash-lite-preview-09-2025Google DeepMind0.3
38claude-sonnet-4-20250514Anthropic0.3
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks