Live
Benchmarks

CAD Eval

Models scored
15
Best score
0.74
Metric
Overall pass (%)

CAD Eval is an AI evaluation tracked by Epoch AI, with 15 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span March 2024 to April 2025. The best score recorded is 0.74, measured as Overall pass (%). Models come from Anthropic, Google DeepMind, Google DeepMind,Google, OpenAI.

Full record
Score metric
Overall pass (%)
Models scored
15
Best score recorded
0.74
Earliest model scored
Mar 7, 2024
Most recent model scored
Apr 16, 2025
Third-party reported
Yes
Leaderboard
01o3-2025-04-16_mediumOpenAI0.74
02gemini-2.5-pro-preview-03-25Google DeepMind0.64
03o4-mini-2025-04-16_mediumOpenAI0.62
04o1-2024-12-17_mediumOpenAI0.56
05o3-mini-2025-01-31_mediumOpenAI0.54
06claude-3-7-sonnet-20250219Anthropic0.54
07claude-3-5-sonnet-20241022Anthropic0.48
08gpt-4.1-2025-04-14OpenAI0.42
09gemini-1.5-pro-002Google DeepMind0.34
10claude-3-5-haiku-20241022Anthropic0.32
11gemini-2.0-flash-001Google DeepMind,Google0.3
12gpt-4o-2024-08-06OpenAI0.26
13ml-elephant0.2
14gpt-4.1-mini-2025-04-14OpenAI0.16
15claude-3-haiku-20240307Anthropic0.12
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks