Онлайн
Benchmarks

Aider Polyglot

Models scored
72
Best score
88
Metric
Percent correct

Aider Polyglot is an AI evaluation tracked by Epoch AI, with 72 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span March 2024 to December 2025. The best score recorded is 88, measured as Percent correct. Models come from OpenAI, Alibaba, DeepSeek, xAI, Moonshot, Google DeepMind, Anthropic, Meta AI and others.

Full record
Score metric
Percent correct
Models scored
72
Best score recorded
88
Earliest model scored
Mar 26, 2024
Most recent model scored
Dec 1, 2025
Third-party reported
Yes
Leaderboard
01gpt-5-2025-08-07_highOpenAI88
02gpt-5-2025-08-07_mediumOpenAI86.7
03o3-pro-2025-06-10_highOpenAI84.9
04gemini-2.5-pro-preview-06-05_32KGoogle DeepMind83.1
05o3-2025-04-16_highOpenAI81.3
06gpt-5-2025-08-07_lowOpenAI81.3
07grok-4-0709_highxAI79.6
08grok-4-0709xAI79.6
09gemini-2.5-pro-preview-06-05Google DeepMind79.1
10gemini-2.5-pro-preview-05-06Google DeepMind76.9
11o3-2025-04-16_mediumOpenAI76.9
12o3-2025-04-16_unknownOpenAI76.9
13deepseek-reasonerDeepSeek74.2
14DeepSeek-V3.2-Exp_thinkingDeepSeek74.2
15gemini-2.5-pro-exp-03-25Google DeepMind72.9
16gemini-2.5-pro-preview-03-25Google DeepMind72.9
17claude-opus-4-20250514_32KAnthropic72
18o4-mini-2025-04-16_highOpenAI72
19DeepSeek-R1-0528DeepSeek71.4
20claude-opus-4-20250514Anthropic70.7
21DeepSeek-V3.2-ExpDeepSeek70.2
22deepseek-chatDeepSeek70.2
23claude-3-7-sonnet-20250219_32KAnthropic64.9
24o1-2024-12-17_highOpenAI61.7
25claude-sonnet-4-20250514_32KAnthropic61.3
26claude-3-7-sonnet-20250219Anthropic60.4
27o3-mini-2025-01-31_highOpenAI60.4
28Qwen3-235B-A22BAlibaba59.6
29Qwen3-235B-A22B-Instruct-2507Alibaba59.6
30moonshotai/kimi-k2-0905Moonshot59.1
31Kimi-K2-InstructMoonshot59.1
32DeepSeek-R1DeepSeek56.9
33claude-sonnet-4-20250514Anthropic56.4
34gemini-2.5-flash-preview-05-20_23KGoogle DeepMind55.1
35DeepSeek-V3-0324DeepSeek55.1
36o3-mini-2025-01-31_mediumOpenAI53.8
37grok-3-betaxAI53.3
38gpt-4.1-2025-04-14OpenAI52.4
39claude-3-5-sonnet-20241022Anthropic51.6
40grok-3-mini-beta_highxAI49.3
41DeepSeek-V3DeepSeek48.4
42gemini-2.5-flash-preview-04-17Google DeepMind47.1
43chatgpt-4o-03-27-2025OpenAI45.3
44gpt-4.5-preview-2025-02-27OpenAI44.9
45gemini-2.5-flash-preview-05-20Google DeepMind44
46openai/gpt-oss-120b_highOpenAI41.8
47gpt-oss-120b_highOpenAI41.8
48Qwen3-32BAlibaba40
49gemini-exp-1206Google DeepMind,Google38.2
50gemini-2.0-pro-exp-02-05Google DeepMind35.6
51grok-3-mini-beta_lowxAI34.7
52o1-mini-2024-09-12_unknownOpenAI32.9
53gpt-4.1-mini-2025-04-14OpenAI32.4
54claude-3-5-haiku-20241022Anthropic28
55chatgpt-4o-01-29-2025OpenAI27.1
56gpt-4o-2024-08-06OpenAI23.1
57gemini-2.0-flash-expGoogle DeepMind,Google22.2
58qwen-max-2025-01-25Alibaba21.8
59QwQ-32BAlibaba20.9
60gemini-2.0-flash-thinking-exp-01-21Google DeepMind,Google18.2
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks