Aider Polyglot is an AI evaluation tracked by Epoch AI, with 72 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span March 2024 to December 2025. The best score recorded is 88, measured as Percent correct. Models come from OpenAI, Alibaba, DeepSeek, xAI, Moonshot, Google DeepMind, Anthropic, Meta AI and others.
| 01 | gpt-5-2025-08-07_high | OpenAI | 88 |
| 02 | gpt-5-2025-08-07_medium | OpenAI | 86.7 |
| 03 | o3-pro-2025-06-10_high | OpenAI | 84.9 |
| 04 | gemini-2.5-pro-preview-06-05_32K | Google DeepMind | 83.1 |
| 05 | o3-2025-04-16_high | OpenAI | 81.3 |
| 06 | gpt-5-2025-08-07_low | OpenAI | 81.3 |
| 07 | grok-4-0709_high | xAI | 79.6 |
| 08 | grok-4-0709 | xAI | 79.6 |
| 09 | gemini-2.5-pro-preview-06-05 | Google DeepMind | 79.1 |
| 10 | gemini-2.5-pro-preview-05-06 | Google DeepMind | 76.9 |
| 11 | o3-2025-04-16_medium | OpenAI | 76.9 |
| 12 | o3-2025-04-16_unknown | OpenAI | 76.9 |
| 13 | deepseek-reasoner | DeepSeek | 74.2 |
| 14 | DeepSeek-V3.2-Exp_thinking | DeepSeek | 74.2 |
| 15 | gemini-2.5-pro-exp-03-25 | Google DeepMind | 72.9 |
| 16 | gemini-2.5-pro-preview-03-25 | Google DeepMind | 72.9 |
| 17 | claude-opus-4-20250514_32K | Anthropic | 72 |
| 18 | o4-mini-2025-04-16_high | OpenAI | 72 |
| 19 | DeepSeek-R1-0528 | DeepSeek | 71.4 |
| 20 | claude-opus-4-20250514 | Anthropic | 70.7 |
| 21 | DeepSeek-V3.2-Exp | DeepSeek | 70.2 |
| 22 | deepseek-chat | DeepSeek | 70.2 |
| 23 | claude-3-7-sonnet-20250219_32K | Anthropic | 64.9 |
| 24 | o1-2024-12-17_high | OpenAI | 61.7 |
| 25 | claude-sonnet-4-20250514_32K | Anthropic | 61.3 |
| 26 | claude-3-7-sonnet-20250219 | Anthropic | 60.4 |
| 27 | o3-mini-2025-01-31_high | OpenAI | 60.4 |
| 28 | Qwen3-235B-A22B | Alibaba | 59.6 |
| 29 | Qwen3-235B-A22B-Instruct-2507 | Alibaba | 59.6 |
| 30 | moonshotai/kimi-k2-0905 | Moonshot | 59.1 |
| 31 | Kimi-K2-Instruct | Moonshot | 59.1 |
| 32 | DeepSeek-R1 | DeepSeek | 56.9 |
| 33 | claude-sonnet-4-20250514 | Anthropic | 56.4 |
| 34 | gemini-2.5-flash-preview-05-20_23K | Google DeepMind | 55.1 |
| 35 | DeepSeek-V3-0324 | DeepSeek | 55.1 |
| 36 | o3-mini-2025-01-31_medium | OpenAI | 53.8 |
| 37 | grok-3-beta | xAI | 53.3 |
| 38 | gpt-4.1-2025-04-14 | OpenAI | 52.4 |
| 39 | claude-3-5-sonnet-20241022 | Anthropic | 51.6 |
| 40 | grok-3-mini-beta_high | xAI | 49.3 |
| 41 | DeepSeek-V3 | DeepSeek | 48.4 |
| 42 | gemini-2.5-flash-preview-04-17 | Google DeepMind | 47.1 |
| 43 | chatgpt-4o-03-27-2025 | OpenAI | 45.3 |
| 44 | gpt-4.5-preview-2025-02-27 | OpenAI | 44.9 |
| 45 | gemini-2.5-flash-preview-05-20 | Google DeepMind | 44 |
| 46 | openai/gpt-oss-120b_high | OpenAI | 41.8 |
| 47 | gpt-oss-120b_high | OpenAI | 41.8 |
| 48 | Qwen3-32B | Alibaba | 40 |
| 49 | gemini-exp-1206 | Google DeepMind,Google | 38.2 |
| 50 | gemini-2.0-pro-exp-02-05 | Google DeepMind | 35.6 |
| 51 | grok-3-mini-beta_low | xAI | 34.7 |
| 52 | o1-mini-2024-09-12_unknown | OpenAI | 32.9 |
| 53 | gpt-4.1-mini-2025-04-14 | OpenAI | 32.4 |
| 54 | claude-3-5-haiku-20241022 | Anthropic | 28 |
| 55 | chatgpt-4o-01-29-2025 | OpenAI | 27.1 |
| 56 | gpt-4o-2024-08-06 | OpenAI | 23.1 |
| 57 | gemini-2.0-flash-exp | Google DeepMind,Google | 22.2 |
| 58 | qwen-max-2025-01-25 | Alibaba | 21.8 |
| 59 | QwQ-32B | Alibaba | 20.9 |
| 60 | gemini-2.0-flash-thinking-exp-01-21 | Google DeepMind,Google | 18.2 |