Enigma Eval is an AI evaluation tracked by Epoch AI, with 42 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span February 2024 to March 2026. The best score recorded is 0.24, measured as Accuracy. Models come from OpenAI, Google DeepMind, Anthropic, Moonshot, Google, Google DeepMind,Google, Mistral AI, Meta AI.
| 01 | gpt-5.4-pro-2026-03-05_unknown | OpenAI | 0.24 |
| 02 | gemini-3.1-pro-preview | Google DeepMind | 0.2 |
| 03 | gpt-5-pro-2025-10-06_unknown | OpenAI | 0.19 |
| 04 | gemini-3-pro-preview | Google DeepMind | 0.18 |
| 05 | gpt-5.4-2026-03-05_xhigh | OpenAI | 0.16 |
| 06 | o3-2025-04-16_medium | OpenAI | 0.13 |
| 07 | o3-2025-04-16_high | OpenAI | 0.12 |
| 08 | claude-opus-4-5-20251101_unknown | Anthropic | 0.12 |
| 09 | gpt-5.1-2025-11-13_unknown | OpenAI | 0.11 |
| 10 | gpt-5-2025-08-07_unknown | OpenAI | 0.1 |
| 11 | gpt-5.2-2025-12-11_unknown | OpenAI | 0.1 |
| 12 | o4-mini-2025-04-16_high | OpenAI | 0.09 |
| 13 | gpt-5-mini-2025-08-07_unknown | OpenAI | 0.08 |
| 14 | claude-opus-4-6_max | Anthropic | 0.08 |
| 15 | claude-opus-4-1-20250805_unknown | Anthropic | 0.07 |
| 16 | claude-opus-4-6 | Anthropic | 0.07 |
| 17 | o4-mini-2025-04-16_medium | OpenAI | 0.07 |
| 18 | o1-pro-2025-03-19 | OpenAI | 0.06 |
| 19 | claude-sonnet-4-5-20250929_unknown | Anthropic | 0.06 |
| 20 | o1-2024-12-17_unknown | OpenAI | 0.06 |
| 21 | gemini-2.5-pro-preview-06-05 | Google DeepMind | 0.06 |
| 22 | claude-opus-4-20250514_unknown | Anthropic | 0.06 |
| 23 | claude-3-7-sonnet-20250219_unknown | Anthropic | 0.04 |
| 24 | gemini-2.5-pro-exp-03-25 | Google DeepMind | 0.04 |
| 25 | kimi-k2.5 | Moonshot | 0.03 |
| 26 | gpt-4.5-preview-2025-02-27 | OpenAI | 0.03 |
| 27 | claude-sonnet-4-20250514_unknown | Anthropic | 0.03 |
| 28 | gemini-3.1-flash-lite | 0.03 | |
| 29 | gemini-2.5-flash-preview-05-20 | Google DeepMind | 0.03 |
| 30 | gemini-2.5-pro-preview-05-06 | Google DeepMind | 0.02 |
| 31 | claude-3-7-sonnet-20250219 | Anthropic | 0.02 |
| 32 | gpt-4.1-2025-04-14 | OpenAI | 0.02 |
| 33 | gpt-5.1-2025-11-13_none | OpenAI | 0.02 |
| 34 | gemini-2.0-flash-thinking-exp-01-21 | Google DeepMind,Google | 0.01 |
| 35 | claude-3-5-sonnet-20241022 | Anthropic | 0.01 |
| 36 | pixtral-large-2411 | Mistral AI | 0.01 |
| 37 | claude-3-opus-20240229 | Anthropic | 0.01 |
| 38 | gpt-4o-2024-11-20 | OpenAI | 0.01 |
| 39 | gemini-2.0-pro-exp-02-05 | Google DeepMind | 0.01 |
| 40 | gemini-2.0-flash-02-05 | Google DeepMind,Google | 0.01 |
| 41 | Llama-4-Maverick-17B-128E-Instruct | Meta AI | 0.01 |
| 42 | Llama-3.2-90B-Vision-Instruct | Meta AI | 0 |