Lech Mazur Writing is an AI evaluation tracked by Epoch AI, with 49 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span July 2024 to August 2025. The best score recorded is 8.6, measured as Mean score. Models come from OpenAI, Alibaba, DeepSeek, Anthropic, Google DeepMind, xAI, Google DeepMind,Google, Mistral AI and others.
| 01 | gpt-5-2025-08-07_medium | OpenAI | 8.6 |
| 02 | Kimi-K2-Instruct | Moonshot | 8.56 |
| 03 | claude-opus-4-1-20250805 | Anthropic | 8.47 |
| 04 | claude-opus-4-1-20250805_16K | Anthropic | 8.45 |
| 05 | o3-pro-2025-06-10_medium | OpenAI | 8.44 |
| 06 | o3-2025-04-16_medium | OpenAI | 8.39 |
| 07 | gemini-2.5-pro | Google DeepMind | 8.38 |
| 08 | claude-opus-4-20250514_16K | Anthropic | 8.36 |
| 09 | claude-opus-4-20250514 | Anthropic | 8.31 |
| 10 | gpt-5-mini-2025-08-07_medium | OpenAI | 8.31 |
| 11 | qwen3-235b-a22b | Alibaba | 8.3 |
| 12 | DeepSeek-R1 | DeepSeek | 8.3 |
| 13 | Qwen3-235B-A22B-Thinking-2507 | Alibaba | 8.24 |
| 14 | DeepSeek-R1-0528 | DeepSeek | 8.19 |
| 15 | gpt-4o-2024-11-20 | OpenAI | 8.18 |
| 16 | claude-sonnet-4-20250514_16K | Anthropic | 8.14 |
| 17 | claude-3-7-sonnet-20250219_16K | Anthropic | 8.11 |
| 18 | claude-sonnet-4-20250514 | Anthropic | 8.09 |
| 19 | gemini-2.5-pro-preview-05-06 | Google DeepMind | 8.09 |
| 20 | gemini-2.5-pro-exp-03-25 | Google DeepMind | 8.05 |
| 21 | claude-3-5-sonnet-20241022 | Anthropic | 8.03 |
| 22 | QwQ-32B (16K thinking) | Alibaba | 8.02 |
| 23 | gemma-3-27b-it | Google DeepMind | 7.99 |
| 24 | claude-3-7-sonnet-20250219 | Anthropic | 7.94 |
| 25 | mistral-medium-2505 | Mistral AI | 7.73 |
| 26 | gpt-oss-120b | OpenAI | 7.71 |
| 27 | DeepSeek-V3-0324 | DeepSeek | 7.7 |
| 28 | grok-4-0709 | xAI | 7.69 |
| 29 | gemini-2.5-flash-preview-04-17 (24K thinking) | Google DeepMind | 7.65 |
| 30 | grok-3-beta | xAI | 7.64 |
| 31 | gpt-4.5-preview-2025-02-27 | OpenAI | 7.56 |
| 32 | Qwen3-30B-A3B | Alibaba | 7.53 |
| 33 | o4-mini-2025-04-16_medium | OpenAI | 7.5 |
| 34 | gemini-2.0-flash-thinking-exp-01-21 | Google DeepMind,Google | 7.38 |
| 35 | grok-3-mini-beta_low | xAI | 7.35 |
| 36 | claude-3-5-haiku-20241022 | Anthropic | 7.35 |
| 37 | glm-4.5 | Z.ai (Zhipu AI),Tsinghua University | 7.34 |
| 38 | qwen2.5-max | Alibaba | 7.29 |
| 39 | gemini-2.0-flash-exp | Google DeepMind,Google | 7.15 |
| 40 | o1-2024-12-17_medium | OpenAI | 7.02 |
| 41 | mistral-large-2407 | Mistral AI | 6.9 |
| 42 | gpt-4o-mini-2024-07-18 | OpenAI | 6.72 |
| 43 | o1-mini-2024-09-12_medium | OpenAI | 6.49 |
| 44 | grok-2-1212 | xAI | 6.36 |
| 45 | phi-4 | Microsoft Research | 6.26 |
| 46 | Llama-4-Maverick-17B-128E-Instruct | Meta AI | 6.2 |
| 47 | o3-mini-2025-01-31_high | OpenAI | 6.17 |
| 48 | o3-mini-2025-01-31_medium | OpenAI | 6.15 |
| 49 | amazon.nova-pro-v1:0 | Amazon | 6.05 |