Live
Benchmarks

Lech Mazur Writing

Models scored
49
Best score
8.6
Metric
Mean score

Lech Mazur Writing is an AI evaluation tracked by Epoch AI, with 49 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.

Scored models span July 2024 to August 2025. The best score recorded is 8.6, measured as Mean score. Models come from OpenAI, Alibaba, DeepSeek, Anthropic, Google DeepMind, xAI, Google DeepMind,Google, Mistral AI and others.

Full record
Score metric
Mean score
Models scored
49
Best score recorded
8.6
Earliest model scored
Jul 18, 2024
Most recent model scored
Aug 7, 2025
Third-party reported
Yes
Leaderboard
01gpt-5-2025-08-07_mediumOpenAI8.6
02Kimi-K2-InstructMoonshot8.56
03claude-opus-4-1-20250805Anthropic8.47
04claude-opus-4-1-20250805_16KAnthropic8.45
05o3-pro-2025-06-10_mediumOpenAI8.44
06o3-2025-04-16_mediumOpenAI8.39
07gemini-2.5-proGoogle DeepMind8.38
08claude-opus-4-20250514_16KAnthropic8.36
09claude-opus-4-20250514Anthropic8.31
10gpt-5-mini-2025-08-07_mediumOpenAI8.31
11qwen3-235b-a22bAlibaba8.3
12DeepSeek-R1DeepSeek8.3
13Qwen3-235B-A22B-Thinking-2507Alibaba8.24
14DeepSeek-R1-0528DeepSeek8.19
15gpt-4o-2024-11-20OpenAI8.18
16claude-sonnet-4-20250514_16KAnthropic8.14
17claude-3-7-sonnet-20250219_16KAnthropic8.11
18claude-sonnet-4-20250514Anthropic8.09
19gemini-2.5-pro-preview-05-06Google DeepMind8.09
20gemini-2.5-pro-exp-03-25Google DeepMind8.05
21claude-3-5-sonnet-20241022Anthropic8.03
22QwQ-32B (16K thinking)Alibaba8.02
23gemma-3-27b-itGoogle DeepMind7.99
24claude-3-7-sonnet-20250219Anthropic7.94
25mistral-medium-2505Mistral AI7.73
26gpt-oss-120bOpenAI7.71
27DeepSeek-V3-0324DeepSeek7.7
28grok-4-0709xAI7.69
29gemini-2.5-flash-preview-04-17 (24K thinking)Google DeepMind7.65
30grok-3-betaxAI7.64
31gpt-4.5-preview-2025-02-27OpenAI7.56
32Qwen3-30B-A3BAlibaba7.53
33o4-mini-2025-04-16_mediumOpenAI7.5
34gemini-2.0-flash-thinking-exp-01-21Google DeepMind,Google7.38
35grok-3-mini-beta_lowxAI7.35
36claude-3-5-haiku-20241022Anthropic7.35
37glm-4.5Z.ai (Zhipu AI),Tsinghua University7.34
38qwen2.5-maxAlibaba7.29
39gemini-2.0-flash-expGoogle DeepMind,Google7.15
40o1-2024-12-17_mediumOpenAI7.02
41mistral-large-2407Mistral AI6.9
42gpt-4o-mini-2024-07-18OpenAI6.72
43o1-mini-2024-09-12_mediumOpenAI6.49
44grok-2-1212xAI6.36
45phi-4Microsoft Research6.26
46Llama-4-Maverick-17B-128E-InstructMeta AI6.2
47o3-mini-2025-01-31_highOpenAI6.17
48o3-mini-2025-01-31_mediumOpenAI6.15
49amazon.nova-pro-v1:0Amazon6.05
More benchmarks
SourceEpoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/benchmarks. Licensed under CC BY 4.0.
← All benchmarks