Live
AI models

GPT-2 Medium (FlashAttention)

Training compute
8.9×10²⁰ FLOP
Parameters
355M
Published
May 27, 2022

GPT-2 Medium (FlashAttention) is an AI model developed by Stanford University and University at Buffalo (United States), first published in May 2022. It works in the language domain, on tasks such as language modeling/generation.

Training it took an estimated 8.9×10²⁰ FLOP of compute (estimation method: comparison with other models,hardware). The model has 355,000,000 parameters. It was trained on roughly 10.1B datapoints. Training ran on 8 NVIDIA A100 SXM4 40 GB for about 166 hours.

Access: Open weights (unrestricted). Its weights are openly available. The reference paper has 4,295 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
Stanford University, University at Buffalo
Country of organization
United States
Domain
Language
Task
Language modeling/generation
Training compute
8.9×10²⁰ FLOP
Compute estimation method
Comparison with other models, Hardware
Parameters
355,000,000
Dataset size
10.1B
Training hardware
NVIDIA A100 SXM4 40 GB
Chips used
8
Training time
166 h
Training power draw
6.4 kW
Model accessibility
Open weights (unrestricted)
Open weights
Yes
Citations
4,295
Epoch confidence
Confident
More from Stanford University,University at Buffalo
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models