Live
AI models

Grover-Mega

Training compute
4.6×10²¹ FLOP
Parameters
1.5B
Published
May 29, 2019

Grover-Mega is an AI model developed by University of Washington (United States), first published in May 2019. It works in the language domain, on tasks such as language modeling/generation. It counts among the frontier models: the systems trained with the most compute of their moment.

Training it took an estimated 4.6×10²¹ FLOP of compute (estimation method: hardware,operation counting). The model has 1,500,000,000 parameters. It was trained on roughly 32B datapoints. Training ran on 128 Google TPU v3 for about 336 hours. The compute alone is estimated at $16K in 2023 dollars.

Access: Open weights (unrestricted). Its weights are openly available. The reference paper has 1,231 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
University of Washington
Country of organization
United States
Domain
Language
Task
Language modeling/generation
Training compute
4.6×10²¹ FLOP
Compute estimation method
Hardware, Operation counting
Parameters
1,500,000,000
Dataset size
32B
Training hardware
Google TPU v3
Chips used
128
Training time
336 h
Training power draw
118.5 kW
Training cost (2023 USD)
$16K
Model accessibility
Open weights (unrestricted)
Open weights
Yes
Citations
1,231
Epoch confidence
Confident
More from University of Washington
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models