Live
AI models

MegaScale (Production)

Training compute
3.9×10²⁴ FLOP
Parameters
530B
Published
Feb 23, 2024

MegaScale (Production) is an AI model developed by ByteDance and Peking University (China), first published in February 2024. It works in the language domain, on tasks such as language modeling/generation.

Training it took an estimated 3.9×10²⁴ FLOP of compute (estimation method: other). The model has 530,000,000,000 parameters. Training ran on 12,288 NVIDIA A100 for about 504 hours. The compute alone is estimated at $3M in 2023 dollars.

Access: Unreleased. Its weights are not openly released. The reference paper has 302 citations. Epoch AI rates the confidence of this record as speculative.

Full record
Organization
ByteDance, Peking University
Country of organization
China
Domain
Language
Task
Language modeling/generation
Training compute
3.9×10²⁴ FLOP
Compute estimation method
Other
Parameters
530,000,000,000
Training hardware
NVIDIA A100
Chips used
12,288
Training time
504 h
Training power draw
9.7 MW
Training cost (2023 USD)
$3M
Model accessibility
Unreleased
Open weights
No
Citations
302
Epoch confidence
Speculative
More from ByteDance,Peking University
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models