Live
AI models

MegaScale (175B)

Training compute
2.7×10²³ FLOP
Parameters
175B
Published
Feb 23, 2024

MegaScale (175B) is an AI model developed by ByteDance and Peking University (China), first published in February 2024. It works in the language domain, on tasks such as language modeling/generation.

Training it took an estimated 2.7×10²³ FLOP of compute (estimation method: operation counting,hardware). The model has 175,000,000,000 parameters. It was trained on roughly 300B datapoints. Training ran on 12,288 NVIDIA A100 for about 42 hours.

Access: Unreleased. Its weights are not openly released. The reference paper has 302 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
ByteDance, Peking University
Country of organization
China
Domain
Language
Task
Language modeling/generation
Training compute
2.7×10²³ FLOP
Compute estimation method
Operation counting, Hardware
Parameters
175,000,000,000
Dataset size
300B
Training hardware
NVIDIA A100
Chips used
12,288
Training time
42 h
Training power draw
9.7 MW
Model accessibility
Unreleased
Open weights
No
Citations
302
Epoch confidence
Confident
More from ByteDance,Peking University
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models