Live
AI models

Megatron-Turing NLG 530B

Training compute
8.6×10²³ FLOP
Parameters
530B
Published
Oct 11, 2021

Megatron-Turing NLG 530B is an AI model developed by Microsoft and NVIDIA (United States), first published in October 2021. It works in the language domain, on tasks such as language modeling, language modeling/generation and question answering. It counts among the frontier models: the systems trained with the most compute of their moment.

Training it took an estimated 8.6×10²³ FLOP of compute (estimation method: third-party estimation). The model has 530,000,000,000 parameters. It was trained on roughly 270B datapoints. Training ran on 4,480 NVIDIA A100 SXM4 80 GB for about 770 hours. The compute alone is estimated at $4M in 2023 dollars.

Access: Unreleased. Its weights are not openly released. The reference paper has 847 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
Microsoft, NVIDIA
Country of organization
United States
Domain
Language
Task
Language modeling, Language modeling/generation, Question answering
Training compute
8.6×10²³ FLOP
Compute estimation method
Third-party estimation
Parameters
530,000,000,000
Dataset size
270B
Training hardware
NVIDIA A100 SXM4 80 GB
Chips used
4,480
Training time
770 h
Chip-hours
3.4M
Training power draw
3.6 MW
Training cost (2023 USD)
$4M
Numerical format
BF16
Model accessibility
Unreleased
Open weights
No
Citations
847
Epoch confidence
Confident
Benchmark results
01LAMBADAScore0.87
02BoolQScore0.85
03PIQAScore0.83
04HellaSwagOverall accuracy0.82
05WinoGrandeAccuracy0.79
06Adversarial NLIScore0.4
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models