Live
AI models

Megatron-BERT

Training compute
2.2×10²² FLOP
Parameters
3.9B
Published
Sep 17, 2019

Megatron-BERT is an AI model developed by NVIDIA (United States), first published in September 2019. It works in the language domain, on tasks such as language modeling/generation. It counts among the frontier models: the systems trained with the most compute of their moment.

Training it took an estimated 2.2×10²² FLOP of compute (estimation method: operation counting,third-party estimation). The model has 3,900,000,000 parameters. It was trained on roughly 7B datapoints. Training ran on 512 NVIDIA Tesla V100S PCIe 32 GB for about 374 hours. The compute alone is estimated at $172K in 2023 dollars.

Access: Unreleased. Its weights are not openly released. The reference paper has 2,766 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
NVIDIA
Country of organization
United States
Domain
Language
Task
Language modeling/generation
Training compute
2.2×10²² FLOP
Compute estimation method
Operation counting, Third-party estimation
Parameters
3,900,000,000
Dataset size
7B
Training hardware
NVIDIA Tesla V100S PCIe 32 GB
Chips used
512
Training time
374 h
Chip-hours
191.5K
Training power draw
262.6 kW
Training cost (2023 USD)
$172K
Numerical format
FP16
Model accessibility
Unreleased
Open weights
No
Citations
2,766
Epoch confidence
Confident
More from NVIDIA
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models