Live
AI models

Megatron-LM (8.3B)

Training compute
9.1×10²¹ FLOP
Parameters
8.3B
Published
Sep 17, 2019

Megatron-LM (8.3B) is an AI model developed by NVIDIA (United States), first published in September 2019. It works in the language domain, on tasks such as language modeling/generation. It counts among the frontier models: the systems trained with the most compute of their moment.

Training it took an estimated 9.1×10²¹ FLOP of compute (estimation method: hardware,operation counting,third-party estimation). The model has 8,300,000,000 parameters. It was trained on roughly 46.4B datapoints. Training ran on 512 NVIDIA Tesla V100 DGXS 32 GB for about 327 hours. The compute alone is estimated at $109K in 2023 dollars.

Access: Unreleased. Its weights are not openly released. The reference paper has 2,766 citations. Epoch AI rates the confidence of this record as likely.

Full record
Organization
NVIDIA
Country of organization
United States
Domain
Language
Task
Language modeling/generation
Training compute
9.1×10²¹ FLOP
Compute estimation method
Hardware, Operation counting, Third-party estimation
Parameters
8,300,000,000
Dataset size
46.4B
Training hardware
NVIDIA Tesla V100 DGXS 32 GB
Chips used
512
Training time
327 h
Chip-hours
167.4K
Training power draw
262.6 kW
Training cost (2023 USD)
$109K
Numerical format
FP16
Model accessibility
Unreleased
Open weights
No
Citations
2,766
Epoch confidence
Likely
More from NVIDIA
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models