Megatron-LM (8.3B) is an AI model developed by NVIDIA (United States), first published in September 2019. It works in the language domain, on tasks such as language modeling/generation. It counts among the frontier models: the systems trained with the most compute of their moment.
Training it took an estimated 9.1×10²¹ FLOP of compute (estimation method: hardware,operation counting,third-party estimation). The model has 8,300,000,000 parameters. It was trained on roughly 46.4B datapoints. Training ran on 512 NVIDIA Tesla V100 DGXS 32 GB for about 327 hours. The compute alone is estimated at $109K in 2023 dollars.
Access: Unreleased. Its weights are not openly released. The reference paper has 2,766 citations. Epoch AI rates the confidence of this record as likely.