MoE-Multi is an AI model developed by Jagiellonian University and Google Brain (Poland and United States), first published in January 2017. It works in the language domain, on tasks such as language modeling and translation.
Training it took an estimated 9.4×10¹⁹ FLOP of compute (estimation method: hardware). The model has 8,700,000,000 parameters. It was trained on roughly 87B datapoints. Training ran on 64 NVIDIA Tesla K40t for about 288 hours. The compute alone is estimated at $4K in 2023 dollars.
Access: Unreleased. Its weights are not openly released. The reference paper has 4,587 citations. Epoch AI rates the confidence of this record as confident.