Live
AI models

MoE-Multi

Training compute
9.4×10¹⁹ FLOP
Parameters
8.7B
Published
Jan 23, 2017

MoE-Multi is an AI model developed by Jagiellonian University and Google Brain (Poland and United States), first published in January 2017. It works in the language domain, on tasks such as language modeling and translation.

Training it took an estimated 9.4×10¹⁹ FLOP of compute (estimation method: hardware). The model has 8,700,000,000 parameters. It was trained on roughly 87B datapoints. Training ran on 64 NVIDIA Tesla K40t for about 288 hours. The compute alone is estimated at $4K in 2023 dollars.

Access: Unreleased. Its weights are not openly released. The reference paper has 4,587 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
Jagiellonian University, Google Brain
Country of organization
Poland, United States
Domain
Language
Task
Language modeling, Translation
Training compute
9.4×10¹⁹ FLOP
Compute estimation method
Hardware
Parameters
8,700,000,000
Dataset size
87B
Training hardware
NVIDIA Tesla K40t
Chips used
64
Training time
288 h
Training power draw
32.9 kW
Training cost (2023 USD)
$4K
Numerical format
FP32
Model accessibility
Unreleased
Open weights
No
Citations
4,587
Epoch confidence
Confident
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models