Live
AI models

Transformer (Adaptive Input Embeddings) WT103

Training compute
4.5×10¹⁹ FLOP
Parameters
247M
Published
Sep 28, 2018

Transformer (Adaptive Input Embeddings) WT103 is an AI model developed by Facebook AI Research (United States and France), first published in September 2018. It works in the language domain, on tasks such as language modeling.

Training it took an estimated 4.5×10¹⁹ FLOP of compute (estimation method: hardware,operation counting). The model has 247,000,000 parameters. It was trained on roughly 100M datapoints. Training ran on 8 NVIDIA V100 for about 67 hours. The compute alone is estimated at $3K in 2023 dollars.

Access: Open weights (unrestricted). Its weights are openly available. The reference paper has 433 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
Facebook AI Research
Country of organization
United States, France
Domain
Language
Task
Language modeling
Training compute
4.5×10¹⁹ FLOP
Compute estimation method
Hardware, Operation counting
Parameters
247,000,000
Dataset size
100M
Training hardware
NVIDIA V100
Chips used
8
Training time
67 h
Chip-hours
4.3K
Training power draw
5.0 kW
Training cost (2023 USD)
$3K
Numerical format
FP16
Model accessibility
Open weights (unrestricted)
Open weights
Yes
Citations
433
Epoch confidence
Confident
More from Facebook AI Research
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models