Live
AI models

Transformer-XL (257M)

Training compute
3.8×10²⁰ FLOP
Parameters
257M
Published
Jan 9, 2019

Transformer-XL (257M) is an AI model developed by Carnegie Mellon University (CMU) and Google Brain (United States), first published in January 2019. It works in the language domain, on tasks such as language modeling/generation.

Training it took an estimated 3.8×10²⁰ FLOP of compute (estimation method: operation counting,hardware). The model has 257,000,000 parameters. It was trained on roughly 103M datapoints. Training ran on 32 Google TPU v3.

Access: Open weights (unrestricted). Its weights are openly available. The reference paper has 4,320 citations. Epoch AI rates the confidence of this record as speculative.

Full record
Organization
Carnegie Mellon University (CMU), Google Brain
Country of organization
United States
Domain
Language
Task
Language modeling/generation
Training compute
3.8×10²⁰ FLOP
Compute estimation method
Operation counting, Hardware
Parameters
257,000,000
Dataset size
103M
Training hardware
Google TPU v3
Chips used
32
Training power draw
29.7 kW
Numerical format
FP32
Model accessibility
Open weights (unrestricted)
Open weights
Yes
Citations
4,320
Epoch confidence
Speculative
More from Carnegie Mellon University (CMU),Google Brain
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models