Live
AI models

Tensorized Transformer (257M)

Training compute
4.8×10¹⁸ FLOP
Parameters
257M
Published
Jun 24, 2019

Tensorized Transformer (257M) is an AI model developed by Tianjin University, Microsoft Research Asia and Beijing Institute of Technology (China), first published in June 2019. It works in the language domain, on tasks such as language modeling/generation.

Training it took an estimated 4.8×10¹⁸ FLOP of compute (estimation method: operation counting). The model has 257,000,000 parameters. It was trained on roughly 103M datapoints. Training ran on 2 NVIDIA P40.

Access: Unreleased. Its weights are not openly released. The reference paper has 194 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
Tianjin University, Microsoft Research Asia, Beijing Institute of Technology
Country of organization
China
Domain
Language
Task
Language modeling/generation
Training compute
4.8×10¹⁸ FLOP
Compute estimation method
Operation counting
Parameters
257,000,000
Dataset size
103M
Training hardware
NVIDIA P40
Chips used
2
Training power draw
1.0 kW
Model accessibility
Unreleased
Open weights
No
Citations
194
Epoch confidence
Confident
More from Tianjin University,Microsoft Research Asia,Beijing Institute of Technology
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models