Live
AI models

SRU++ Large only 2 attention layers (k=5) (WT103)

Training compute
3.6×10¹⁹ FLOP
Parameters
225M
Published
Feb 24, 2021

SRU++ Large only 2 attention layers (k=5) (WT103) is an AI model developed by ASAPP (United States), first published in February 2021. It works in the language domain, on tasks such as language modeling.

Training it took an estimated 3.6×10¹⁹ FLOP of compute (estimation method: hardware). The model has 225,000,000 parameters. Training ran on 8 NVIDIA V100 for about 33 hours.

Access: Unreleased. Its weights are not openly released. The reference paper has 54 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
ASAPP
Country of organization
United States
Domain
Language
Task
Language modeling
Training compute
3.6×10¹⁹ FLOP
Compute estimation method
Hardware
Parameters
225,000,000
Training hardware
NVIDIA V100
Chips used
8
Training time
33 h
Training power draw
4.9 kW
Model accessibility
Unreleased
Open weights
No
Citations
54
Epoch confidence
Confident
More from ASAPP
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models