SRU++ Large only 2 attention layers (k=5) (WT103) is an AI model developed by ASAPP (United States), first published in February 2021. It works in the language domain, on tasks such as language modeling.
Training it took an estimated 3.6×10¹⁹ FLOP of compute (estimation method: hardware). The model has 225,000,000 parameters. Training ran on 8 NVIDIA V100 for about 33 hours.
Access: Unreleased. Its weights are not openly released. The reference paper has 54 citations. Epoch AI rates the confidence of this record as confident.