T2R 75% + Pretrain (WT-103) is an AI model developed by University of Washington, Microsoft and DeepMind (United States and United Kingdom), first published in March 2021. It works in the language domain, on tasks such as language modeling.
Training it took an estimated 1.4×10¹⁹ FLOP of compute (estimation method: operation counting,hardware). The model has 668,893,184 parameters. Training ran on 8 NVIDIA V100 for about 11.8 hours.
Access: Unreleased. Its weights are not openly released. The reference paper has 94 citations. Epoch AI rates the confidence of this record as confident.