GPT-2 Medium (FlashAttention) is an AI model developed by Stanford University and University at Buffalo (United States), first published in May 2022. It works in the language domain, on tasks such as language modeling/generation.
Training it took an estimated 8.9×10²⁰ FLOP of compute (estimation method: comparison with other models,hardware). The model has 355,000,000 parameters. It was trained on roughly 10.1B datapoints. Training ran on 8 NVIDIA A100 SXM4 40 GB for about 166 hours.
Access: Open weights (unrestricted). Its weights are openly available. The reference paper has 4,295 citations. Epoch AI rates the confidence of this record as confident.