GBERT-Large is an AI model developed by deepset and Bayerische Staatsbibliothek Muenchen (Germany), first published in October 2020. It works in the language domain, on tasks such as document classification and named entity recognition (ner).
Training it took an estimated 2.2×10²¹ FLOP of compute (estimation method: hardware). The model has 335,000,000 parameters. It was trained on roughly 5.5B datapoints. Training ran on 64 Google TPU v3 for about 264 hours. The compute alone is estimated at $4K in 2023 dollars.
Access: Open weights (unrestricted). Its weights are openly available. The reference paper has 333 citations. Epoch AI rates the confidence of this record as likely.