Live
AI models

GBERT-Large

Training compute
2.2×10²¹ FLOP
Parameters
335M
Published
Oct 21, 2020

GBERT-Large is an AI model developed by deepset and Bayerische Staatsbibliothek Muenchen (Germany), first published in October 2020. It works in the language domain, on tasks such as document classification and named entity recognition (ner).

Training it took an estimated 2.2×10²¹ FLOP of compute (estimation method: hardware). The model has 335,000,000 parameters. It was trained on roughly 5.5B datapoints. Training ran on 64 Google TPU v3 for about 264 hours. The compute alone is estimated at $4K in 2023 dollars.

Access: Open weights (unrestricted). Its weights are openly available. The reference paper has 333 citations. Epoch AI rates the confidence of this record as likely.

Full record
Organization
deepset, Bayerische Staatsbibliothek Muenchen
Country of organization
Germany
Domain
Language
Task
Document classification, Named entity recognition (NER)
Training compute
2.2×10²¹ FLOP
Compute estimation method
Hardware
Parameters
335,000,000
Dataset size
5.5B
Training hardware
Google TPU v3
Chips used
64
Training time
264 h
Chip-hours
16.9K
Training power draw
58.6 kW
Training cost (2023 USD)
$4K
Model accessibility
Open weights (unrestricted)
Open weights
Yes
Citations
333
Epoch confidence
Likely
More from deepset,Bayerische Staatsbibliothek Muenchen
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models