Live
AI models

GShard (dense)

Training compute
4.8×10²² FLOP
Parameters
2.3B
Published
Jun 30, 2020

GShard (dense) is an AI model developed by Google (United States), first published in June 2020. It works in the language domain, on tasks such as translation. It counts among the frontier models: the systems trained with the most compute of their moment.

Training it took an estimated 4.8×10²² FLOP of compute (estimation method: operation counting,hardware). The model has 2,300,000,000 parameters. It was trained on roughly 346.7B datapoints. Training ran on 1,024 Google TPU v3 for about 1K hours. The compute alone is estimated at $266K in 2023 dollars.

Access: Unreleased. Its weights are not openly released. The reference paper has 1,992 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
Google
Country of organization
United States
Domain
Language
Task
Translation
Training compute
4.8×10²² FLOP
Compute estimation method
Operation counting, Hardware
Parameters
2,300,000,000
Dataset size
346.7B
Training hardware
Google TPU v3
Chips used
1,024
Training time
1,008 h
Chip-hours
1M
Training power draw
939.5 kW
Training cost (2023 USD)
$266K
Numerical format
FP32
Model accessibility
Unreleased
Open weights
No
Citations
1,992
Epoch confidence
Confident
More from Google
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models