Live
AI models

DeepSpeech2 (English)

Training compute
2.6×10¹⁹ FLOP
Parameters
38M
Published
Dec 8, 2015

DeepSpeech2 (English) is an AI model developed by Baidu Research - Silicon Valley AI Lab (United States), first published in December 2015. It works in the speech domain, on tasks such as speech recognition (asr). It counts among the frontier models: the systems trained with the most compute of their moment.

Training it took an estimated 2.6×10¹⁹ FLOP of compute (estimation method: operation counting,third-party estimation). The model has 38,000,000 parameters. It was trained on roughly 716.4M datapoints. Training ran on 16 NVIDIA GeForce GTX TITAN X for about 120 hours. The compute alone is estimated at $214 in 2023 dollars.

The reference paper has 3,150 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
Baidu Research - Silicon Valley AI Lab
Country of organization
United States
Domain
Speech
Task
Speech recognition (ASR)
Training compute
2.6×10¹⁹ FLOP
Compute estimation method
Operation counting, Third-party estimation
Parameters
38,000,000
Dataset size
716.4M
Training hardware
NVIDIA GeForce GTX TITAN X
Chips used
16
Training time
120 h
Chip-hours
1.9K
Training power draw
8.5 kW
Training cost (2023 USD)
$214
Numerical format
FP32
Citations
3,150
Epoch confidence
Confident
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models