DeepSpeech2 (English) is an AI model developed by Baidu Research - Silicon Valley AI Lab (United States), first published in December 2015. It works in the speech domain, on tasks such as speech recognition (asr). It counts among the frontier models: the systems trained with the most compute of their moment.
Training it took an estimated 2.6×10¹⁹ FLOP of compute (estimation method: operation counting,third-party estimation). The model has 38,000,000 parameters. It was trained on roughly 716.4M datapoints. Training ran on 16 NVIDIA GeForce GTX TITAN X for about 120 hours. The compute alone is estimated at $214 in 2023 dollars.
The reference paper has 3,150 citations. Epoch AI rates the confidence of this record as confident.