LSTM (Hebbian, Cache, MbPA) is an AI model developed by DeepMind and University College London (UCL) (United Kingdom), first published in March 2018. It works in the language domain, on tasks such as language modeling.
Training it took an estimated 3.3×10¹⁹ FLOP of compute (estimation method: hardware,operation counting). The model has 530,442,240 parameters. It was trained on roughly 175.2M datapoints. Training ran on 8 NVIDIA P100 for about 144 hours. The compute alone is estimated at $591 in 2023 dollars.
Access: Unreleased. Its weights are not openly released. The reference paper has 47 citations. Epoch AI rates the confidence of this record as confident.