Verbatim Memory Transformer (108M) is an AI model developed by Johns Hopkins University and New York University (NYU) (United States), first published in October 2022. It works in the language domain, on tasks such as language modeling.
Training it took an estimated 9.9×10¹⁷ FLOP of compute. The model has 107,700,000 parameters. It was trained on roughly 40M datapoints. Training ran on 1 NVIDIA Quadro RTX 8000 for about 12 hours.
Access: Unreleased. Its weights are not openly released. The reference paper has 9 citations. Epoch AI rates the confidence of this record as confident.