Live
AI models

Verbatim Memory Transformer (108M)

Training compute
9.9×10¹⁷ FLOP
Parameters
107.7M
Published
Oct 24, 2022

Verbatim Memory Transformer (108M) is an AI model developed by Johns Hopkins University and New York University (NYU) (United States), first published in October 2022. It works in the language domain, on tasks such as language modeling.

Training it took an estimated 9.9×10¹⁷ FLOP of compute. The model has 107,700,000 parameters. It was trained on roughly 40M datapoints. Training ran on 1 NVIDIA Quadro RTX 8000 for about 12 hours.

Access: Unreleased. Its weights are not openly released. The reference paper has 9 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
Johns Hopkins University, New York University (NYU)
Country of organization
United States
Domain
Language
Task
Language modeling
Training compute
9.9×10¹⁷ FLOP
Parameters
107,700,000
Dataset size
40M
Training hardware
NVIDIA Quadro RTX 8000
Chips used
1
Training time
12 h
Training power draw
286 W
Model accessibility
Unreleased
Open weights
No
Citations
9
Epoch confidence
Confident
More from Johns Hopkins University,New York University (NYU)
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models