Live
AI models

VideoMAE V2

Training compute
9.7×10²¹ FLOP
Parameters
1B
Published
Mar 29, 2023

VideoMAE V2 is an AI model developed by Nanjing University, Shenzhen Institute of Advanced Technology and Shanghai AI Lab (China), first published in March 2023. It works in the video domain, on tasks such as action recognition.

Training it took an estimated 9.7×10²¹ FLOP of compute (estimation method: hardware). The model has 1,000,000,000 parameters. It was trained on roughly 1.2B datapoints. Training ran on 64 NVIDIA A100 SXM4 80 GB for about 336 hours. The compute alone is estimated at $18K in 2023 dollars.

Access: Open weights (unrestricted). Its weights are openly available. It is built on top of ViT-G/14. The reference paper has 664 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
Nanjing University, Shenzhen Institute of Advanced Technology, Shanghai AI Lab
Country of organization
China
Domain
Video
Task
Action recognition
Training compute
9.7×10²¹ FLOP
Compute estimation method
Hardware
Parameters
1,000,000,000
Dataset size
1.2B
Training hardware
NVIDIA A100 SXM4 80 GB
Chips used
64
Training time
336 h
Training power draw
51.0 kW
Training cost (2023 USD)
$18K
Numerical format
FP16
Model accessibility
Open weights (unrestricted)
Open weights
Yes
Base model
ViT-G/14
Citations
664
Epoch confidence
Confident
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models