VideoMAE V2 is an AI model developed by Nanjing University, Shenzhen Institute of Advanced Technology and Shanghai AI Lab (China), first published in March 2023. It works in the video domain, on tasks such as action recognition.
Training it took an estimated 9.7×10²¹ FLOP of compute (estimation method: hardware). The model has 1,000,000,000 parameters. It was trained on roughly 1.2B datapoints. Training ran on 64 NVIDIA A100 SXM4 80 GB for about 336 hours. The compute alone is estimated at $18K in 2023 dollars.
Access: Open weights (unrestricted). Its weights are openly available. It is built on top of ViT-G/14. The reference paper has 664 citations. Epoch AI rates the confidence of this record as confident.