VASA-1 is an AI model developed by Microsoft Research Asia (China), first published in October 2024. It works in the video and audio domain, on tasks such as video generation.
Training it took an estimated 4×10¹⁹ FLOP of compute (estimation method: hardware). The model has 229,000,000 parameters. Training ran on 4 NVIDIA RTX A6000 for about 240 hours.
Access: Unreleased. Its weights are not openly released. Epoch AI rates the confidence of this record as confident.