Step-Video-T2V is an AI model developed by StepFun (China), first published in February 2025. It works in the video domain, on tasks such as video generation and text-to-video.
Training it took an estimated 4.1×10²⁴ FLOP of compute (estimation method: hardware). The model has 30,000,000,000 parameters. Training ran on 5,000 NVIDIA H800 SXM5 for about 720 hours.
Access: Open weights (unrestricted). Its weights are openly available. Epoch AI rates the confidence of this record as likely.