Live
AI models

Step-Video-T2V

Training compute
4.1×10²⁴ FLOP
Parameters
30B
Published
Feb 24, 2025

Step-Video-T2V is an AI model developed by StepFun (China), first published in February 2025. It works in the video domain, on tasks such as video generation and text-to-video.

Training it took an estimated 4.1×10²⁴ FLOP of compute (estimation method: hardware). The model has 30,000,000,000 parameters. Training ran on 5,000 NVIDIA H800 SXM5 for about 720 hours.

Access: Open weights (unrestricted). Its weights are openly available. Epoch AI rates the confidence of this record as likely.

Full record
Organization
StepFun
Country of organization
China
Domain
Video
Task
Video generation, Text-to-video
Training compute
4.1×10²⁴ FLOP
Compute estimation method
Hardware
Parameters
30,000,000,000
Training hardware
NVIDIA H800 SXM5
Chips used
5,000
Training time
720 h
Training power draw
6.9 MW
Model accessibility
Open weights (unrestricted)
Open weights
Yes
Epoch confidence
Likely
More from StepFun
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models