Step-Omni is an AI model developed by StepFun (China), first published in February 2025. It works in the speech, language, multimodal and vision domain, on tasks such as speech synthesis, speech recognition (asr), speech-to-text and 6 more.
Training it took an estimated 2.5×10²⁴ FLOP of compute (estimation method: operation counting). The model has 130,000,000,000 parameters. It was trained on roughly 2.5T datapoints. Training ran on NVIDIA H800 SXM5.
Access: Unreleased. Its weights are not openly released. It is built on top of Step-1. Epoch AI rates the confidence of this record as confident.