Live
AI models

Step-Omni

Training compute
2.5×10²⁴ FLOP
Parameters
130B
Published
Feb 18, 2025

Step-Omni is an AI model developed by StepFun (China), first published in February 2025. It works in the speech, language, multimodal and vision domain, on tasks such as speech synthesis, speech recognition (asr), speech-to-text and 6 more.

Training it took an estimated 2.5×10²⁴ FLOP of compute (estimation method: operation counting). The model has 130,000,000,000 parameters. It was trained on roughly 2.5T datapoints. Training ran on NVIDIA H800 SXM5.

Access: Unreleased. Its weights are not openly released. It is built on top of Step-1. Epoch AI rates the confidence of this record as confident.

Full record
Organization
StepFun
Country of organization
China
Domain
Speech, Language, Multimodal, Vision
Task
Speech synthesis, Speech recognition (ASR), Speech-to-text, Text-to-speech (TTS), Audio question answering, Audio generation, Image captioning, Visual question answering, Speech-to-speech
Training compute
2.5×10²⁴ FLOP
Compute estimation method
Operation counting
Parameters
130,000,000,000
Dataset size
2.5T
Training hardware
NVIDIA H800 SXM5
Model accessibility
Unreleased
Open weights
No
Base model
Step-1
Epoch confidence
Confident
More from StepFun
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models