OpenOmni is an AI model developed by Chinese Academy of Sciences, Shenzhen Institute of Advanced Technology, University of Chinese Academy of Sciences, National University of Singapore and University of Science and Technology of China (USTC) (China and Singapore), first published in May 2025. It works in the multimodal, language, vision and speech domain, on tasks such as speech-to-text, speech recognition (asr), image captioning and 5 more.
Epoch AI has no training-compute estimate for this model. The model has 7,000,000,000 parameters. Training ran on 8 NVIDIA A100.
Access: Open weights (unrestricted). Its weights are openly available. It is built on top of Qwen2.5 Instruct (7B),CLIP (ViT L/14@336px),Whisper v3. Epoch AI rates the confidence of this record as confident.