Live
AI models

LLaVA-OV-72B

Training compute
3×10²⁴ FLOP
Parameters
72B
Published
Aug 6, 2024

LLaVA-OV-72B is an AI model developed by ByteDance, Nanyang Technological University, Chinese University of Hong Kong (CUHK) and Hong Kong University of Science and Technology (HKUST) (China, Singapore and Hong Kong), first published in August 2024. It works in the multimodal, vision, language and video domain, on tasks such as image captioning, visual question answering, video description and 3 more.

Training it took an estimated 3×10²⁴ FLOP of compute (estimation method: operation counting). The model has 72,000,000,000 parameters. It was trained on roughly 38.3B datapoints.

Access: Open weights (unrestricted). Its weights are openly available. It is built on top of Qwen2-72B. The reference paper has 2,428 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
ByteDance, Nanyang Technological University, Chinese University of Hong Kong (CUHK), Hong Kong University of Science and Technology (HKUST)
Country of organization
China, Singapore, Hong Kong
Domain
Multimodal, Vision, Language, Video
Task
Image captioning, Visual question answering, Video description, Object recognition, Action recognition, Language modeling/generation
Training compute
3×10²⁴ FLOP
Compute estimation method
Operation counting
Parameters
72,000,000,000
Dataset size
38.3B
Numerical format
BF16
Model accessibility
Open weights (unrestricted)
Open weights
Yes
Base model
Qwen2-72B
Citations
2,428
Epoch confidence
Confident
More from ByteDance,Nanyang Technological University,Chinese University of Hong Kong (CUHK),Hong Kong University of Science and Technology (HKUST)
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models