Live
AI models

LLaVA-OV-7B

Parameters
7.6B
Published
Aug 6, 2024

LLaVA-OV-7B is an AI model developed by ByteDance, Nanyang Technological University, Chinese University of Hong Kong (CUHK) and Hong Kong University of Science and Technology (HKUST) (China, Singapore and Hong Kong), first published in August 2024. It works in the multimodal, video, vision and language domain, on tasks such as image captioning, visual question answering, video description and 3 more.

Epoch AI has no training-compute estimate for this model. The model has 7,600,000,000 parameters. It was trained on roughly 945.1M datapoints.

Access: Open weights (unrestricted). Its weights are openly available. It is built on top of Qwen2-7B. The reference paper has 2,428 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
ByteDance, Nanyang Technological University, Chinese University of Hong Kong (CUHK), Hong Kong University of Science and Technology (HKUST)
Country of organization
China, Singapore, Hong Kong
Domain
Multimodal, Video, Vision, Language
Task
Image captioning, Visual question answering, Video description, Object recognition, Action recognition, Language modeling/generation
Parameters
7,600,000,000
Dataset size
945.1M
Numerical format
BF16
Model accessibility
Open weights (unrestricted)
Open weights
Yes
Base model
Qwen2-7B
Citations
2,428
Epoch confidence
Confident
More from ByteDance,Nanyang Technological University,Chinese University of Hong Kong (CUHK),Hong Kong University of Science and Technology (HKUST)
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models