LLaVA-OV-7B is an AI model developed by ByteDance, Nanyang Technological University, Chinese University of Hong Kong (CUHK) and Hong Kong University of Science and Technology (HKUST) (China, Singapore and Hong Kong), first published in August 2024. It works in the multimodal, video, vision and language domain, on tasks such as image captioning, visual question answering, video description and 3 more.
Epoch AI has no training-compute estimate for this model. The model has 7,600,000,000 parameters. It was trained on roughly 945.1M datapoints.
Access: Open weights (unrestricted). Its weights are openly available. It is built on top of Qwen2-7B. The reference paper has 2,428 citations. Epoch AI rates the confidence of this record as confident.