LLaVA-OV-72B is an AI model developed by ByteDance, Nanyang Technological University, Chinese University of Hong Kong (CUHK) and Hong Kong University of Science and Technology (HKUST) (China, Singapore and Hong Kong), first published in August 2024. It works in the multimodal, vision, language and video domain, on tasks such as image captioning, visual question answering, video description and 3 more.
Training it took an estimated 3×10²⁴ FLOP of compute (estimation method: operation counting). The model has 72,000,000,000 parameters. It was trained on roughly 38.3B datapoints.
Access: Open weights (unrestricted). Its weights are openly available. It is built on top of Qwen2-72B. The reference paper has 2,428 citations. Epoch AI rates the confidence of this record as confident.