LLaVA is an AI model developed by University of Wisconsin Madison, Microsoft Research and Columbia University (United States), first published in April 2023. It works in the multimodal, vision and language domain, on tasks such as chat, question answering and visual question answering.
Training it took an estimated 7.8×10²² FLOP of compute (estimation method: hardware). The model has 13,000,000,000 parameters. Training ran on 8 NVIDIA A100 for about 10 hours. The compute alone is estimated at $42 in 2023 dollars.
Access: Open weights (unrestricted). Its weights are openly available. It is built on top of Vicuna-13B v0. The reference paper has 9,391 citations. Epoch AI rates the confidence of this record as confident.