Live
AI models

LLaVA

Training compute
7.8×10²² FLOP
Parameters
13B
Published
Apr 17, 2023

LLaVA is an AI model developed by University of Wisconsin Madison, Microsoft Research and Columbia University (United States), first published in April 2023. It works in the multimodal, vision and language domain, on tasks such as chat, question answering and visual question answering.

Training it took an estimated 7.8×10²² FLOP of compute (estimation method: hardware). The model has 13,000,000,000 parameters. Training ran on 8 NVIDIA A100 for about 10 hours. The compute alone is estimated at $42 in 2023 dollars.

Access: Open weights (unrestricted). Its weights are openly available. It is built on top of Vicuna-13B v0. The reference paper has 9,391 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
University of Wisconsin Madison, Microsoft Research, Columbia University
Country of organization
United States
Domain
Multimodal, Vision, Language
Task
Chat, Question answering, Visual question answering
Training compute
7.8×10²² FLOP
Compute estimation method
Hardware
Parameters
13,000,000,000
Training hardware
NVIDIA A100
Chips used
8
Training time
10 h
Chip-hours
80
Training power draw
6.4 kW
Training cost (2023 USD)
$42
Numerical format
BF16
Model accessibility
Open weights (unrestricted)
Open weights
Yes
Base model
Vicuna-13B v0
Citations
9,391
Epoch confidence
Confident
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models