Live
AI models

LLaVA-CoT

Parameters
11B
Published
Nov 15, 2024

LLaVA-CoT is an AI model developed by Peking University, Tsinghua University, Peng Cheng Laboratory, Alibaba DAMO Academy and Lehigh University (China and United States), first published in November 2024. It works in the language, vision and multimodal domain, on tasks such as visual question answering, language modeling/generation and quantitative reasoning.

Epoch AI has no training-compute estimate for this model. The model has 11,000,000,000 parameters. It was trained on roughly 24M datapoints. Training ran on 8 NVIDIA H100 SXM5 80GB.

Access: Open weights (non-commercial). Its weights are openly available. It is built on top of Llama 3.2 11B. The reference paper has 459 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
Peking University, Tsinghua University, Peng Cheng Laboratory, Alibaba DAMO Academy, Lehigh University
Country of organization
China, United States
Domain
Language, Vision, Multimodal
Task
Visual question answering, Language modeling/generation, Quantitative reasoning
Parameters
11,000,000,000
Dataset size
24M
Training hardware
NVIDIA H100 SXM5 80GB
Chips used
8
Training power draw
11.0 kW
Model accessibility
Open weights (non-commercial)
Open weights
Yes
Base model
Llama 3.2 11B
Citations
459
Epoch confidence
Confident
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models