Live
AI models

VILA-13B

Training compute
2.3×10²¹ FLOP
Parameters
13.4B
Published
Dec 12, 2023

VILA-13B is an AI model developed by NVIDIA and Massachusetts Institute of Technology (MIT) (United States), first published in December 2023. It works in the multimodal, language and vision domain, on tasks such as chat, visual question answering, image captioning and 2 more.

Training it took an estimated 2.3×10²¹ FLOP of compute. The model has 13,350,839,296 parameters. It was trained on roughly 32.4B datapoints. Training ran on 128 NVIDIA A100 SXM4 80 GB.

Access: Open weights (non-commercial). Its weights are openly available. It is built on top of Llama 2-13B,CLIP (ViT L/14@336px). The reference paper has 827 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
NVIDIA, Massachusetts Institute of Technology (MIT)
Country of organization
United States
Domain
Multimodal, Language, Vision
Task
Chat, Visual question answering, Image captioning, Language modeling/generation, Question answering
Training compute
2.3×10²¹ FLOP
Parameters
13,350,839,296
Dataset size
32.4B
Training hardware
NVIDIA A100 SXM4 80 GB
Chips used
128
Training power draw
101.5 kW
Numerical format
BF16
Model accessibility
Open weights (non-commercial)
Open weights
Yes
Base model
Llama 2-13B, CLIP (ViT L/14@336px)
Citations
827
Epoch confidence
Confident
More from NVIDIA,Massachusetts Institute of Technology (MIT)
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models