Live
AI models

VILA-7B

Parameters
7B
Published
May 15, 2024

VILA-7B is an AI model developed by NVIDIA and Massachusetts Institute of Technology (MIT) (United States), first published in May 2024. It works in the multimodal, language, vision and video domain, on tasks such as chat, visual question answering, image captioning and 2 more.

Epoch AI has no training-compute estimate for this model. The model has 7,000,000,000 parameters. Training ran on 128 NVIDIA A100.

Access: Unreleased. Its weights are not openly released. It is built on top of Llama 2-7B. The reference paper has 827 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
NVIDIA, Massachusetts Institute of Technology (MIT)
Country of organization
United States
Domain
Multimodal, Language, Vision, Video
Task
Chat, Visual question answering, Image captioning, Language modeling/generation, Question answering
Compute estimation method
Hardware
Parameters
7,000,000,000
Training hardware
NVIDIA A100
Chips used
128
Chip-hours
5.1K
Training power draw
101.1 kW
Model accessibility
Unreleased
Open weights
No
Base model
Llama 2-7B
Citations
827
Epoch confidence
Confident
More from NVIDIA,Massachusetts Institute of Technology (MIT)
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models