VILA-7B is an AI model developed by NVIDIA and Massachusetts Institute of Technology (MIT) (United States), first published in May 2024. It works in the multimodal, language, vision and video domain, on tasks such as chat, visual question answering, image captioning and 2 more.
Epoch AI has no training-compute estimate for this model. The model has 7,000,000,000 parameters. Training ran on 128 NVIDIA A100.
Access: Unreleased. Its weights are not openly released. It is built on top of Llama 2-7B. The reference paper has 827 citations. Epoch AI rates the confidence of this record as confident.