VILA1.5-13B is an AI model developed by NVIDIA and Massachusetts Institute of Technology (MIT) (United States), first published in May 2024. It works in the multimodal, language, vision and video domain, on tasks such as chat, visual question answering, image captioning and 2 more.
Training it took an estimated 2.3×10²¹ FLOP of compute. The model has 13,493,916,736 parameters. It was trained on roughly 32.4B datapoints. Training ran on 128 NVIDIA A100.
Access: Open weights (non-commercial). Its weights are openly available. It is built on top of SigLIP 400M,Vicuna-13B-v1.5. The reference paper has 827 citations. Epoch AI rates the confidence of this record as confident.