CLIP (ViT L/14@336px) is an AI model developed by OpenAI (United States), first published in January 2021. It works in the multimodal, vision, language and video domain, on tasks such as zero-shot image classification, character recognition (ocr) and video description.
Training it took an estimated 10²² FLOP of compute (estimation method: third-party estimation). The model has 370,000,000 parameters. It was trained on roughly 400M datapoints. Training ran on 256 NVIDIA V100 for about 288 hours. The compute alone is estimated at $25K in 2023 dollars.
Access: Open weights (unrestricted). Its weights are openly available. The reference paper has 48,743 citations. Epoch AI rates the confidence of this record as confident.