Live
AI models

CLIP (ViT L/14@336px)

Training compute
10²² FLOP
Parameters
370M
Published
Jan 5, 2021

CLIP (ViT L/14@336px) is an AI model developed by OpenAI (United States), first published in January 2021. It works in the multimodal, vision, language and video domain, on tasks such as zero-shot image classification, character recognition (ocr) and video description.

Training it took an estimated 10²² FLOP of compute (estimation method: third-party estimation). The model has 370,000,000 parameters. It was trained on roughly 400M datapoints. Training ran on 256 NVIDIA V100 for about 288 hours. The compute alone is estimated at $25K in 2023 dollars.

Access: Open weights (unrestricted). Its weights are openly available. The reference paper has 48,743 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
OpenAI
Country of organization
United States
Domain
Multimodal, Vision, Language, Video
Task
Zero-shot image classification, Character recognition (OCR), Video description
Training compute
10²² FLOP
Compute estimation method
Third-party estimation
Parameters
370,000,000
Dataset size
400M
Training hardware
NVIDIA V100
Chips used
256
Training time
288 h
Training power draw
155.9 kW
Training cost (2023 USD)
$25K
Numerical format
FP16
Model accessibility
Open weights (unrestricted)
Open weights
Yes
Citations
48,743
Epoch confidence
Confident
More from OpenAI
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models