Live
AI models

BLIP-2 (Q-Former)

Training compute
1.2×10²¹ FLOP
Parameters
1.5B
Published
Jan 30, 2023

BLIP-2 (Q-Former) is an AI model developed by Salesforce Research (United States), first published in January 2023. It works in the vision and language domain, on tasks such as visual question answering and image captioning.

Training it took an estimated 1.2×10²¹ FLOP of compute (estimation method: hardware). The model has 1,480,000,000 parameters. It was trained on roughly 2.3B datapoints. Training ran on 16 NVIDIA A100 SXM4 40 GB for about 200 hours. The compute alone is estimated at $2K in 2023 dollars.

Access: Open weights (unrestricted). Its weights are openly available. The reference paper has 8,019 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
Salesforce Research
Country of organization
United States
Domain
Vision, Language
Task
Visual question answering, Image captioning
Training compute
1.2×10²¹ FLOP
Compute estimation method
Hardware
Parameters
1,480,000,000
Dataset size
2.3B
Training hardware
NVIDIA A100 SXM4 40 GB
Chips used
16
Training time
200 h
Chip-hours
3.2K
Training power draw
12.8 kW
Training cost (2023 USD)
$2K
Numerical format
FP16
Model accessibility
Open weights (unrestricted)
Open weights
Yes
Citations
8,019
Epoch confidence
Confident
More from Salesforce Research
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models