Live
AI models

DeepSeek-VL-1.3B

Training compute
8.7×10²¹ FLOP
Parameters
1.3B
Published
Mar 8, 2024

DeepSeek-VL-1.3B is an AI model developed by DeepSeek (China), first published in March 2024. It works in the multimodal, vision and language domain, on tasks such as character recognition (ocr), language modeling/generation, visual question answering and question answering.

Training it took an estimated 8.7×10²¹ FLOP of compute (estimation method: operation counting,hardware). The model has 1,300,000,000 parameters. It was trained on roughly 400M datapoints. Training ran on 128 NVIDIA A100 for about 168 hours.

Access: Open weights (restricted use). Its weights are openly available. It is built on top of DeepSeek-LLM-1.3b-base. The reference paper has 797 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
DeepSeek
Country of organization
China
Domain
Multimodal, Vision, Language
Task
Character recognition (OCR), Language modeling/generation, Visual question answering, Question answering
Training compute
8.7×10²¹ FLOP
Compute estimation method
Operation counting, Hardware
Parameters
1,300,000,000
Dataset size
400M
Training hardware
NVIDIA A100
Chips used
128
Training time
168 h
Training power draw
101.3 kW
Model accessibility
Open weights (restricted use)
Open weights
Yes
Base model
DeepSeek-LLM-1.3b-base
Citations
797
Epoch confidence
Confident
More from DeepSeek
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models