Live
AI models

DeepSeek-V3 (Mar 2025)

Training compute
3.3×10²⁴ FLOP
Parameters
671B
Published
Mar 24, 2025

DeepSeek-V3 (Mar 2025) is an AI model developed by DeepSeek (China), first published in March 2025. It works in the language domain, on tasks such as language modeling/generation, code generation, quantitative reasoning and question answering.

Training it took an estimated 3.3×10²⁴ FLOP of compute (estimation method: operation counting,hardware). The model has 671,000,000,000 parameters. It was trained on roughly 14.8T datapoints. Training ran on 2,048 NVIDIA H800 SXM5. The compute alone is estimated at $5M in 2023 dollars.

Access: Open weights (restricted use). Its weights are openly available. Epoch AI rates the confidence of this record as confident.

Full record
Organization
DeepSeek
Country of organization
China
Domain
Language
Task
Language modeling/generation, Code generation, Quantitative reasoning, Question answering
Training compute
3.3×10²⁴ FLOP
Compute estimation method
Operation counting, Hardware
Parameters
671,000,000,000
Dataset size
14.8T
Training hardware
NVIDIA H800 SXM5
Chips used
2,048
Chip-hours
2.8M
Training power draw
2.8 MW
Training cost (2023 USD)
$5M
Numerical format
FP8
Model accessibility
Open weights (restricted use)
Open weights
Yes
Epoch confidence
Confident
Benchmark results
01Epoch Capabilities IndexECI Score137.1
More from DeepSeek
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models