Live
AI models

DeepSeek-V2 (MoE-236B)

Training compute
10²⁴ FLOP
Parameters
236B
Published
May 7, 2024

DeepSeek-V2 (MoE-236B) is an AI model developed by DeepSeek (China), first published in May 2024. It works in the language domain, on tasks such as language modeling/generation, chat and code generation.

Training it took an estimated 10²⁴ FLOP of compute (estimation method: operation counting). The model has 236,000,000,000 parameters. It was trained on roughly 8.1T datapoints. Training ran on NVIDIA H800 SXM5.

Access: Open weights (restricted use). Its weights are openly available. Epoch AI rates the confidence of this record as confident.

Full record
Organization
DeepSeek
Country of organization
China
Domain
Language
Task
Language modeling/generation, Chat, Code generation
Training compute
10²⁴ FLOP
Compute estimation method
Operation counting
Parameters
236,000,000,000
Dataset size
8.1T
Training hardware
NVIDIA H800 SXM5
Chip-hours
172.8K
Model accessibility
Open weights (restricted use)
Open weights
Yes
Epoch confidence
Confident
Benchmark results
01Epoch Capabilities IndexECI Score124.6
More from DeepSeek
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models