En vivo
Head to head

Grok 3 vs Llama Nemotron Ultra 253B

xAI
Grok 3
February 2025
vs
3.5×10²⁶Training compute (FLOP)3.9×10²⁵
$218MTraining cost--

Grok 3 (xAI) and Llama Nemotron Ultra 253B (NVIDIA) are both frontier AI models. Grok 3 was published in February 2025 and Llama Nemotron Ultra 253B in March 2025.

Grok 3 was trained on 3.5×10²⁶ FLOP, about 8.9x the compute of Llama Nemotron Ultra 253B at 3.9×10²⁵ FLOP. Training compute is the closest available proxy for how much was invested in a model, though it says nothing on its own about how well that compute was spent.

These two models share no benchmark on which both have been scored, so no direct performance comparison is possible here. The specification table below is a comparison of inputs, not of results.

Specifications
xAI
Organization
NVIDIA
Feb 17, 2025
Published
Mar 18, 2025
3.5×10²⁶ FLOP
Training compute8.9x
3.9×10²⁵ FLOP
3T
Parameters12x
253B
--
Dataset size
603B
NVIDIA H100 SXM5 80GB
Training hardware
--
80,000
Chips used
--
2,160 h
Training time
--
$218M
Training cost (2023 USD)
--
109.9 MW
Training power draw
--
API access
Accessibility
Open weights (restricted use)
No
Open weights
Yes
United States
Country
United States
Related comparisons
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.