Grok 4 (xAI) and Llama Nemotron Ultra 253B (NVIDIA) are both frontier AI models. Grok 4 was published in July 2025 and Llama Nemotron Ultra 253B in March 2025.
Grok 4 was trained on 5×10²⁶ FLOP, about 12.8x the compute of Llama Nemotron Ultra 253B at 3.9×10²⁵ FLOP. Training compute is the closest available proxy for how much was invested in a model, though it says nothing on its own about how well that compute was spent.
These two models share no benchmark on which both have been scored, so no direct performance comparison is possible here. The specification table below is a comparison of inputs, not of results.