Grok 4 (xAI) and Llama 3.1-405B (Meta AI) are both frontier AI models. Grok 4 was published in July 2025 and Llama 3.1-405B in July 2024.
Grok 4 was trained on 5×10²⁶ FLOP, about 13.2x the compute of Llama 3.1-405B at 3.8×10²⁵ FLOP. Training compute is the closest available proxy for how much was invested in a model, though it says nothing on its own about how well that compute was spent.
Grok 4 wins the single benchmark both were scored on. Benchmark counts are a crude scoreboard — the evaluations differ wildly in what they measure and in how saturated they are — so the per-benchmark table below matters more than the tally.