GPT-4.5 (OpenAI) and Grok 3 (xAI) are both frontier AI models. GPT-4.5 was published in February 2025 and Grok 3 in February 2025.
GPT-4.5 was trained on 3.8×10²⁶ FLOP, essentially the same compute as Grok 3 at 3.5×10²⁶ FLOP. Training compute is the closest available proxy for how much was invested in a model, though it says nothing on its own about how well that compute was spent.
Grok 3 wins the single benchmark both were scored on. Benchmark counts are a crude scoreboard — the evaluations differ wildly in what they measure and in how saturated they are — so the per-benchmark table below matters more than the tally.