Falcon-180B (Technology Innovation Institute) and Llama 3.1-405B (Meta AI) are both frontier AI models. Falcon-180B was published in September 2023 and Llama 3.1-405B in July 2024.
Llama 3.1-405B was trained on 3.8 x 10^25 FLOP, about 10.1x the compute of Falcon-180B at 3.8 x 10^24 FLOP. Training compute is the closest available proxy for how much was invested in a model, though it says nothing on its own about how well that compute was spent.
Llama 3.1-405B leads on 6 of the 6 benchmarks they were both scored on, against 0 for Falcon-180B. Benchmark counts are a crude scoreboard — the evaluations differ wildly in what they measure and in how saturated they are — so the per-benchmark table below matters more than the tally.
| Field | Falcon-180B | Llama 3.1-405B |
|---|---|---|
| Organization | Technology Innovation Institute | Meta AI |
| Published | Sep 6, 2023 | Jul 23, 2024 |
| Training compute | 3.8 x 10^24 FLOP | 3.8 x 10^25 FLOP |
| Parameters | 180B | 405B |
| Dataset size | 3.5T | 15.6T |
| Training hardware | NVIDIA A100 SXM4 40 GB | NVIDIA H100 SXM5 80GB |
| Chips used | 4,096 | 16,384 |
| Training time | 4.3K h | 2.1K h |
| Training cost (2023 USD) | $11M | $53M |
| Training power draw | 3.3 MW | 22.6 MW |
| Accessibility | Open weights (restricted use) | Open weights (restricted use) |
| Open weights | Yes | Yes |
| Country | United Arab Emirates | United States |
| Benchmark | Falcon-180B | Llama 3.1-405B |
|---|---|---|
| ARC AI2(Challenge score) | 0.68 | 0.95 |
| Epoch Capabilities Index(ECI Score) | 111.9 | 129 |
| HellaSwag(Overall accuracy) | 0.89 | 0.89 |
| MMLU(EM) | 0.71 | 0.84 |
| PIQA(Score) | 0.85 | 0.86 |
| WinoGrande(Accuracy) | 0.87 | 0.89 |