Falcon-180B (Technology Innovation Institute) and PaLM (540B) (Google Research) are both frontier AI models. Falcon-180B was published in September 2023 and PaLM (540B) in April 2022.
Falcon-180B was trained on 3.8 x 10^24 FLOP, about 1.5x the compute of PaLM (540B) at 2.5 x 10^24 FLOP. Training compute is the closest available proxy for how much was invested in a model, though it says nothing on its own about how well that compute was spent.
Falcon-180B leads on 6 of the 9 benchmarks they were both scored on, against 3 for PaLM (540B). Benchmark counts are a crude scoreboard — the evaluations differ wildly in what they measure and in how saturated they are — so the per-benchmark table below matters more than the tally.
| Field | Falcon-180B | PaLM (540B) |
|---|---|---|
| Organization | Technology Innovation Institute | Google Research |
| Published | Sep 6, 2023 | Apr 4, 2022 |
| Training compute | 3.8 x 10^24 FLOP | 2.5 x 10^24 FLOP |
| Parameters | 180B | 540.4B |
| Dataset size | 3.5T | 780B |
| Training hardware | NVIDIA A100 SXM4 40 GB | Google TPU v4 |
| Chips used | 4,096 | 6,144 |
| Training time | 4.3K h | 1.5K h |
| Training cost (2023 USD) | $11M | $3M |
| Training power draw | 3.3 MW | 4.2 MW |
| Accessibility | Open weights (restricted use) | Unreleased |
| Open weights | Yes | No |
| Country | United Arab Emirates | United States |
| Benchmark | Falcon-180B | PaLM (540B) |
|---|---|---|
| ARC AI2(Challenge score) | 0.68 | 0.85 |
| BoolQ(Score) | 0.89 | 0.89 |
| GSM8K(EM) | 0.54 | 0.56 |
| HellaSwag(Overall accuracy) | 0.89 | 0.84 |
| LAMBADA(Score) | 0.8 | 0.78 |
| MMLU(EM) | 0.71 | 0.69 |
| OpenBookQA(Accuracy) | 0.64 | 0.68 |
| PIQA(Score) | 0.85 | 0.82 |
| WinoGrande(Accuracy) | 0.87 | 0.85 |