LAMBADA is an AI evaluation tracked by Epoch AI, with 53 scored model versions on record. The scores here are reported by third parties rather than produced by Epoch AI running the evaluation itself.
Scored models span December 2021 to November 2023. The best score recorded is 0.87, measured as Score. Models come from MosaicML, Baichuan, Meta AI, Stability AI, Alibaba, OpenAI, DeepMind, Microsoft,NVIDIA and others.
| 01 | Megatron-Turing NLG 530B | Microsoft,NVIDIA | 0.87 |
| 02 | text-davinci-001 | OpenAI | 0.86 |
| 03 | falcon-180B | Technology Innovation Institute | 0.8 |
| 04 | Llama-2-70b-hf | Meta AI | 0.79 |
| 05 | Inflection-1 | Inflection AI | 0.79 |
| 06 | PaLM 540B | Google Research | 0.78 |
| 07 | LLaMA-65B | Meta AI | 0.78 |
| 08 | Chinchilla (70B) | DeepMind | 0.77 |
| 09 | falcon-40b | Technology Innovation Institute | 0.77 |
| 10 | LLaMA-33B | Meta AI | 0.77 |
| 11 | Llama-2-13b | Meta AI | 0.77 |
| 12 | LLaMA-13B | Meta AI | 0.75 |
| 13 | falcon-7b | Technology Innovation Institute | 0.75 |
| 14 | Gopher (280B) | DeepMind | 0.74 |
| 15 | Baichuan-2-13B-Base | Baichuan | 0.74 |
| 16 | Baichuan-2-7B-Base | Baichuan | 0.73 |
| 17 | LLaMA-7B | Meta AI | 0.73 |
| 18 | Llama-2-7b | Meta AI | 0.73 |
| 19 | internlm-20b | 0.72 | |
| 20 | StableBeluga2 | Stability AI | 0.71 |
| 21 | Qwen-14B | Alibaba | 0.71 |
| 22 | mpt-7b | MosaicML | 0.7 |
| 23 | Qwen-7B | Alibaba | 0.68 |
| 24 | internlm-7b | 0.67 | |
| 25 | Qwen-1_8B | 0.58 | |
| 26 | chatglm2-6b | 0.54 |