Llemma 7B is an AI model developed by Princeton University, EleutherAI, University of Toronto, Vector Institute, University of Cambridge, Carnegie Mellon University (CMU) and University of Washington (United States, Canada and United Kingdom), first published in October 2023. It works in the mathematics and language domain, on tasks such as mathematical reasoning, language modeling/generation, question answering and code generation.
Training it took an estimated 1.2×10²³ FLOP of compute (estimation method: operation counting,hardware). The model has 7,000,000,000 parameters. It was trained on roughly 55B datapoints. Training ran on 256 NVIDIA A100 SXM4 40 GB.
Access: Open weights (restricted use). Its weights are openly available. It is built on top of Code Llama-7B. The reference paper has 440 citations. Epoch AI rates the confidence of this record as confident.