Live
AI models

Llemma 7B

Training compute
1.2×10²³ FLOP
Parameters
7B
Published
Oct 16, 2023

Llemma 7B is an AI model developed by Princeton University, EleutherAI, University of Toronto, Vector Institute, University of Cambridge, Carnegie Mellon University (CMU) and University of Washington (United States, Canada and United Kingdom), first published in October 2023. It works in the mathematics and language domain, on tasks such as mathematical reasoning, language modeling/generation, question answering and code generation.

Training it took an estimated 1.2×10²³ FLOP of compute (estimation method: operation counting,hardware). The model has 7,000,000,000 parameters. It was trained on roughly 55B datapoints. Training ran on 256 NVIDIA A100 SXM4 40 GB.

Access: Open weights (restricted use). Its weights are openly available. It is built on top of Code Llama-7B. The reference paper has 440 citations. Epoch AI rates the confidence of this record as confident.

Full record
Organization
Princeton University, EleutherAI, University of Toronto, Vector Institute, University of Cambridge, Carnegie Mellon University (CMU), University of Washington
Country of organization
United States, Canada, United Kingdom
Domain
Mathematics, Language
Task
Mathematical reasoning, Language modeling/generation, Question answering, Code generation
Training compute
1.2×10²³ FLOP
Compute estimation method
Operation counting, Hardware
Parameters
7,000,000,000
Dataset size
55B
Training hardware
NVIDIA A100 SXM4 40 GB
Chips used
256
Chip-hours
23K
Training power draw
203.3 kW
Numerical format
BF16
Model accessibility
Open weights (restricted use)
Open weights
Yes
Base model
Code Llama-7B
Citations
440
Epoch confidence
Confident
SourceEpoch AI, 'AI Models'. Published online at epoch.ai. Retrieved 2026-07-29 from https://epoch.ai/data/ai-models. Licensed under CC BY 4.0.
← All ai models