Mamba-2.8B is an AI model developed by Carnegie Mellon University (CMU) and Princeton University (United States), first published in December 2023. It works in the language domain, on tasks such as language generation and question answering.
Training it took an estimated 5.4×10²¹ FLOP of compute (estimation method: operation counting). The model has 2,800,000,000 parameters. It was trained on roughly 300B datapoints.
Access: Open weights (unrestricted). Its weights are openly available. The reference paper has 7,063 citations. Epoch AI rates the confidence of this record as likely.