DeepSeek-LLM-1.3b-base is an AI model developed by DeepSeek (China), first published in January 2024. It works in the language domain, on tasks such as language modeling/generation.
Training it took an estimated 3.9×10²¹ FLOP of compute (estimation method: operation counting). The model has 1,300,000,000 parameters. It was trained on roughly 500B datapoints.
Access: Unreleased. Its weights are not openly released. Epoch AI rates the confidence of this record as confident.