
DeepSeek-R1-Zero is a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step. It's 671B parameters in size, with 37B active in an inference pass.
It demonstrates remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors.
DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. See DeepSeek R1 for the SFT model.
Modalities
Context
164K
Released
Mar 6, 2025
Knowledge Cutoff
Jul 2024
DeepSeek-R1-Zero is a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step. It's 671B parameters in size, with 37B active in an inference pass. It demonstrates remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors.
DeepSeek R1 Zero has a 163,840 token context window.
DeepSeek V4 Pro 0813, DeepSeek V4 Flash 0731, DeepSeek V4 Pro 0423 and 9 more are other text models from DeepSeek.
DeepSeek R1 Zero was released on March 6, 2025. Its knowledge cutoff is July 31, 2024.
Token volume and request traffic to this model over time.