Sebastian Raschka
Read post

A Technical Tour of the DeepSeek Models from V3 to V3.2

DeepSeek V3.2 represents a significant evolution in open-weight language models, introducing sparse attention mechanisms (DSA) for improved efficiency and self-verification techniques from DeepSeekMath V2 for enhanced reasoning. The model maintains the Multi-Head Latent Attention (MLA) and Mixture-of-Experts architecture from V3 while adding learned token selection instead of fixed sliding windows. Training improvements include domain-specific KL tuning in GRPO, off-policy sequence masking, and hybrid RLVR combining symbolic verification with LLM-as-judge rewards. The release achieves GPT-5 level performance through architectural efficiency gains and refined reinforcement learning methods, with V3.2-Speciale offering extended thinking capabilities via reduced length penalties during training.

    #machine-learning#llm#reinforcement-learning#deepseek
Dec 03, 2025•27m read time•From sebastianraschka.com
Post cover image
Table of contents
1. The DeepSeek Release Timeline2. Hybrid Versus Dedicated Reasoning Models3. From DeepSeek V3 to V3.14. DeepSeek V3.2-Exp and Sparse Attention5. DeepSeekMath V2 with Self-Verification and Self-Refinement6. DeepSeek V3.2 (Dec 1, 2025)7. Conclusion
1.4K Impressions
Sebastian Raschka's image
Sebastian Raschka

Sebastian Raschka's Blog offers insights, tutorials, and research updates on machine learning, deep ...

176 Followers

•

1.3K Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard