DeepSeek-V3.2-Exp Release Notes

DeepSeek has launched DeepSeek-V3.2-Exp, an experimental model designed to optimize long-context performance and reduce operational costs. This release introduces DeepSeek Sparse Attention (DSA), a mechanism that enables faster training and inference while maintaining the output quality of its predecessor, V3.1-Terminus.

DeepSeek Sparse Attention (DSA) and Efficiency

DeepSeek Sparse Attention (DSA) implements fine-grained sparse attention to reduce compute costs and boost performance during long-context processing. According to DeepSeek, DSA achieves these efficiency gains with minimal impact on output quality. Benchmarks indicate that DeepSeek-V3.2-Exp performs on par with V3.1-Terminus.

API Pricing and Availability

DeepSeek has reduced API prices by more than 50% effective immediately. The new model is available across the DeepSeek App, Web interface, and API. To facilitate comparison testing, the V3.1-Terminus model remains available via a temporary API until October 15, 2025, at 15:59 UTC.

Open Source and Technical Resources

DeepSeek-V3.2-Exp is available as an open-source release. The company has provided the following resources for developers and researchers:

  • Model Weights: Available on Hugging Face.
  • Technical Report: Detailed documentation is hosted on GitHub.
  • GPU Kernels: Key GPU kernels are provided in TileLang and CUDA, with TileLang recommended for rapid research prototyping.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch