DeepSeek-V3.2-Exp introduces a novel Sparse Attention mechanism for efficient long-context inference, reducing costs by up to 50% for long-context API calls. With immediate support on cutting-edge NVIDIA architectures, the model is open-sourced, allowing experimentation and enterprise deployment through Red Hat AI systems. This post details the model's deployment process, scalability with llm-d, and ongoing optimization plans.