LLM Compressor 0.9.0: Attention quantization, MXFP4 support, and more

14 · Red Hat · Jan. 16, 2026, 7:34 a.m.
Summary
The LLM Compressor 0.9.0 release introduces significant features like attention quantization, support for MXFP4, a model_free_ptq pathway, and enhanced calibration capabilities. Designed for better quantization of large language models, this update enhances the tools available for developers, ensuring improved performance and flexibility in model deployment. Key advancements include compatibility with Hugging Face, new hooks for attention quantization, and accelerated calibration processes, making this version crucial for optimizing LLMs.