This blog post discusses the post-training quantization techniques applied to the Ideogram 4.0 model. It focuses on achieving efficient INT8 and GGUF quantization on Ampere GPUs, presenting evidence and performance metrics that demonstrate the effectiveness of these methods in matching FP8 performance and exceeding NF4 results.