Building Pinterest’s VLM Serving Stack on NVIDIA Dynamo

· Pinterest · Sept. 10, 2026, 11:37 p.m.
Summary
The blog post details Pinterest's development of its Vision-Language Model (VLM) serving stack built on NVIDIA Dynamo. It addresses the challenges and optimizations in serving AI systems that process both language and visual content, focusing on low-latency inference, improved performance, and the architectural innovations of NVIDIA Blackwell GPUs. The post also highlights various use cases, including Pinterest Assistant, and discusses the importance of a flexible serving platform for multimodal AI across the company.
AUTHOR