In-House LLM Serving at Netflix

· Netflix, Inc. · July 17, 2026, 10:08 p.m.
Summary
This blog post discusses Netflix's in-house LLM serving architecture, which runs the full stack from model deployment to inference within the production environment. It details the technology choices made during the development of the platform, including engine selection (vLLM), model packaging, API design, and deployment strategies. The article emphasizes lessons learned from production that guided decisions in designing a robust and scalable LLM serving system, alongside operational challenges faced and future improvements considered. It reflects a collaborative effort to enhance the machine learning workflow at Netflix.
AUTHOR
Sponsored
Zulip logo Zulip
Organized team chat for people who take work seriously. Topic-based threading keeps conversations focused.
Try Zulip
Become a sponsor →
BLOG POST FEATURED ON

Add this plugin to your blog