This post discusses the challenges posed by the increasing complexity of LLM inference workloads when using a monolithic serving process, and introduces strategies for deploying disaggregated workloads on Kubernetes to improve scalability and efficiency.