This blog post explores the intricacies of how a large language model (LLM) server handles multiple simultaneous requests. It details the processes of memory management, including the use of KV caches, the difference between prefill and decode tasks, and the implications of continuous vs. static batching for performance. Through experimental observations, the author shares insights on how memory limitations, rather than computational speed, primarily dictate the efficiency of serving LLM requests. This analysis is supported by empirical data and references to relevant research papers.