This post discusses the importance of evaluation in the development of production-grade LLM (Language Model) systems, highlighting the gap between demo systems and trustworthy applications. It includes a narrative relevant to current engineering challenges and practical insights into building evaluation platforms from scratch.