This blog post delves into the author's journey of training a GPT-2-small-style large language model (LLM) based on interventions, specifically focusing on updated results from instruction fine-tuning (IFT) tests. The author critiques previous methodologies and highlights the impact of training epochs on model performance while exploring various models. Through detailed evaluations, the author aims to identify the effectiveness of their models but observes inconsistencies in results, suggesting the relationship between loss minimization and instruction quality isn't straightforward. The conclusion emphasizes that chasing lower loss should not be the sole objective – practical performance also matters. This post is part of an ongoing series on building a language model from scratch.