This blog post provides an analysis of the knowledge cutoffs and pre-training timelines of language models such as Claude and GPT. It discusses what these models understand and what that reveals about their training processes, although the insights are described as only mildly interesting.