This blog post details the author's experience porting a Linux CLI subset of the Agents' Last Exam (ALE) benchmark to Prime Intellect's Verifiers v1, highlighting numerous challenges encountered, including dependency issues and benchmark defects. The author discusses the evolution of the project into ALE-Gold, identifying both weaknesses and strengths in the models tested during evaluation runs, and ultimately reflects on the implications of benchmark limitations and the nature of performance measures in AI models.