RefineBench introduces a benchmark to evaluate the self-refinement abilities of language models (LMs) through 1,000 challenging problems and a checklist-based approach. It examines both guided and self-refinement modes, revealing that while frontier models show modest improvement in self-refinement, they excel in guided refinement, suggesting the need for advancements in their self-correcting capabilities.