RefineBench: Evaluating Refinement Capability of Language Models via Checklists

230 · Research Nvidia · Aug. 11, 2026, 7:19 a.m.
Summary
RefineBench introduces a benchmark to evaluate the self-refinement abilities of language models (LMs) through 1,000 challenging problems and a checklist-based approach. It examines both guided and self-refinement modes, revealing that while frontier models show modest improvement in self-refinement, they excel in guided refinement, suggesting the need for advancements in their self-correcting capabilities.