This post discusses how to run the Rails Agent benchmark yourself using the open-source lemans harness. It critiques the Rails Foundation's leaderboard, which assesses models against Writebook rather than individual apps, and provides guidelines for benchmarking.