This post critiques the inadequacy of current safety benchmarks for large language models, arguing that they fail to capture the models' behavior under continuous, multi-turn interactions during adversarial attacks, suggesting a need for revised evaluation metrics.