Where are the token-level LLM kill-switches?

· Boydkane · Sept. 8, 2026, 12:54 p.m.
Summary
The blog post discusses the concept of 'poisoned strings' in the context of large language models (LLMs) and proposes a method for training LLMs to recognize certain strings that induce them to emit end-of-sequence tokens prematurely. This concept introduces a potential mechanism for creating kill-switches in LLMs, enhancing control and safety in their usage.
AUTHOR