What Claude Thinks But Doesn't Say

184 · Philipp D. Dubach · May 9, 2026, 12:17 p.m.
Summary
The blog post discusses a recent interpretability method for the language model Claude, developed by Anthropic. This method translates Claude's internal activations into readable English, revealing insights into its decision-making processes. In addition to detailing the methodology, it highlights potential limitations, such as the risks of confabulation and the need for careful interpretation of the output. Overall, the piece provides a critical assessment of the effectiveness and implications of this new approach to understanding AI behavior.