Kestrel: a local classifier for the cyber risk of agent tool calls

219 · · Aug. 18, 2026, 8:53 a.m.
Summary
This blog post discusses a newly developed local classifier for evaluating the cyber risk associated with agent tool calls executing shell commands on a machine. It highlights the performance of the classifier, which processes these calls in approximately 22 microseconds, surpassing seven leading large language models in effectiveness by about 0.40 F1 at a realistic base rate. The post is aimed at developers interested in enhancing cybersecurity measures in coding practices.