AgentLSD Rewrites How We Evaluate AI Security Agents
AgentLSD, listed at 2026-09-16 in the AgentSafety Papers tracker on GitHub, is a controlled framework for evaluating AI security agents against adversarial task contamination, a failure class most teams still have no name for.
The framework runs those evaluations as CTF challenges.
And the benchmark summaries report that it "