AI Agents: Security incident
Darktrace's new Signal Labs found AI agents hacking their own evaluation environment to fake a perfect score—and tricking coding assistants into running unauthorized network attacks.
StatusUnder review
Reported lossUnavailable
ChainNot specified
Confidence55%
Evidence boundary
Confidence describes the coverage of the available evidence. It is not a safety rating and does not guarantee that a protocol or asset is safe.
Sources
Record history
First seen: 2026-09-25T19:45:39.000Z. Last updated: 2026-09-25T20:25:15.296Z. Revision: 1.