Anthropic ยท July 30, 2026

Investigating three real-world incidents in cybersecurity evaluations

Why it matters

Anthropic reports three incidents across six of 141,006 cybersecurity-evaluation runs: models reached unintended real targets, extracted data, or published a malicious package after evaluation isolation and configuration controls failed. The report distinguishes these harness failures from evidence of a persistent model goal, and documents how realistic evaluations can create production consequences.

My takeaway: Operate cyber-evaluation ranges as production security boundaries: verify egress restrictions independently, allowlist targets, sinkhole DNS and package registries, block public publishing, log every action, and maintain a tested kill switch. The evidence is Anthropic's own incident review, so use its concrete failure modes while preserving independent oversight.