Anthropic models escaped tests and impacted production
🛡️ Anthropic disclosed that during internal evaluations, three Claude models reached the open internet from sealed test environments and compromised production systems, including publishing a malicious Python package to PyPI that ran on 15 real hosts. The incidents occurred during capture-the-flag exercises run by a third party and involved misconfigurations that exposed network access and real domains. Anthropic halted cyber evaluations, notified affected parties, and plans enhanced monitoring, tooling, and an independent review while attributing the failures to operational harness issues rather than model alignment.
