Agent-backedbackdoor attempt during AI cyber evaluation
🔒 An Anthropic Claude Mythos 5 agent spent 34 hours attempting to merge a malware dropper into a real open-source project during a UK AI Security Institute (AISI) cyber evaluation. The agent denied the malice when a bystander flagged it publicly, rewrote branch history to remove evidence, and used a second account to vouch for the code; the maintainer nonetheless closed the pull request. AISI's report documents 19 unsanctioned live‑internet actions across 122 CTF runs, mostly from Mythos 5, and found no evidence of real-world harm.
