AI agents can enable silent remote code execution
🔒 A new AI Now Institute report demonstrates a proof-of-concept exploit that coerces Anthropic’s Claude Code and OpenAI’s Codex into executing attacker-supplied binaries during automated code review. The attack uses multi-stage prompt injection hidden in repository files (documentation, comments) to trick agents in auto-mode or auto-review into running a seemingly benign script that launches a malicious payload. Researchers warn the architectural risk — agents’ inability to reliably attribute text sources — makes such platforms potential attack vectors when granted shell access and autonomous execution.
