PuzzleMask prompt injection bypasses LLM gatekeepers
🔍 PuzzleMask embeds policy-violating instructions inside natural prose to evade input classifiers. Researchers tested the technique against four lightweight gatekeepers and found a 100% bypass rate; a downstream, high-capability target model recovered and executed the hidden payload in roughly 94% of trials. Anthropic’s Opus models resisted the approach, suggesting detection that monitors reasoning rather than input alone. Defenses include paraphrasing inputs, hardening gatekeeper policies, monitoring model output/reasoning, or raising gatekeeper capability.
