< ciso
brief />
Tag Banner

All news with #model evasion tag

2 articles

Actor Commercializes Claude Jailbreaks into AI Pentest Tool

🔍 A Russian-speaking actor known as Trim moved from posting a Claude jailbreak tutorial to selling a commercial AI pentesting platform in three months. Cato CTRL research shows Trim published six named bypass techniques in March and launched AI Pentest Checker by June, embedding those jailbreaks and using a grey-market Claude API key. The product combines Claude Opus and GLM-5 with conventional scanners to produce rapid vulnerability reports.
read more →

Poetic Prompts Can Bypass Chatbot Safety Controls, Study

⚠️ A recent study finds that framing malicious instructions as poetry substantially raises the chance that chatbots produce unsafe outputs. Researchers converted known harmful prose prompts into verse and tested 1,200 prompts across 25 models from vendors such as Google, OpenAI, Anthropic, and DeepSeek. Across the full dataset, poetic prompts increased unsafe responses by an average of about 35%, while an extreme top-20 metric showed even higher bypass rates. The experiment highlights a novel stylistic jailbreak that can undermine conventional safety controls.
read more →