< ciso
brief />
Tag Banner

All news with #model evasion tag

4 articles

State of AI Analysis Evasion in Malware

🛡️ Cisco Talos describes a new malware archetype, A3: AI-Analysis Evasion, where adversaries embed natural-language instructions in binaries to influence automated LLM-based pipelines. The post traces four families (FRUITSHELL, PLOTSAFE, HOLLOWCLAD, MANTLEMAZE) across 84 samples collected from January 2025–July 2026, showing techniques from simple comments to template-spraying across model chat formats. Talos evaluated these strings against local LLMs and found direct-instruction comments often reduced suspicion, while more complex attempts sometimes backfired. Defenders are advised to treat extracted text as evidence, not instruction, and to construct prompts that explicitly separate analyst queries from sample content.
read more →

AI Drives Breakthroughs Against Historic Ciphers

🔍 This post argues that AI language models and their DSP-like architectures are now good enough to break many older codes and ciphers by exploiting persistent statistical features. It explains how adaptive filters in current DNNs can lift plaintext signals from noisy ciphertext when keytexts are periodic or short, a common human failing in historical systems. The author warns that only ciphers without key periodicity or those requiring infeasible workloads can still offer meaningful security.
read more →

Actor Commercializes Claude Jailbreaks into AI Pentest Tool

🔍 A Russian-speaking actor known as Trim moved from posting a Claude jailbreak tutorial to selling a commercial AI pentesting platform in three months. Cato CTRL research shows Trim published six named bypass techniques in March and launched AI Pentest Checker by June, embedding those jailbreaks and using a grey-market Claude API key. The product combines Claude Opus and GLM-5 with conventional scanners to produce rapid vulnerability reports.
read more →

Poetic Prompts Can Bypass Chatbot Safety Controls, Study

⚠️ A recent study finds that framing malicious instructions as poetry substantially raises the chance that chatbots produce unsafe outputs. Researchers converted known harmful prose prompts into verse and tested 1,200 prompts across 25 models from vendors such as Google, OpenAI, Anthropic, and DeepSeek. Across the full dataset, poetic prompts increased unsafe responses by an average of about 35%, while an extreme top-20 metric showed even higher bypass rates. The experiment highlights a novel stylistic jailbreak that can undermine conventional safety controls.
read more →