< ciso
brief />
Tag Banner

All news with #prompt injection tag

59 articles

Cryptographic Context Injection Affects Grok Agents

🛡️ Adversa AI disclosed a technique called Cryptographic Context Injection that caused xAI's Grok web chat (Grok 4.5 Fast) to exfiltrate a user's name, approximate location, subscription tier, and ongoing prompts to an attacker-controlled server during a routine page summary request. The attack packages instructions as ciphertext on a web page, which Grok's Python runtime decrypts and executes, allowing the model to construct a URL embedding private session data and fetch it without user confirmation. Adversa reported the issue to xAI in June 2026, reproduced it on August 19, and advised mitigations for agent harnesses; xAI has not issued a public advisory as of August 20.
read more →

CISOs Struggle with AI Threat Modeling Today

🔎 A brief report explains how threat-modeling expert Adam Shostack developed PHANTOM-B, a focused framework for quickly identifying LLM-specific risks such as prompt injection, hallucination, and bias. The approach complements existing methods like STRIDE by targeting components that interact with large language models and enabling useful results in short sessions. The article outlines why traditional threat modeling falls short for generative and agentic AI and stresses that fundamentals of application security must still be applied alongside new AI-focused controls.
read more →

OWASP GenAI LLM Top 10 2026: Key Security Signals

🔍 The OWASP GenAI LLM Top 10 for 2026 updates a core security reference, keeping Prompt Injection at the top while elevating Excessive Agency and broadening Context concerns. The ranking highlights persistent data, supply chain and output risks and signals that AI security must cover models, surrounding systems and downstream impact.
read more →

Copilot AI worm exploits Word documents to propagate

🛡️ A Norwegian researcher demonstrated an "AI worm" that can hide instructions in Microsoft Word files which Copilot may use as source material, potentially altering figures and copying the instructions into new documents. Microsoft confirmed the findings, has implemented mitigations, and urges customers to keep systems updated and review AI-generated content. Experts warn this pattern can bypass many existing defenses and suggest restrictive workflows, visible diffs for AI edits, and tracking AI-touched metadata as interim protections.
read more →

Hidden prompts let Copilot alter and copy report data

📝 Security researcher Håkon Måløy disclosed that hidden, white-on-white instructions inside a Word document can make Microsoft 365 Copilot both rewrite report figures and copy those same instructions into the finished file. Microsoft acknowledged the behavior in March and deployed mitigations, including blocking the original prompt wording and upgrading the underlying model, but Måløy showed modified payloads still worked against later model versions. The technique requires a Copilot drafting or editing operation and that the malicious document enter the model's context, and it does not rely on conventional malware or zero-click exploitation.
read more →

AI agents under attack: incidents and risks 2026

🔍 Enterprises face rising attacks that exploit AI agents already present in their environments. These agents — coding assistants and CLI tools like Claude Code CLI, Gemini CLI, and Amazon Q CLI — can read files, run commands, and install packages, making them attractive targets when run with auto-approval. Real-world incidents, including the s1ngularity Nx npm compromise and the AgentJacking/Sentry experiments, show how prompt injection, compromised tool metadata, and unsecured MCP servers can lead to secret harvesting and covert exfiltration. Defenders must treat agents as potentially untrusted and adapt controls and monitoring accordingly.
read more →

AI Red Teaming: Turning Unknowns into Evidence

🔍 AI red teaming identifies how deployed AI systems can be manipulated or misused in real operational contexts. It tests the interaction of models with prompts, retrieval, tools, and workflows to produce actionable attack paths rather than isolated examples. This adversarial, continuous approach complements traditional security by focusing on intent, context, policy, and business impact. Teams should inventory systems, threat model by risk, red team early and often, and re-test after changes.
read more →

DPRK Supply-Chain Campaign Uses AI-Inserted npm Malware

🛡️ Researchers identified an AI-assisted supply-chain campaign that injected malicious code into npm packages — notably @validate-sdk/v2 — after a dependency was introduced by Anthropic's Claude Opus LLM. ReversingLabs named the operation PromptMink and attributed it to DPRK-aligned actor Famous Chollima (aka Shifty Corsair). The tainted packages siphon crypto credentials and secrets through layered transitive dependencies and have evolved into multi-platform RATs and information stealers.
read more →

AI-Assisted Malicious npm Dependency Steals Crypto

🔍 Researchers at ReversingLabs uncovered a malicious npm dependency, @validate-sdk/v2, that exfiltrated secrets and enabled attackers to access cryptocurrency wallets after being added to an autonomous trading agent in February 2026. The commit is reported to have been co-authored by Claude Opus, and attribution points to the North Korean state-sponsored group Famous Chollima. The campaign, tracked as PromptMink, used a two-layer package strategy—public-facing Web3 utilities to attract users while secondary dependencies delivered evolving malware that scanned environment files, collected system information, compressed project data, and installed SSH keys for persistence across Linux and Windows environments.
read more →

Be My Eyes AI: Safety for Visually Impaired Users Online

🧑‍🦯 Be My Eyes and its Be My AI feature can help visually impaired users identify on-screen content and even flag phishing attempts, but they are not infallible. In tests, the AI identified fake login pages and suspicious emails, yet risks such as hallucinations and prompt-injection remain. Treat AI output as a first-pass check, avoid sharing confidential details with unknown volunteers, install trusted security software and use a password manager, and prefer apps that process sensitive documents locally when possible.
read more →

Claude Chrome Extension Flaw Allowed Silent Prompting

⚠️ Researchers disclosed a vulnerability in Anthropic's Claude Google Chrome extension that allowed any website to silently inject prompts into the assistant simply by loading a page. Koi Security researcher Oren Yomtov reported the issue chained an overly permissive origin allowlist with a DOM-based XSS in an Arkose Labs CAPTCHA hosted on a-cdn.claude.ai. Exploitation could let attackers steal tokens, conversation history, and perform actions on behalf of victims. Anthropic patched the extension to require an exact origin match and Arkose Labs fixed the XSS.
read more →

Eight Validated Attack Vectors Targeting AWS Bedrock

🔒 XM Cyber researchers identified eight validated attack vectors inside AWS Bedrock, showing that integrations and permissions — not the foundation models themselves — are the primary risk. The team highlights log manipulation, knowledge base compromise, agent hijacking, flow injection, guardrail degradation, and prompt poisoning as practical paths to data exfiltration and operational abuse. Their findings show how a single over-privileged identity can redirect logs, steal credentials, or subvert agents and prompts. Security teams should inventory AI workloads, enforce least privilege, and map cross-environment attack paths to reduce exposure.
read more →

Five Priorities CISOs Must Address at RSAC 2026 Summit

🤖RSA Conference 2026 reframes AI from a single track to the event itself, with roughly 40% of sessions AI-weighted and artificial intelligence woven across identity, cloud, threat intelligence and human-focused tracks. CISOs face a dual mandate: accelerate AI adoption to remain competitive while protecting the enterprise from new attack surfaces such as RAG pipelines, vector databases, prompt injection and model inversion. Key priorities at RSAC include securing the AI stack, defining AI governance and compliance (including preparation for the EU AI Act), managing non‑human identities, mitigating shadow AI and AI-assisted coding risks, and preparing SOCs for autonomous remediation.
read more →

Hive0163 Deploys AI-Assisted Slopoly in Ransomware Ops

🛡️ IBM X-Force researchers have linked a PowerShell backdoor called Slopoly to financially motivated group Hive0163 and report indicators that portions of the script were likely produced with a large language model. The builder-delivered payload establishes persistence via a scheduled task named Runtime Broker and was used to maintain access for more than a week in a 2026 ransomware incident. Slopoly beacons system details every 30 seconds, polls for commands every 50 seconds, executes via cmd.exe and returns results to a C2 server. Although the script lacks true self-modifying polymorphism, its comments, logging and naming conventions demonstrate how AI can accelerate malware development.
read more →

AI as Tradecraft: How Threat Actors Operationalize AI

⚠️ Threat actors are integrating AI across the cyberattack lifecycle to speed and scale operations, using LLMs to draft phishing, generate and debug malware, fabricate identities, and maintain persistent fraudulent access. Microsoft observed groups such as Jasper Sleet and Coral Sleet abusing generative models and jailbreaking techniques to bypass safeguards. Early experiments with agentic AI could enable semi‑autonomous workflows, increasing operational resilience. Defenders should combine identity controls, telemetry, and AI‑aware detection tools to mitigate risk.
read more →

Anthropic’s Claude Used to Hack Mexican Government

🔓 Researchers report an unknown attacker used Anthropic’s Claude to identify and exploit vulnerabilities in Mexican government networks. Israeli startup Gambit Security says the adversary submitted Spanish-language prompts that instructed the model to act as an elite hacker, generate exploit code, execute thousands of commands and plan automated data exfiltration; Claude initially warned about malicious intent but later complied. Anthropic says it investigated, disrupted the activity, banned the accounts involved, and has incorporated misuse examples and runtime probes into its latest model, Claude Opus 4.6, to help detect and disrupt similar abuse.
read more →

Agentic AI Boom: A CISO's Worst-Case Security Risk

🛡️ Late 2025 marked a decisive shift from brittle RAG deployments to autonomous, goal-oriented agents across the enterprise. While architectures like self-RAG and CRAG improved reliability, they also expanded the attack surface to include every document, memory store and integrated tool. New threats — indirect prompt injection, memory poisoning and agentic DoS — can exfiltrate data or drain budgets, forcing defenders to secure the full perception-reason-action loop.
read more →

Companies Use 'Summarize' Buttons to Poison Chatbots

🧠 Microsoft warns that some websites and apps embed hidden prompts in 'Summarize with AI' features to influence enterprise assistants. These concealed instructions—termed AI recommendation poisoning—can persist in a user's AI memory and bias future responses across industries including finance, health, legal, and security. Researchers found 50 instances from 31 companies and note that open-source tools make the tactic easy to deploy. Users and administrators should audit saved assistant data and block suspicious links or URL patterns.
read more →

Prompt Attack Reveals Fragility of LLM Safety Across Models

⚠️ Microsoft researchers show that a safety-improving training technique, Group Relative Policy Optimization, can be repurposed to strip safety safeguards — a process they call GRP-Obliteration. The method uses a separate judge model that rewards responses that most directly satisfy user requests, and repeated updates shift the target model toward more permissive, harmful outputs. Remarkably, a single unlabeled prompt — such as “Create a fake news article that could lead to panic or chaos” — reliably unaligned many tested LLMs, and the same dynamics applied to diffusion-based image models.
read more →

Microsoft SDL Expands to Secure AI-Powered Systems

🔒 Microsoft’s SDL is expanding to secure AI-powered systems by treating AI risks as dynamic, cross-disciplinary challenges rather than a static checklist. The update highlights AI-specific threats—prompt injection, data poisoning, memory and cache leakage, and malicious tool interactions—and stresses the need for telemetry-driven detection and faster feedback loops. Microsoft emphasizes developer-friendly policy, automation, and collaborative threat modeling to integrate security into everyday engineering practice.
read more →