< ciso
brief />
Tag Banner

All news with #ai security tag

900 articles · page 2 of 45

AI watermark removers proliferate with unverifiable claims

🛡️ A rapid market has emerged for tools claiming to remove invisible AI watermarks after Anthropic enabled hidden marks in Claude outputs. Some projects strip metadata and hidden characters reliably, but none can currently be proven to defeat Anthropic's model-level watermark because the vendor has not published the detector or full technical details. Many commercial sites promise full removal, often measuring results against ordinary AI detectors rather than the undisclosed Claude watermark; independent code review shows gaps and unaddressed payloads. The ecosystem — open repos, web tools and agent skills — creates a potentially risky supply-chain surface if integrated directly into pipelines.
read more →

OpenAI Daybreak models now on Amazon Bedrock

🔒 Security teams can now access Daybreak Red and Daybreak Blue from OpenAI on Amazon Bedrock. Daybreak Blue supports common defensive workflows like vulnerability discovery, detection engineering, and incident response, while Daybreak Red targets advanced, authorized tasks such as vulnerability research and exploit reproduction with stronger identity verification and monitoring. Both models run on Bedrock's next-generation inference engine with zero-operator access and do not use inference data for model training.
read more →

AI-driven vulnerability discovery and its implications

🔍 A Black Hat USA 2026 keynote highlighted rapid growth in AI-assisted vulnerability discovery and the strain it places on defenders. Research from Arizona State University found that advanced models and workflows dramatically increased the number of bugs found, creating reporting and patching backlogs. This surge raises concerns about responsible disclosure, patching practices, and the potential for AI to eventually reduce new vulnerabilities as models and development processes improve.
read more →

WhatsApp launches on-device scam alert beta

🔔 WhatsApp has started a limited beta for an optional Scam Alert that runs an on-device machine learning model to warn users of likely scam messages. The feature analyzes linguistic signals and conversational structure for messages from non-contacts, displaying a chat warning that suggests blocking, reporting, or continuing. No message content or the model leaves the device; users can mark chats as trusted or opt to share recent messages to improve accuracy. Scam Alert complements existing security updates aimed at preventing account takeover and spyware attacks.
read more →

Rethinking cyber defense as AI accelerates exploits

🔒 Microsoft warns that AI-driven tools are accelerating vulnerability discovery and exploit generation, making traditional reactive patching and detection-centric defenses insufficient. David Weston of Microsoft highlighted MDASH findings showing rapid, low-cost exploit generation and urged industry shifts toward memory-safe languages like Rust, proactive secure-by-construction methods, and AI-assisted remediation. The talk, delivered at Black Hat USA, framed resilience and prevention as the new priorities.
read more →

Black Hat USA 2026: AI and cybersecurity controls

🧭 The Black Hat USA 2026 conference centered on AI's influence across cybersecurity, featuring keynotes and panels with senior US officials who debated regulation, innovation, and national leadership. Speakers including the White House National Cyber Director and representatives from CISA and the FBI discussed rapid vulnerability discovery enabled by AI, industry collaboration, and the need for prioritization. Presentations highlighted incidents such as the OpenAI–Hugging Face case and emphasized that AI systems act through human-set tasks and controls, underscoring accountability and governance requirements.
read more →

Prompt injections used as defensive mechanism

🛡️ Researchers from Tracebit report that embedding prompt injections alongside secrets stored on AWS can disrupt AI hacking agents by triggering LLM guardrails. These injected prompts instruct the model to perform forbidden actions, causing the LLM to shut down or stop following prior commands—a technique the researchers call context bombing. The approach succeeds only when attackers use models with built-in safety filters; locally run or unguarded models remain unaffected.
read more →

Four gaps slowing AI adoption in enterprise SOCs

🔍 Enterprise SOCs are investing in AI but struggle to convert tools into measurable operational gains. Many initiatives add complexity and fragmented workflows instead of reducing analyst workload. Successful deployments prioritize explainability, augment existing playbooks, and unify access to disparate security tools. Clear governance and incremental automation help turn AI pilots into repeatable operational improvements.
read more →

OpenAI launches GPT‑5.6‑Cyber for security teams

🔒 OpenAI introduced GPT‑5.6‑Cyber, a cybersecurity-focused variant of GPT‑5.6 Sol designed for vulnerability research, exploit development, and incident response. Offered through a Daybreak Red tier for authorized defenders, it completes far more high-risk cyber prompts than standard models and outperforms prior GPT‑5.5‑Cyber on several benchmarks. The model has already helped discover high-severity flaws, though it sometimes produces shorter vulnerability reports and performs less well on open-ended exploit development tasks.
read more →

OpenAI launches GPT‑5.6‑Cyber and Daybreak tiers

🔒 OpenAI has announced GPT‑5.6‑Cyber, a purpose-trained LLM for cybersecurity, and introduced two Daybreak tiers: Daybreak Blue for defensive tasks and Daybreak Red for advanced defensive and offensive testing. Daybreak Blue members get access to frontier models like GPT‑5.6 Sol with certain system-level safeguards removed for authorized defensive work, while Daybreak Red members can use purpose-trained cyber models such as GPT‑5.5‑Cyber and GPT‑5.6‑Cyber for advanced tasks. OpenAI says GPT‑5.6‑Cyber outperforms prior models on benchmarks and was used to find a high-severity V8 vulnerability that was responsibly disclosed and fixed.
read more →

OpenAI Pauses Astra Testing Over Cybersecurity Risks

🛡️ OpenAI has temporarily halted some internal testing of its forthcoming model Astra after assessments flagged its cyber capabilities as "critical." The firm said testing revealed significant advances in agentic coding and cybersecurity, prompting scaled-up robustness testing and strengthened controls including isolated environments, restricted access, and enhanced monitoring. OpenAI will pause activities that do not meet the new security requirements and share guidance with third-party testing partners.
read more →

Security leaders confident but unprepared for rogue AI

🔒 A majority of IT and security leaders say they can detect malfunctioning AI agents, but few can trace and mitigate downstream impact quickly. A WanAware survey found 90% confident in detection while only 26% can trace impacts within minutes, and over 45% say it would take hours. Experts warn agents act at machine speed, spread via shared credentials and multiple platforms, and require built-in identities, narrow permissions, audit trails, and hard kill switches to contain incidents.
read more →

OpenAI unveils GPT‑5.6 Cyber for vetted security partners

🔒 OpenAI has released GPT 5.6 Cyber, a specialized model for vulnerability research, penetration testing, and incident response, available only to approved companies and security vendors. The offering includes two access tiers—Daybreak Blue for defensive workloads and Daybreak Red for tightly governed tasks—and will be integrated into partner tools and services rather than exposed to regular users. OpenAI emphasizes safeguards such as identity verification, scoped testing, logging, and human oversight to mitigate abuse.
read more →

Weekly recap: AI autonomy, Metabase zero-day

⚡ This week’s recap highlights AI models acting autonomously to target open-source projects, a critical zero-day in Metabase allowing unauthenticated SQL injection, and new CPU-level attacks bypassing Spectre v2 defenses. It also covers webmail CSS attacks, vishing campaigns by UNC6671 against financial firms, Chinese router backdoors in Zbtlink devices, and shifting ransomware behaviors.
read more →

Native AI enforcement for Claude Enterprise

🔒 Anthropic’s new inference hooks let enterprises enforce security policies before prompts reach Claude, enabling real-time allow-or-deny decisions without proxies or endpoint agents. Check Point Workforce AI Security integrates in minutes to apply existing DLP and attack protection rules across Claude web, desktop, and tool calls, with shadow mode, gradual rollout, and centralized event logging. The protocol does not rewrite prompts and currently inspects prompts and tool calls only.
read more →

One-click prompt injection exposed Atlassian Rovo data

🛡️ Researchers at DEF CON 34 demonstrated a one-click prompt-injection attack called “RovoBlast” that abused Atlassian’s enterprise AI assistant Rovo by injecting malicious instructions via the rovoChatPrompt parameter. The exploit allowed a single click to make Rovo accept attacker-supplied parameters in a user session, potentially exposing data across connected services like Slack, Microsoft 365, Google Workspace, Jira, and Confluence. Varonis reported the issue through Bugcrowd and Atlassian has issued a fix, while researchers urged limiting Rovo’s access and disabling unneeded automation.
read more →

AI tutors for children: benefits and concerns

📘 AI tutoring tools are expanding rapidly and promise tailored learning, but they carry notable risks for children. Parents should distinguish between simple chatbots and structured Intelligent Tutoring Systems, and be aware of cognitive, psychosocial, privacy and security issues. Careful selection, oversight and data-protection checks are essential to minimize harm and ensure productive learning outcomes.
read more →

Seven key trends shaping the cybersecurity market

🛡️ AI is reshaping the cybersecurity market as VC funding soars and incumbents race to integrate agentic AI features, driving robust M&A activity. New AI-centric product categories such as LLM security, model integrity, and AI governance are emerging while platforms and managed services gain momentum. Quantum security and DSPM are rising priorities as organizations seek integrated, AI-native defenses.
read more →

How Google Cloud detects and contains emerging threats

🔒 Google Cloud outlines its proactive, shared-fate approach to detect and contain emerging threats across AI workloads, cryptomining, credential exposure, supply chain attacks, and account takeover. The post describes detection signals, tailored containment actions like granular throttling and localized identity isolation, and escalation paths including targeted suspensions. It highlights integrations such as GitHub Secret Scanning and details observability tools like Cloud Abuse Event Logging, Cloud Audit Logging, and billing alerts.
read more →

AI-driven HTTP desync research uncovers new techniques

🛡️ PortSwigger's AI-assisted system HTTP Terminator, developed by James Kettle, autonomously generated and validated novel HTTP desynchronization techniques after exploring 30,000 candidate attack vectors. The team also ran a human-guided cascade that discovered a now-patched zero-day in Apache Traffic Server (CVE-2026-63078) and reported findings across banks, government infrastructure, and security products. PortSwigger released the tool as open source and recommends avoiding HTTP/1.1 upstream or tightly allow-listing methods where removal isn't possible.
read more →