< ciso
brief />
Tag Banner

All news with #resilience tag

95 articles

OT Resilience Becomes a Boardroom Imperative

πŸ”’ The long-standing assumption that operational technology (OT) is safe due to physical isolation is no longer valid. Integration of OT with IT and cloud has exposed nearly 20 million OT assets online, increasing risk and making operational resilience a strategic concern. Modern industrial environments require continuous verification, micro-segmentation, and governance to protect safety, availability, and business continuity.
read more β†’

AWS European Sovereign Cloud independent operation exercise

πŸ›‘οΈ The AWS European Sovereign Cloud will run an exercise on October 24, 2026, demonstrating it can operate without using the AWS Global Network backbone or any non-EU infrastructure. The EU-based operational team will execute the test using only in-EU hardware and software, routing traffic over European ISPs and disabling dedicated global systems to validate independent operation. Customers should expect brief convergence-related connectivity interruptions, but no service availability impact within the sovereign cloud or other AWS Regions.
read more β†’

Architecture Diagrams Aren’t Proof of Resilience

πŸ” This article, based on a conversation with Mark Russinovich, explains why an architecture diagram is only an intent and not evidence of resilience. It highlights how drift, human and machine authorship, and AI model dependencies introduce new failure modes. The post argues for continuous validation, deterministic checks where possible, and automated, measurable practices to keep resilience real over time.
read more β†’

DynamoDB MRSC Now in Five More Regions

πŸ”” Amazon DynamoDB global tables with multi-Region strong consistency (MRSC) are now available in five additional AWS Regions: Canada (Central), Europe (Stockholm), Europe (Spain), Asia Pacific (Mumbai), and Asia Pacific (Singapore). This expands MRSC support to 15 Regions and allows creation of three-Region configurations with either three replicas or two replicas plus a witness Region. MRSC provides zero RPO and supports strongly consistent reads across geographies, helping build highly available, globally consistent applications.
read more β†’

AWS Resilience Hub adds EKS labels, insights, sharing

πŸ”§ AWS Resilience Hub introduces three capabilities: EKS labels as service input sources, generative AI-powered dependency insights, and resilience policy sharing via AWS Organizations. EKS labels let teams scope discovery using Kubernetes labeling conventions. Dependency insights analyze discovered dependencies to surface new links, cross-Region dependencies, and unusual patterns. Policy sharing enables centralized policy distribution and organization-wide observability.
read more β†’

Architecting resilient authentication with Cognito MRR

πŸ”’ Amazon Cognito now supports multi-Region replication (MRR) to automatically replicate user pools across AWS Regions with near-real-time synchronization, built-in failover, and interoperable JWT sessions. Replica user pools support sign-in and token operations but are read-only for configuration and attribute writes, which must be performed in the primary Region. To use MRR you must configure a symmetric multi-Region AWS KMS customer managed key and consider adopting the updated multi-Region OIDC issuer to ensure consistent discovery and JWKS endpoints. Cognito supports automatic domain and OAuth failover via Route 53 health checks and recommends using infrastructure-as-code and JWKS caching strategies for smooth migration and operational continuity.
read more β†’

AWS enhances regional resiliency for root sign-ins

πŸ”’ AWS now serves root user sign-in across US East (N. Virginia), US East (Ohio), and US West (Oregon), distributing traffic across these Regions to reduce dependency on a single Region and improve resiliency during disruptions. Sign-in routing is automatic and requires no action from customers. CloudTrail records ConsoleLogin events in the Region that processed the sign-in; update monitoring and alerts to include all three Regions. This change is available now for all AWS accounts.
read more β†’

AWS Lambda recursion protection reaches EU sovereign cloud

πŸ”’ AWS Lambda recursive loop detection is now supported in Europe Sovereign Cloud, helping prevent unintended billing by automatically detecting and stopping recursive invocations between Lambda functions and supported services. The feature covers common event sources like Amazon S3, Amazon SQS, and Amazon SNS, and halts processing when a loop is detected while sending an AWS Health Dashboard notification with troubleshooting guidance. Recursive loop detection is enabled by default for Lambda functions using a supported SDK version, and can be disabled per function using the PutFunctionRecursionConfig API for intentional recursion.
read more β†’

Building Modern Infrastructure Resiliency with Azure

πŸ”’ Microsoft Azure outlines how resiliency must be integrated into modernization efforts, emphasizing design-time planning, continuous assessment, and recovery readiness. The post highlights tools like Azure Infrastructure Resiliency Manager, Azure Chaos Studio, and Azure Backup to reduce blast radius, validate failover, and protect recovery points. It positions resiliency as a continuous operational practice aligned to workload criticality and business outcomes.
read more β†’

Design framework for zone-resilient Azure workloads

πŸ›‘οΈ This post argues that zone resiliency should be decided per component rather than applied as a single setting across a workload. It explains Azure availability zones, contrasts service-managed zone redundancy with user-managed zonal designs, and outlines common component categories and when two or three zones are appropriate. The guidance emphasizes validating service-specific behavior, modeling cost and capacity tradeoffs, and documenting failover, capacity, and operational responsibilities.
read more β†’

Amazon Connect adds cross-region routing for resiliency

πŸ” Amazon Connect Global Resiliency now supports cross-region routing, enabling contacts arriving in one AWS region to be offered to the longest-available matching agent across two linked regions. Managers can view analytics and search contacts across both regions as a unified contact center, and both regions carry live traffic at all times to ensure configurations and integrations are continuously exercised. Customers retain control over traffic distribution and can shift new contacts and agents to a single region when needed.
read more β†’

Google Cloud introduces Fault Injection Testing (Preview)

πŸ› οΈ Google Cloud announces Fault Injection Testing (FIT) in public preview to help teams automate failure testing and validate application resilience. FIT lets you create experiment templates to inject targeted faults such as Cloud SQL failovers and degraded Layer 7 traffic, with an automated dry run to verify affected resources and permissions. Experiments run for a defined duration with stop-and-revert controls, and Google recommends using FIT in non-production during preview. Access is via the Cloud console, gcloud CLI, or REST API; request preview through your account team.
read more β†’

Ransomware Forces Shift Toward Enterprise Resilience

πŸ”’ Ransomware has evolved from simple encryption schemes into multifaceted campaigns that combine data theft, extortion, and operational disruption. Attackers increasingly leverage AI and target third parties, expanding the attack surface and complicating detection. CISOs must now prioritize business continuity, vendor risk, and AI governance alongside traditional security controls to maintain trust and operational resilience.
read more β†’

MediaConnect Router Adds Configurable Latency Modes

πŸ”§ AWS Elemental MediaConnect Router now lets customers set internal recovery latency per output. Customers can choose between balanced mode (default behavior) and low-latency mode to optimize recovery time for latency-sensitive workflows. The setting is configurable via the MediaConnect API, AWS Management Console, or AWS CLI, and a new CloudWatch metric, RouteFabricRecoveryLatency, exposes recovery latency per route. This feature is available in all regions where MediaConnect Router is deployed.
read more β†’

Agentic AI Enhances Operational Resilience in Banking

πŸ€– Deutsche Bank partnered with Google Cloud to build an agentic AI-driven resilience platform that modernizes regulatory tabletop exercises and incident analysis. The platform ingests architecture, logs, data flows, and telemetry to generate context-aware scenarios, audit-ready evidence, and structured session records. Using Gemini Enterprise Agent Platform, LangGraph, and Google ADK, the bank achieves both deterministic, traceable execution and adaptive investigation. This approach supports continuous, regulator-aligned resilience across complex, distributed systems.
read more β†’

Sharded Hub-and-Spoke to Mitigate Noisy Neighbors

πŸ”Ž This article explains how shifting from a monolithic data pipeline to a sharded hub-and-spoke architecture reduces the impact of "noisy neighbor" tenants. The Hub acts as a lightweight router while Spokes provide isolated processing with Pub/Sub buffers between them. The design enables independent scaling, fault isolation, tiered pipelines for priority tenants, and spoke-level best practices such as DLQs, strict connection pooling, and asynchronous I/O.
read more β†’

AWS Resilience Hub adds recommended resilience tests

πŸ› οΈ AWS Resilience Hub now offers recommended resilience tests to help platform engineering and site reliability teams validate service behavior under known failure scenarios. The feature uses AWS Fault Injection Service (FIS) to inject controlled faults, evaluate recovery against defined objectives, and produce pass/fail results with detailed reports. Tests cover Availability Zone and Regional impairments and dependency failures, and are available in multiple global regions.
read more β†’

NCSC issues guidance for disruptive cyber incidents

πŸ›‘οΈ The UK's National Cyber Security Centre (NCSC) has published What To Do When Cyber-Attacks Disrupt Your Organisation, outlining three chronological stages for response: immediate hours and days, recovery to minimum viable operations, and longer-term restoration to business as usual. The guidance emphasizes preparing in advance, practicing realistic simulations, and engaging NCSC-vetted incident response firms to build resilience against escalating threats such as AI-accelerated attacks.
read more β†’

CISOs Rising to Lead Business Resilience

πŸ”’ CISOs are increasingly acting as de facto chief resilience officers, expanding from prevention to incident response and recovery. Experts recommend framing resilience in business terms β€” uptime, data protection, and financial impact β€” to secure board-level buy-in and funding. Practical steps include defining minimum viable operations, rehearsing recovery through a "ResOps" approach, and partnering with GRC, finance, and operations to share responsibility.
read more β†’

Jurassic Park and the Myth of Cyber Control

πŸ¦– The article compares Jurassic Park’s failed containment to modern cybersecurity, arguing that visibility is often mistaken for control. It asserts that tooling, dashboards, and backups provide friction but not guaranteed survivability, and that dynamic cloud and AI-driven change invalidate static recovery assumptions. The piece recommends continuous resilience engineering, dependency awareness, and validation to operate through inevitable disruptions rather than assume they can be prevented.
read more β†’