< ciso
brief />
Tag Banner

All news with #nvidia tag

119 articles · page 3 of 6

P4de Instances Expand to SageMaker Studio Notebooks

🚀 Amazon Web Services has announced the general availability of EC2 P4de instances on SageMaker Studio notebooks in Asia Pacific (Tokyo, Singapore) and Europe (Frankfurt). Each P4de packs eight NVIDIA A100 GPUs with 80GB HBM2e (640GB total), offering 2× the per‑GPU memory versus P4d. AWS reports up to 60% faster ML training and roughly 20% lower training cost compared to P4d, benefiting large high‑resolution datasets and reducing model training time. Developers can follow the SageMaker JupyterLab and CodeEditor guides and consult pricing for cost planning.
read more →

NVIDIA Confirms GeForce NOW Data Breach in Armenia

🔒 NVIDIA confirmed that GeForce NOW user information was exposed in a breach limited to Armenia after a regional partner's infrastructure was compromised. The company said its own network and NVIDIA-operated services were not affected and it is assisting the partner. Regional operator GFN.am said the incident occurred March 20–26 and that impacted users will be notified. Exposed fields reportedly include names, emails, phone numbers, dates of birth and usernames; no passwords were exposed.
read more →

Amazon EC2 G6 with NVIDIA L4 Now in Germany Sovereign Cloud

🚀 Amazon Web Services now offers Amazon EC2 G6 instances powered by NVIDIA L4 GPUs in the AWS European Sovereign Cloud (Germany). These instances address graphics-intensive and machine learning workloads — including natural language processing, translation, video and image analysis, speech recognition, and personalization — with up to 8 L4 Tensor Core GPUs (24 GB each), third-generation AMD EPYC processors, up to 192 vCPUs, 100 Gbps networking, and 7.52 TB local NVMe SSD. G6 instances are available as On-Demand, Spot, and Savings Plans and can be launched via the AWS Management Console, CLI, or SDKs. They expand AWS's GPU capabilities for customers with sovereignty and compliance needs.
read more →

Amazon EC2 G7e with NVIDIA Blackwell GPUs in London

🚀 Amazon EC2 G7e instances with NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs are now available in the Europe (London) region. G7e delivers up to 2.3x inference performance over G6e and supports up to eight GPUs with 96 GB each, 5th Gen Intel Xeon processors, 192 vCPUs, and up to 1600 Gbps networking. G7e includes NVIDIA GPUDirect P2P and GPUDirect RDMA with EFA for lower-latency multi-GPU and multi-node workloads, and is available as On‑Demand, Spot, and via Savings Plans.
read more →

Amazon EC2 P6-B200 (Blackwell GPUs) Arrive in GovCloud

🚀 Amazon EC2 P6-B200 instances with NVIDIA Blackwell GPUs are now available in AWS GovCloud (US-West) region as of May 6, 2026. They offer up to 2x performance versus P5en for AI training and inference and include eight Blackwell GPUs with 1440 GB of high-bandwidth GPU memory and a 60% increase in memory bandwidth. P6-B200 runs on 5th Generation Intel Xeon processors (Emerald Rapids), supports up to 3.2 terabits per second of EFAv4 networking, and is delivered in the p6-b200.48xlarge size.
read more →

Amazon EC2 P6-B300 Instances Available in US East Region

🚀 Amazon Web Services announced that Amazon EC2 P6-B300 instances are now available in the US East (N. Virginia) Region. The p6-b300.48xlarge ships with 8x NVIDIA Blackwell Ultra GPUs, 2.1 TB high-bandwidth GPU memory, 6.4 Tbps EFA networking, 300 Gbps ENA throughput and 4 TB system memory. Compared with P6-B200, P6-B300 delivers 2x networking bandwidth and 1.5x GPU memory and TFLOPS (FP4, without sparsity), making it suited for training and serving large trillion-parameter foundation models and LLMs with improved token throughput and faster distributed training.
read more →

Rowhammer GPU Attacks Grant Full Control of NVIDIA CPUs

⚠️ Two independent research teams disclosed new Rowhammer-style attacks against NVIDIA Ampere GPUs that induce GDDR bitflips to gain arbitrary read/write access to host memory, enabling full system compromise when IOMMU is disabled by default in many BIOS settings. The proofs of concept — GDDRHammer and GeForge — manipulate GPU page tables and page directories to escalate privileges and, in demonstrations, open root shells on affected machines. A subsequent variant was shown to succeed even with IOMMU enabled; tested cards include RTX 3060, RTX A6000, and RTX 6000.
read more →

Amazon ECS Managed Instances Adds NVIDIA GPU Metrics

🖥️ Amazon ECS Managed Instances now exposes NVIDIA GPU metrics through CloudWatch Container Insights with enhanced observability. Customers can monitor GPU capacity, utilization, memory usage, device-level hardware health, and thermal conditions for containerized workloads. The metrics are available in all commercial AWS Regions; to use them, enable Container Insights with enhanced observability and launch GPU-accelerated EC2 instance types via an ECS Managed Instances capacity provider.
read more →

Amazon Bedrock Adds OpenAI GPT OSS and NVIDIA Nemotron

🚀 Amazon Bedrock now includes OpenAI GPT OSS (120B and 20B) and NVIDIA Nemotron models (Nano 9B v2, Nano 12B v2, Nano 30B, Super 120B), enabling developers to access open-weight foundation models through a single API. The integration is powered by Mantle, a distributed inference engine that provides serverless, high-performance inference, unified capacity pools, automated quota management, and OpenAI API compatibility. These models are available on AWS GovCloud (US) for compliant, enterprise-grade deployments.
read more →

Amazon SageMaker HyperPod adds G7e and r5d.16xlarge

🚀 Amazon SageMaker HyperPod now supports G7e and r5d.16xlarge instances to improve large-model development, training, and deployment at scale. G7e uses NVIDIA RTX PRO 6000 Blackwell GPUs, offering up to 2.3x better inference performance than G6e and up to 768 GB GPU memory for larger LLMs, multimodal, and agentic AI workloads. The r5d.16xlarge provides 64 vCPUs, 512 GB RAM, and NVMe storage for distributed preprocessing, feature engineering, and memory-heavy orchestration; G7e is available in select US and Asia Pacific regions while r5d.16xlarge is available across all HyperPod regions.
read more →

Amazon ECS Adds NVIDIA GPU Health Monitoring & Repair

🔧 Amazon Elastic Container Service now includes NVIDIA GPU health monitoring and auto repair for ECS Managed Instances. The capability leverages NVIDIA Data Center GPU Manager (DCGM) to detect critical GPU hardware failures and proactively replace impaired instances to maintain availability for GPU-accelerated container workloads. You can view GPU health via the DescribeContainerInstances API and receive notifications through Amazon EventBridge. Auto repair is enabled by default on supported instances at no additional cost and is available in all AWS Commercial Regions.
read more →

Building the AI Foundation for Public Sector Partners

🚀 Google Public Sector is launching coordinated initiatives to help partners build, certify, and bring AI solutions to government customers faster. The program includes a federal startup accelerator in collaboration with NVIDIA for AI-focused ISVs, an expanded ISV ATO Accelerator offering up to $1M in funding, and a new Distributor Channel Private Offer with Carahsoft. These efforts target procurement, compliance, and legacy environment barriers to speed deployment of mission-critical AI.
read more →

Amazon EC2 G7e Instances Reach Los Angeles Local Zones

🚀 AWS has made Amazon EC2 G7e instances generally available in the Los Angeles Local Zone (us-west-2-lax-1b), bringing high-performance GPU compute closer to end users. G7e instances combine NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs with 5th-generation Intel Xeon Scalable (Emerald Rapids) processors to support creative workflows such as studio workstations, VFX editing, color correction, and enhanced real-time rendering. They are also targeted at AI workloads including LLM deployment, inference, and agentic AI at the edge; availability is via On Demand and Savings Plans and you can enable the Local Zone through AWS Global View or launch instances from the EC2 console, AWS CLI, and SDKs.
read more →

AWS Neuron SDK 2.29 Released with Stable NKI expanded tools

🚀 AWS released Neuron SDK 2.29.0, promoting the Neuron Kernel Interface (NKI) to Stable (v0.3.0) and adding a Standard Library plus a CPU Simulator for local kernel development. The update introduces ISA-level features, DMA priority controls, and variable-length collectives, along with seven new experimental kernels and improvements to existing ones. NxD Inference and the vLLM Neuron Plugin receive vision-language optimizations. Neuron Explorer moves to Stable and is available on the VS Code marketplace.
read more →

WPP Accelerates Humanoid Robot Training with G4 VMs

🤖 WPP leveraged Google Cloud G4 VM instances powered by the NVIDIA RTX PRO 6000 Blackwell and the NVIDIA Isaac Sim image to cut humanoid robot training from hours or days to under an hour, achieving more than 10x speedups. Their pipeline combines OptiTrack motion capture, OpenUSD digital twins, and MuJoCo-based validation to retarget complex human motion to constrained robot kinematics. Training ran at scale using GPU P2P topology and the AI Hypercomputer, condensing learned policies into ONNX for real-time deployment while preserving safety and robustness.
read more →

Amazon EC2 P6-B300 Instances Now in GovCloud (US-East)

🚀 Amazon has added EC2 P6-B300 instances to the AWS GovCloud (US‑East) Region. The p6-b300.48xlarge configuration provides 8x NVIDIA Blackwell Ultra GPUs with 2.1 TB of high-bandwidth GPU memory, 6.4 Tbps EFA networking, 300 Gbps ENA throughput, and 4 TB of system memory. P6-B300 delivers ~2x networking, 1.5x GPU memory and 1.5x FP4 TFLOPS vs P6-B200, targeting training and deployment of large trillion-parameter foundation models and LLMs.
read more →

Rowhammer Attacks Targeting GDDR6 GPUs and Servers

🔒 Three recent academic studies — GDDRHammer, GeForge, and GPUBreach — describe Rowhammer-style attacks that target GDDR6 on modern GPUs. The first two demonstrate memory-access patterns that can bypass TRR and corrupt GPU page tables, enabling arbitrary reads and writes in video memory and potential escalation into system RAM. GPUBreach goes further by chaining driver flaws to defeat IOMMU-based isolation. While enabling ECC, using HBM, and applying IOMMU mitigations reduce risk, these findings highlight a credible threat to shared GPU/cloud environments.
read more →

Nemotron-3-Super-120B and Qwen3.5 Models Added to SageMaker

🚀 Amazon SageMaker JumpStart now includes NVIDIA’s Nemotron-3-Super-120B and the Qwen3.5 family (9B and 27B), giving customers turnkey access to foundation models optimized for agentic reasoning, multilingual coding, and advanced instruction following. Nemotron-3-Super-120B employs a hybrid LatentMixture-of-Experts architecture with Mamba-2 and MoE layers to support collaborative agents and high-volume automation such as IT ticket triage and cybersecurity workflows. The Qwen3.5-9B prioritizes efficiency for resource-constrained environments, while Qwen3.5-27B offers deeper contextual and multimodal reasoning for large-scale document processing and complex scenarios. Users can deploy these models directly from the JumpStart catalog or programmatically via the SageMaker Python SDK.
read more →

Are $30,000 AI GPUs Better at Cracking Passwords Today?

🔒 Specops compared two flagship AI accelerators, the Nvidia H200 and AMD MI300X, against the consumer RTX 5090 using Hashcat benchmarks for MD5, NTLM, bcrypt, SHA-256 and SHA-512. The RTX 5090 outperformed both AI GPUs across all tested algorithms, often by wide margins, meaning the expensive AI hardware does not translate to superior password-cracking performance. Price-to-performance was stark: the H200 costs at least ten times an RTX 5090 yet delivers lower hash rates. The practical risk remains weak or reused credentials; long passphrases, breached-password detection, and MFA are the recommended mitigations.
read more →

Experimenting with GPUs, GKE DRANET and Inference Gateway

🔧 This post walks through deploying and serving a large model on Google Kubernetes Engine using managed DRANET and NVIDIA B200 GPUs. It explains how RDMA networking is provisioned as an isolated regional VPC for low-latency GPU-to-GPU communication and how to provision A4 nodes and reservations for RoCEv2-capable accelerators. The author provides example gcloud and kubectl commands to create the cluster, a GPU node pool with DRA labels, a ResourceClaimTemplate for mrdma workloads, and steps to serve a DeepSeek model privately via GKE Inference Gateway and a regional internal Application Load Balancer.
read more →