< ciso
brief />
Tag Banner

All news with #google kubernetes engine tag

58 articles

AI21 Accelerates Model Training with AI Hypercomputer

🔧 AI21 Labs adopted Google Cloud AI Hypercomputer and Kueue on GKE to run foundation models like the Jamba family at scale. They pooled thousands of A3 and A3 Ultra GPU instances into a shared GKE cluster to maximize utilization and replaced manual capacity negotiation in Slack with automated scheduling. The change reduced high-priority job wait times from 72 hours to 12, cut manual scheduling interventions from 20 per week to zero, and lowered fragmentation.
read more →

GKE CPU Startup Boost for Faster Pod Starts

🚀 GKE introduces CPU startup boost in preview, integrated into the Vertical Pod Autoscaler (VPA), to temporarily elevate CPU allocation during container initialization and then in-place scale it back to steady-state levels without restarting containers. The feature leverages Kubernetes In-place Pod Resize (IPPR) and operates across admission, startup, and unboosting phases. It targets CPU-intensive startup workloads—such as Java JVMs, Node.js, and Python/AI microservices—so teams can avoid overprovisioning while reducing cold starts and costs. Available on GKE 1.36.0-gke.4447000+ for Standard and Autopilot clusters.
read more →

Networking architectures for AI inference model serving

🔒 This post compares two reference networking architectures for AI inference model serving: one tailored for Google Kubernetes Engine (GKE) and one for mixed or alternative backends. It explains a common control-plane pattern using Private Service Connect, optional Apigee, and Model Armor as a centralized entry point for secure, private inference calls. The GKE design adds a specialized GKE Inference Gateway, inference pools, and replica sets for GPU/TPU workloads. The multi-backend design uses a regional internal Application Load Balancer, a Cloud Run payload processor service extension to inject model headers, and Network Endpoint Groups to route to heterogeneous backends.
read more →

Google Named a Leader in 2026 Container Management

🚀 Google Cloud was named a Leader in the 2026 Gartner® Magic Quadrant™ for Container Management, ranking highest for Ability to Execute. The accompanying Gartner Critical Capabilities report placed Google Cloud first across all evaluated use cases, including AI training and inference. Google highlights recent GKE and Cloud Run innovations—like predictive latency boosts, rapid startup times, serverless GPU scale-to-zero, and Agent Substrate/Sandbox—to support enterprise and AI workloads at scale. The post invites users to try the open-source Agent projects and explore Cloud Run and GKE enhancements.
read more →

GKE agentic migration for safer cloud modernizations

🔧 Google Cloud released the open-source GKE agentic migration plugin to simplify migrations from AWS EKS to Google Kubernetes Engine. The tool combines LLM-driven translations with deterministic server-side validation and GitOps PR workflows to avoid direct changes to live clusters. It persists long-running migration state for multi-persona handoffs and produces runbooks for stateful transport, enabling human-in-the-loop approvals and safer, auditable migrations.
read more →

GKE Adds Native Prometheus Metrics for Autoscaling

🔔 Google Cloud announced built-in Prometheus metrics processing for GKE, enabling HPA to use PromQL queries directly via Google Managed Service for Prometheus. This removes the need for third-party adapters, reduces latency, and simplifies autoscaling configuration. The controller runs in the control plane and only deploys pods on nodes when PromQL metrics are actively requested, minimizing resource use. The feature is in preview with plans to add self-hosted Prometheus support before GA.
read more →

GKE Adds Native Scale-to-Zero Capabilities

🚀 GKE 1.37 introduces native scale-to-zero so workloads can fully scale down to zero replicas and stop consuming compute while idle. The feature uses HPA with the new AutoscalingMetric CRD and KEP-2021 support for minReplicas: 0 to wake workloads based on external signals such as Pub/Sub or Cloud Monitoring. Capacity buffers provide pooled warm capacity to eliminate cold-start latency and balance cost with instant responsiveness.
read more →

Kubernetes YAML can hand over a GCP organization

🛡️ When developers declare cloud resources via Kubernetes GitOps, controllers like Google Kubernetes Config Connector (KCC) act on their behalf using a platform service account. KCC authenticates with Google Cloud through a single service account that may have broad project, folder, or organization-level roles. If a user can create IAM-related resources in a namespace KCC watches, they can escalate privileges by having KCC apply bindings using its powerful account.
read more →

Global multi-cluster GKE inference for GPUs and TPUs

🧭 This post describes a layered routing architecture that makes globally scattered accelerator capacity behave like a single pool behind one entry point. The multi-cluster GKE Inference Gateway performs global traffic distribution and high availability while an LLM-d router applies memory-aware scheduling to maximize utilization across GPUs and TPUs. Benchmarks across three regions and 17,000 nodes showed near-linear throughput scaling and <1% routing overhead while maintaining ~99.9% success rates under heavy concurrency.
read more →

GKE Pod Snapshots: Faster Startup for AI Workloads

🚀 GKE Pod snapshots let you capture and restore a running Pod’s full state, including CPU and GPU memory, to dramatically reduce cold-start times for AI inference and sandboxed agent workloads. By persisting snapshots to high-throughput Cloud Storage, new replicas restore directly from a warmed state, cutting startup latency by up to 89% in benchmarks. This enables aggressive autoscaling, lowers idle GPU costs, and replaces complex custom caching solutions with a simple declarative CRD-driven workflow.
read more →

SeaVerse selects GKE Agent Sandbox for isolation

🎮 SeaVerse, a SeaArt gaming startup, built a platform for playable AI experiences and needed infrastructure that could run dynamic, multi-tenant sandboxes with low latency, strong isolation, and better observability. By adopting Google Kubernetes Engine (GKE) and GKE Agent Sandbox, SeaVerse achieved kernel-level isolation with microVMs and gVisor options, improved runtime visibility, and scaled sandbox allocations rapidly. The change reduced infrastructure costs by up to 60% while preserving a fast creator experience.
read more →

Agent Substrate now available on GKE clusters

🛡️ Agent Substrate is now available on Google Kubernetes Engine (GKE). This open-source agent execution runtime is engineered for high-density sandboxing, delivering sub-500ms resume times and hundreds of suspend/resume activations per second with native kernel and network isolation. Optimized for GKE but portable to any Kubernetes cluster, it supports hardware-isolated microVMs or gVisor sandboxes and integrates with existing agent frameworks like Hermes, Claude Code, and OpenClaw.
read more →

Best practices for dynamic capacity management

🚀 This post outlines practical strategies for dynamic capacity management to support large-scale AI and agent workloads, emphasizing predictable cost and performance. It describes three implementations: scheduling capacity for planned events, creating automated fallback plans with managed instance groups, and automating the lifecycle with GKE and Custom ComputeClasses. The guidance highlights tools like Dynamic Workload Scheduler, instance flexibility, Spot VMs, Hyperdisk, and dynamic resource allocation to improve utilization and resilience.
read more →

GKE introduces ClusterNetworkPolicy for cluster-wide control

🔒 ClusterNetworkPolicy (CNP) is a new cluster-scoped network policy API added to GKE to let administrators enforce deterministic, non-bypassable network guardrails across namespaces. CNP introduces a hierarchical tier model—admin, network policy, and baseline—with top-to-bottom evaluation and an explicit Pass action to delegate final decisions. Built with the Kubernetes SIG-Policy WG and implemented with Cilium, CNP is open source and intended to improve scalability, compliance, and portability of network security controls in multi-tenant clusters.
read more →

Filestore now runs on Colossus for scalable NFS

🚀 Filestore, Google Cloud’s first-party NFS file service, now runs on Colossus, Google’s distributed storage system, to deliver greater scalability, flexibility, and operational efficiency. The update decouples capacity from performance so IOPS can be provisioned independently, and offers deep GKE integration including CSI driver support and multishares for smaller persistent volumes. Backed by Colossus, Filestore targets high-concurrency AI and agentic workflows with NFS-based shared workspaces, improved failure recovery, and integrated security via IAM, UIDs/GIDs, and IP ACLs.
read more →

GKE Agent Sandbox boosts agent density and efficiency

🧭 This article explains how Google Kubernetes Engine (GKE) Agent Sandbox helps platform teams run more AI agents per node by replacing heavy microVMs with lightweight gVisor sandboxes. It summarizes testing on an n2-standard-48 VM showing Agent Sandbox increased agent density from 61 to 88 in a baseline scenario and enabled higher-density modes up to 274 agents with suspend/resume and warm-pool strategies. The piece highlights cost and performance trade-offs and orchestration patterns like pod snapshots for freezing idle agents.
read more →

GKE Blueprint for Securing AI Workloads at Scale

🔒 This article presents a blueprint for securing AI workloads on Google Kubernetes Engine (GKE), consolidating controls across Google Cloud services and GKE features to create a secure-by-default platform. It covers three layers—infrastructure, supply chain, and application—and details capabilities such as Confidential GKE Nodes, Workload Identity Federation, k8s-aibom for AI SBOMs, Model Armor, and the GKE Inference Gateway. The blueprint recommends a three-phase rollout: Deploy, Operate, and Govern, and emphasizes integrating Google Cloud controls to maintain security at enterprise scale.
read more →

GKE Autopilot Clusters with Managed DRANET

🛠️ This blog explains how to configure GKE Autopilot clusters to use GKE managed DRANET for GPU and TPU workloads. It outlines the setup flow: create a VPC, deploy an Autopilot cluster, define a custom ComputeClass, create a ResourceClaimTemplate for RDMA (GPUs) or netdev (TPUs), and deploy workloads that reference those resources. Examples and YAML snippets demonstrate GPU and TPU ComputeClasses, resource claim templates, and a deployment that binds pods to accelerators.
read more →

Google Cloud unveils C4N network and storage VMs

🚀 C4N is Google Cloud’s new network- and block-storage-optimized Compute Engine instance family, now generally available after its Next ’26 preview. Built on a custom Titanium offload architecture and 5th Gen Intel Xeon CPUs, C4N delivers up to 400 Gbps network bandwidth, 95M PPS, and up to 25 GiB/s with Hyperdisk Extreme. It targets network-intensive, storage-heavy, and latency-sensitive workloads without requiring premium add-ons.
read more →

Ray Serve LLM on GKE: Major performance gains

🚀 Developers using Ray Serve for LLM inference on Google Kubernetes Engine (GKE) now get significantly better performance thanks to a joint effort with Anyscale. Three architectural changes — HAProxy integration for internal routing, a direct token streaming path, and a v2 Ray executor backend for vLLM — reduce overhead and latency. Benchmarks on A4 VMs with NVIDIA HGX B200 hardware show up to 5x higher throughput and 8x lower latency, while preserving Ray's developer-friendly features.
read more →