< ciso
brief />
Tag Banner

All news with #cloud run tag

25 articles

Serverless Apache Spark on Google Cloud: Architecture

🚀 This technical guide explains Google Cloud’s Managed Service for Apache Spark, contrasting traditional managed clusters with serverless deployment modes and execution models (interactive sessions and batches). It covers resource and cost optimization techniques including history-based autotuning, tuning cores/memory, dynamic allocation caps, and shuffle partition sizing. The article also demonstrates integrated troubleshooting using Gemini Cloud Assist to diagnose runtime failures and generate resilient PySpark fixes.
read more →

Google Cloud launches Developer Device Platform preview

📱 Google Cloud announced the public preview of Developer Device Platform (DDP), a fully managed service offering on-demand access to real physical devices and high-concurrency virtual emulators. DDP provides interactive debugging via Device Streaming and parallel CI/CD testing via Device Run, enabling faster iteration, smarter sharding, and auto-retries. The platform supports integration with coding agents and will integrate with Android Studio and CLI, charging users on a pay-per-minute public preview model.
read more →

BigQuery Autonomous Performance and Cost Optimizations

🧭 BigQuery introduces autonomous, history-based query optimizations and an upgraded advanced runtime to improve performance and reduce compute costs without user intervention. These capabilities include enhanced vectorization, short query optimizations, and support for open formats like Apache Iceberg, delivering up to 35% faster queries and 40% lower slot usage in 2025. The platform’s fluid scaling autoscaler enables per-second billing and average cost reductions up to 34%, with built-in safety guardrails to prevent regressions.
read more →

Checkout.com boosts reliability with Managed Airflow

🚀 Checkout.com migrated from a self-hosted Apache Airflow to Google Cloud’s Managed Service for Apache Airflow (Gen 3) to reduce operational overhead and improve reliability. The move eliminated manual patching and scaling, introduced DAG isolation and managed upgrades, and integrated with Cloud Monitoring and Cloud Logging for better visibility. As a result, the team saw faster deployments, higher stability, and estimated monthly cost savings of ~30%.
read more →

Cloud Run multi-region enhancements for high availability

🚀 Cloud Run now supports one-command multi-region deployments with automatic failover when paired with a global or internal application load balancer. Readiness probes give instance-level health checks and Service health aggregates those checks per region, exposed via serverless NEGs to enable rapid failover. These features target both public internet and private VPC applications and are intended to reduce downtime and simplify HA architectures.
read more →

Hardening Public Serverless Functions on Google Cloud

🛡️ This post from Mandiant highlights how publicly exposed serverless applications — often unauthenticated by design — are frequent targets for application-level attacks like LFI/RFI and command injection. It explains exploitation paths including file retrieval and service account token exfiltration, and demonstrates attack examples against Cloud Run Python functions. The article provides actionable hardening guidance such as using dedicated service accounts with least privilege, isolating public services in separate projects, enforcing S‑SDLC practices, and deploying Cloud Armor WAF and Layer 7 load balancing for centralized protection.
read more →

Google Cloud Run Sandboxes Enter Public Preview

🛡️ Cloud Run sandboxes are now in public preview, offering a native, secure, and ultra-fast runtime to execute untrusted code and agent workloads in milliseconds. These lightweight, isolated execution boundaries can spawn within existing Cloud Run service instances and enforce credential isolation, deny-by-default network egress, and a safe read-only filesystem overlay. Enabling sandboxes requires a single deployment flag and integrates with the Agent Development Kit and ComputeSDK for streamlined use.
read more →

Context-Aware Polymorphic Schema Validation

🛠️ This post outlines an architecture using Google's ADK and Gemini Flash to replace static prompt-driven agents with a just-in-time, metadata-driven orchestration. It externalizes JSON schema descriptors to a Central Metadata Registry and employs a lightweight discovery prompt plus dynamic validation hooks (Cloud Run) to ensure deterministic, schema-compliant payloads. The pattern reduces context bloat, lowers token costs, and prevents attention diffusion in multi-agent workflows.
read more →

AI-focused innovations in Dataflow platform

🧭 Google describes how innovations from its internal Flume platform power Dataflow, a fully managed batch and streaming service supporting large-scale ML workloads. The post outlines features like liquid sharding for dynamic rebalancing, global compute for cross-region scaling, automatic pipeline optimization, and rate-limiting for external API calls. It also highlights TPU-focused efficiencies such as heterogeneous worker pools, TPU-aware autoscaling, duty-cycle enforcement, and TPU fungibility. The article notes developer conveniences—multi-language SDKs, unified batch/streaming, ML framework integration, observability, and advanced workflows—and cites customer use cases and ongoing platform enhancements.
read more →

Guide to Reducing AI Cold Starts on Cloud Run

🧭 This article examines practical strategies to reduce AI cold-start latency on Cloud Run when serving GPU-backed models. It outlines the four-phase cold-start process, highlights storage and model-format choices (Cloud Storage, container images, GGUF, Safetensors, quantization), and explains Cloud Run features like image streaming, temporary CPU boosts, and concurrency tuning. The piece also shares operational tactics—warmup endpoints, startup probe tuning, regional deployment choices—and production patterns used by Elastic to treat GPUs as fungible compute.
read more →

Deploy a Multi-Agent System on Cloud Run with Terraform

📣 This article describes how the Dev Signal team transitioned a multi-agent prototype into production on Google Cloud by combining a FastAPI service, a Vertex AI memory bank, and the Agent Developer Kit. It highlights production-ready concerns including OpenTelemetry traces exported to Cloud Trace for visibility into agent reasoning, and secure secret handling via Secret Manager so credentials never appear in environment variables. The guide also demonstrates reproducible infrastructure using Terraform to provision Artifact Registry, service accounts, Cloud Run, and related APIs, and outlines containerization and Cloud Build steps to deploy new revisions.
read more →

Cloud Run Worker Pools at Estée Lauder Companies: Use Cases

🔁 Google Cloud's Cloud Run worker pools provide an always-on, pull-based execution model that Estée Lauder Companies used to scale LLM-powered services. The company's Rostrum platform migrated from a request-driven service to a producer-consumer architecture: a FastAPI web tier publishes user messages to Pub/Sub and worker pools consume them for LLM inference. This decoupling improved message durability, UI latency SLAs, and reduced operational overhead while enabling GPU-backed distributed workloads and cost improvements for long-running background tasks.
read more →

Orchestrator Pattern for Distributed AI Agents at Scale

🤖 The post proposes the orchestrator pattern to turn monolithic AI scripts into a team of specialized, distributed microservices that integrate directly with existing frontends. It demonstrates using Google's Agent Development Kit (ADK), the Agent-to-Agent (A2A) protocol, and Cloud Run to host separate researcher, judge, and orchestrator services. The design enables independent scaling, strict JSON contracts for reliable decision-making, and language-agnostic implementations. The authors emphasize production hardening: secure agent endpoints, mitigate latency across hops, and implement robust retries and error handling.
read more →

Updated Spend-Based Committed Use Discounts Guide Overview

💡 Google Cloud updated its spend-based Committed Use Discounts (CUDs), moving from a credit-based model to a direct discounted price model that makes net costs and savings visible at a glance. The rollout began in July 2025 and is now generally available, expanding SKU coverage to include Cloud Run and H3/M-series VMs and correcting reporting gaps for mixed Flex CUD environments. The unified CUD Analysis provides hourly granularity (up to 30 days), CSV exports, and a metadata export for programmatic joins with Billing BigQuery Export datasets. Enhanced recommendation and scenario modeling let FinOps teams size commitments, tune coverage thresholds, and validate pre/post migration savings.
read more →

Cloud Run Adds NVIDIA RTX PRO 6000 Blackwell GPUs for AI

🚀 Cloud Run now supports NVIDIA RTX PRO 6000 Blackwell GPUs in preview, enabling serverless deployment of large inference models such as Gemma 3 27B and Llama 3.1 70B. The GPUs provide 96GB vGPU memory, 1.6 TB/s bandwidth and support for FP4 and FP6 precision. Cloud Run pre-installs drivers, offers rapid GPU startup and autoscaling to zero, and integrates with Cloud Storage and IAP for production use.
read more →

Full-Stack Dart Architecture: Flutter on Cloud Run

🚀 This article demonstrates a full-stack architecture that uses Flutter for the web frontend and Dart for the backend, enabling shared models and business logic across client and server. It walks through a To-Do example that places the domain model in a shared package, uses Shelf to serve both API routes and static web files, and compiles the server to a native executable for fast startup. Deployment options include Cloud Run's OS-only runtime for mounting precompiled artifacts or a Dockerfile-based multi-stage build for portable containers, and the article includes CI guidance using GitHub Actions to automate analysis, tests, and web builds.
read more →

Deploy Gemini 3 Apps Quickly with Google Cloud Run

🚀 This guide demonstrates how to create and deploy a public web app using Gemini 3 Flash Preview via Google AI Studio and Google Cloud Run. In Build mode you describe the application in natural language and let the model "vibe code" a complete app, which appears instantly in the Preview panel for testing. When satisfied, a single Deploy App action pushes the app to Cloud Run, exports your API key as an environment variable, and provides a shareable URL. Note that deployment requires a Google Cloud project with billing enabled.
read more →

Hands-on with Gemma 3: Deploying Open Models on GCP

🚀 Google Cloud introduces hands-on labs for Gemma 3, a family of lightweight open models offering multimodal (text and image) capabilities and efficient performance on smaller hardware footprints. The labs present two deployment paths: a serverless approach using Cloud Run with GPU support, and a platform approach using GKE for scalable production environments. Choose Cloud Run for simplicity and cost-efficiency or GKE Autopilot for control and robust orchestration to move models from local testing to production.
read more →

Deploy n8n on Cloud Run for Serverless AI Workflows

🚀 Deploy the official n8n Docker image to Cloud Run in minutes to run scalable, serverless AI workflows. Cloud Run scales from zero and persists data in Cloud SQL while you only pay for active usage. The post shows how to call Gemini as the agent LLM and optionally connect workflows to Google Workspace via OAuth for Gmail, Calendar, and Drive. For production, follow the n8n docs to add Secrets Manager, Cloud SQL, and Terraform-based deployment.
read more →

Giles AI on Google Cloud: Transforming Medical Research

🚀 Giles AI migrated its healthcare-focused platform to Google Cloud to reduce latency, improve scalability, and accelerate developer velocity. Using Google Kubernetes Engine, Cloud Run, and Compute Engine, the company orchestrates complex clinical data flows and routes prompts through Vertex AI and Model Garden to remain model-agnostic. Data storage and extraction are handled with Cloud SQL, Cloud Storage, and Document AI, while Cloud Armor and Security Command Center bolster security and compliance. Early customer results include dramatic reductions in research time and improvements in response accuracy.
read more →