< ciso
brief />
Tag Banner

All news with #vertex ai tag

107 articles

Google Cloud Modernize: AI-Driven Enterprise Transformation

๐Ÿš€ Google Cloud announces Google Cloud Modernize, an end-to-end portfolio that consolidates migration and modernization tools to accelerate enterprise transformation with AI. The offering centers on Modernization Hub, an in-console experience for analyzing code and mapping dependencies for Java, .NET, and mainframe apps. New Gemini-powered capabilities in Migration Center provide rapid TCO estimates and interactive cost modeling. Purpose-built compute, VMware support, and agentic migration tools (EKS-to-GKE) help organizations modernize infrastructure and application estates with enterprise-grade controls.
read more โ†’

Distributed GraphFlow: Scalable GNNs for Telco Networks

๐Ÿš€ Google Cloud introduces Distributed GraphFlow (DGF), an open-source Python library and framework designed to train and deploy Graph Neural Networks (GNNs) at scale for telecommunications. The post outlines an Autonomous Network Operations architecture built around a real-time network digital twin hosted in Spanner Graph, and explains how DGF integrates with that twin to enable anomaly detection, root cause analysis, predictive maintenance, and what-if simulations. DGF offers composable primitives and a high-level API to simplify GNN lifecycle management and production inference via Gemini Enterprise endpoints.
read more โ†’

KDDI Optimizes RAG Performance with ADK

๐Ÿ“˜ KDDI developed Buffmee, a consumer-facing Retrieval-Augmented Generation (RAG) app, to provide grounded, trustworthy AI-assisted learning across books, magazines, and web media. Facing latency and hallucination challenges, KDDI worked with KDDI iret and Google Cloud teams to implement automated evaluation using Gemini Enterprise, BigQuery Agent Analytics, and the Agent Development Kit (ADK). These optimizations reduced response latency by 38% and improved TTFT by nearly 18%, while boosting groundedness by 25% through systematic testing and prompt modularization. The result is a faster, more reliable experience that preserves content trust and compliance.
read more โ†’

TPU Performance: Gemma 3 on v6e for Real Workloads

๐Ÿ” This post benchmarks Gemma 3 12B and 27B on Google Cloud TPU v6e to compare classification (prefill-heavy) and generation (decode-heavy) workloads. It highlights that the 27B model saturates in high-concurrency generation beyond 64 users, while the 12B scales substantially better. For classification, both models show similar scaling up to 128 users. The article recommends E2E latency-based autoscaling and vLLM padding optimizations.
read more โ†’

Flexible billing and cost controls for AI agents

๐Ÿงพ Google Cloud announces expanded billing flexibility and cost controls for agent workloads across Gemini Enterprise and developer tools like Google Antigravity and Android Studio. New options include pay-as-you-go consumption, pooled quotas, Flexible Savings Plans with 10โ€“20% token discounts, and deferred execution pricing for off-peak discounts. Admins gain consolidated spend guardrails, hard monthly caps, anomaly detection, and centralized billing visibility to align FinOps with agent-driven innovation.
read more โ†’

Mirendil Chooses Google Cloud AI Hypercomputer

๐Ÿ” Mirendil will leverage Google Cloudโ€™s AI Hypercomputer, combining TPU accelerators and NVIDIA full-stack AI infrastructure to support model pre-training and post-training workloads. Google Cloud partnered closely with Mirendil on design and deployment across compute, storage, networking, and control planes. Managed training clusters run in Gemini Enterprise Agent Platform, and Mirendil is already live with TPU v5P chips while NVIDIA systems come online soon.
read more โ†’

Google Cloud Gemini Enterprise Agent Platform Updates

๐Ÿงญ Google Cloud announces broader availability of key features in the Gemini Enterprise Agent Platform, including Agent Memory Bank, Agent Runtime, Agent Identity, Agent Gateway, and Agent Registry. These additions enable long-running, personalized agents with enterprise-grade security, governance, and centralized discovery. The platform also adds unified observability and evaluation tools to monitor agent behavior and performance in production.
read more โ†’

Automate agent lifecycles with Gemini Enterprise

๐Ÿ› ๏ธ This deep dive shows how to build a production-ready agent using the Agents CLI and Gemini Enterprise. It walks developers through six stagesโ€”Setup, Build, Deploy, Govern, Evaluate, and Publishโ€”using an Industry Watch agent that reconciles press coverage with SEC filings. The tutorial emphasizes deterministic tools, managed runtime, memory, identity controls, and automated evaluations to prevent hallucination and ensure grounded, auditable results.
read more โ†’

Voicify and Google Cloud: AI Calling Transformation

๐Ÿค– Voicify partnered with Google Cloud to transform phone calls into reliable, AI-driven interactions for restaurants and healthcare. By adopting Gemini Enterprise and Vertex AI, the company improved latency, reduced costs, and achieved enterprise-grade security and compliance. Their orchestration platform validates orders against POS systems, handles traffic spikes with provisioned throughput and pay-as-you-go, and shortened client onboarding dramatically. The architecture emphasizes scalability, data integrity, and multicloud availability.
read more โ†’

AlloyDB enables accurate multilingual search with AI

๐Ÿงญ AlloyDB introduces native AI Functions to solve tokenization issues for logographical languages like Chinese, Japanese, and Korean. By calling Gemini models from SQL via ai.generate(), developers can perform in-database segmentation, stop-word removal, and embedding generation without ETL pipelines or external services. The approach uses stored-procedure batching, generated columns for search vectors and embeddings, and RUM plus ScaNN indexes to enable fast hybrid lexical and semantic search.
read more โ†’

Ray Serve LLM on GKE: Major performance gains

๐Ÿš€ Developers using Ray Serve for LLM inference on Google Kubernetes Engine (GKE) now get significantly better performance thanks to a joint effort with Anyscale. Three architectural changes โ€” HAProxy integration for internal routing, a direct token streaming path, and a v2 Ray executor backend for vLLM โ€” reduce overhead and latency. Benchmarks on A4 VMs with NVIDIA HGX B200 hardware show up to 5x higher throughput and 8x lower latency, while preserving Ray's developer-friendly features.
read more โ†’

Google Vertex AI SDK bucket squatting enables RCE

๐Ÿ”’ A design flaw in the Vertex AI SDK for Python allowed attackers to hijack model staging buckets across projects by predicting bucket names derived from project ID and region. Unit 42 researchers called this class of issue Bucket Squatting, where global bucket name uniqueness enabled pre-creation and silent takeover. The flaw could lead to cross-tenant model poisoning and remote code execution via pickle deserialization. Google issued fixes in SDK versions 1.144.0 and 1.148.0 and users should upgrade.
read more โ†’

Google Vertex AI SDK bucket-squatting flaw patched

๐Ÿ›ก๏ธ Palo Alto Networks Unit 42 disclosed a flaw in the Google Cloud Vertex AI Python SDK that let an attacker with only their own Google Cloud project and a victim's project ID hijack model uploads and execute code in Vertex AI serving containers. Google fixed the issue; users must update to google-cloud-aiplatform version 1.148.0 or later and explicitly set a staging_bucket. The bug arose from predictable default bucket names and lack of ownership checks, enabling an attacker to precreate the bucket, swap uploaded model files (often pickled), and run malicious code when Vertex AI loaded the model.
read more โ†’

Vertex AI SDK bucket-squatting enables RCE

๐Ÿ›ก๏ธ We discovered a vulnerability in the Google Cloud Vertex AI Python SDK that allowed an attacker to hijack a model upload and poison it, enabling remote code execution (RCE) in a victim's serving infrastructure. The issue stems from a predictable default staging bucket name and a missing ownership check in the SDK. By creating the same deterministic bucket in their own project and granting broad permissions, an attacker could replace uploaded model artifacts within a short window before Vertex AI reads them. Google fixed the issue in google-cloud-aiplatform v1.148.0 released April 15, 2026; developers should upgrade to the patched SDK.
read more โ†’

Agentic AI Bridges Dental Manufacturing Gaps

๐Ÿฆท Movix built a custom agentic AI platform to address a severe shortage of skilled dental technicians and reduce costly remakes in aligner and appliance manufacturing. Using Google Cloud infrastructure, including Gemini Enterprise Agent Platform, Cloud Run with L4 GPUs, and Compute Engine, Movix developed deep learning, computer vision, and 3D mesh models to automate quality control and data entry. The solution integrates with legacy lab systems, anonymizes PHI for compliance, and targets large-volume labs to improve accuracy, speed, and cost savings.
read more โ†’

AI Studio expands database choices and Starter Tier

๐Ÿ› ๏ธ At Google I/O 2026, Google announced expanded integration between AI Studio and Google Cloud, allowing new users to deploy up to two full-stack apps on the Starter Tier without a billing account. Developers can now choose between Firestore (non-relational) and Cloud SQL (relational) with Firebase Auth for unified authentication. The AI agent can infer or provision the appropriate database, provision resources, generate schema and code, and deploy apps directly to Cloud Run for rapid prototyping.
read more โ†’

Public Sector Embraces Agentic AI: Highlights from Next '26

๐Ÿค– At Google Cloud Next, public sector leaders showcased how they are using AI agents to boost productivity and mission impact across government and research organizations. Google introduced the Gemini Enterprise Agent Platformโ€”an evolution of Vertex AIโ€”plus the Gemini Enterprise App with Gemini 3.1 Pro and an Agent Designer for inspectable, scheduleโ€‘based workflows. The announcement also covered AI infrastructure (TPU 8 series), an Agentic Data Cloud, enhanced security and Agentic Defense, partner initiatives, and upskilling through the GEAR program.
read more โ†’

Google Cloud Next '26 Day 1: Gemini and the Agentic Stack

๐Ÿš€ At Google Cloud Next โ€™26, Google presented a unified stack to move AI into enterprise production, anchored by Gemini Enterprise as the connective tissue between data, people, and goals. Key launches include the Gemini Enterprise Agent Platform for building, scaling, governing, and optimizing agents, and the AI Hypercomputer with next-generation TPU 8 chips. Google also outlined the Agentic Data Cloud to ground agents in enterprise context, expanded security agents in Agentic Defense, Workspace Intelligence enhancements, and cross-cloud data capabilities to accelerate real-world deployment.
read more โ†’

Partner-Built Agents Now Available in Gemini Enterprise

๐Ÿš€ Google Cloud has integrated partner-built agents from its Agent Marketplace into the Agent Gallery inside the Gemini Enterprise app, creating a centrally governed hub for discovering and managing specialist, role-specific AI. Featured partners โ€” including Accenture, Adobe, Atlassian, Palo Alto Networks, Salesforce and others โ€” must pass a four-step evaluation to earn the Google Cloud Ready - Gemini Enterprise badge. Built-in safeguards such as cryptographic agent identities, Agent Gateway, and Model Armor protect data and prevent use for model training. Customers can trial the Gallery, while partners can apply to the AI Agents Program and access a rapid deployment framework.
read more โ†’

Gemini Enterprise: One Platform for Agent Development

๐Ÿš€ Gemini Enterprise is an end-to-end system for the agentic era, combining access to frontier models, a developer platform, a collaborative app, and a partner ecosystem to build and deploy agent fleets. The offering centers on the Gemini Enterprise Agent Platform โ€” an evolution of Vertex AI โ€” with an enhanced Agent Development Kit (ADK), graph-based orchestration, persistent Memory Bank, and fast Agent Runtime for multi-step work. IT teams gain a unified control plane for identity, governance, Model Armor, and auditing, while knowledge workers use a no-code Agent Designer, Inbox, Projects, and Canvas to create and monitor agents.
read more โ†’