< ciso
brief />
Tag Banner

All news with #vertex ai tag

102 articles · page 2 of 6

Anthropic's Claude Mythos Preview Now on Vertex AI

🔒 Anthropic’s newest and most capable model, Claude Mythos Preview, is available in Private Preview to a select group of Google Cloud customers through Project Glasswing. Its placement on Vertex AI provides enterprises access to a frontier model integrated with Google Cloud’s tools to build, scale, and govern AI applications and agents. The announcement emphasizes high performance across use cases and a renewed focus on reducing cybersecurity risk in enterprise deployments.
read more →

Rightmove modernizes property search with unified cloud data

🏠 Rightmove migrated from siloed on-premises databases to Google Cloud to build a unified analytics and AI platform it calls the data hive. Using BigQuery, Vertex AI, and Looker, the company extracts metadata from listings and images to deliver personalized search, agent-assist messaging, and an Automated Valuation Model. The hub-and-spoke architecture centralizes governance while enabling business units to run tailored forecasting and ML use cases. Around 300 staff now use the platform to convert data into operational and commercial value.
read more →

Ultimate Prompting Guide for Lyria 3 and Lyria 3 Pro

🎵 This guide outlines best practices for prompting Lyria 3 and Lyria 3 Pro, Google’s music generation models that deliver granular control over vocals, instrumentation, arrangement, and timing. It highlights technical details—track lengths from rapid 30-second prototypes to three‑minute compositions, multi‑vocal support in eight languages, timed-lyrics and tempo conditioning—and includes a concise prompting framework. The post also covers advanced workflows such as timestamped segment instructions and multimodal generation using images or PDFs, plus integration paths through Vertex AI and the Gen AI SDK.
read more →

Google launches Lyria 3 and Lyria 3 Pro on Vertex AI

🎵 Google has made Lyria 3 and Lyria 3 Pro available on Vertex AI in public preview, bringing high-fidelity music generation to the Vertex AI API and Media Studio. Lyria 3 Pro composes studio-quality tracks up to three minutes with structural elements (intros, verses, choruses, bridges), while Lyria 3 produces 30-second tracks for rapid prototyping. Both models accept multi-modal inputs (text or images), support vocal generation with timed lyrics or user-provided lyrics, and can produce purely instrumental pieces. Outputs are embedded with SynthID watermarking and filtered for policy and IP compliance.
read more →

Google Cloud unveils Veo 3.1 Lite and Upscaling on Vertex AI

🚀 Google Cloud has launched Veo 3.1 Lite, a cost‑effective video generation model available now on Vertex AI, and introduced a new standalone Veo upscaling capability currently in private preview. The Veo 3.1 family now includes three tiers—Veo 3.1, Veo 3.1 Fast, and Veo 3.1 Lite—all with native audio generation. The upscaling tool enhances existing low‑resolution videos to 1080p and 4K, regardless of source, and access is provided via the Vertex AI API and Vertex AI Media Studio. Developer documentation and a sample video editor agent are available to help teams get started.
read more →

Gemma 4 Now Available Across Google Cloud Ecosystem

🔒 Gemma 4 is now available on Google Cloud as an open, commercially permissive (Apache 2.0) family of models with context windows up to 256K, native vision and audio processing, and support for over 140 languages. Enterprises can deploy and fine-tune Gemma 4 via Vertex AI, serve inference serverlessly on Cloud Run, or run production workloads on GKE and TPUs. The release highlights data residency and compliance through Sovereign Cloud options and open weights, enabling secure, controlled AI deployments.
read more →

Vertex AI P4SA Permissions Flaw Exposes Google Cloud Data

🔒 Unit 42 disclosed a permissions flaw in Vertex AI where the default Per-Project, Per-Product Service Agent (P4SA) can expose credentials and OAuth scopes via the metadata service. Researchers showed attackers could use those credentials to pivot into customer projects, read Google Cloud Storage buckets, and download images from restricted Artifact Registry repositories. Google updated docs and advises using BYOSA and least-privilege scopes; organizations should validate agent permissions before deployment.
read more →

Double Agents: Security Blind Spots in Vertex AI on GCP

🔒 Unit 42 researchers discovered that AI agents deployed with Google Cloud’s Vertex AI ADK can inherit overly broad default permissions, enabling a deployed agent to leak service‑agent credentials and act as a “double agent.” By exploiting the Per‑Project, Per‑Product Service Agent (P4SA), the team pivoted into consumer projects and downloaded restricted Artifact Registry images from Google‑managed producer projects. Google collaborated with Unit 42, updated documentation, and recommended Bring Your Own Service Account (BYOSA) as a mitigation. Palo Alto Networks highlights protection via Prisma AIRS, Cortex Cloud Identity Security, and Cortex AI‑SPM.
read more →

Multi-Agent Architecture and Long-Term Memory with ADK

🤖 Dev Signal is a multi-agent system designed to turn raw community signals into reliable technical guidance by automating the path from trend discovery to expert content creation. It relies on the Model Context Protocol (MCP) to standardize integrations with Reddit, Google Cloud Docs, and a custom Nano Banana Pro MCP server, all coordinated by a Root Orchestrator that manages three specialist agents. A dual-layer memory model uses Vertex AI for long-term embeddings while the Session Service preserves short-term state, with automated callbacks and tools (save_session_to_memory_callback, PreloadMemoryTool, LoadMemoryTool) to persist and fetch user preferences and stylistic signals.
read more →

Five Techniques to Optimize LLM Inference Efficiency

⚡ Karl Weinmeister frames LLM inference as an efficient frontier that trades latency against throughput and argues production systems often sit below this curve. He presents five actionable optimizations—semantic model routing, prefill/decode disaggregation, modern quantization, context-aware L7 routing with prefix caching, and speculative decoding—and explains their practical tradeoffs. A Vertex AI case study reports 35% faster time-to-first-token and doubled prefix cache hit rates after deploying GKE Inference Gateway.
read more →

Why Context Matters for AI Data Security with SDP Now

🔒 Google Cloud’s Sensitive Data Protection (SDP) now applies advanced AI context classifiers and image object detectors to identify and redact sensitive content across text and images. It detects medical and financial contexts, faces, passports, credit cards, and other PII, and can generate redacted versions so organizations keep valuable training data while protecting privacy. SDP supports both Vertex AI tuning and live agent interactions and integrates with Model Armor, Security Command Center, and contact center solutions.
read more →

Reduce 429 Errors and Build Resilient Vertex AI Apps

⚠️ Building LLM applications on Vertex AI can trigger 429 errors when request rates exceed available throughput, degrading user experience and increasing retries. This article explains consumption options—Standard and Priority PayGo, Provisioned Throughput, Flex PayGo, and Batch—and prescribes five operational practices: smart retries, global model routing, context caching, prompt optimization, and traffic shaping. Combining these approaches (for example PT for critical real-time traffic and Batch for latency-tolerant jobs) helps preserve performance and control costs.
read more →

BMW and Google Cloud Build Automated SLM Optimization

🚗 BMW Group and Google Cloud present a proof-of-concept pipeline to compress, fine-tune, evaluate, and deploy domain-specific small language models (SLMs) for in-vehicle voice commands. They position SLMs as a practical compromise between full cloud-based LLMs and constrained onboard hardware, reducing latency and network dependence. Using Vertex AI Pipelines, the automated workflow explores quantization, pruning, distillation, LoRA fine-tuning, and RL-based alignment, and validates models on Android/AOSP head-unit environments. The team publishes the pipeline code to encourage reuse and reproducible experimentation.
read more →

Agentic Autonomous Networks at MWC 2026 — Platform Advances

🚀 At MWC Barcelona, Google Cloud outlines a shift from AI-driven insights to agentic telco operations, showcasing tools that embed AI into network control to achieve Level 4–5 autonomy. The company highlights a dynamic network digital twin, a unified graph data layer using Spanner Graph and BigQuery, and real-time GNN predictions in Vertex AI. New open-source telco data pipelines and two proof-of-value agents — a data steward and autonomous network agents — aim to accelerate trials and reduce legacy bottlenecks.
read more →

Nano Banana 2 Brings Pro-Level Image AI to Enterprise

🖼️ Nano Banana 2 is Google’s latest image-generation and editing model, delivering Pro-level image quality and fast iteration for enterprise creative workflows. Powered by real-time web search and integrated with Gemini API in Vertex AI, it provides accurate, localized visuals plus premium features like text rendering, translations, and upscaling to 2K/4K. Enterprise-ready provenance is supported via SynthID and interoperable C2PA Content Credentials to surface how AI was used.
read more →

Google Releases Gemini 3.1 Pro for Enhanced Reasoning

🚀 Google announced Gemini 3.1 Pro, an upgraded foundation model in the Gemini 3 series that emphasizes deeper reasoning and complex problem solving. The model is available in preview in Vertex AI and Gemini Enterprise, and developers can access it through Google AI Studio, the Gemini API, Android Studio, Google Antigravity, and the Gemini CLI. Early customers report meaningful gains in speed, efficiency, and accuracy across code, 3D transformations, and product design workflows.
read more →

Provisioned Throughput on Vertex AI: Expanded Capacity

⚙️ Provisioned Throughput on Vertex AI standardizes reserved capacity across first-party, third-party, and open-source models, adding multimodal and operational enhancements to support production-scale AI agents. The update introduces Anthropic integration (private preview), PT for popular open models such as Llama 4, Qwen3, and GLM-4.7, and native support for high-bandwidth modalities including Gemini 3, Nano Banana, and Gemini Live API. Operational improvements — one-week PT terms, scheduled change orders, and explicit caching for long contexts — enable predictable latency, flexible commitments, and lower input costs for peak events and high-concurrency workloads.
read more →

Mastering Model Adaptation: Fine-Tuning on Google Cloud

🔧 This guide explains how to adapt foundation models on Google Cloud by fine-tuning both managed and self-managed workflows. It contrasts a fully managed Vertex AI Supervised Fine-Tuning path for models like Gemini with a customizable GKE approach using LoRA on open-source models such as Llama. The labs walk through data preparation, baseline evaluation, tuning, and automated evaluation metrics, as well as GKE infrastructure, GPU provisioning, security with Workload Identity, and containerized training for production readiness.
read more →

Seven Technical Lessons from Using Gemini at Scale

🧰 The Google Cloud samples team describes building a specialized end-to-end system that uses Gemini on Vertex AI and Genkit to produce production-ready educational code samples across many languages and products. Their architecture separates generation, validation, and delivery so LLM outputs are combined with deterministic automations, linters, unit tests, and human review. The post presents seven practical technical takeaways—decomposition, determinism, precise prompts, vetted evaluation, scaled downstream processes, end-to-end testing, and solid engineering practices—that drove reliable, scalable sample generation.
read more →

GKE Inference Gateway Cuts Latency for Vertex AI Performance

🚀 The Vertex AI team deployed the GKE Inference Gateway, built on the Kubernetes Gateway API, to reduce inference latency and improve cache efficiency without a custom scheduler. The gateway applies load-aware routing—scraping Prometheus metrics like KV cache utilization and queue depth—and content-aware routing that inspects request prefixes to send traffic to pods with warm context. In production this cut Time to First Token by ~35% for Qwen3-Coder, improved P95 by ~52% for a bursty chat model, and doubled prefix-cache hit rates from 35% to 70%.
read more →