< ciso
brief />
Tag Banner

All news with #bigquery tag

118 articles

Lakehouse runtime catalog powered by Spanner

๐Ÿš€ The Lakehouse runtime catalog is a fully serverless, Spanner-backed implementation of the Apache Iceberg REST catalog designed to provide high availability, strong consistency, and horizontal scale for metadata management. It decouples metadata discovery from compute engines to enable multi-engine interoperability and supports credential vending, governance integration, and bi-directional federation. The catalog aims to reduce operational overhead and support agent-scale workloads with enterprise-grade features.
read more โ†’

PayPal modernizes analytics with managed Apache Spark

๐Ÿ” PayPal migrated its legacy, on-premise Hadoop analytics platform to Googleโ€™s Managed Service for Apache Spark to streamline operations and accelerate insights. The move delivered rapid provisioning, elastic scaling, and native integration with Google Cloud Storage and BigQuery, reducing operational overhead and data silos. Results included 25% faster processing, 30% better SLA adherence, lower costs, and more engineering time for innovation.
read more โ†’

Cloudflare publishes BEACON: real-world web speed data

๐Ÿ“Š Cloudflare today announced the release of BEACON, an anonymized dataset of billions of real-user performance measurements across 10,000 top websites. BEACON follows the RUM Archive standard, reports Core Web Vitals as full histograms, and is updated daily in Google BigQuery. The dataset breaks down metrics by browser engine, device, country, and industry, and includes queries and examples to aid researcher analysis.
read more โ†’

Borderless Lakehouse adds cross-cloud caching

๐Ÿ”’ Today Google Cloud announced enhancements to the borderless Lakehouse to let data engineers, analysts, and AI agents query governed data in place across clouds. The update introduces preview cross-cloud caching for BigQuery to reduce remote data transfer by caching columnar blocks locally and preview cross-cloud connections to query non-Iceberg data. The features use the Apache Iceberg REST catalog spec, Partner Cross-Cloud Interconnect, and default encryption to improve performance, security, and TCO for multi-cloud analytics.
read more โ†’

BigQuery Adds Augmented Analytics Table Functions

๐Ÿ”Ž BigQuery introduces six augmented analytics Table-Valued Functions (TVFs) to automate insight discovery and explain patterns using AI, ML and statistical methods. These functions run where data resides, produce structured SQL outputs, and can be chained to diagnose metric shifts, identify drivers, and estimate causal effects. They are integrable into conversational analytics and AI agent workflows for rapid, scalable investigation.
read more โ†’

BigQuery Identity Columns Generate Sequential IDs

๐Ÿ“ฃ BigQuery now supports identity columns that automatically generate sequential 64-bit integer values for table rows. This feature simplifies unique identifier management by letting BigQuery handle ID generation natively, reducing ETL complexity and boilerplate SQL. Identity columns integrate with standard DML such as INSERT and MERGE and can be defined in CREATE TABLE statements with options for fully managed sequences or allow manual overrides when needed.
read more โ†’

Google unveils TabFM for predictive ML in BigQuery

๐Ÿงญ TabFM is a new pre-trained foundation model in BigQuery for regression and classification on tabular data, now available in preview. Exposed through SQL functions AI.PREDICT and AI.EVALUATE, it delivers instant predictions and evaluation without explicit training or feature engineering. TabFM uses in-context learning and distributed inference to handle large-scale prediction workloads and outperforms many classic models on benchmarks.
read more โ†’

Data Agent Kit simplifies enterprise data pipelines

๐Ÿ› ๏ธ The Data Agent Kit brings the Orchestration Pipelines framework directly into your IDE or CLI, offering a unified, open-source toolkit for building, deploying, and troubleshooting data pipelines. It provides a Data Engineering tab and an agentic skill that authors production-grade Apache Airflowยฎ DAGs via natural language, paired with a declarative YAML DSL to remove Python boilerplate. The kit integrates with BigQuery, Managed Spark, Gemini Enterprise Agent Platform, and dbt to enable reproducible MLOps workflows and automated CI/CD deployment.
read more โ†’

BigQuery Graph: Native graph analytics at scale

๐Ÿ“ˆ BigQuery Graph brings native graph capabilities to the data warehouse, unifying ISO-standard GQL with SQL to run traversals natively at petabyte scale without ETL. Built on BigQuery, it inherits row- and column-level security, integrates with BigQuery ML and Gemini models, and supports cross-cloud Iceberg tables for virtual graphs. GA improves traversal performance, adds CALL and extended subquery support, and includes agentic tooling for modeling, conversational analytics, and auditable context graphs.
read more โ†’

Scaling OKF Bundles Using Knowledge Catalog

๐Ÿ“˜ This article explains how to publish and govern Open Knowledge Format (OKF) bundles across an organization by mapping OKF concepts onto Google Cloud's Knowledge Catalog. It outlines a one-time setup and a single push workflow using sample code and the kcmd CLI to register EntryGroups, EntryTypes, and an okf AspectType carrying OKF v0.2 signals. The approach makes bundles discoverable, searchable, and governed by existing IAM policies alongside BigQuery, Cloud Storage, and other cataloged resources.
read more โ†’

Automating data governance with lineage and automation

๐Ÿงญ This post describes Google's Governance Agent project that automates metadata propagation using column-level lineage, Knowledge Catalog, and BigQuery. It explains how the agent propagates descriptions, glossary terms, policy tags, and trust scores from upstream sources while applying confidence thresholds and conservative grounding. The project provides both a Gradio dashboard and a CLI to support steward review and automated pipelines, and emphasizes that automation is meant to reduce repetitive work, not remove steward oversight.
read more โ†’

BigQuery Graphs with Measures for Agentic Workloads

๐Ÿงญ BigQuery Graph introduces measures to unify governed metrics with relationship mapping, enabling agents to reason across complex, multi-hop dependencies without ETL. By mapping tables to an in-place property graph and defining MEASURE in the Property Graph DDL, BigQuery resolves graph paths before computing aggregations using GRAPH_EXPAND and AGG. The release includes a visual graph modeler in BigQuery Studio and native Looker integration to keep business metrics at the data layer.
read more โ†’

Malachyte Reinvents Retail Recommendations

๐Ÿ” Malachyte applies attention-based neural networks and LLM-inspired sequence modeling to address the retail "cold start" problem, updating user vectors in real time to personalize search and product pages. By streaming every interaction through Managed Service for Apache Kafka into Bigtable and combining multimodal embeddings, the platform refines predictions and privacy-friendly personalization within 100 milliseconds. Built on Google Cloud's AI stack, the solution leverages GKE, GCE, and Cloud Pub/Sub to enable continuous learning across retailers.
read more โ†’

BigQuery DTS expands integrations and features

๐Ÿš€ BigQuery Data Transfer Service (DTS) reduces engineering overhead by automating zero-code data ingestion into BigQuery, enabling teams to shift focus from pipeline maintenance to analytics. Recent additions include Open Lakehouse ingestion to Apache Iceberg, a managed Model Context Protocol (MCP) Server, expanded database connectors (PostgreSQL, MySQL, SQL Server), SaaS connectors (Shopify, Klaviyo, HubSpot, Mailchimp), and a Snowflake migration path. DTS emphasizes free ingestion for many first-party sources, low consumption-based pricing for third-party SaaS, integrated Cloud IAM security, and a 99.99% SLA for resilient data pipelines.
read more โ†’

BigQuery Autonomous Performance and Cost Optimizations

๐Ÿงญ BigQuery introduces autonomous, history-based query optimizations and an upgraded advanced runtime to improve performance and reduce compute costs without user intervention. These capabilities include enhanced vectorization, short query optimizations, and support for open formats like Apache Iceberg, delivering up to 35% faster queries and 40% lower slot usage in 2025. The platformโ€™s fluid scaling autoscaler enables per-second billing and average cost reductions up to 34%, with built-in safety guardrails to prevent regressions.
read more โ†’

Cortex Framework v7 Enables Agentโ€‘Ready SAP Data

๐Ÿ” Google Cloud announces general availability of Cortex Framework v7, designed to convert SAP transactional data into AI-ready, semantically rich data products deployed in BigQuery and registered in Knowledge Catalog. The release introduces purpose-built accelerators for SAP ERP and SAP Business Data Cloud, modular Dataform-powered deployments, and incremental, costโ€‘efficient processing to scale without extra infrastructure. New agent skills include an agentic data product builder to automate custom data product creation and preserve separation between vendor-delivered content and customer customizations.
read more โ†’

Google Cloud Conversational Analytics Expanded in Q3

๐Ÿ—‚๏ธ Google Cloud has advanced Conversational Analytics from experiments into enterprise-ready offerings across BigQuery, Looker, and preview support for AlloyDB, Cloud SQL, and Spanner. The platform supports querying data across clouds, Lakehouse and Iceberg catalogs, and integrates into tools like BigQuery Studio, Looker, and Gemini Enterprise. Enterprises gain governance features such as CMEK, VPC, DRZ, and row- and column-level access controls, plus cost and observability tools via OpenTelemetry. Agentic Workflows, anomaly detection, and APIs/SDKs enable embedding conversational agents across applications and workflows.
read more โ†’

SAP and Google Cloud launch BDC Connect for BigQuery

๐Ÿš€ SAP and Google Cloud announced general availability of SAP Business Data Cloud Connect for BigQuery, enabling zero-copy, bi-directional access between SAP Business Data Cloud and BigQuery. The integration exposes SAP tables, metadata, and business semantics directly in BigQuery and Knowledge Catalog to accelerate analytics and agentic AI while reducing data replication and costs. Early adopters report faster data pipelines and improved operational insights.
read more โ†’

Preparing Infrastructure for the Agentic Data Cloud

๐Ÿš€ In the agentic era, organizations must move from passive data stores to proactive systems of action by providing AI agents with trusted business context. Google introduces the Agentic Data Cloud to unify data, models, and operational databases on an AI-native stack, leveraging BigQuery, Spanner, and open standards like Apache Iceberg. The approach reduces latency, operational overhead, and integration gaps that hinder production-grade agentic AI.
read more โ†’

BigQuery advances unify structured and unstructured data

๐Ÿ” BigQuery announced GA for Autonomous Embedding Generation and AI.SEARCH(), plus a public preview of Hybrid Search to simplify retrieval and analytics over unstructured data. It automates embedding creation (including images), offers large single-query performance gains, and combines semantic and lexical techniques for more precise results. These features integrate into a broader end-to-end document analytics workflow in BigQuery.
read more โ†’