< ciso
brief />
Tag Banner

All news with #data governance tag

123 articles

Serverless Lakehouse Catalog Modernizes Apache Hive Metastore

🛠️ The blog explains how legacy Apache Hive Metastores become bottlenecks as enterprises scale their data lakes and adopt multiple query engines. It introduces the Google Cloud Lakehouse runtime catalog, a serverless metadata registry built on the Apache Iceberg REST Catalog specification that supports both legacy Hive tables and modern table formats. The post outlines common pain points — scaling, governance, and operational TCO — and describes a migration path that extracts Hive table definitions and registers them into the serverless catalog. The result is unified governance, zero-data-copy access across engines, and reduced operational overhead.
read more →

SageMaker Unified Studio adds data profiling

🔍 Amazon SageMaker Unified Studio now integrates data profiling and anomaly detection powered by AWS Glue Data Quality. Data stewards, engineers, and analysts can generate dataset- and column-level statistics on catalog tables and Visual ETL job results to understand data shape and completeness. A dedicated Data profile tab supports on-demand and scheduled profiling while anomaly detection flags drift without predefined thresholds. These capabilities are available in all Regions where SageMaker Unified Studio is offered.
read more →

Automating data governance with lineage and automation

🧭 This post describes Google's Governance Agent project that automates metadata propagation using column-level lineage, Knowledge Catalog, and BigQuery. It explains how the agent propagates descriptions, glossary terms, policy tags, and trust scores from upstream sources while applying confidence thresholds and conservative grounding. The project provides both a Gradio dashboard and a CLI to support steward review and automated pipelines, and emphasizes that automation is meant to reduce repetitive work, not remove steward oversight.
read more →

UK Legal Regulator Issues AI Safety Warning

🛡️ The Solicitors Regulation Authority (SRA) has issued a warning to solicitors and law firms about using AI responsibly after spotting hallucinations and data leaks. The notice emphasizes that regulated individuals remain accountable for AI outputs and must maintain appropriate human oversight, governance and secure handling of client data. The SRA highlighted risks including false case citations, potential contempt of court and breaches of client confidentiality when information is entered into public AI tools.
read more →

Amazon S3 Metadata and Annotations Reach GovCloud

🔎 Amazon S3 Metadata and annotations are now available in AWS GovCloud (US-East) and AWS GovCloud (US-West), enabling faster discovery, understanding, and enrichment of S3 data. S3 Metadata captures system-defined details like object size and source and stores them in Amazon S3 Tables for near real-time tabular queries. Annotations let you attach rich business context in JSON, XML, or YAML (up to 1 GB per object) that travels with the object and follows durability and consistency properties.
read more →

BigQuery DTS expands integrations and features

🚀 BigQuery Data Transfer Service (DTS) reduces engineering overhead by automating zero-code data ingestion into BigQuery, enabling teams to shift focus from pipeline maintenance to analytics. Recent additions include Open Lakehouse ingestion to Apache Iceberg, a managed Model Context Protocol (MCP) Server, expanded database connectors (PostgreSQL, MySQL, SQL Server), SaaS connectors (Shopify, Klaviyo, HubSpot, Mailchimp), and a Snowflake migration path. DTS emphasizes free ingestion for many first-party sources, low consumption-based pricing for third-party SaaS, integrated Cloud IAM security, and a 99.99% SLA for resilient data pipelines.
read more →

AWS Glue Schema Registry expands to 10 regions

📢 The AWS Glue Schema Registry is now available in ten additional regions, including New Zealand, Thailand, Hyderabad, Osaka, Malaysia, Melbourne, Mexico (Central), Israel (Tel Aviv), Taipei, and Canada West (Calgary). The serverless, free registry supports Apache Avro, JSON, and Protobuf formats to validate and manage streaming data evolution. It serves as a centralized repository that reduces validation logic and cross-team coordination, improving data quality and lowering downstream failures. The registry integrates with C# and Java apps for Apache Kafka/MSK, Amazon Kinesis Data Streams, Apache Flink/Managed Flink, and AWS Lambda.
read more →

AWS Glue Data Quality adds smarter anomaly detection

🛠️ AWS Glue Data Quality introduces a new observation mode that reduces false positives and better handles irregular data arrival intervals for notebook and exploratory workflows. The mode uses a constant baseline instead of linear trend extrapolation to avoid over-alerting, improving accuracy and reducing noise. Additionally, anomaly detection for Glue ETL jobs is now offered at no extra charge across all AWS commercial and GovCloud (US) regions.
read more →

AWS Glue Data Catalog adds S3 Tables export

🔔 Today AWS Glue Data Catalog preview adds two features: exporting catalog metadata to S3 Tables and enabling semantic search for catalogs encrypted with customer managed AWS KMS keys. Exported metadata — including glossary terms, custom attachments, and descriptions — is written to the managed aws-catalog S3 table bucket in Apache Iceberg format, enabling SQL queries, auditing, and time travel in engines like Amazon Athena and third-party tools. These capabilities are available in US East (N. Virginia), US East (Ohio), US West (Oregon), and Europe (Ireland); S3 Tables usage follows S3 pricing.
read more →

AWS Glue Data Quality adds Catalog anomaly detection

🛠️ AWS Glue Data Quality now supports anomaly detection for Catalog-based evaluations and can write evaluation results to AWS Glue Data Catalog (GDC) tables. These features apply to both ETL jobs and Catalog evaluations, enabling ML-driven time-series forecasting to surface unexpected changes in table statistics without manual thresholds. Evaluation outcomes, profiling metrics, and anomaly predictions (with confidence bounds) are persisted to GDC tables and can be queried via standard SQL. The capabilities are available in all AWS commercial regions and AWS GovCloud (US).
read more →

AWS Glue Data Quality adds Distribution Analyzer

📊 AWS announced a Distribution Analyzer for AWS Glue Data Quality that produces frequency distribution profiles for datasets. The feature generates histograms for numeric columns and value distributions for categorical, date, and boolean columns, with support for custom bin counts to tune granularity. Distribution statistics integrate with existing DQDL rulesets, are stored in Amazon S3, and are accessible via APIs for querying and visualization.
read more →

SAP and Google Cloud launch BDC Connect for BigQuery

🚀 SAP and Google Cloud announced general availability of SAP Business Data Cloud Connect for BigQuery, enabling zero-copy, bi-directional access between SAP Business Data Cloud and BigQuery. The integration exposes SAP tables, metadata, and business semantics directly in BigQuery and Knowledge Catalog to accelerate analytics and agentic AI while reducing data replication and costs. Early adopters report faster data pipelines and improved operational insights.
read more →

AWS Wickr launches managed Data Retention Service

🗂️ AWS Wickr introduces a managed Data Retention Service for Premium customers, enabling retention of conversations across networks—including direct messages, Groups, Rooms, and federated teams—for archival and audit purposes. The serverless, cloud-native solution replaces container-based approaches with simplified deployment, managed infrastructure, automatic scaling, and monitoring while preserving Wickr's end-to-end encryption. The feature is opt-in for Premium networks and is available in multiple AWS regions including US, Canada, Asia Pacific, Europe, and AWS GovCloud.
read more →

AWS HealthOmics private workflows expand to regions

🧬 AWS HealthOmics private workflows are now available in the Asia Pacific (Tokyo) and US East (Ohio) Regions, extending access to fully managed bioinformatics pipelines for research, drug discovery, and agricultural science with regional compliance. The HIPAA-eligible service supports domain-specific languages like Nextflow, WDL, and CWL, and includes Git integrations and Amazon ECR container support to simplify migration and maintain data provenance.
read more →

Amazon SageMaker Unified Studio adds custom transforms

🔧 With Amazon SageMaker Unified Studio, teams can now create, share, and reuse custom visual transforms within visual ETL flows. This feature enables data engineers to encapsulate business-specific transformation logic—such as phone number standardization, PII masking, or data-quality checks—into reusable components that non-coders can apply across ETL jobs. The capability is available in all AWS Regions where Unified Studio is offered.
read more →

Global IAM Data Governance Tags for BigQuery

🔒 This post introduces the preview of IAM data governance tags for BigQuery column-level security. Built on Google Cloud Resource Manager tags with purpose=DATA_GOVERNANCE, these tags are global, support hierarchical classification up to five levels, and are replicated for disaster recovery. The article explains creating tag keys/values, attaching tags to columns via JSON or SQL, and defining regional BigQuery data policies for masking or raw access. It highlights decoupled governance, regional policy enforcement, and layered security requirements.
read more →

AWS Glue zero-ETL and SAP OData now in GovCloud

🛡️ AWS Glue now offers the SAP OData connector and zero-ETL integrations in AWS GovCloud (US-West) and AWS GovCloud (US-East). These integrations support Amazon DynamoDB, Salesforce, and SAP OData as sources and let regulated customers replicate data into Amazon Redshift, Amazon S3, or other destinations without custom pipelines. The fully managed zero-ETL option reduces operational overhead and engineering effort by providing a no-code interface to set up continuous data replication and maintain up-to-date replicas for analytics.
read more →

Fixing data architecture vs. upgrading detection models

🔍 Security teams often default to retraining AI models when detections fail, but the real root cause is usually upstream data issues. Fragmented telemetry, inconsistent schemas and stale baselines degrade ML effectiveness long before models see events. Standardizing schemas, monitoring data quality at ingestion and applying governance to security telemetry are practical priorities that restore detection reliability without wholesale platform replacements.
read more →

SageMaker adds OpenLineage for IAM-based domains

📊 Amazon SageMaker Unified Studio now supports OpenLineage-compatible data lineage in IAM-based domains, capturing events from Apache Spark on Amazon EMR, AWS Glue, SageMaker Visual ETL, and notebooks. The interactive lineage graph shows data flow with configurable depth, timestamp modes for column-level detail, and a dataset-only view. You can programmatically publish, query, manage, and delete lineage events via OpenLineage APIs and the DeleteLineageEvent API.
read more →

Flock’s Vehicle Fingerprinting Enables Plateless Surveillance

🚨 A 2024 company presentation reveals that Flock uses a so-called “Vehicle Fingerprint” combining decals, bumper stickers, racks and temporary tags to identify cars when license plates are incomplete or absent. The system enables officers to search that dataset, perform multi-geo queries and locate vehicles believed to be traveling together. Bruce Schneier notes this capability echoes older surveillance practices and warns that similar outcomes are possible with broad access to cell phone location data.
read more →