< ciso
brief />
Tag Banner

All news with #data governance tag

137 articles

AWS Glue adds optimization for Iceberg V3 tables

πŸ”§ AWS Glue Data Catalog now supports table optimization, statistics, and crawlers for Apache Iceberg Version 3 (V3) tables. With table optimization you can compact V3 tables using binpack, sort, or z-order strategies and remove expired snapshots and orphan files. Glue can generate NDV statistics for V3 types and Glue crawlers can discover V3 tables in Amazon S3 and register them in the Glue Data Catalog.
read more β†’

AWS introduces system-managed Iceberg materialized views

πŸ›ˆ AWS announces system-managed materialized views for Apache Iceberg that lock write access to view data and definitions to AWS Glue. These views let you precompute expensive queries and store results as standard Iceberg tables in Amazon S3 while preventing other writers from altering the computed outcome. You define the view with SQL and an optional refresh schedule; AWS Glue computes and maintains the results in the Glue Data Catalog. The feature is available in all supported regions, providing governed, consistent datasets readable by Iceberg-compatible engines.
read more β†’

Amazon S3 Tables add Iceberg V3 data types

πŸ†• Amazon S3 Tables now support the Apache Iceberg Version 3 data types including geometry, geography, unknown, and nanosecond-precision timestamps, plus column default values. These additions let you store geospatial coordinates and nanosecond event times natively, reducing the need to encode values as strings or integers and improving efficiency and query performance. S3 Tables continue to provide automatic maintenance and compaction for Iceberg tables and are available in all Regions where S3 Tables are offered.
read more β†’

Most Organizations Face Microsoft 365 Governance Incidents

πŸ“Š ShareGate's State of Microsoft 365 report found 77% of global organizations experienced at least one Microsoft 365 governance incident in the past year. The survey of nearly 1,800 IT professionals across nine countries highlights failures such as lingering access for former users, audit and compliance gaps, and sensitive data reaching unintended recipients. Rapid AI adoption and overconfidence in AI controls, plus limited proactive monitoring, are cited as key drivers of increased risk.
read more β†’

EU fines Google €403M for mishandling location data

πŸ“Œ The Irish Data Protection Commission fined Google €403 million for GDPR breaches in how three features handled location data between May 2018 and February 2020. The DPC found issues with Web & App Activity, Location History and the Location Accuracy feature, citing failures in lawful processing, transparency and accountability, and excessive data retention. Google says the case concerns historical policies and notes it has updated practices, including introducing auto-delete controls and changing defaults since 2019.
read more β†’

Irish DPC Fines Google €403M Over Location Data

πŸ“The Irish Data Protection Commission has fined Google €403m for GDPR breaches related to its handling of users' location data across features such as Web & App Activity, Location History and Location Accuracy. The inquiry, covering May 25, 2018 to February 4, 2020, found failures in lawfulness, transparency, accountability and retention practices. The DPC said Google must rectify its processing within six months, while Google contends policies have since changed and tools improved.
read more β†’

What the 3M ChatGPT case reveals about AI governance

πŸ“ The Watson Grinding litigation involving 3M highlighted how AI chat histories can become central to discovery and governance. The article explains that prompts and interaction logs may capture assumptions, preferred outcomes and rejected alternatives that don't appear in final artifacts. It argues organizations must treat consequential AI interactions as part of the information lifecycle, defining retention, ownership and access. Practical governance should scale with risk to preserve sufficient provenance without indiscriminate retention.
read more β†’

Amazon Connect Customer Profiles: Segment Events

πŸ”” Amazon Connect Customer Profiles now emits segment membership events to notify when customer profiles enter or exit segments such as high-value or low-satisfaction cohorts. This removes the prior need to export full segments and run differential scripts, which caused delays and errors. Events stream to Amazon Kinesis Data Streams and include profile ID, segment name, operation type, and whether the change was real-time or from a scheduled snapshot. Real-time detection applies to standard-condition segments, while enhanced segments (Spark SQL) use configurable periodic snapshots.
read more β†’

Amazon S3 Object Lock Adds Variable Retention

πŸ”’ Amazon S3 Object Lock now supports variable retention via event holds that start on a future triggering event, enabling WORM protection that begins when a contract closing, audit completion, or similar event occurs. You can apply event holds to individual objects, set them as a bucket default, or scale them with S3 Batch Operations, while new IAM and bucket policy condition keys let you control who can set or release holds and enforce duration limits. CloudTrail logs hold operations and S3 Inventory reports hold status; the feature is available in all AWS Regions at no extra charge and has relevant regulatory assessments.
read more β†’

Amazon Redshift adds Apache Iceberg v3 support

πŸ”” Amazon Redshift now reads from and writes to Apache Iceberg v3 tables in your data lake, adding support for default column values, row lineage, and deletion vectors. Default column values simplify schema evolution by providing initial values when none are supplied. Row lineage exposes pseudo-columns for row identity and update sequence, enabling incremental and CDC workflows. Deletion vectors use compact compressed bitmaps to replace positional delete files, improving read/write performance for frequent updates and deletes.
read more β†’

AWS Glue 5.1 Arrives in European Sovereign Cloud

πŸ”” AWS Glue 5.1 is now available in the AWS European Sovereign Cloud Region, bringing core engine upgrades to Apache Spark 3.5.6, Python 3.11, and Scala 2.12.18 for improved performance and security. The release updates open table format support for Apache Hudi 1.0.2, Apache Iceberg 1.10.0, and Delta Lake 3.3.2, and adds Iceberg format 3.0 features such as default column values and row lineage tracking. Lake Formation now enforces fine-grained access control for write DML and DDL operations in Spark DataFrames and Spark SQL, and full-table access control is added for Hudi and Delta Lake tables.
read more β†’

Scaling OKF Bundles Using Knowledge Catalog

πŸ“˜ This article explains how to publish and govern Open Knowledge Format (OKF) bundles across an organization by mapping OKF concepts onto Google Cloud's Knowledge Catalog. It outlines a one-time setup and a single push workflow using sample code and the kcmd CLI to register EntryGroups, EntryTypes, and an okf AspectType carrying OKF v0.2 signals. The approach makes bundles discoverable, searchable, and governed by existing IAM policies alongside BigQuery, Cloud Storage, and other cataloged resources.
read more β†’

Surveillance and AI Targeting Infant Monitoring

🍼 The New York Times reports on companies developing continuous surveillance systems for infants that increasingly incorporate AI. These devices, led by vendors like Nanit, collect extensive data to monitor health, development, and behavior. Recent funding rounds, including Nanit’s $50 million, aim to expand analytics into speech, motor skills, and longer-term tracking into childhood. The article raises concerns about privacy, data use, and the implications of pervasive monitoring.
read more β†’

AWS Glue adds Iceberg catalog federation in GovCloud

πŸ”’ AWS Glue now supports catalog federation for remote Iceberg REST catalogs in AWS GovCloud (US) Regions, enabling direct, secure access to Iceberg tables stored in Amazon S3 and cataloged remotely. With this feature, analytics engines can query remote Iceberg tables without copying data, and metadata synchronizes in real time between AWS Glue Data Catalog and the remote catalog. The capability integrates with AWS Lake Formation for fine-grained access controls and cross-account sharing, and is supported by Amazon Redshift, Amazon EMR, Amazon Athena, AWS Glue, and third-party engines. Catalog federation is available via the Lake Formation console and the AWS Glue and Lake Formation SDKs/APIs in GovCloud (US-East) and GovCloud (US-West).
read more β†’

Serverless Lakehouse Catalog Modernizes Apache Hive Metastore

πŸ› οΈ The blog explains how legacy Apache Hive Metastores become bottlenecks as enterprises scale their data lakes and adopt multiple query engines. It introduces the Google Cloud Lakehouse runtime catalog, a serverless metadata registry built on the Apache Iceberg REST Catalog specification that supports both legacy Hive tables and modern table formats. The post outlines common pain points β€” scaling, governance, and operational TCO β€” and describes a migration path that extracts Hive table definitions and registers them into the serverless catalog. The result is unified governance, zero-data-copy access across engines, and reduced operational overhead.
read more β†’

SageMaker Unified Studio adds data profiling

πŸ” Amazon SageMaker Unified Studio now integrates data profiling and anomaly detection powered by AWS Glue Data Quality. Data stewards, engineers, and analysts can generate dataset- and column-level statistics on catalog tables and Visual ETL job results to understand data shape and completeness. A dedicated Data profile tab supports on-demand and scheduled profiling while anomaly detection flags drift without predefined thresholds. These capabilities are available in all Regions where SageMaker Unified Studio is offered.
read more β†’

Automating data governance with lineage and automation

🧭 This post describes Google's Governance Agent project that automates metadata propagation using column-level lineage, Knowledge Catalog, and BigQuery. It explains how the agent propagates descriptions, glossary terms, policy tags, and trust scores from upstream sources while applying confidence thresholds and conservative grounding. The project provides both a Gradio dashboard and a CLI to support steward review and automated pipelines, and emphasizes that automation is meant to reduce repetitive work, not remove steward oversight.
read more β†’

UK Legal Regulator Issues AI Safety Warning

πŸ›‘οΈ The Solicitors Regulation Authority (SRA) has issued a warning to solicitors and law firms about using AI responsibly after spotting hallucinations and data leaks. The notice emphasizes that regulated individuals remain accountable for AI outputs and must maintain appropriate human oversight, governance and secure handling of client data. The SRA highlighted risks including false case citations, potential contempt of court and breaches of client confidentiality when information is entered into public AI tools.
read more β†’

Amazon S3 Metadata and Annotations Reach GovCloud

πŸ”Ž Amazon S3 Metadata and annotations are now available in AWS GovCloud (US-East) and AWS GovCloud (US-West), enabling faster discovery, understanding, and enrichment of S3 data. S3 Metadata captures system-defined details like object size and source and stores them in Amazon S3 Tables for near real-time tabular queries. Annotations let you attach rich business context in JSON, XML, or YAML (up to 1 GB per object) that travels with the object and follows durability and consistency properties.
read more β†’

BigQuery DTS expands integrations and features

πŸš€ BigQuery Data Transfer Service (DTS) reduces engineering overhead by automating zero-code data ingestion into BigQuery, enabling teams to shift focus from pipeline maintenance to analytics. Recent additions include Open Lakehouse ingestion to Apache Iceberg, a managed Model Context Protocol (MCP) Server, expanded database connectors (PostgreSQL, MySQL, SQL Server), SaaS connectors (Shopify, Klaviyo, HubSpot, Mailchimp), and a Snowflake migration path. DTS emphasizes free ingestion for many first-party sources, low consumption-based pricing for third-party SaaS, integrated Cloud IAM security, and a 99.99% SLA for resilient data pipelines.
read more β†’