< ciso
brief />
Tag Banner

All news with #data governance tag

137 articles · page 2 of 7

AWS Glue Schema Registry expands to 10 regions

📢 The AWS Glue Schema Registry is now available in ten additional regions, including New Zealand, Thailand, Hyderabad, Osaka, Malaysia, Melbourne, Mexico (Central), Israel (Tel Aviv), Taipei, and Canada West (Calgary). The serverless, free registry supports Apache Avro, JSON, and Protobuf formats to validate and manage streaming data evolution. It serves as a centralized repository that reduces validation logic and cross-team coordination, improving data quality and lowering downstream failures. The registry integrates with C# and Java apps for Apache Kafka/MSK, Amazon Kinesis Data Streams, Apache Flink/Managed Flink, and AWS Lambda.
read more →

AWS Glue Data Quality adds smarter anomaly detection

🛠️ AWS Glue Data Quality introduces a new observation mode that reduces false positives and better handles irregular data arrival intervals for notebook and exploratory workflows. The mode uses a constant baseline instead of linear trend extrapolation to avoid over-alerting, improving accuracy and reducing noise. Additionally, anomaly detection for Glue ETL jobs is now offered at no extra charge across all AWS commercial and GovCloud (US) regions.
read more →

AWS Glue Data Catalog adds S3 Tables export

🔔 Today AWS Glue Data Catalog preview adds two features: exporting catalog metadata to S3 Tables and enabling semantic search for catalogs encrypted with customer managed AWS KMS keys. Exported metadata — including glossary terms, custom attachments, and descriptions — is written to the managed aws-catalog S3 table bucket in Apache Iceberg format, enabling SQL queries, auditing, and time travel in engines like Amazon Athena and third-party tools. These capabilities are available in US East (N. Virginia), US East (Ohio), US West (Oregon), and Europe (Ireland); S3 Tables usage follows S3 pricing.
read more →

AWS Glue Data Quality adds Catalog anomaly detection

🛠️ AWS Glue Data Quality now supports anomaly detection for Catalog-based evaluations and can write evaluation results to AWS Glue Data Catalog (GDC) tables. These features apply to both ETL jobs and Catalog evaluations, enabling ML-driven time-series forecasting to surface unexpected changes in table statistics without manual thresholds. Evaluation outcomes, profiling metrics, and anomaly predictions (with confidence bounds) are persisted to GDC tables and can be queried via standard SQL. The capabilities are available in all AWS commercial regions and AWS GovCloud (US).
read more →

AWS Glue Data Quality adds Distribution Analyzer

📊 AWS announced a Distribution Analyzer for AWS Glue Data Quality that produces frequency distribution profiles for datasets. The feature generates histograms for numeric columns and value distributions for categorical, date, and boolean columns, with support for custom bin counts to tune granularity. Distribution statistics integrate with existing DQDL rulesets, are stored in Amazon S3, and are accessible via APIs for querying and visualization.
read more →

SAP and Google Cloud launch BDC Connect for BigQuery

🚀 SAP and Google Cloud announced general availability of SAP Business Data Cloud Connect for BigQuery, enabling zero-copy, bi-directional access between SAP Business Data Cloud and BigQuery. The integration exposes SAP tables, metadata, and business semantics directly in BigQuery and Knowledge Catalog to accelerate analytics and agentic AI while reducing data replication and costs. Early adopters report faster data pipelines and improved operational insights.
read more →

AWS Wickr launches managed Data Retention Service

🗂️ AWS Wickr introduces a managed Data Retention Service for Premium customers, enabling retention of conversations across networks—including direct messages, Groups, Rooms, and federated teams—for archival and audit purposes. The serverless, cloud-native solution replaces container-based approaches with simplified deployment, managed infrastructure, automatic scaling, and monitoring while preserving Wickr's end-to-end encryption. The feature is opt-in for Premium networks and is available in multiple AWS regions including US, Canada, Asia Pacific, Europe, and AWS GovCloud.
read more →

AWS HealthOmics private workflows expand to regions

🧬 AWS HealthOmics private workflows are now available in the Asia Pacific (Tokyo) and US East (Ohio) Regions, extending access to fully managed bioinformatics pipelines for research, drug discovery, and agricultural science with regional compliance. The HIPAA-eligible service supports domain-specific languages like Nextflow, WDL, and CWL, and includes Git integrations and Amazon ECR container support to simplify migration and maintain data provenance.
read more →

Amazon SageMaker Unified Studio adds custom transforms

🔧 With Amazon SageMaker Unified Studio, teams can now create, share, and reuse custom visual transforms within visual ETL flows. This feature enables data engineers to encapsulate business-specific transformation logic—such as phone number standardization, PII masking, or data-quality checks—into reusable components that non-coders can apply across ETL jobs. The capability is available in all AWS Regions where Unified Studio is offered.
read more →

Global IAM Data Governance Tags for BigQuery

🔒 This post introduces the preview of IAM data governance tags for BigQuery column-level security. Built on Google Cloud Resource Manager tags with purpose=DATA_GOVERNANCE, these tags are global, support hierarchical classification up to five levels, and are replicated for disaster recovery. The article explains creating tag keys/values, attaching tags to columns via JSON or SQL, and defining regional BigQuery data policies for masking or raw access. It highlights decoupled governance, regional policy enforcement, and layered security requirements.
read more →

AWS Glue zero-ETL and SAP OData now in GovCloud

🛡️ AWS Glue now offers the SAP OData connector and zero-ETL integrations in AWS GovCloud (US-West) and AWS GovCloud (US-East). These integrations support Amazon DynamoDB, Salesforce, and SAP OData as sources and let regulated customers replicate data into Amazon Redshift, Amazon S3, or other destinations without custom pipelines. The fully managed zero-ETL option reduces operational overhead and engineering effort by providing a no-code interface to set up continuous data replication and maintain up-to-date replicas for analytics.
read more →

Fixing data architecture vs. upgrading detection models

🔍 Security teams often default to retraining AI models when detections fail, but the real root cause is usually upstream data issues. Fragmented telemetry, inconsistent schemas and stale baselines degrade ML effectiveness long before models see events. Standardizing schemas, monitoring data quality at ingestion and applying governance to security telemetry are practical priorities that restore detection reliability without wholesale platform replacements.
read more →

SageMaker adds OpenLineage for IAM-based domains

📊 Amazon SageMaker Unified Studio now supports OpenLineage-compatible data lineage in IAM-based domains, capturing events from Apache Spark on Amazon EMR, AWS Glue, SageMaker Visual ETL, and notebooks. The interactive lineage graph shows data flow with configurable depth, timestamp modes for column-level detail, and a dataset-only view. You can programmatically publish, query, manage, and delete lineage events via OpenLineage APIs and the DeleteLineageEvent API.
read more →

Flock’s Vehicle Fingerprinting Enables Plateless Surveillance

🚨 A 2024 company presentation reveals that Flock uses a so-called “Vehicle Fingerprint” combining decals, bumper stickers, racks and temporary tags to identify cars when license plates are incomplete or absent. The system enables officers to search that dataset, perform multi-geo queries and locate vehicles believed to be traveling together. Bruce Schneier notes this capability echoes older surveillance practices and warns that similar outcomes are possible with broad access to cell phone location data.
read more →

Google Cloud cleared for Dutch public sector use

🔒 Google Cloud announced completion of a Dutch data protection impact assessment (DPIA) by SLM Rijk, confirming there are no known high data protection risks when recommended measures are applied. The outcome enables the Dutch central public sector to adopt Google Cloud from a privacy-assessment perspective and builds on earlier DPIA work for Google Workspace. Google emphasizes continued investment in privacy-enhancing technologies and support resources for customers.
read more →

Papa John’s Uses Shopping Data to Target Ads

🍕Papa John’s partnered with NBCUniversal, Instacart, and media agency Carat to target consumers when they’re likely low on groceries by analyzing Instacart purchase patterns. The campaign creates custom audiences based on purchases of staples like eggs, milk, and produce, then serves tailored creatives on NBCU streaming with prompts such as “Light on groceries?” and QR codes. Carat framed the approach as learning what’s in consumers’ fridges without being “too creepy.” The author notes historical parallels and ethical concerns about such predictive advertising.
read more →

AWS Clean Rooms adds intermediate tables for SQL

🧩 AWS Clean Rooms now supports writing SQL query results to intermediate tables within a collaboration, enabling multi-step analytical workflows between partners. These intermediate tables allow reuse of complex joins and creation of shared ID mapping tables for downstream analyses, all within the collaboration’s privacy boundary. The feature helps reduce costs and improve performance for subsequent analyses such as reach, frequency, and attribution.
read more →

Google expands privacy controls for Search and Play

🔒 Google announced new privacy controls that separate saved history and personalization for Search services and Google Play, rolling out in users' Google Accounts in the coming days. The update creates distinct Search Services History and Personalized Recommendations settings, and similarly splits Play History and Personalization in Play. If Web & App Activity is on, the new Search Services History and its Save Media subsetting will be enabled after transition, but users can disable or delete saved media later.
read more →

AWS Glue Data Catalog adds semantic search preview

🔍 Today AWS announced a preview of business context and semantic search for AWS Glue Data Catalog, enabling discovery of data by semantic meaning. You can enrich catalog tables with glossary terms, custom metadata fields, and add skills that provide agents with additional context. The new Glue Search API lets you find tables by both structure and attached business meaning, and MCP-compatible agents can use the aws-data-analytics plugin to integrate with minimal setup.
read more →

Plan and Migrate Data with Azure Storage

📌 This blog explains a structured approach to enterprise storage migration using Microsoft tools. It emphasizes planning, assessment, and choosing the right migration path based on data volume, connectivity, and downtime tolerance. Key solutions covered include Azure Migrate, Azure Storage Mover, Azure Data Box, and a preview Azure Copilot Migration Agent. The post illustrates phased strategies, real customer examples, and guidance for regulated and AI use cases.
read more →