Click any tag below to further narrow down your results
Links
Marquez is an open-source metadata server that collects and visualizes data lineage across your organization's pipelines. It works with major tools like Airflow, Spark, dbt, and Dagster to show you where data comes from, where it goes, and how jobs depend on each other.
- Real-time metadata collection through an OpenLineage-compatible API endpoint that integrates with existing data orchestration tools
- Web UI displays data dependencies and lineage as a visual graph, letting you trace datasets back through pipelines and see job inputs/outputs
- Lineage API enables automation for tasks like backfills and root cause analysis by letting you query dependencies across multiple platforms
Organizations face three recurring data problems—inconsistent metric definitions, fragmented access controls, and metric changes that don't propagate everywhere. A semantic layer solves this by centralizing metric definitions and governance in one place, so all tools pull the same numbers and changes cascade automatically. It won't fix bad data at the source, but it shrinks the surface area you need to manage and makes self-service analytics actually work.
- The same metric getting defined differently across Tableau, Power BI, and Python isn't a minor annoyance—it's a root cause of bad executive decisions.
- Hiring more BI analysts as gatekeepers just creates ticket queues and bottlenecks; it doesn't fix fragmented governance across tools.
- A semantic layer centralizes metric definitions and access controls so a single change (e.g., redefining ARR) propagates automatically everywhere instead of requiring manual updates across systems.
- Centralizing definitions and logic in one place also makes the data self-documenting, enabling real self-service instead of ticket-based requests.
Apache Ossie is an Apache Incubator project that defines a vendor-neutral YAML spec for semantic data models. It lets teams declare metrics, dimensions and joins once and share them across BI, analytics and AI tools. This prevents metric drift, cuts integration debt and creates a single source of truth.
- Apache Ossie standardizes metric/dimension definitions in YAML so BI, analytics, and AI tools all reference one shared source of truth instead of redefining metrics per dashboard.
- Over 50 organizations, including Snowflake, Databricks, Oracle, Salesforce, dbt Labs, Qlik, and ThoughtSpot, have joined the effort.
- The spec was renamed from Open Semantic Interchange to Apache Ossie in July 2026, shortly after launching a Financial Services Semantic Working Group in June 2026.
- It embeds AI context instructions so LLMs can ground answers in official business logic, aiming to cut reconciliation costs and eliminate conflicting dashboards.
Databricks is launching a Software-Defined Storage ecosystem that uses the open-source OpenSharing protocol to link on-premises, edge, and private-cloud systems directly into its Data Intelligence Platform. This zero-copy approach lets teams run serverless compute and train models on local datasets under Unity Catalog governance without migrating any data.
- Databricks now lets you query on-prem/edge/private-cloud data (via MinIO, Everpure, Qumulo, VAST Data) directly through Unity Catalog with zero-copy access and no egress fees, using the open-source OpenSharing protocol
- MinIO's AIStor integration is already GA, letting users run live queries on on-prem Iceberg and Delta tables; Everpure and Qumulo are in private preview, VAST Data joins in August
- Targeted at regulated/data-sovereignty-heavy industries (banks, healthcare, semiconductors, trading firms) facing GDPR/HIPAA/NIS2 rules and untenable cloud egress costs at exabyte scale
- Connections must pass Databricks's Partner Well-Architected Framework for security/certification before going live, preserving full lineage, access controls, and audit logs across hybrid environments
The article argues that companies are increasingly recording every meeting by default to feed AI systems the living context of their culture, decisions, and conversations. This shift turns unstructured voice data into a searchable, structured system of record that boosts individual productivity and executive oversight, making meeting recording inevitable.
- Meeting recording became default not through decision but because tools shipped with it on and nobody turned it off
- Companies like Bridgewater, OpenAI, and a16z (via Granola) already treat recorded conversation as a living system of record smarter than any wiki
- Skipping recording costs both individual productivity (no context-aware AI assistant) and executive oversight (no early warning system for risks)
- Verbal-culture companies (Shopify, OpenAI) have more to gain from this shift than document-heavy cultures (Stripe, Anthropic), since their key knowledge was previously unrecorded