Click any tag below to further narrow down your results
Links
Marquez is an open-source metadata server that collects and visualizes data lineage across your organization's pipelines. It works with major tools like Airflow, Spark, dbt, and Dagster to show you where data comes from, where it goes, and how jobs depend on each other.
- Real-time metadata collection through an OpenLineage-compatible API endpoint that integrates with existing data orchestration tools
- Web UI displays data dependencies and lineage as a visual graph, letting you trace datasets back through pipelines and see job inputs/outputs
- Lineage API enables automation for tasks like backfills and root cause analysis by letting you query dependencies across multiple platforms
Nvidia's dominance is shifting from raw GPU competition to controlling the entire data center infrastructure around compute. As AI systems scale to gigawatt levels, the company's specialized hardware for data orchestration—CPUs, networking, storage—is becoming harder to replicate than the chips themselves.
- Nvidia's Vera CPU and supporting hardware deliver 3x performance improvements by optimizing data movement to GPUs, addressing a critical bottleneck as companies optimize for tokens-per-watt efficiency.
- Competitors like OpenAI are tackling the same data movement problem differently (integrated chips like Jalapeño), but the underlying challenge shows the competition has shifted from GPU design to full-system efficiency.
- Operating megascale data centers at peak efficiency is still incredibly difficult, creating a new competitive layer where system integration matters more than individual chip superiority.