This is the lightboard from my Datadog Illuminated episode with Julien Le Dem on OpenLineage.

Before OpenLineage, if your data broke somewhere in the pipeline, you had no way to trace it back to where it went wrong. OpenLineage fixes that by giving every tool a shared language for where data came from, so you can find the exact source of the problem instead of guessing.

Read the full board notes here: gist.github.com/wiggitywh…

A chalkboard-style lightboard diagram titled "OpenLineage - A common observability framework for data pipelines." The left side lists problems: isolated teams, bad data propagates (errors, privacy), finding bad data sources is hard, no common language. The center shows a diagram of Org A, Org B, and Org C connected by data flow arrows feeding into a "Data Lineage Repo," along with definitions for Job, Run, and Datasets (input/output), and a note that facets are metadata specs specific to a technology. Below that, a box lists what OpenLineage unlocks: trying different tools easily, easier onboarding, root cause identification, data reliability, and impact analysis and compliance. The right side shows OpenLineage as a hub connecting lineage producers (Airflow, Spark, Snowflake) to lineage consumers (Marquez, DataHub, Datadog), with a circled note reading "OpenLineage at Datadog" listing data job monitoring, data quality monitoring, and lineage graph use cases.