One workflow, two graphs
The example below has three jobs and one view. JOB_A loadsstg_orders, JOB_REPORT writes report_daily fromv_customer_orders, and the view also reads customers — a table loaded by another team. Walk the scenarios and watch what each graph can and cannot see.
Baseline: How does the workflow look before anything goes wrong?
Not in this graph: customers is read by the view but has no declared depends_on edge to JOB_REPORT.
v_customer_ordersThe scheduler knows JOB_REPORT waits for JOB_A. The data lineage also shows that report_daily depends on customers through v_customer_orders. The two graphs already describe different facts.
Why the two graphs disagree
- A view hides its SQL inputs from the scheduler. The scheduler only sees the jobs and the dependencies someone declared. The view reads
stg_ordersandcustomers, but only the job relationship is declared. - A dependency can exist without data flow. A job may wait for an upstream job that writes nothing it reads — for example a cleanup, a gate or a notification job.
- Data can flow without a dependency. Tables loaded by another team, external sources and manually maintained tables have lineage but no scheduler edge.
- Dynamic SQL hides the edges. A query assembled at runtime can read different tables per execution, so neither static lineage nor declared dependencies fully describe it.
“What runs next?” is an execution question. “Where did this value come from?” is a data question. A single graph cannot answer both without becoming ambiguous.
What this does to impact analysis
Impact analysis starts from a change — a late table, a corrected value, a failed job — and asks what else is affected. The answer depends on which graph you query:
| Change | Scheduler graph | Data lineage graph |
|---|---|---|
customers arrives late | No impact path — no declared dependency | Finds report_daily as affected |
JOB_A must be rerun | Shows which jobs to re-trigger | Shows which tables are rebuilt |
JOB_REPORT is skipped | Shows the missed execution | Shows which data becomes stale |
The practical rule: use lineage to enumerate affected data objects, then use the scheduler graph to plan the execution — which jobs to rerun and in what order. Neither is a substitute for the other.
Keeping the graphs aligned without over-declaring
- Derive candidate dependencies from lineage. If a job reads a table that another job writes, that is a candidate
depends_onedge. Review it, do not declare it blindly. - Mark external or shared tables explicitly. If
customersis loaded outside the pipeline, record its owner and freshness expectation so the missing scheduler edge is a known decision. - Do not declare every read. Adding a dependency for every table access can create cycles and unnecessary waits. Declare what must be ready before the job starts.
- Test the disagreement. Periodically compare the declared dependencies with the lineage of the queries. The differences are the places where impact analysis will be wrong.
Edge cases and trade-offs
- View chains. Each view adds a layer the scheduler does not see. Lineage must expand through the view definition to find the real tables.
- Temporary and staging tables. They create lineage edges that do not map to a durable scheduler dependency; decide whether they belong in the lineage graph.
- Cross-team loads. A table loaded by another team may have an SLA but no job edge in this scheduler; a sensor or freshness check can make the wait explicit.
- Column-level lineage. Table-level lineage answers “is this table affected”. Column-level lineage answers “is this field affected” and is more expensive to maintain.