Home › Interview Questions › How would you design immutable lineage for data pi…

How would you design immutable lineage for data pipelines, ensuring that each run gets a new versioned snapshot and rollback is handled by pointing the 'current' flag to the previous snapshot?

🟡 Medium Conceptual Mid level
1Times asked
Sep 2026Last seen
Sep 2026First seen

💡 Model Answer

Immutable lineage treats metadata as an append‑only log. For each pipeline run, generate a unique run_id and capture the full graph of source tables, transformations, and destination tables. Store this graph in a versioned metadata store (e.g., a relational DB or a graph database like Neo4j) with a timestamp and run_id. The 'current' view is a materialized view or a flag column that points to the latest run_id. When a rollback is required, simply update the flag to the previous run_id; no data is altered, only the pointer changes. This approach guarantees auditability, as every change is recorded. To support queries, expose lineage via a REST API or a UI that can traverse the graph. For performance, index the run_id and use incremental snapshots to avoid full graph rewrites. Additionally, store lineage in a separate immutable store (e.g., S3 objects with versioning) so that even if the database is corrupted, the lineage history remains intact. This design ensures that lineage is tamper‑proof, auditable, and easy to rollback.

This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.

🎤 Get questions like this answered in real-time

Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.

Get Assisting AI — Starts at ₹500