Home › Interview Questions › Suppose the source system contains duplicate recor…

Suppose the source system contains duplicate records. How would you design a scalable pipeline to handle them, and can you describe your end‑to‑end experience from design to deployment?

🟡 Medium Conceptual Senior level
1Times asked
Sep 2026Last seen
Sep 2026First seen

💡 Model Answer

I would start by defining a unique key for the entity and capturing changes via CDC or incremental snapshots. The ingestion layer uses Kafka or a cloud queue to buffer records, ensuring at‑least‑once delivery. In the processing layer, a Spark Structured Streaming job reads the stream, applies a dedupe window using a stateful aggregation keyed on the unique key, and writes the cleaned data to a Delta Lake or Snowflake. Partitioning by date or key range keeps the job scalable. For the target, I use a MERGE statement to upsert into the final table, guaranteeing idempotence. I prototype the job in a dev cluster, run unit tests with sample data, and use CI/CD pipelines (GitHub Actions + Airflow) to deploy to production. Monitoring is set up with Prometheus and Grafana to track duplicate rates and job latency. My experience includes designing such pipelines for a financial services client, where we reduced duplicate processing by 95% and achieved 99.9% data freshness.

This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.

🎤 Get questions like this answered in real-time

Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.

Get Assisting AI — Starts at ₹500