How would you design a job that processes one million records, ensuring that after a restart it resumes from the last successfully processed record instead of reprocessing all records?
💡 Model Answer
To resume processing after a restart, you need a checkpointing or offset mechanism. Store the last successfully processed record ID or timestamp in a durable store (database, file, or distributed key‑value store). When the job starts, read the checkpoint and skip records before that point. Use idempotent operations so reprocessing the same record does not cause side effects. In frameworks like Spark or Flink, enable checkpointing and set the checkpoint interval. For custom batch jobs, write the checkpoint after each batch or after a set number of records. Also consider using a message queue with consumer offsets (Kafka) or a database transaction log. This approach guarantees that after a crash or restart, the job continues from the last successful point without reprocessing the entire dataset.
This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.
🎤 Get questions like this answered in real-time
Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.
Get Assisting AI — Starts at ₹500