You have a Spark job that normally runs in 15 minutes but today it ran for two hours. How would you debug this performance regression?
💡 Model Answer
First, open the Spark UI and look at the stages that took the longest. Check the number of tasks, their durations, and any failed or long‑running tasks. Inspect the shuffle read/write metrics to see if a particular stage has a huge shuffle size or many tasks. Look for data skew: a single partition may be much larger than others. Examine the executor logs for GC pauses or memory spills. Compare the current job's configuration (spark.sql.shuffle.partitions, broadcast thresholds) with the previous run. If the schema changed, verify that the new column caused a broadcast join to fail or a larger shuffle. Use the spark.sql.adaptive.enabled flag to let Spark adjust partitions. If the job reads from a new source, check network latency or source throughput. Finally, profile the code: use explain(true) to see the physical plan, and add cache() or persist() where appropriate. By systematically narrowing down the culprit—skew, configuration drift, or external factors—you can restore the 15‑minute runtime.
This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.
🎤 Get questions like this answered in real-time
Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.
Get Assisting AI — Starts at ₹500