What is shuffling in Spark?
💡 Model Answer
In Spark, shuffling is the process of redistributing data across the cluster based on a key, typically triggered by operations like groupBy, join, or reduceByKey. During a shuffle, each executor writes its partitioned data to disk, then other executors read the relevant partitions from the network. This incurs significant I/O and network overhead, so minimizing shuffles is key to performance. Techniques such as broadcast joins, partitioning data to match join keys, or using map-side combine can reduce shuffle volume. Understanding when Spark triggers a shuffle helps you design more efficient transformations.
This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.
🎤 Get questions like this answered in real-time
Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.
Get Assisting AI — Starts at ₹500