What is broadcast join in Spark, and why would you broadcast a small table across all executors?
💡 Model Answer
Broadcast join is a join strategy in Spark where the smaller dataset is sent to every executor so that each executor can perform the join locally with its partition of the larger dataset. This eliminates the shuffle that would normally be required for a regular join, reducing network I/O and improving performance. You would use broadcast join when one side of the join is small enough to fit comfortably in the memory of each executor, such as a lookup or dimension table that is a few megabytes or less. The trade‑off is that the broadcasted data is replicated on every executor, so you must ensure that the total size of the broadcasted dataset plus the executor memory overhead does not exceed the available memory. Complexity is O(n + m) where n and m are the sizes of the two datasets, compared to O(n log n + m log m) for a shuffle join.
This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.
🎤 Get questions like this answered in real-time
Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.
Get Assisting AI — Starts at ₹500