Home › Interview Questions › How would you solve the small file problem?

How would you solve the small file problem?

🟡 Medium Conceptual Junior level
1Times asked
Sep 2026Last seen
Sep 2026First seen

💡 Model Answer

The small file problem arises in Hadoop/Spark when a dataset is split into many tiny files (often < 64 MB). This leads to high metadata overhead, increased job scheduling time, and poor compression. To solve it, you can: 1) Combine small files into larger ones using coalesce or repartition before writing; 2) Use mapreduce.fileoutputcommitter.algorithm.version=2 to merge files during write; 3) Enable spark.sql.files.maxPartitionBytes to control partition size; 4) Use partitioning schemes that produce fewer files per partition; 5) Employ tools like Hadoop DistCp or HDFS concat to merge files post‑write; 6) Store data in columnar formats (Parquet, ORC) that support efficient compression and predicate pushdown. Additionally, consider using a data lakehouse approach where small files are periodically compacted by a background job. These strategies reduce the number of files, lower metadata load, and improve query performance.

This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.

🎤 Get questions like this answered in real-time

Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.

Get Assisting AI — Starts at ₹500