What file formats did you use for raw data ingestion, and why did you choose JSON over Parquet for the raw layer?
💡 Model Answer
In a typical data lake architecture, the raw layer is the first landing zone for data. It should preserve the original format to avoid data loss. Common formats include CSV, JSON, Parquet, Avro, and Delta Lake. CSV is simple but lacks schema enforcement and is inefficient for large volumes. JSON is semi‑structured, easy to ingest, and supports nested structures, making it a good choice for raw ingestion when the downstream system can handle schema evolution. Parquet is columnar, compressed, and efficient for analytical workloads, but it requires a defined schema upfront. Delta Lake adds ACID transactions and schema enforcement on top of Parquet. Choosing JSON for the raw layer allows you to keep the data as‑is, capture all fields, and later transform it into a more efficient format like Parquet or Delta for downstream processing. It also simplifies debugging and data lineage. In contrast, storing raw data in Parquet would require you to define a schema early, which might not be feasible if the source schema is evolving. Therefore, JSON is often preferred for raw ingestion, while Parquet/Delta is used in the curated layer.
This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.
🎤 Get questions like this answered in real-time
Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.
Get Assisting AI — Starts at ₹500