You have an upstream producer that recently added a new column to the dataset. How would you handle this to ensure your downstream job continues to run without breaking?
💡 Model Answer
When an upstream source adds a new column, the downstream job must be resilient to schema changes. First, use a schema registry (e.g., Confluent Schema Registry, AWS Glue Schema Registry) so the producer publishes the updated schema and consumers can retrieve it. In the consumer, enable schema evolution support in the data format (Avro, Parquet, ORC) so that new fields are treated as optional. If the job is written in Spark, you can use the spark.sql.columnNameOfAlias or schema option to read the latest schema and use DataFrame.selectExpr to ignore unknown columns. For legacy code that expects a fixed schema, add a fallback: read the raw data into a generic structure, then map to the expected schema, filling missing columns with nulls or defaults. Finally, test the pipeline with the new schema in a staging environment before promoting to production. This approach keeps the job running while the new column is gradually adopted.
This answer was generated by AI for study purposes. Use it as a starting point — personalize it with your own experience.
🎤 Get questions like this answered in real-time
Assisting AI listens to your interview, captures questions live, and gives you instant AI-powered answers on a discreet on-screen overlay.
Get Assisting AI — Starts at ₹500