Data Transformation and Processing Logic Questions

Implementing transformation logic: joins, aggregations, deduplication, pivoting/reshaping, and business-rule application over datasets. Covers writing correct and maintainable transformation code, handling edge cases in the transform layer, and preparing data for downstream consumption. Focuses on the logic of turning raw data into analytics-ready outputs.

HardTechnical
54 practiced

A join between a very large fact table and a much smaller dimension or lookup table is suffering from severe key skew: a handful of hot keys account for a large share of the rows, causing straggler tasks and out-of-memory failures in a distributed engine like Spark. Explain and compare the standard mitigation techniques: broadcasting the small side, salting the skewed keys, pre-aggregating the hot keys before the join, and letting the engine's adaptive query execution handle it automatically. Discuss when each is the right choice and what can go wrong with a broadcast join if your assumption about which side is 'small' turns out to be wrong.

MediumTechnical
42 practiced

Implement a Python generator that reads a large CSV file in fixed-size chunks, applies a transformation to each chunk, and yields (or persists) the transformed results without ever loading the whole file into memory. Handle header rows, malformed lines, and different encodings, and describe how the job could resume from the last successfully processed chunk after a failure.

MediumTechnical
34 practiced

A dataset has missing values scattered across several columns. Walk through how you would decide what to do about each column's missingness, and how you would document and communicate that decision to stakeholders who will consume the resulting dashboard or model.

MediumTechnical
30 practiced

A Spark job must process JSON records with deeply nested arrays and optional fields, producing a wide, flat analytical table. Show PySpark code that safely explodes the nested arrays and selects the needed fields, handling optional/missing nested keys and empty arrays without failing the job.

MediumTechnical
28 practiced

Design and implement a streaming deduplication component that consumes a stream of (id, timestamp) events and reports whether each event is new or a duplicate within a bounded time window, using bounded memory (an LRU cache or a Bloom filter, your choice). Discuss the correctness trade-off of the approach you chose: can it produce false positives or false negatives, and what does that mean for events that get silently dropped or double-counted?

Unlock Full Question Bank

Get access to all 42 Data Transformation and Processing Logic interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.