Data Transformation and Processing Logic Questions

Implementing transformation logic: joins, aggregations, deduplication, pivoting/reshaping, and business-rule application over datasets. Covers writing correct and maintainable transformation code, handling edge cases in the transform layer, and preparing data for downstream consumption. Focuses on the logic of turning raw data into analytics-ready outputs.

MediumTechnical
60 practiced

Implement a defensive type-casting utility that attempts to cast a column (or Series) to a target type, returning both the successfully cast values and the indices that failed to cast, rather than raising an exception or silently producing bad data. Discuss common pitfalls of casting strings to numbers or dates in real-world data, and how your function behaves on an empty input or an input where every value fails to cast.

MediumTechnical
37 practiced

A numeric column contains both NULLs and extreme outliers. Describe a pragmatic decision process for handling both before analytics: statistical detection methods (IQR, z-score), business-rule thresholds, when to impute versus flag versus drop, and how you would document and test these decisions so downstream analysts understand what happened to their data.

HardTechnical
34 practiced

During feature or dataset preparation for a supervised model, a join can accidentally introduce label leakage: for example, joining in a table that contains information only available after the outcome is known, or a join that looks ahead in time. Describe common sources of this kind of leakage, how you would detect it automatically, and how you would design a rollback or re-computation process once leakage is found.

MediumTechnical
34 practiced

You receive timestamps in several inconsistent formats and time zones from different sources (for example '2024-01-02 13:00', '01/02/2024 1pm PST', an epoch value, or a date string with an ambiguous day/month order). Design a robust approach to parse and normalize all of these into timezone-aware UTC timestamps. Discuss how you detect and handle ambiguous dates (02/03/2024), missing timezone information, and DST transitions, and how you would detect signs of clock skew or duplicated timestamps between sources before trusting the normalized result.

MediumTechnical
42 practiced

Implement a Python generator that reads a large CSV file in fixed-size chunks, applies a transformation to each chunk, and yields (or persists) the transformed results without ever loading the whole file into memory. Handle header rows, malformed lines, and different encodings, and describe how the job could resume from the last successfully processed chunk after a failure.

Unlock Full Question Bank

Get access to all 34 Data Transformation and Processing Logic interview questions and detailed answers.

Sign in to Continue

Join thousands of developers preparing for their dream job.