Interview Prep11 min read

Data Engineer Data Retention Interview: Delete Raw, Keep Funnels

A TTL on one table is not a retention design. See where a prepared mid-level Data Engineer loses points in a 30-minute retention interview, then practice live.

IT
InterviewStack TeamResearch
|

A TTL on raw_events Is Only Step One

You have a clean answer ready: set a 90-day TTL on the event table, schedule a nightly delete, done. Then the interviewer asks what happens to the funnel dashboard, the fraud cases, the free-text support tickets, and last month's backup. Four places hold the same user, and your design touched one.

This post walks one 30-minute, mid-level Data Engineer mock interview turn by turn, built from a real AI interview blueprint. The candidate answers are dramatized and illustrative, never a transcript of a real person or a real company's questions.

Key Findings

  • The interview is scored out of 100 points; Interviewer Objectives Alignment and Level-Specific Expectations are worth 30 points each, so 60 points reward judgment over tooling.
  • Technical Proficiency is only 20 of the 100 points, the same weight as Communication and Problem Solving.
  • The blueprint runs 30 minutes in 3 phases: 8 minutes of policy mapping, 12 minutes of enforcement design, and 10 minutes of edge cases and trade-offs.
  • The scenario has 4 source tables and 4 teams (product, fraud, support, analytics) with competing retention needs.
  • Phase 1 expects you to call out 5 high-risk fields: IP address, message body, investigation notes, payload JSON, and attachment references.
  • Across the 3 phases the blueprint lists 14 checklist items (4, 5, and 5), and the deletion follow-up alone spans 4 layers: raw, curated, aggregates, and backups.

What Is the Data Retention Interview Really Testing?

It tests whether you can turn a vague policy ("keep only what is necessary") into field-level decisions, enforcement points, and proof, while four teams argue for their own data.

The interview question

Your team supports a consumer app with web and mobile clients. Product, fraud, support, and analytics teams rely on a shared event platform and a user profile store. The sources are an append-only event stream, current user profiles, support tickets, and fraud cases.

raw_events(event_id, user_id, session_id, event_name, event_ts,
  ip_address, user_agent, device_id, page_url, referrer,
  search_query, payload JSON)
user_profiles(user_id, email, phone, country, birth_date,
  marketing_opt_in, created_at, deleted_at)
support_tickets(ticket_id, user_id, created_at, channel, subject,
  message_body, attachment_url, resolution_code)
fraud_cases(case_id, user_id, opened_at, closed_at, risk_score,
  investigation_notes, linked_ip, linked_device_id)

A new company policy says teams must keep only data necessary for defined purposes, assign retention periods by data category, and automatically delete or de-identify expired data. Analytics still needs trend reporting, product wants funnel metrics, fraud wants longer access to certain identifiers, and support needs recent conversation history. Design how you would update the data platform to enforce data minimization and retention while still supporting those needs.

Under the surface, the interviewer is probing whether you can translate requirements into data models, retention rules, and deletion enforcement across raw, curated, and downstream datasets, and whether you reason honestly about analytics value versus minimization, late-arriving data, and backups.

Interviewer scoring weights for the Data Engineer data minimization and retention interview

Judgment and level signals account for 60 of the 100 points, so a flawless DELETE statement cannot carry the score alone.

Four Follow-Ups, Four Places Points Leak

Tomas is a prepared mid-level candidate. The mistakes below are ones the blueprint's checklist penalizes, not a recording of anyone.

Turn 1: Drop or Keep Briefly

Interviewer: "How would you decide which fields should be dropped at collection time versus retained briefly and deleted later?"

COMMON MISTAKE
Tomas says to collect everything and rely on a delete job later, with one blanket retention period for the whole platform. That misses Phase 1's last checklist item, which asks you to cut fields at collection rather than only relying on downstream deletion, and it costs points on Interviewer Objectives Alignment.
STRONGER MOVE
Start from purpose: for each field, name who needs it and why, then sort fields into drop now, reduce (truncate the IP, drop the user agent detail), keep briefly, and keep as an aggregate. Say your assumptions out loud, then flag the riskiest fields first: message body, investigation notes, payload JSON.

Turn 2: Year-Over-Year Funnels

Interviewer: "If product says they need year-over-year funnel reporting, how would you support that without keeping all raw event-level data indefinitely?"

COMMON MISTAKE
Tomas either refuses the request outright or agrees to keep raw events for two years "just in case". Neither picks a side when business needs conflict with minimization, the last checklist item, so Level-Specific Expectations suffer.
STRONGER MOVE
Recommend a curated aggregate or de-identified summary table that outlives the raw rows: counts by funnel step, cohort, and week, with no user-level identifiers. Then state the trade-off plainly: you lose ad hoc re-slicing of old data, and product should list the dimensions up front.

Turn 3: Free Text and JSON

Interviewer: "How would you deal with fields like IP address, free-form support messages, and JSON payloads that may contain unexpected sensitive data?"

COMMON MISTAKE
Tomas treats the payload column as one opaque blob with one expiry date and never mentions what may be hiding inside it. That skips the checklist item on nested and unstructured data, a hit to Technical Proficiency.
STRONGER MOVE
Prefer an allowlist: parse the payload at ingestion and keep only approved keys, quarantining anything unknown. For free text, scan and redact before it lands in long-lived tables, and give it a shorter retention class than structured fields.

Turn 4: One Deletion, Four Layers

Interviewer: "How would your design handle user deletion requests when the same user's data appears in raw events, curated tables, downstream aggregates, and backups?"

COMMON MISTAKE
Tomas deletes the user from the raw table and calls it done, never mentioning derived tables, late-arriving records, or backups. That leaves out the deletion-request and backup items in the final phase and costs points on Interviewer Objectives Alignment, since the interviewer's stated objectives name late-arriving data and backups.
STRONGER MOVE
Walk the lineage layer by layer, and keep a deletion log so late-arriving events for that user get dropped on arrival. For backups, propose time-bounded retention plus a restore-time step that replays the deletion log, and admit you will not rewrite old snapshots.

Why Isn't Reading This Enough?

Spotting these four mistakes on the page is easy. Avoiding them live is the real skill: the clock is running, the interviewer interrupts, and the follow-up you have not seen asks about the exact layer you skipped. Knowing the right design is not the same as delivering it: under the clock it is easy to ramble, forget backups, or skip the trade-off when the interviewer pushes back. That habit only forms through reps, saying the plan out loud and recovering when a follow-up lands somewhere unplanned.

The Blueprint a Strong Candidate Hits

This is the blueprint a strong candidate covers across the 30 minutes, and exactly what the AI mock interview tracks you against in real time. Each phase has a time range, and the checklist fills in as you hit each item.

The 30-minute interview blueprint timeline for the Data Engineer data minimization and retention interview

The first phase is the shortest at 8 minutes, so classify fields and state assumptions fast to leave the 12-minute design phase room for enforcement.

Blueprinta strong 30-minute interview, phase by phase
1
Problem framing and policy mapping 0-8
  • ✓Asks or states assumptions about business purposes for events, support, fraud, and profiles
  • ✓Separates raw identifiers, quasi-identifiers, free text, and aggregate data into different handling categories
  • ✓Calls out high-risk fields such as IP address, message_body, investigation_notes, payload JSON, and attachment references
  • ✓Suggests that some fields should be removed or reduced at collection rather than only relying on downstream deletion
2
Technical design for enforcement 8-20
  • ✓Proposes a policy model such as retention metadata by dataset/field/purpose or tagged data classes
  • ✓Describes where enforcement occurs, for example ingestion filters, transformation jobs, table TTLs, partition deletes, row-level delete workflows, or scheduled compaction/rewrite
  • ✓Preserves analytics value via curated aggregate tables or de-identified summaries with longer retention than raw event-level records
  • ✓Addresses nested or unstructured data by restricting payload contents, parsing and whitelisting fields, or scanning/quarantining unexpected sensitive content
  • ✓Considers downstream dependencies so deletions or expirations propagate beyond source tables
3
Edge cases, operations, and tradeoffs 20-30
  • ✓Explains how to handle deletion requests, late-arriving records, reprocessing, and historical backfills
  • ✓Mentions audit evidence such as deletion logs, policy versioning, counts of expired records removed, and periodic validation queries
  • ✓Acknowledges backup constraints and gives a reasonable approach such as time-bounded backup retention and restore-time deletion controls
  • ✓Identifies failure modes like orphaned derived data, missed partitions, schema drift, or job failures and proposes alerts or reconciliation checks
  • ✓Makes a clear recommendation when business needs conflict with minimization, such as keeping aggregated metrics instead of raw identifiers

Run This Interview Live

The fastest way to find out whether your deletion design survives a follow-up is to be asked, out loud, by someone who will not let you skip backups.

Start the AI mock interview for this exact scenario and get scored against the same four dimensions, with turn-by-turn coaching afterward. To warm up first, drill Data Engineer data minimization and retention questions in the question bank, skim the preparation guides, or browse Data Engineer openings on the job board.

FAQ

Q. How long is the Data Engineer data minimization and retention mock interview?

The blueprint runs 30 minutes in three phases: problem framing and policy mapping (minutes 0-8), technical design for enforcement (8-20), and edge cases, operations, and tradeoffs (20-30).

Q. How is a Data Engineer data retention interview scored?

Out of 100 points: 30 for Interviewer Objectives Alignment, 30 for Level-Specific Expectations, 20 for Technical Proficiency, and 20 for Communication and Problem Solving. Judgment and seniority signals carry 60 of the 100 points.

Q. Should a Data Engineer keep raw event data after the retention period ends?

No. Strong answers keep longer-lived aggregates or de-identified summaries for trend and funnel reporting, and expire the identifiable event-level records on a shorter schedule tied to a defined purpose.

Q. How do you handle a user deletion request across raw events, curated tables, aggregates, and backups?

Treat it as a propagation problem. Delete or de-identify in the raw layer, push the same delete through curated and derived tables, confirm aggregates no longer carry identifiable rows, and bound backup retention so restored data is re-deleted from a deletion log.

Q. How do you prove that a retention job actually deleted the data?

Keep audit evidence: deletion logs, policy versions, counts of expired records removed, and periodic validation queries that look for rows past their retention date. A job that reports success is not proof.

Q. What separates a mid-level answer from a junior one on data retention?

A mid-level candidate structures the problem into data classes, retention rules, and enforcement points without being led, and makes field-level choices to aggregate, tokenize, hash, or drop. Naming the legal framework perfectly is not required.

Where Your Next Rep Goes

Retention interviews reward the candidate who starts with purpose and ends with proof. Pick one field, say why it exists, say when it dies, and say how you would show that it did.

Topics

Data Engineerdata retentiondata minimizationmock interviewinterview walkthroughdata deletion

Ready to practice?

Put what you've learned into practice with AI mock interviews and structured preparation guides.