A TTL on raw_events Is Only Step One
You have a clean answer ready: set a 90-day TTL on the event table, schedule a nightly delete, done. Then the interviewer asks what happens to the funnel dashboard, the fraud cases, the free-text support tickets, and last month's backup. Four places hold the same user, and your design touched one.
This post walks one 30-minute, mid-level Data Engineer mock interview turn by turn, built from a real AI interview blueprint. The candidate answers are dramatized and illustrative, never a transcript of a real person or a real company's questions.
Key Findings
- The interview is scored out of 100 points; Interviewer Objectives Alignment and Level-Specific Expectations are worth 30 points each, so 60 points reward judgment over tooling.
- Technical Proficiency is only 20 of the 100 points, the same weight as Communication and Problem Solving.
- The blueprint runs 30 minutes in 3 phases: 8 minutes of policy mapping, 12 minutes of enforcement design, and 10 minutes of edge cases and trade-offs.
- The scenario has 4 source tables and 4 teams (product, fraud, support, analytics) with competing retention needs.
- Phase 1 expects you to call out 5 high-risk fields: IP address, message body, investigation notes, payload JSON, and attachment references.
- Across the 3 phases the blueprint lists 14 checklist items (4, 5, and 5), and the deletion follow-up alone spans 4 layers: raw, curated, aggregates, and backups.
What Is the Data Retention Interview Really Testing?
It tests whether you can turn a vague policy ("keep only what is necessary") into field-level decisions, enforcement points, and proof, while four teams argue for their own data.
The interview question
Your team supports a consumer app with web and mobile clients. Product, fraud, support, and analytics teams rely on a shared event platform and a user profile store. The sources are an append-only event stream, current user profiles, support tickets, and fraud cases.
raw_events(event_id, user_id, session_id, event_name, event_ts, ip_address, user_agent, device_id, page_url, referrer, search_query, payload JSON) user_profiles(user_id, email, phone, country, birth_date, marketing_opt_in, created_at, deleted_at) support_tickets(ticket_id, user_id, created_at, channel, subject, message_body, attachment_url, resolution_code) fraud_cases(case_id, user_id, opened_at, closed_at, risk_score, investigation_notes, linked_ip, linked_device_id)
A new company policy says teams must keep only data necessary for defined purposes, assign retention periods by data category, and automatically delete or de-identify expired data. Analytics still needs trend reporting, product wants funnel metrics, fraud wants longer access to certain identifiers, and support needs recent conversation history. Design how you would update the data platform to enforce data minimization and retention while still supporting those needs.
Under the surface, the interviewer is probing whether you can translate requirements into data models, retention rules, and deletion enforcement across raw, curated, and downstream datasets, and whether you reason honestly about analytics value versus minimization, late-arriving data, and backups.

Judgment and level signals account for 60 of the 100 points, so a flawless DELETE statement cannot carry the score alone.
Four Follow-Ups, Four Places Points Leak
Tomas is a prepared mid-level candidate. The mistakes below are ones the blueprint's checklist penalizes, not a recording of anyone.
Turn 1: Drop or Keep Briefly
Interviewer: "How would you decide which fields should be dropped at collection time versus retained briefly and deleted later?"
Turn 2: Year-Over-Year Funnels
Interviewer: "If product says they need year-over-year funnel reporting, how would you support that without keeping all raw event-level data indefinitely?"
Turn 3: Free Text and JSON
Interviewer: "How would you deal with fields like IP address, free-form support messages, and JSON payloads that may contain unexpected sensitive data?"
Turn 4: One Deletion, Four Layers
Interviewer: "How would your design handle user deletion requests when the same user's data appears in raw events, curated tables, downstream aggregates, and backups?"
Why Isn't Reading This Enough?
Spotting these four mistakes on the page is easy. Avoiding them live is the real skill: the clock is running, the interviewer interrupts, and the follow-up you have not seen asks about the exact layer you skipped. Knowing the right design is not the same as delivering it: under the clock it is easy to ramble, forget backups, or skip the trade-off when the interviewer pushes back. That habit only forms through reps, saying the plan out loud and recovering when a follow-up lands somewhere unplanned.
The Blueprint a Strong Candidate Hits
This is the blueprint a strong candidate covers across the 30 minutes, and exactly what the AI mock interview tracks you against in real time. Each phase has a time range, and the checklist fills in as you hit each item.

The first phase is the shortest at 8 minutes, so classify fields and state assumptions fast to leave the 12-minute design phase room for enforcement.
- ✓Asks or states assumptions about business purposes for events, support, fraud, and profiles
- ✓Separates raw identifiers, quasi-identifiers, free text, and aggregate data into different handling categories
- ✓Calls out high-risk fields such as IP address, message_body, investigation_notes, payload JSON, and attachment references
- ✓Suggests that some fields should be removed or reduced at collection rather than only relying on downstream deletion
- ✓Proposes a policy model such as retention metadata by dataset/field/purpose or tagged data classes
- ✓Describes where enforcement occurs, for example ingestion filters, transformation jobs, table TTLs, partition deletes, row-level delete workflows, or scheduled compaction/rewrite
- ✓Preserves analytics value via curated aggregate tables or de-identified summaries with longer retention than raw event-level records
- ✓Addresses nested or unstructured data by restricting payload contents, parsing and whitelisting fields, or scanning/quarantining unexpected sensitive content
- ✓Considers downstream dependencies so deletions or expirations propagate beyond source tables
- ✓Explains how to handle deletion requests, late-arriving records, reprocessing, and historical backfills
- ✓Mentions audit evidence such as deletion logs, policy versioning, counts of expired records removed, and periodic validation queries
- ✓Acknowledges backup constraints and gives a reasonable approach such as time-bounded backup retention and restore-time deletion controls
- ✓Identifies failure modes like orphaned derived data, missed partitions, schema drift, or job failures and proposes alerts or reconciliation checks
- ✓Makes a clear recommendation when business needs conflict with minimization, such as keeping aggregated metrics instead of raw identifiers
Run This Interview Live
The fastest way to find out whether your deletion design survives a follow-up is to be asked, out loud, by someone who will not let you skip backups.
Start the AI mock interview for this exact scenario and get scored against the same four dimensions, with turn-by-turn coaching afterward. To warm up first, drill Data Engineer data minimization and retention questions in the question bank, skim the preparation guides, or browse Data Engineer openings on the job board.
FAQ
Q. How long is the Data Engineer data minimization and retention mock interview?
The blueprint runs 30 minutes in three phases: problem framing and policy mapping (minutes 0-8), technical design for enforcement (8-20), and edge cases, operations, and tradeoffs (20-30).
Q. How is a Data Engineer data retention interview scored?
Out of 100 points: 30 for Interviewer Objectives Alignment, 30 for Level-Specific Expectations, 20 for Technical Proficiency, and 20 for Communication and Problem Solving. Judgment and seniority signals carry 60 of the 100 points.
Q. Should a Data Engineer keep raw event data after the retention period ends?
No. Strong answers keep longer-lived aggregates or de-identified summaries for trend and funnel reporting, and expire the identifiable event-level records on a shorter schedule tied to a defined purpose.
Q. How do you handle a user deletion request across raw events, curated tables, aggregates, and backups?
Treat it as a propagation problem. Delete or de-identify in the raw layer, push the same delete through curated and derived tables, confirm aggregates no longer carry identifiable rows, and bound backup retention so restored data is re-deleted from a deletion log.
Q. How do you prove that a retention job actually deleted the data?
Keep audit evidence: deletion logs, policy versions, counts of expired records removed, and periodic validation queries that look for rows past their retention date. A job that reports success is not proof.
Q. What separates a mid-level answer from a junior one on data retention?
A mid-level candidate structures the problem into data classes, retention rules, and enforcement points without being led, and makes field-level choices to aggregate, tokenize, hash, or drop. Naming the legal framework perfectly is not required.
Where Your Next Rep Goes
Retention interviews reward the candidate who starts with purpose and ends with proof. Pick one field, say why it exists, say when it dies, and say how you would show that it did.
Topics
Ready to practice?
Put what you've learned into practice with AI mock interviews and structured preparation guides.