Applied Scientist (Junior Level) Interview Preparation Guide - FAANG Standards
Applied Scientist interviews at FAANG companies follow a rigorous multi-stage process designed to evaluate your ability to conduct applied research, develop ML/AI algorithms, prototype solutions, and collaborate across teams. The process typically spans 4-6 weeks and includes initial recruiter screening, technical phone rounds assessing ML fundamentals and coding proficiency, and comprehensive onsite rounds covering research problem-solving, system design for ML systems, statistical analysis, and behavioral/cultural fit. For junior-level candidates, the focus is on demonstrating solid fundamentals, hands-on experience with real projects, ability to work independently with occasional guidance, and strong learning potential.
Interview Rounds
Recruiter Screening
What to Expect
Your first conversation with a recruiter focused on assessing basic qualifications, understanding your career aspirations, and evaluating cultural fit. The recruiter will review your background, discuss why you're interested in the Applied Scientist role, and provide an overview of the interview process ahead. This is an opportunity to showcase your soft skills, enthusiasm for the company, and alignment with their mission.
Tips & Advice
Research the company's ML/AI initiatives and product roadmap before this call. Be prepared to discuss your resume in detail, especially projects involving machine learning, research, or algorithm development. Clearly articulate why you want to join this specific company and how your interests align with their work in AI/ML. Prepare 2-3 brief examples of projects you've worked on that demonstrate your applied research capabilities. Ask thoughtful questions about the role, team structure, and research focus areas. Show enthusiasm for learning and solving real-world problems.
Focus Topics
Communication and Soft Skills
Ability to communicate clearly, listen actively, and demonstrate professionalism in conversation
Practice Interview
Study Questions
Resume and Project Discussion
Ability to clearly articulate your background, projects, and experiences relevant to applied research and ML development
Practice Interview
Study Questions
Career Motivation and Company Fit
Understanding of why you're pursuing an Applied Scientist role and alignment with the company's mission and values
Practice Interview
Study Questions
Technical Phone Screen - Machine Learning Fundamentals and Statistics
What to Expect
This technical phone interview assesses your foundational knowledge of machine learning concepts and statistical reasoning. You'll be asked to explain core ML algorithms, their applications, probability theory, hypothesis testing, and statistical inference. The interviewer will evaluate your understanding of when and why you'd choose specific approaches, your ability to explain concepts clearly, and your grasp of bias-variance tradeoff and model evaluation. Expect questions that start with fundamentals but may probe deeper into your understanding.
Tips & Advice
Focus on understanding the 'why' behind algorithms, not just the 'what'. Be prepared to derive or explain key equations for common models. Practice explaining ML concepts at different levels of complexity - you should be able to explain logistic regression to both a technical audience and a product manager. Use whiteboard or shared document to sketch diagrams when explaining concepts. When asked about your projects, focus on the ML methodology: problem definition, feature engineering, model selection, evaluation metrics, and results. For junior level, it's acceptable to say 'I wasn't sure about X so I researched it' - this shows good learning habits. Always validate assumptions and clarify ambiguous questions before diving into answers.
Focus Topics
Unsupervised Learning and Dimensionality Reduction
Clustering methods (K-means, hierarchical), PCA, t-SNE, and when to use each for exploratory analysis and feature engineering
Practice Interview
Study Questions
Supervised Learning Algorithms
Linear regression, logistic regression, decision trees, ensemble methods (random forests, gradient boosting), and when to apply each
Practice Interview
Study Questions
Bias-Variance Tradeoff and Model Evaluation
Understanding overfitting vs underfitting, regularization techniques, cross-validation, and appropriate evaluation metrics for different problem types
Practice Interview
Study Questions
Hypothesis Testing and Statistical Inference
p-values, confidence intervals, significance levels, Type I/II errors, power analysis, and interpreting experimental results
Practice Interview
Study Questions
Probability and Bayesian Reasoning
Conditional probability, Bayes' theorem, prior and posterior probability, and application in ML contexts like classification
Practice Interview
Study Questions
Technical Phone Screen - Coding and Data Structures
What to Expect
This round evaluates your programming proficiency in the context of data science and machine learning. You'll solve coding problems that may involve data manipulation, algorithm implementation, or optimization. Problems are typically medium difficulty and focus on your ability to write clean, efficient code; think through edge cases; and communicate your approach. You may be asked to implement ML-adjacent algorithms, work with data structures efficiently, or optimize solutions. The interviewer is assessing problem-solving approach as much as the final solution.
Tips & Advice
Write pseudocode or outline your approach before jumping into implementation. Explain your thought process as you code - this helps the interviewer follow your logic and provide hints if needed. Start with a clear, straightforward solution even if it's not optimal - then discuss optimization. For junior level, an O(n²) solution that you can explain is better than stumbling through an O(n) approach you don't fully understand. Test your code with examples including edge cases. If you get stuck, ask clarifying questions and discuss your approach with the interviewer rather than sitting in silence. Practice on platforms like LeetCode (medium difficulty) and focus on problems involving arrays, hashmaps, sorting, and basic dynamic programming. Be comfortable coding in Python, as it's the standard for ML roles.
Focus Topics
Sorting and Searching
Common sorting algorithms (merge sort, quicksort conceptually), binary search, and when to apply each
Practice Interview
Study Questions
Basic Dynamic Programming
Recognizing DP problems, memoization vs tabulation, solving problems like climbing stairs, coin change, and longest subsequence
Practice Interview
Study Questions
Arrays and String Manipulation
Working with arrays, lists, string operations, slicing, and common patterns like two-pointer techniques and prefix sums
Practice Interview
Study Questions
Hash Maps and Sets
Hash table operations, collision handling conceptually, and using maps/sets for efficient lookups and deduplication
Practice Interview
Study Questions
Coding Best Practices
Writing readable code, handling edge cases, debugging systematically, and explaining your reasoning clearly
Practice Interview
Study Questions
Onsite Round 1 - Applied Research Problem and Algorithm Design
What to Expect
This comprehensive onsite round focuses on your ability to approach real-world applied research problems similar to what you'd encounter in the role. You'll be presented with a problem statement (often ambiguous) and asked to design a solution from first principles. This may involve proposing novel algorithms, selecting appropriate techniques, defining metrics, or planning an experimental approach. The interviewer will assess your problem-solving methodology, creativity, technical depth, and ability to make reasonable trade-offs with incomplete information. At the junior level, you're expected to ask good clarifying questions and propose reasonable solutions with clear justification.
Tips & Advice
Start by clarifying the problem: what are the constraints, what does success look like, what data is available, what computational resources matter? Don't jump to solutions. Think out loud and involve the interviewer in your reasoning - they may guide you toward productive directions. For junior level, it's perfectly acceptable to say 'I would research X' or 'I'm not sure about Y but here's my thinking'. Propose multiple approaches and discuss trade-offs before settling on one. Focus on end-to-end thinking: problem definition, approach, implementation strategy, evaluation methodology, and potential challenges. Use your real project experience as anchors - relate the problem to things you've actually done. Draw diagrams to explain your approach. Be prepared for the interviewer to poke holes in your solution and adjust gracefully.
Focus Topics
From Research to Production
Thinking about how research ideas translate to production systems, scalability considerations, and collaboration with engineering teams
Practice Interview
Study Questions
Handling Trade-offs and Constraints
Making informed decisions when optimizing for accuracy vs speed vs interpretability, considering computational constraints and business requirements
Practice Interview
Study Questions
Problem Decomposition and Clarification
Ability to break down ambiguous problems into well-defined components, ask clarifying questions, and establish clear success criteria
Practice Interview
Study Questions
Evaluation Metrics and Experimental Design
Selecting appropriate metrics for different objectives, designing validation strategies, and planning how to measure success
Practice Interview
Study Questions
Feature Engineering and Data Representation
Designing effective feature representations, transformations, and selection strategies that capture relevant information for the problem
Practice Interview
Study Questions
Algorithm and Approach Selection
Knowledge of when to apply different ML techniques, data structures, and algorithmic approaches based on problem constraints and characteristics
Practice Interview
Study Questions
Onsite Round 2 - Machine Learning System Design
What to Expect
This round assesses your ability to design machine learning systems at a higher level. You'll be asked to design the architecture of an ML system that solves a real-world problem at scale (e.g., 'Design a recommendation system for millions of users' or 'Design a fraud detection system'). The focus is on how you think about system architecture, data pipelines, model training and serving, monitoring, and deployment considerations. This differs from coding interviews - you're drawing high-level architecture, not writing code. You should discuss trade-offs between different design choices, scalability challenges, and how to iterate on the system. At junior level, you're expected to think about basic system design - data pipeline, training infrastructure, and serving considerations.
Tips & Advice
Start by clarifying requirements and constraints: scale (users, data volume, queries per second), latency requirements, accuracy requirements, infrastructure constraints. Ask about the problem context before diving into design. Use a structured approach: define the problem, propose a high-level architecture, discuss data pipeline, model training approach, serving strategy, and monitoring. Draw diagrams to illustrate components and data flow. For junior level, don't be expected to know every detail - focus on logical, reasonable design choices with clear justification. Discuss trade-offs: should you use a simple model or complex one? Batch or real-time predictions? This shows thinking beyond 'just make it accurate'. Consider feedback loops and how to monitor model performance in production. Use examples from your experience or things you've researched. Be comfortable saying 'I would need to learn more about X' while proposing reasonable approaches.
Focus Topics
Scalability and Performance Optimization
Designing systems that scale to millions of users/requests, handling large datasets, and optimizing for latency and throughput
Practice Interview
Study Questions
Monitoring, Evaluation, and Feedback Loops
Setting up monitoring for model performance in production, defining success metrics, and designing feedback loops for continuous improvement
Practice Interview
Study Questions
Model Training and Serving Trade-offs
Batch vs online training, batch vs real-time serving, model serving frameworks, and latency/accuracy trade-offs
Practice Interview
Study Questions
Data Pipeline and Feature Store Considerations
Designing data collection, storage, preprocessing, and feature engineering pipelines that support model development and serving
Practice Interview
Study Questions
ML System Architecture and Design
Designing end-to-end ML systems covering data pipeline, feature engineering, model training, serving, and monitoring components
Practice Interview
Study Questions
Onsite Round 3 - SQL and Data Analysis
What to Expect
This technical round evaluates your SQL proficiency and ability to extract insights from data. You'll write SQL queries to answer business questions, perform exploratory data analysis, and demonstrate understanding of data at scale. Problems typically involve joins, aggregations, window functions, and require thinking about query efficiency. This is practical skills assessment - can you extract what you need from production databases? You may also discuss approach to analyzing experimental results or debugging data issues. For junior level, focus on writing correct, readable queries and explaining your analysis approach clearly.
Tips & Advice
Write SQL queries that are readable first - use clear table aliases, meaningful column names, and comments. Start with a simple query and optimize if needed. Explain your approach before writing code: what tables do you need, what joins, what aggregations? Test your logic with examples. For complex problems, break them into steps: first get the subset of data you need, then aggregate, then order. Be comfortable with window functions, subqueries, and CTEs (Common Table Expressions) - these are common in real work. When analyzing results, calculate metrics properly and think about edge cases (null values, zero divisions, etc.). If asked about data quality issues, think systematically: missing data, duplicates, outliers, schema mismatches. Practice on platforms like LeetCode SQL section or StrataScratch. Focus on medium-difficulty problems that require multiple joins and aggregations.
Focus Topics
Exploratory Data Analysis and Insight Generation
Using SQL to understand data distributions, identify anomalies, validate hypotheses, and generate insights for research
Practice Interview
Study Questions
Query Optimization and Performance
Understanding indexing, query execution plans, and writing queries that run efficiently on large datasets
Practice Interview
Study Questions
SQL Fundamentals and Query Writing
Writing correct, efficient SQL queries using SELECT, WHERE, JOIN, GROUP BY, ORDER BY, and HAVING clauses
Practice Interview
Study Questions
Joins and Complex Query Logic
Understanding different join types (INNER, LEFT, RIGHT, FULL), subqueries, and combining data from multiple tables correctly
Practice Interview
Study Questions
Aggregations and Window Functions
GROUP BY operations, window functions (ROW_NUMBER, RANK, LAG, LEAD), and calculating running totals or rankings
Practice Interview
Study Questions
Onsite Round 4 - Behavioral and Leadership
What to Expect
This final round assesses cultural fit, collaboration style, and how you approach challenges beyond technical execution. The interviewer will ask behavioral questions focused on your past experiences, how you've handled conflicts, learned from failures, collaborated with others, and contributed to team success. At the junior level, the focus is on demonstrating ability to work independently with occasional guidance, positive attitude toward learning, collaboration with team members, and ownership of assigned work. You'll discuss your past projects, challenges you overcame, and how you interact with colleagues. The interviewer is assessing whether you'll fit the team and grow into a strong contributor.
Tips & Advice
Prepare 3-4 concrete project stories from your experience and use them to answer multiple behavioral questions. Use the STAR method: Situation (context), Task (what you were responsible for), Action (what you did), Result (outcome and lessons). Be specific and quantify results when possible. For junior level, it's appropriate to discuss scenarios where you sought help or learned from senior colleagues - this shows good judgment. Talk about technical challenges you solved and collaboration experiences. Discuss how you handle feedback and disagreement - show you're coachable. Prepare for questions about motivations, growth areas, and what success looks like for you. Research the company's values and weave them into your stories naturally. Practice with a friend or mentor. Be authentic - interviewers can tell when you're reciting prepared answers vs. having genuine conversations.
Focus Topics
Handling Ambiguity and Setbacks
Navigating unclear requirements, dealing with failed experiments, and adapting when plans change
Practice Interview
Study Questions
Communication and Stakeholder Engagement
Explaining technical work to non-technical stakeholders, presenting findings, and gaining buy-in for research directions
Practice Interview
Study Questions
Project Ownership and Problem-Solving
Demonstrating ability to independently own parts of projects, overcome technical and non-technical challenges, and drive toward completion
Practice Interview
Study Questions
Collaboration and Teamwork
Working effectively with team members, communicating progress, integrating feedback, and contributing to team success beyond your individual work
Practice Interview
Study Questions
Learning Ability and Growth Mindset
Examples of learning new skills, adapting to new challenges, seeking feedback, and growing from failures or mistakes
Practice Interview
Study Questions
Frequently Asked Applied Scientist Interview Questions
Your spot instance training jobs are frequently interrupted, and rerunning from scratch is too expensive. How would you design checkpointing and restart behavior so that recovery is fast, state is consistent, and the training run remains reproducible?
Sample Answer
Approach
I would make checkpointing a consistent snapshot of everything needed to resume training, not just model weights. That means saving model parameters, optimizer state, learning-rate scheduler state, random number generator seeds, current epoch and batch cursor, and any data sampler state.
Design
- Write checkpoints atomically: first write to temporary storage, validate checksums, then publish a manifest pointer.
- Keep checkpoints incremental when possible, but always make the latest one self-contained for fast restart.
- Trigger an immediate checkpoint on spot interruption notice, then resume from the last good manifest.
- Store code version, config, and data version so the run is reproducible.
Worked example
If I checkpoint every 10 minutes and a preemption happens at minute 37, the restart only redoes at most 7 minutes of work, not the whole job.
Why this works
The manifest guarantees consistency, the RNG and sampler state keep replay deterministic, and the atomic publish prevents half-written checkpoints from being used.
You have only 1,000 labeled examples for a 10-class classification problem. Describe a pragmatic model-selection process given how little data you have to both train and validate on.
Sample Answer
Direct answer
With only 1,000 labeled examples for 10 classes (about 100 per class on average, likely less for some), favor simpler, lower-variance models, use nested or repeated cross-validation rather than a single held-out split (since a single split leaves too little data for reliable validation), and consider transfer learning or a pre-trained representation if any exists for your domain rather than training a model from scratch.
Structured elaboration
At this scale, a complex model (a deep network trained from scratch) is very likely to overfit; simpler model families (regularized logistic regression, a shallow tree ensemble) or a pre-trained feature extractor (if your data is images, text, or another domain with strong existing pre-trained models) fine-tuned lightly are more reliable choices. Validation strategy matters more than usual here: a single 80/20 split leaves only ~200 examples for validation across 10 classes (~20 per class), which is barely enough to distinguish real improvements from noise; repeated k-fold CV (or even repeated random splits, averaged) gives a much more stable estimate of true performance from the same limited data. Stratification (ensuring each fold/split has proportional representation of all 10 classes) is essential given how few examples per class you have, an unstratified split risks a fold with zero examples of a rare class.
Worked example
Rather than train a CNN from scratch on 1,000 labeled images across 10 classes (likely to overfit badly), extract features from a pre-trained model (fine-tuning only a small classification head, or even just the head with the backbone frozen) and combine that with heavily-regularized logistic regression on top, validated with 5-times-repeated stratified 5-fold CV to get a stable performance estimate from the limited data.
Trade-offs & pitfalls
Repeated CV costs more compute (many more fits than a single split) but that's rarely the binding constraint at this small a data scale; the bigger risk is over-interpreting small differences between models when your validation estimate itself, however carefully computed, still carries real uncertainty from having so few labeled examples to work with.
Finance wants a month-to-date revenue trend by product from a daily sales fact table, but some product-day combinations are missing because there were no sales. The report still needs to show zero-revenue days and reset correctly at each month boundary. How would you structure the query and what reference data, if any, would you need?
Sample Answer
Approach
I would build a date spine, which is a table of every calendar date in the reporting range, then cross join it to the product list. That guarantees zero-sales days are present. After that, I would left join the daily sales fact and use a running sum partitioned by product and month.
WITH spine AS (
SELECT d::date AS dt
FROM generate_series(date '2025-01-01', date '2025-01-31', interval '1 day') AS g(d)
), daily AS (
SELECT product_id, sale_date, SUM(revenue) AS revenue
FROM sales_fact
GROUP BY product_id, sale_date
), base AS (
SELECT p.product_id, s.dt
FROM products p
CROSS JOIN spine s
)
SELECT
b.product_id,
b.dt,
COALESCE(d.revenue, 0) AS revenue,
SUM(COALESCE(d.revenue, 0)) OVER (
PARTITION BY b.product_id, DATE_TRUNC('month', b.dt)
ORDER BY b.dt
) AS mtd_revenue
FROM base b
LEFT JOIN daily d
ON d.product_id = b.product_id
AND d.sale_date = b.dt;
Reference data needed
- A product dimension.
- A calendar or date spine table, or a generated date series.
Why it works
Partitioning by month resets the running total at each boundary, and the spine ensures missing product-day combinations still show as zero.
Tell me about an experiment or attempt of yours that did not work out. How long did you keep at it before deciding, how did you make that call, and what did you do with what you had learned by then?
Sample Answer
Direct answer
I ran a six-week test of a new onboarding email sequence, hypothesizing that adding a short personalized video would raise activation, and by week four the data was inconclusive rather than clearly negative, which is the harder call: deciding whether to keep running for a real signal or stop because the result had stopped being informative. I stopped at week five, explained the decision and the reasoning to the two stakeholders who had sunk real time into producing the videos, and made sure what we'd learned about the underlying segment behavior carried into the next attempt instead of being lost with the failed one.
The hypothesis, design, and timeline
The hypothesis was that a short, personalized video early in onboarding would raise activation among users who had signed up but not completed setup, based on a pattern we'd seen in a smaller pilot. I designed a six-week A/B test with a defined minimum sample size calculated up front, specifically so I wouldn't be tempted to call it early or late based on how the numbers happened to be trending on a given day.
How I made the stop-or-continue call
By week four, the treatment group's activation rate wasn't meaningfully different from control, but the sample was also smaller than planned because a tracking issue had silently dropped a portion of the treatment group's data for the first ten days, which meant the result was underpowered (we didn't have enough clean data left to trust a negative result either way, not that the result was actually bad), not simply negative. I spent part of week four determining whether that was an environmental problem, the tracking gap, rather than a genuine sign the video didn't work. Extending the test to compensate was one option; I decided against it, because even a clean extension wouldn't have told us anything about the actual hypothesis with confidence by a reasonable date, and continuing mainly to avoid calling it a failure would have been the wrong reason to keep going.
What I did with what I'd learned
I stopped at week five and told the two people who had built the videos directly: the specific reason, an underpowered and contaminated dataset rather than a clear negative result, and that the honest conclusion was "inconclusive," not "the idea doesn't work." Rather than letting the attempt just end there, I salvaged what was usable: the clean portion of the data still showed a real behavioral pattern in how users engaged with onboarding content at all, which fed directly into redesigning the next attempt's tracking and targeting before we tried a similar idea again.
Trade-offs and pitfalls
The trade-off in a stop-or-continue call like this is sunk cost against real signal: the video work represented real time from real people, and there's pressure to keep going just to justify that investment rather than to actually learn something. The pitfall I watch for is treating "inconclusive" and "failed" as the same thing when explaining the decision, since conflating them either overstates how wrong the idea was or understates how little the test actually proved either way.
Given orders(order_id, order_date, channel, total_amount), write a query producing monthly revenue with separate online_revenue and in_store_revenue columns, using conditional aggregation (SUM with CASE WHEN).
Sample Answer
Conditional aggregation lets one query produce several columns, each summing only the rows matching a specific condition, instead of running separate queries per category and stitching the results together.
Structured elaboration
SELECT strftime('%Y-%m', order_date) AS month,
SUM(CASE WHEN channel = 'online' THEN total_amount ELSE 0 END) AS online_revenue,
SUM(CASE WHEN channel = 'in_store' THEN total_amount ELSE 0 END) AS in_store_revenue
FROM orders
GROUP BY month;
Each SUM has its own CASE expression that zeroes out rows not matching that column's condition, so every row contributes to exactly one of the conditional sums (or none, if some third channel value exists and isn't covered) while all rows still participate in the single GROUP BY pass. This is meaningfully more efficient than running two separate queries (one filtered to online, one to in_store) and joining the results, and it guarantees the two numbers come from a single consistent snapshot of the data.
Worked example
Given orders on 2025-01-05 (online, 100), 2025-01-10 (in_store, 50), 2025-02-01 (online, 30): January shows online_revenue 100, in_store_revenue 50; February shows online_revenue 30, in_store_revenue 0 (correctly showing zero rather than NULL, since the ELSE 0 branch contributes to every row's sum even when the CASE condition is never true for that month).
Trade-offs and pitfalls
Using ELSE 0 (not leaving it implicit/NULL) is what guarantees a clean zero instead of NULL for months with no matching rows in one category; if ELSE is omitted, CASE implicitly returns NULL for non-matching rows, and SUM correctly still gives 0 in that case too (since SUM ignores NULLs), but it's clearer and safer to be explicit rather than rely on that behavior. This same technique is the foundation for building a full pivot/cross-tab with many columns, covered in a following question.
Legal or compliance flags that something you're about to ship may violate a regulation in a key market and asks for a freeze, but the business wants to proceed. How do you work through that?
Sample Answer
Direct answer
When legal or compliance flags a possible regulatory problem on something about to ship, that flag is new information, not an attack on the project. The first move is to separate the specific risk from the whole feature: find out exactly what triggers the concern, then look for a way to ship everything outside that blast radius (the specific data, users, or markets the flagged concern actually touches) while the risky piece gets handled properly. Treating the flag as either a full block to fight or a formality to route around are both weak answers; the senior move is to make the freeze as small as the actual risk.
Structured elaboration
1. Turn the flag into a scoped, written finding
Ask for the specific clause or regulation, the specific data flow or behavior it applies to, and which markets or user segments are affected. A flag that sounds like 'this violates a regulation' often narrows down to 'this one data field, in these two markets.' Until that scoping happens, nobody can reason about mitigation, they can only argue about the abstract freeze.
2. Sort what's actually blocked from what's just slow
Once scoped, most flags fall into three buckets: genuinely unsafe to ship anywhere (rare, but real, treat it as a hard stop); unsafe in specific markets or for specific data (the common case, often scoped out with a flag or market-level rule); or unsafe as currently designed but fixable with a smaller change than a full freeze (needs a scoped rework, not a blanket delay).
3. Bring a mitigation, not just a constraint
Offer a concrete option: disable the flagged behavior for the affected markets, gate it behind a feature flag (a toggle that turns a piece of functionality on or off without a new deployment), or ship a version that omits the specific data flow while the rest proceeds. This turns the conversation from 'can we go or not' into 'does this mitigation satisfy the concern,' which moves much faster.
4. Get joint, written sign-off before proceeding
Both the business owner and compliance need to agree in writing on what shipped, what did not, the remaining risk, and who owns closing it. This protects everyone if the interpretation is questioned later and prevents the same argument from recurring next release.
5. If a real freeze can't be avoided, negotiate the timeline explicitly
Sometimes there is no safe scoped path and the freeze has to hold for the affected piece. Here the negotiation shifts to: what's the minimum change needed to clear the concern, who is assigned to it, and can the review be fast-tracked with a dedicated reviewer instead of sitting in a general queue. A freeze with a committed, shrinking timeline is a very different conversation from an open-ended one.
Worked example
A team is about to ship a feature that logs a new field for product analytics, and legal flags that collecting that field may violate a data-protection rule in one region. Scoping the flag shows the issue is narrow: one field, one region. Instead of freezing the whole release, the team ships everywhere else immediately, and for the flagged region ships the same feature with that one field's collection disabled behind a config switch. Legal signs off on the scoped version in writing. The team opens a follow-up item, with an owner and a target date, to redesign how that field is collected (for example, aggregating it instead of storing it per user), so the region isn't stuck without the feature indefinitely.
Trade-offs and pitfalls
- Treating every compliance flag as either a full block or a nuisance to route around is the most common mistake here; both extremes erode trust with the compliance function over time.
- Scoped mitigations (flags, market gating, field exclusions) are good short-term tools but can quietly become permanent if nobody owns the follow-up fix. The sign-off should name an owner and a date, not just describe a workaround.
- Escalating past compliance to force a ship date, without addressing the underlying concern, tends to resurface later as a bigger problem: a real violation or a regulator inquiry. Speed gained by skipping the process rarely survives contact with the risk it was protecting against.
- The strongest signal of seniority isn't how fast the team got to yes, it's whether the final decision is something both sides would still defend the same way months later.
What is the difference between event time and processing time in a streaming system? Give a concrete example where using processing time would produce an incorrect feature value, and explain how you would correct for it.
Sample Answer
Direct answer: Event time is when something actually happened in the real world (recorded on the event itself); processing time is when the stream-processing system happens to observe and process that event. They can diverge because of network delays, retries, batching, or a client being offline and syncing later, and computing features on processing time instead of event time can silently corrupt time-based features.
Structured elaboration: Consider ad-impression logging: a mobile client logs an impression at 2:00pm but is offline and only successfully sends the event at 2:45pm because of a poor connection. If a feature like "impressions in the last hour" is computed using processing time, that impression gets counted as if it happened at 2:45pm, which both undercounts the 2:00-2:45 window it actually belongs to and overcounts the 2:45-3:45 window it gets misattributed to. Over many such delayed events, this systematically skews any time-windowed feature, and the skew is worse for exactly the segments most likely to have delayed or bursty connectivity (e.g. users on poor networks), which can introduce a subtle bias into the feature.
Correcting for it: use the event's own recorded timestamp (event time) as the basis for windowing, not the time the processing system received it. This requires a watermarking strategy, since event-time processing means the system has to decide when it is safe to close a window despite not knowing in advance whether more delayed events for that window will still arrive. Most modern stream processors (Flink, Spark Structured Streaming) support event-time windowing natively, so the fix is largely a matter of configuring the job to key off the event's timestamp field rather than the system's ingestion clock, plus setting an appropriate allowed-lateness/watermark policy.
Trade-offs & pitfalls: Event-time processing requires trusting the timestamp on the event itself, which assumes the producing client's clock is reasonably accurate; a client with a badly skewed clock can inject events with timestamps far in the future or past, which either get dropped by sanity-check bounds or corrupt a window if not filtered. Processing time is simpler and requires no watermark logic at all, which is why naive implementations default to it, but that simplicity is exactly why it silently produces incorrect features whenever delivery is not instantaneous and uniform, which in practice it never is at scale.
Explain what a hash function is and list the properties that make a good general-purpose hash function for hash tables used in data pipelines. Cover determinism, uniform distribution, speed, low collision rate, avalanche effect, and non-adversarial guarantees. Give concrete examples of a poor hash choice and a good hash choice for string keys and why.
Sample Answer
Direct answer
A hash function is good for hash-table use if it is deterministic, distributes keys uniformly across
buckets, computes quickly, and has the avalanche property: changing even one bit of the input flips
roughly half the bits of the output, so similar-looking keys don't cluster together.
Structured elaboration
Determinism. The same key must always produce the same hash within one process/run, or the table
could never find something it already inserted. Note this is compatible with hashing being
randomized across different runs (Python salts string hashing per-process specifically to defeat
hash-flooding attacks), as long as it is stable within a run.
Uniform distribution. Across the actual key population you expect, outputs should spread evenly
over the output range. A hash function can be "random-looking" in general and still distribute one
particular workload badly, e.g. a hash that is uniform over arbitrary strings can still collide badly
if your real keys are all short numeric-looking strings that share a common prefix.
Speed. A hash function runs on every insert, lookup, and delete, so it needs to be cheap,
typically a handful of multiply/xor/shift operations, not something like SHA-256 (see below).
Avalanche effect. A well-mixed hash function should make output bits maximally sensitive to
input bits, so that keys differing by one character (e.g. "user1" vs "user2") land in unrelated
buckets rather than adjacent ones.
Cryptographic vs non-cryptographic choice. Non-cryptographic hash families (FNV, MurmurHash,
xxHash, CityHash) are built purely to satisfy the four properties above as cheaply as possible.
Cryptographic hashes (the SHA family) additionally guarantee collision resistance against an
adversary who is trying to find two inputs with the same output, and preimage resistance
(you can't work backward from the hash to the input). Those extra guarantees cost 10 to 100 times
more CPU per hash and are unnecessary for an ordinary hash table where the only failure mode you
care about is accidental clustering, not a cryptographic attacker trying to forge a signature.
The one place a cryptographic-strength hash earns its cost inside a hash table is exactly the
hash-flooding defense: SipHash (a keyed, non-cryptographic-speed but attack-resistant function) is
used specifically because an attacker who can choose keys should not be able to predict or engineer
collisions even knowing the algorithm, unless they also know the per-process secret key.
Worked example
A concrete case of "uniform in general, bad on your workload": if your keys are internal object IDs
like "obj_00001", "obj_00002", ..., "obj_00999", and your hash function only uses the last
few characters heavily (a genuinely weak, easy-to-write-by-accident hash such as
sum(ord(c) for c in s[-2:])), then IDs sharing the same last two digits collide constantly even
though the hash "looks fine" on random strings. A well-mixed hash (Python's built-in hash(),
MurmurHash, xxHash) processes the whole string and avalanches, so "obj_00001" and "obj_00002"
land in unrelated buckets.
Trade-offs and pitfalls
The most common pitfall is picking a hash function by its name or its Wikipedia reputation
("MurmurHash is a good hash") without checking it against the actual key distribution the table
will see, uniform-on-paper is not the same as uniform-on-your-data. The second common pitfall is
reaching for a cryptographic hash "to be safe," which pays a real latency cost across every
operation for a property (adversarial collision resistance under an unknown-algorithm attacker)
that a table almost never actually needs, since the algorithm is usually knowable and the real
defense against adversarial keys is a per-process secret seed (SipHash-style), not raw
cryptographic strength.
Explain what a RIGHT JOIN does, then rewrite a RIGHT JOIN query as an equivalent LEFT JOIN by swapping the table order. Why do many teams avoid RIGHT JOIN in their codebase even though it's standard SQL?
Sample Answer
Direct answer. A RIGHT JOIN keeps every row from the right-hand table and NULL-pads the left side where there's no match, the exact mirror of a LEFT JOIN; swapping which table is written first and changing RIGHT JOIN to LEFT JOIN produces an identical result, which is why many teams simply avoid RIGHT JOIN and standardize on LEFT JOIN for everything.
Structured elaboration. Because RIGHT JOIN and LEFT JOIN are true mirrors of each other, any RIGHT JOIN can be rewritten by physically swapping the two table references and changing the keyword, with no change in meaning at all. Teams that ban RIGHT JOIN do it for pure readability and consistency: a codebase where every join is either INNER or LEFT is easier to scan (a reader always knows "the important, must-appear table comes first"), and RIGHT JOIN's rarity means readers spend an extra beat parsing it every time it shows up, even though it's perfectly valid SQL.
Worked example. employees(employee_id, name, dept_id): (1, 'Alice', 10). departments(dept_id, dept_name): (10, 'Eng'), (20, 'Sales').
-- RIGHT JOIN: keep every department, even ones with no employees
SELECT e.name, d.dept_name
FROM employees e RIGHT JOIN departments d ON e.dept_id = d.dept_id
ORDER BY d.dept_id;
-- returns (Alice, Eng), (NULL, Sales)
-- identical result via LEFT JOIN, with the tables swapped
SELECT e.name, d.dept_name
FROM departments d LEFT JOIN employees e ON e.dept_id = d.dept_id
ORDER BY d.dept_id;
-- returns (Alice, Eng), (NULL, Sales)
Both queries return the identical two rows: Sales has no employees, so it appears with a NULL name either way.
Trade-offs and pitfalls. There's no correctness or performance difference between the two forms, so this is purely a house-style convention, not a technical constraint; some optimizers may even normalize a RIGHT JOIN into the equivalent LEFT JOIN internally before planning it. The main practical risk with RIGHT JOIN isn't the join itself, it's that a query mixing LEFT and RIGHT JOINs across a longer chain becomes genuinely hard for a reader to trace, since they can no longer assume "the must-preserve table is always written first"; standardizing on LEFT JOIN everywhere removes that ambiguity by convention rather than by any property of the SQL itself.
Write a function remove_duplicates_preserve_order(items: List[T]) -> List[T] that removes duplicate elements while preserving original order. Implement it in Python using only built-in data structures (no third-party libs). Explain time and space complexity and why an ordinary set alone is insufficient to preserve order.
Sample Answer
Direct answer
Scan the list once, keep a hash set of values already seen for O(1) membership checks, and only append an item to the output when it is not yet in that set. This is O(n) time and O(n) space (the seen-set plus the output), and it preserves order because the output is built by appending in the same order items are encountered, never by any operation that would reorder them.
Why a set alone is insufficient
A Python set (or, in other languages, an unordered hash set) has no defined iteration order tied to insertion; converting a list to a set and back, list(set(items)), removes duplicates but does not guarantee the original order is preserved (in CPython, set iteration order is an artifact of hash values and internal table layout, not an ordering guarantee, and it can even change between runs or Python versions for certain types). The fix is to use the set only for membership testing, never as the container that determines output order; a separate list, appended to in scan order, is what actually preserves order.
Implementation
def remove_duplicates_preserve_order(items):
seen = set()
result = []
for item in items:
if item not in seen:
seen.add(item)
result.append(item)
return result
Worked example
cases = [
[1, 2, 1, 3, 2, 4],
["b", "a", "b", "a", "c"],
[3, 1, 2], # already unique, order must stay untouched
]
for items in cases:
print(f"remove_duplicates_preserve_order({items}) = {remove_duplicates_preserve_order(items)}")
Output (verified by execution):
remove_duplicates_preserve_order([1, 2, 1, 3, 2, 4]) = [1, 2, 3, 4]
remove_duplicates_preserve_order(['b', 'a', 'b', 'a', 'c']) = ['b', 'a', 'c']
remove_duplicates_preserve_order([3, 1, 2]) = [3, 1, 2]
Tracing [1, 2, 1, 3, 2, 4]: 1 is new (seen becomes {1}, result [1]), 2 is new ({1,2}, [1,2]), the second 1 is already seen and skipped, 3 is new ({1,2,3}, [1,2,3]), the second 2 is already seen and skipped, 4 is new ({1,2,3,4}, [1,2,3,4]). The final order, [1,2,3,4], is exactly the order of first appearance in the input, which is the property "preserve order" means here: it is not sorted, and it is not the reverse-of-last-occurrence, it is first-occurrence order.
Trade-offs and pitfalls
- The
list(set(items))shortcut is the single most common wrong answer to this question. It is O(n) and does remove duplicates, but it does not preserve order, and candidates who reach for it are often not aware thatsetordering is unspecified rather than merely "seems to work in my test." - This requires items to be hashable. A list of dicts or other unhashable objects breaks
item not in seenoutright (raisesTypeError: unhashable type); the fix there is to derive a hashable key from each item (for example, a canonical JSON string with sorted keys) and check/store that key instead of the raw item, while still appending the original item to the output. dict.fromkeys(items)is a cleaner idiom for this exact problem in modern Python (dicts preserve insertion order since Python 3.7, andfromkeysnaturally dedupes on the keys), andlist(dict.fromkeys(items))is equivalent in behavior and complexity to the explicit loop above; the explicit version is worth knowing because it generalizes directly to the unhashable-items case, while thedict.fromkeysshortcut does not.- Space is unavoidable at O(n) in the worst case (no duplicates at all), which is worth stating rather than implying this can be done with less than linear extra space in general; achieving less space would require either sorting the input (which changes the order, defeating the purpose) or an external, more complex structure that trades memory for something else entirely.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Applied Scientist jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs