Forecasting and Time-Series Analysis Questions
Analyzing and projecting data that moves over time. Covers trend and seasonality decomposition, forecasting approaches, demand modeling, and anomaly detection on time series. Emphasizes reasoning about baselines, drivers, and forecast reliability.
Sales leaders argue the statistical forecast underestimates next quarter. How would you handle this disagreement? Walk through your steps: validating the data, reconciling assumptions, producing side-by-side scenarios, proposing compromise approaches, and restoring trust for future forecasting cycles.
Sample Answer
Direct answer
When sales leaders dispute a statistical forecast, the right response is a structured process, not a unilateral defense of the model or a silent capitulation: validate the underlying data first, surface and reconcile the specific assumptions being disputed, produce side-by-side scenarios so both views are visible rather than argued in the abstract, propose a documented compromise where warranted, and use the resolution to rebuild trust for future cycles rather than treating it as a one-off argument to win.
Structured elaboration
- Validating the data: before engaging on the substance of the disagreement, confirm the forecast's inputs are actually correct (no stale data, no broken pipeline, no missed recent event) - a meaningful fraction of forecast disputes turn out to be legitimate data-quality catches, not a genuine model-vs-judgment disagreement, so ruling this out first is both fast and often the actual resolution.
- Reconciling assumptions: identify precisely WHERE the statistical model and the sales leaders' intuition diverge - is it a different read on a specific deal, a market condition the model has no way to see (a competitor's recent move), or genuine overconfidence on one side? Naming the specific assumption in dispute turns a vague "the number feels wrong" into something concrete enough to actually resolve.
- Producing side-by-side scenarios: rather than picking one number, present the model's forecast AND a scenario reflecting the sales leaders' adjustment, with the specific assumption difference driving the gap made explicit - this respects both inputs and lets the eventual decision-maker see the trade-off rather than a black-box disagreement.
- Proposing compromise approaches: a blended forecast (weighting the statistical model and human judgment, as in collaborative forecasting more broadly) is often defensible when both sides have genuine signal the other lacks; document the blend and the reasoning, not just the final number.
- Restoring trust for future cycles: track which side (model or human adjustment) was closer to the eventual actual, and feed that back explicitly into the NEXT cycle's process - if human overrides have systematically been closer to actuals recently, that's a real signal the model may be missing something structural; if the model has been closer, that's useful evidence too, and either way, closing the loop with actual outcomes (rather than re-litigating the same disagreement from scratch each cycle) is what actually builds durable trust in the forecasting process.
Worked example
This is the same underlying discipline as any collaborative forecasting workflow where human adjustments (sales judgment, marketing input) are layered onto a statistical base: the productive version tracks whether those adjustments, on average, have actually improved on the model's own accuracy over time, rather than treating every override as automatically valuable or automatically suspect - a defensible process measures this explicitly (comparing overridden vs. non-overridden forecast accuracy against eventual actuals) instead of relying on anecdote or seniority to settle the disagreement.
Trade-offs & pitfalls
The most damaging failure mode here is NOT reconciling anything and instead quietly picking whichever number is more politically convenient in the moment - that erodes the forecasting process's credibility either way (with the modelers if judgment always wins with no accountability, or with the business side if the model always wins with no acknowledgment of real on-the-ground signal it can't see). A durable process needs to be able to point to a track record, not just a single episode, to settle these disputes over time.
Discuss responsible AI and governance considerations specific to forecasting systems. Cover detection and mitigation of bias across regions or product lines, fairness when forecasts drive allocation decisions, data retention and privacy of training data, and what operational governance practices you would put in place to keep the system auditable and correctable over time.
Sample Answer
Direct answer
Responsible AI and governance for forecasting systems means checking whether the model treats different regions or product lines fairly (not just accurately on average), protecting the privacy of training data, and running the operational governance machinery - model cards, periodic audits, defined remediation processes - that make the whole system auditable and correctable rather than a black box that only gets scrutinized after something visibly goes wrong.
Structured elaboration
- Detecting bias across regions or product lines: check forecast accuracy and, separately, forecast BIAS (systematic over- or under-prediction) broken out BY segment (region, product line), not just in aggregate - a model that's unbiased on average can still systematically under-forecast one region while over-forecasting another, which is invisible to an aggregate accuracy check but very real in its downstream consequences.
- Mitigating detected bias: once a segment-specific bias is confirmed (not just suspected from a single period, but validated the way any bias-detection exercise should be, with a proper statistical check), remediation options range from a segment-specific recalibration correction to retraining with segment-balanced data or segment-aware features, chosen based on WHY the bias exists (a data-representation issue vs a genuine, harder-to-model difference in that segment's underlying dynamics).
- Fairness when forecasts drive allocation decisions: when a forecast feeds directly into an ALLOCATION decision (inventory, staffing, incentive dollars distributed across regions), a systematic forecasting bias against one region translates directly into that region being under-served - this is the concrete mechanism by which a "purely technical" forecasting bias becomes a real fairness issue, and it's the reason bias detection here needs to be checked specifically against WHATEVER downstream decision the forecast drives, not evaluated as an abstract accuracy statistic alone.
- Data retention and privacy: forecasting models trained on customer-level or location-level transaction data inherit the same retention-limits and privacy-handling obligations as any other system using that data - define and enforce a retention policy for training data, and ensure the model itself (and any cached intermediate features) doesn't become an unintended long-term store of data that should have been deleted under the organization's stated retention policy.
- Operational governance practices: model cards (a standardized, versioned document describing a model's intended use, known limitations, training-data characteristics, and validated performance across segments) make a model's fairness/bias characteristics legible to anyone who needs to evaluate or approve its use, rather than requiring a fresh investigation each time; periodic audits (a scheduled, recurring re-check of segment-level bias and accuracy, not just a one-time pre-launch check) catch drift into unfairness that emerges only after deployment; a defined remediation process (what happens, and who's accountable, once an audit finds a problem) ensures a detected issue actually gets fixed rather than merely documented.
Worked example
An automated demand-allocation system that systematically under-forecasts demand in lower-income neighborhoods (perhaps because those areas have historically been served by fewer marketing dollars, and the model has learned that pattern as if it were an accurate reflection of true underlying demand rather than a historical resourcing artifact) would, if left unchecked, perpetuate and even reinforce that under-service through the forecast-driven allocation decision itself - detecting this requires deliberately checking bias BY neighborhood income level (not something an aggregate accuracy metric would surface on its own), and the mitigation needs to address the root cause (the historical resourcing pattern baked into the training data) rather than just a numeric recalibration that treats the symptom.
Trade-offs & pitfalls
The most consequential governance gap is treating fairness/bias review as a one-time, pre-launch checklist item rather than an ongoing, periodically-repeated audit - a model that was checked and cleared at launch can still drift into segment-specific bias over time as the underlying data and business context evolve, and only a recurring audit cadence (not a single point-in-time check) catches that drift before it compounds into a real, sustained allocation harm.
Given weekly retail sales data, explain how you would choose between ARIMA, ETS (Holt-Winters), Prophet, and machine learning models (e.g., gradient boosting). Discuss considerations such as amount of data, seasonality complexity, external regressors, interpretability, and deployment/maintenance trade-offs.
Sample Answer
Direct answer
Choosing between ARIMA, ETS (Holt-Winters), Prophet, and gradient boosting for a forecasting problem comes down to five practical considerations: how much history you have, how complex the seasonality is, whether you have useful external regressors, how much interpretability you need, and how much ongoing maintenance the approach demands at your deployment scale.
Structured elaboration
- Amount of data: ARIMA/ETS need enough history to reliably estimate their (few) parameters and at least a couple of full seasonal cycles; they work fine on a single series with a year or more of clean history. Gradient boosting typically needs MORE data, or many related series pooled together, because it's learning a much more flexible function; it's the natural choice once you have hundreds of similar series to train a single shared model on ("global" forecasting).
- Seasonality complexity: ETS handles one clean, fixed seasonal period well. ARIMA/SARIMA can add a seasonal term but the search space grows fast with multiple seasonalities. Prophet and gradient boosting (via Fourier or calendar features) both handle MULTIPLE overlapping seasonalities (daily + weekly + yearly) more naturally.
- External regressors: ARIMA can incorporate them (SARIMAX), Prophet supports them as additional regressors, gradient boosting handles them trivially as extra feature columns - plain ETS does not support exogenous regressors at all, which rules it out the moment promotions, price, or weather matter.
- Interpretability: ARIMA and ETS have few, well-understood parameters that a statistically-literate stakeholder can reason about directly. Prophet's decomposition (trend + seasonality + holidays, each separately plottable) is unusually interpretable for a more flexible model. Gradient boosting is the least directly interpretable of the four, though feature importances and partial dependence can partially compensate.
- Deployment/maintenance trade-offs: ARIMA/SARIMA require re-identifying orders as the series' structure drifts and can be finicky to auto-fit reliably at scale across thousands of series; ETS auto-fits more reliably and cheaply; Prophet is designed to be a low-maintenance "fit it and mostly forget it" option, tolerant of missing data and outliers; gradient boosting needs a feature-engineering pipeline (lags, calendar features) to maintain but scales best to "one model, many series" architectures with heavy compute already in place.
Worked example
For weekly retail sales with clean weekly-cycle-only seasonality, a couple of years of history, no meaningful external regressors, and a small analytics team: ETS (Holt-Winters) is the pragmatic default - simple, auto-fits well, and easy to explain. If that same retailer wants to incorporate promotion calendars and multiple seasonalities (daily transactions with weekly AND yearly effects) across thousands of SKUs, gradient boosting on lag/calendar/promo features (trained once across all SKUs) or Prophet (fit per-SKU, tolerant of the sparser individual histories) becomes the more defensible choice, trading some interpretability for the ability to use exogenous signal and scale to many series with shared infrastructure.
Trade-offs & pitfalls
The single biggest mistake is picking the most sophisticated option by default: a gradient-boosted model trained on one thin, noisy series with no exogenous signal will often be beaten by a well-fit ETS model, because there's nothing for the extra flexibility to learn from. Conversely, sticking with per-series ARIMA at a scale of thousands of series usually breaks down operationally (auto-fitting SARIMA reliably at that scale is genuinely hard) well before it breaks down statistically - the deployment/maintenance dimension often decides the choice as much as raw forecast accuracy does.
A pandemic caused a sudden 70% drop in ride volume and changed weekly seasonality. Present a prioritized plan to adapt forecasting, dashboards, and stakeholder communications: rapid diagnostics to quantify the change, short-term heuristic fallbacks for operations, a retraining strategy, scenario planning for recovery paths, and how to communicate uncertainty and recommended actions to operations and executives.
Sample Answer
Direct answer
Responding to a sudden, dramatic demand shock (like a 70% drop with a changed seasonal pattern) means moving fast through rapid diagnostics, short-term operational fallbacks that don't depend on a not-yet-retrained model, a deliberate retraining strategy once the new regime is better understood, explicit scenario planning for recovery, and continuous, honest uncertainty communication throughout - in roughly that priority order.
Structured elaboration
- Rapid diagnostics to quantify the change: first establish the basic facts fast - how large is the drop, is it uniform across segments/regions or concentrated, and has the WEEKLY seasonal pattern itself genuinely changed shape (not just shifted level) - this initial read determines how radically the forecasting approach needs to change, versus a milder response being sufficient.
- Short-term heuristic fallbacks for operations: a model trained entirely on pre-shock history will actively mislead if used naively during the shock itself (extrapolating a pattern that no longer holds); a simple, transparent heuristic (e.g. "current week's actuals, or a short recent-window average, rather than the stale model's output") is often a more honest and more useful stopgap than a sophisticated model confidently extrapolating an outdated pattern.
- Retraining strategy: once enough post-shock data has accumulated to characterize the NEW pattern (not just the initial shock transient), retrain - but be deliberate about the training window: including too much pre-shock history dilutes the new pattern, too little post-shock history produces an unstable, premature re-estimate; a staged approach (retrain on a growing post-shock window as it becomes available, rather than waiting for a fixed amount before doing anything) balances this.
- Scenario planning for recovery paths: since the actual recovery trajectory (fast rebound, slow recovery, a new permanent lower baseline, a different shape than pre-shock) is genuinely uncertain during the shock itself, present multiple named scenarios rather than a single point forecast - this is exactly the same "acknowledge genuine deep uncertainty explicitly, rather than false point-forecast precision" discipline used for any severely disrupted or unprecedented forecasting situation.
- Adapting dashboards: any dashboard panel built on WoW/YoY percentage-change comparisons, or on alert thresholds calibrated against the pre-shock baseline, will either look absurdly extreme (a 70% YoY drop swamping every other number on the page) or silently mislead (thresholds firing constantly, or not firing at all, against a baseline that's no longer valid). Annotate the shock period explicitly on any historical-trend chart so viewers don't misread the discontinuity as a data error or pipeline bug; temporarily suspend or clearly flag any YoY/WoW comparison metric whose window spans the shock boundary, since the underlying comparison is no longer meaningful; and recalibrate alert thresholds against the NEW short-term-heuristic baseline rather than the stale pre-shock one, revisiting both the annotations and the thresholds again once retraining and scenario tracking are underway.
- Communicating uncertainty and recommended actions: to operations, translate directly into concrete near-term actions under the current best-guess scenario (staffing, inventory); to executives, communicate the RANGE of plausible recovery scenarios and what data signal would indicate which one is actually materializing, framed as "here's what we're watching to tell these scenarios apart" rather than false confidence in one specific path.
Worked example
In the first days of such a shock: report the magnitude and shape of the drop (rapid diagnostics), recommend operations use a short recent-window heuristic rather than the stale model's now-invalid output (short-term fallback), and present 2-3 named recovery scenarios (e.g. "sharp V-shaped rebound within weeks," "gradual multi-month recovery," "a new, lower permanent baseline") with the specific signals (weekly growth rate, whether the pre-shock weekly seasonal shape has resumed) that would distinguish which scenario is actually unfolding as more data arrives, updating the recommendation as that evidence accumulates rather than committing to one scenario prematurely.
Trade-offs & pitfalls
The most damaging mistake in a disruption like this is continuing to trust (or silently keep serving) a model's output as if nothing had changed, simply because retraining hasn't happened yet - an explicit, fast decision to fall back to a transparent heuristic during the gap between "the old model is now known to be wrong" and "a properly retrained model is ready" is what protects operational decisions during exactly the period they're most vulnerable.
Explain the difference between forecasting and causal inference in a business context. Give examples of when each approach is appropriate (for example, demand planning versus measuring marketing lift), describe core assumptions required for causal claims, and explain how a BI analyst should decide which approach to use when stakeholders ask for 'what will happen' versus 'what caused it'.
Sample Answer
Direct answer
Forecasting predicts WHAT WILL HAPPEN by extrapolating patterns, while causal inference asks WHAT CAUSED an observed change and requires a different set of assumptions (a valid comparison group, or an experiment) to answer credibly; demand planning is a forecasting problem, while measuring the lift from a marketing campaign is a causal-inference problem, and conflating the two - using a forecasting model's output as if it answered a causal question - is a common and consequential mistake.
Structured elaboration
- Forecasting: uses historical patterns (trend, seasonality, correlated signals) to project forward, with NO requirement that the model understand mechanism or causation - a model that has never heard of "why" can still forecast well if the pattern is stable. Appropriate when the question is purely "what's likely to happen if nothing unusual intervenes."
- Causal inference: asks what would have happened in a counterfactual world WITHOUT a specific intervention, which forecasting alone cannot answer, because a plain extrapolation of history has no way to represent "the same period, but without the campaign." Appropriate when the question is "did X cause Y," and answering it credibly requires either a genuine randomized experiment, or a quasi-experimental design (difference-in-differences, synthetic control, regression discontinuity) with its own specific, checkable assumptions.
- Core assumptions required for causal claims: at minimum, some form of comparability between what happened and what would have happened otherwise - either through randomization (an A/B test, the gold standard) or through an argued similarity between a treated group and a comparison group that wasn't exposed to the intervention (with assumptions like "parallel trends" for difference-in-differences, or "no other simultaneous change" for a simple before/after comparison), each of which can fail in ways that are not always obvious from the data alone.
- Deciding which approach to use: when a stakeholder asks "what will next quarter look like" - forecasting. When a stakeholder asks "did our marketing campaign work" or "would revenue have grown anyway" - causal inference, and a pure forecast (even an accurate one) does not answer that question, because an accurate forecast of the ACTUAL future doesn't tell you what the counterfactual (no-campaign) future would have looked like.
Worked example
A revenue series shows +15% growth the same quarter a new marketing campaign launched. A forecasting model extrapolating the pre-campaign trend might have predicted +8% growth anyway from ordinary momentum and seasonality - the DIFFERENCE between the forecast's counterfactual-like extrapolation and the observed +15% (roughly +7 percentage points) is a rough causal-impact ESTIMATE, but only a credible one if the forecasting model's pre-period fit was genuinely reliable and nothing else changed at the same time; a rigorous causal-impact method (like BSTS-based causal impact analysis, or a proper control-group comparison) formalizes exactly this idea with an explicit counterfactual and uncertainty interval, rather than treating a forecast's extrapolation as automatically causal.
Trade-offs & pitfalls
The most consequential version of this mistake is presenting a forecast-vs-actual gap as if it were automatically a measured causal effect, without checking whether anything ELSE changed at the same time that could explain the gap - seasonality not fully captured by the forecasting model, a competitor's simultaneous move, or a broader macro shift are all alternative explanations a rigorous causal analysis has to rule out (or explicitly account for) that a bare forecast comparison does not.
Unlock Full Question Bank
Get access to all 20 Forecasting and Time-Series Analysis interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.