Forecasting and Time-Series Analysis Questions
Analyzing and projecting data that moves over time. Covers trend and seasonality decomposition, forecasting approaches, demand modeling, and anomaly detection on time series. Emphasizes reasoning about baselines, drivers, and forecast reliability.
Architect a low-latency forecasting service that must serve per-zone 15-minute demand predictions on request for 10k zones with 500 QPS and 100ms p95 latency. Describe overall architecture (model store, feature store, online features, caching), how you would serve batched predictions, and strategies for availability and rollback.
Sample Answer
Direct answer
A low-latency (100ms p95, 500 QPS, 10k zones) forecasting service needs pre-computed or cached predictions wherever possible (rather than computing a fresh forecast per request), an online feature store fast enough to serve the remaining real-time inputs within the latency budget, a model store supporting fast model loading/versioning, and explicit availability/rollback strategies given the operational stakes of a customer-facing latency SLA.
Structured elaboration
- Model store: hosts versioned, trained models ready for fast loading into the serving layer; at 10k zones, likely one (or a small number of clustered) GLOBAL models rather than 10k independently-loaded per-zone models, both for the same statistical pooling reasons that favor global models at scale and for the practical serving reason that loading 10k separate model artifacts per request path is operationally unworkable within a 100ms budget.
- Feature store, online features: the features needed at inference time (recent lag values, calendar features, any real-time exogenous signal) must be servable with very low latency - a dedicated online feature store (backed by a fast key-value store, not a query against a data warehouse) is standard here, since a warehouse query would blow the 100ms budget on its own.
- Caching: for a request pattern where many requests ask for the SAME zone's forecast within a short window (a very plausible pattern at 500 QPS across "only" 10k zones), caching a recently-computed forecast (invalidated on a schedule matching how often the underlying forecast is genuinely refreshed, e.g. every 15 minutes) avoids recomputing an identical answer repeatedly, and is often the single biggest lever for hitting a tight latency SLA without needing every request to hit the full model-inference path.
- Serving batched predictions: even for on-request serving, batching multiple zones' feature vectors into one model inference call (rather than one model call per zone) amortizes fixed per-call overhead and is usually meaningfully faster in aggregate than fully independent per-request inference, especially for a model architecture (e.g. a global gradient-boosted or neural model) that scales well with batch size.
- Availability and rollback: run multiple redundant serving instances behind a load balancer so a single instance failure doesn't violate the latency SLA for other requests; deploy new model versions via a canary (a small slice of traffic first) with automated rollback if the new version's latency or accuracy regresses, rather than a full instantaneous cutover.
Worked example
A concrete architecture: features refreshed into a fast online store on a schedule (e.g. every few minutes) rather than computed synchronously per-request; a background job pre-computes and caches each zone's forecast at the SAME refresh cadence; the actual request-serving path becomes largely a cache lookup (very fast, easily within 100ms even at 500 QPS) with a fallback to synchronous on-demand inference only for a cache miss (a zone whose forecast hasn't been pre-computed recently, ideally a rare case) - this architecture trades a small amount of freshness (forecasts are as fresh as the last refresh cycle, not literally real-time) for a much more reliable and cheaper latency profile than computing everything synchronously per request.
Trade-offs & pitfalls
The most consequential design trade-off here is freshness vs latency/cost: pure synchronous per-request inference gives the freshest possible forecast but is the most expensive and riskiest way to hit a tight SLA at this request volume; a pre-compute-and-cache architecture is cheaper and more reliably fast but serves a slightly staler forecast - the right balance depends on how much genuine value fresher-than-cache-refresh-cycle forecasts actually provide for THIS specific use case (per-zone 15-minute demand predictions), which is worth validating rather than assuming synchronous freshness is automatically worth its cost.
Explain quantile regression forests and conformal prediction as two approaches to uncertainty quantification for forecasts. Discuss strengths and weaknesses and how you would check calibration of the resulting prediction intervals in a production forecasting system.
Sample Answer
Direct answer
Quantile regression forests and conformal prediction are both distribution-free (no Gaussian assumption) ways to quantify forecast uncertainty: quantile regression forests directly predict specific quantiles of the outcome from the input features (naturally capturing heteroscedasticity), while conformal prediction wraps ANY underlying point-forecast model with a calibration step that produces intervals with a formal, finite-sample coverage guarantee - different mechanisms, complementary strengths, both needing empirical calibration checks before trusting them in production.
Structured elaboration
- Quantile regression forests (QRF): an extension of random forests that, instead of averaging leaf predictions to a single point estimate, retains the full distribution of training-target values that land in each leaf, letting you read off any desired quantile directly from that empirical leaf-level distribution - naturally captures heteroscedasticity (different inputs can produce different-width predicted distributions) without any explicit distributional assumption, at the cost of the same general tree-ensemble weaknesses (can struggle to extrapolate cleanly beyond the range of training data).
- Conformal prediction: a more general wrapper technique - given ANY underlying point-forecast model, hold out a calibration set, measure the model's actual errors (or a chosen nonconformity score) on that set, and use the empirical distribution of those errors to construct an interval around any new prediction with a formal guarantee (under an exchangeability assumption) that the true coverage rate matches the target, in finite samples, without needing the underlying model to be probabilistic at all.
- Strengths and weaknesses: QRF gives input-conditional interval width naturally (wider for genuinely more uncertain inputs) as a direct byproduct of the tree structure, but is tied to a tree-ensemble model family specifically; conformal prediction is model-agnostic (wrap literally any point-forecast model, including a deep-learning one) and has the strongest FORMAL coverage guarantee of the methods discussed in this space, but its exchangeability assumption is a real one to check for time series specifically - data isn't strictly exchangeable when there's genuine temporal structure (a value's error isn't independent of nearby-in-time errors the way conformal prediction's classic guarantee assumes), so time-series-specific conformal variants (adapting the calibration to respect temporal order, e.g. via a rolling/adaptive calibration window) are needed rather than the vanilla i.i.d.-assuming version.
- Checking calibration in production: for either method, the practical check is the same - track EMPIRICAL coverage on genuinely new, out-of-sample data over time (does the stated 90% interval actually contain the truth ~90% of the time, tracked on a rolling basis) rather than trusting either method's theoretical guarantee blindly, since QRF's quality depends on how well the forest itself fits, and conformal prediction's formal guarantee depends on an exchangeability assumption that a shifting, non-stationary business time series can genuinely violate.
Worked example
For a production forecasting system already using a gradient-boosted point-forecast model, wrapping it with a TIME-SERIES-ADAPTED conformal prediction layer (using a rolling recent calibration window rather than a fixed historical one, to respect that the error distribution itself may drift over time) is often more practical than switching the underlying model to a QRF, since it adds calibrated intervals to a model you've already validated and deployed, without needing to retrain or replace the point-forecast model itself.
Trade-offs & pitfalls
The most consequential mistake with conformal prediction specifically is applying the standard (i.i.d.-exchangeability-assuming) version naively to time-series data without checking whether that assumption is reasonable - a rolling or adaptive calibration window, re-estimated as new data arrives, is the standard fix, and skipping it risks a formal-sounding "guarantee" that doesn't actually hold for your genuinely non-stationary series.
Design a probabilistic deep-learning approach for multi-step demand forecasting (e.g., 1 to 24 steps ahead). Specify model architecture (seq2seq, Transformer, or DeepAR), loss functions for probabilistic outputs, how to generate quantile or full predictive distributions, and which probabilistic metrics (e.g., CRPS) you would report.
Sample Answer
Direct answer
A probabilistic deep-learning approach for multi-step demand forecasting (seq2seq, Transformer, or DeepAR) produces a full predictive DISTRIBUTION at each future step rather than a single point value, trained with a distribution-appropriate loss (negative log-likelihood of an assumed distribution, or a quantile/pinball loss for a quantile-based output), and evaluated with distribution-aware metrics like CRPS rather than plain point-forecast error metrics.
Structured elaboration
- Model architecture choices: seq2seq (an encoder processing history, a decoder generating the multi-step forecast) is a flexible general framework; DeepAR specifically models each future step's output as parameters of a chosen probability distribution (e.g. a Negative Binomial for count data) conditioned on a learned recurrent state, trained across MANY related series jointly (a global model, which helps sparse individual series borrow strength from similar ones); a Transformer-based architecture can be adapted similarly by having its output heads predict distribution parameters or quantiles rather than a single point per step.
- Loss functions for probabilistic outputs: negative log-likelihood of the assumed output distribution (DeepAR's native approach - the model literally predicts distribution parameters like a Negative Binomial's mean and dispersion, and the loss is how well the actual outcome fits that distribution); pinball/quantile loss (train multiple output heads, one per target quantile, as in the GBM-quantile approach but inside a neural architecture) is an alternative that makes no parametric distributional assumption at all.
- Generating quantile or full predictive distributions: a parametric approach (DeepAR-style) gives a full distribution directly from its predicted parameters, from which any quantile can be derived analytically or by sampling; a quantile-head approach directly outputs the specific quantiles you trained it on, with the crossing/monotonicity caveat that applies to any independently-trained quantile heads (worth an explicit check, the same one relevant to GBM quantile forecasting).
- Probabilistic metrics to report: CRPS is the standard single-number summary of full-distribution forecast quality (generalizes MAE to a full distribution); pinball loss per quantile if using a quantile-head approach; and empirical interval coverage (does the stated interval actually contain the truth at the stated rate) as the calibration check that complements whichever sharpness-oriented metric (CRPS/pinball) you're using.
Worked example
For 1-to-24-step-ahead demand forecasting across many related zones, a DeepAR-style model trained jointly across all zones (global model) with a Negative Binomial output distribution (appropriate for overdispersed count demand) directly gives, for each future step and each zone, both a point estimate AND the full predictive distribution needed to compute a safety-stock-style quantile or a CRPS-based accuracy check - the joint training across zones is specifically what lets sparse/low-volume zones borrow statistical strength from the more data-rich ones, the same pooling logic that helps any cold-start or thin-history forecasting problem.
Trade-offs & pitfalls
A parametric approach (DeepAR-style) is more data-efficient (fewer independent things to estimate per step) but only as good as its distributional assumption - a mismatched distribution family (assuming Negative Binomial when the true generating process is meaningfully different) will produce a systematically miscalibrated predictive distribution even with a well-trained model, so validate the assumed distribution against the data's actual shape before committing to it, and always run the empirical-coverage calibration check on a genuine holdout rather than trusting the model's own stated confidence.
You need to forecast weekly demand for a newly launched product with only 3 months of sales history. Describe modeling strategies to generate forecasts and quantify uncertainty: hierarchical Bayesian pooling across similar products, transfer learning from related SKUs, using external regressors (search trends, category sales), and scenario-based simulation for inventory planning. Explain advantages, assumptions, and validation strategies for each.
Sample Answer
Direct answer
With only 3 months of history, a single-series model is unreliable, so the practical strategies all borrow strength from OUTSIDE the target series itself: hierarchical Bayesian pooling across similar products, transfer learning from related SKUs, using external/leading-indicator regressors, and scenario-based simulation to bound the range of plausible outcomes for planning.
Structured elaboration
- Hierarchical Bayesian pooling: treat the new product's parameters (e.g. its baseline demand level, seasonal shape) as drawn from a distribution shared with similar products, rather than estimated from scratch. Early on, the model leans heavily on the shared "prior" (what similar products typically do); as the new product accumulates its own data, the posterior shifts toward its own observed pattern. This partial-pooling behavior is exactly what you want: it prevents a noisy 3-month history from producing wild extrapolations.
- Transfer learning from related SKUs: fit (or fine-tune) a model on a broader set of comparable products and apply it to the new one, either directly (a "global" model trained across many series with the new SKU's ID/features as inputs) or as a warm start that gets fine-tuned once enough of the new product's own data arrives.
- External regressors: features that don't depend on the new product's own sales history - category-level sales trend, related search-interest trends, or a comparable "analogue" product's early trajectory - give the model signal in the earliest weeks when the target's own history is too short to be informative.
- Scenario-based simulation: rather than a single point forecast, run the demand model under a small set of named scenarios (e.g. "tracks the median of the 3 closest analogue launches", "tracks the top-quartile analogue", "tracks the bottom-quartile analogue") to give inventory/staffing planners an explicit range rather than false point-forecast precision.
Worked example
Say a new product looks most similar to 4 analogue products launched in the past two years. A hierarchical Bayesian approach would fit a shared model across those 4 analogues (and other same-category products) to get a prior distribution on early-life weekly demand shape (e.g. a typical launch-week spike followed by decay to a steady state around week 6), then update that prior with the new product's actual 12 weeks of data. If the new product's first 3 weeks are running 20% above the analogue median, the posterior shifts upward, but by less than a naive "just extrapolate the last 3 weeks" approach would, because the model is also weighting the shared prior - exactly the shrinkage behavior that keeps a 3-month-old series from producing an overconfident forecast.
Trade-offs & pitfalls
The main assumption every one of these methods leans on is that "similar" products really are similar in the way that matters (comparable category, price point, launch channel) - a poorly chosen analogue set will systematically bias the pooled prior. Validate by holding out a recent, already-mature product as if it were new (use only its first 3 months) and checking whether the pooled/transfer approach would have predicted its actual mature-demand trajectory; if it consistently over- or under-shoots on that backtest, the analogue set needs revisiting before trusting it on the live new product.
Explain the ARIMA model components: AR(p), I(d), MA(q). For each component give intuition about what it captures, how you would identify appropriate orders using ACF/PACF and stationarity tests, and when to include seasonal terms (SARIMA).
Sample Answer
Direct answer
ARIMA(p,d,q) models a series as a combination of three pieces: AR(p), autoregression on the series' own past p values; I(d), the number of times you difference the series to make it stationary; and MA(q), a moving average of the past q forecast errors. You identify p and q by reading the ACF and PACF plots of the (differenced) series, and you add seasonal terms (SARIMA) when the ACF/PACF still show a repeating spike pattern at the seasonal lag after ordinary differencing.
Structured elaboration
- AR(p): today's value is a linear function of the last p values plus noise. A PACF that cuts off sharply after lag p (with earlier lags significant) points to an AR(p) term, because the PACF isolates the direct effect of each lag after removing the effect of the lags in between.
- I(d): the number of times you difference yt (i.e. work with yt−yt−1, or the second difference) before the series is stationary. You determine d with a stationarity test (ADF/KPSS) rather than by eye: difference until the test says stationary, and stop as soon as it does. Over-differencing (d too large) introduces artificial negative autocorrelation and inflates forecast variance.
- MA(q): today's value depends on the last q forecast errors, not raw values. An ACF that cuts off sharply after lag q (while the PACF decays slowly) points to an MA(q) term.
- Order identification in practice: plot ACF and PACF of the differenced series. AR signature = PACF cuts off, ACF tails off. MA signature = ACF cuts off, PACF tails off. Mixed ARMA signatures (both tail off) are common in practice, so analysts usually also compare a small grid of candidate (p,d,q) by AIC/BIC rather than trusting the plots alone.
- When to add seasonal terms (SARIMA): if, after taking the ordinary difference, the ACF/PACF still show significant spikes at the seasonal period (e.g. lag 7 for daily-with-weekly-seasonality data, lag 12 for monthly-with-yearly-seasonality), a plain ARIMA hasn't captured the seasonal structure. SARIMA adds seasonal AR/MA/differencing terms (P,D,Q)s operating at multiples of the seasonal period s, on top of the ordinary (p,d,q) terms.
Worked example
Take daily retail sales with a weekly pattern. Fitting an ARIMA(1,1,1) leaves a residual ACF with a clear spike at lag 7 (and 14, 21) - the model has removed the trend (via d=1) but not the weekly repetition. Adding a seasonal term, e.g. SARIMA(1,1,1)(1,0,1)7, lets the seasonal AR/MA terms absorb that lag-7 structure; after refitting, the residual ACF should show no significant spikes at multiples of 7. Concretely: raising d from 0 to 1 changes what "capturing p and q" even means, since AR(p)/MA(q) are now fit on the differenced series, not the raw one - a common early mistake is reading ACF/PACF on the raw series and getting orders that don't apply once differencing is applied.
Trade-offs & pitfalls
Adding seasonal terms multiplies the number of parameters and the search space (p,d,q,P,D,Q,s); over-specifying any one of them (especially D, seasonal differencing) can remove real signal along with the seasonality. A senior candidate will also flag that pure ACF/PACF reading is a starting point, not a final answer: automated order search guided by AIC/BIC (or pmdarima.auto_arima) is standard practice once the visual signature is ambiguous, and the final choice should always be checked with a residual diagnostic (no significant autocorrelation left, e.g. a Ljung-Box test) rather than trusted from the identification step alone.
Unlock Full Question Bank
Get access to all Forecasting and Time-Series Analysis interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.