Direct answer
A persistent, one-directional error in one city is a bias problem, not a noise problem, so I would treat it as a hunt for something systematically different about that city: its data, its map, its label definition, or a change in how people and drivers behave there. I would (1) stop the bleeding with a measured, capped bias correction and wider arrival windows for that city, (2) decompose the error to find which part of the trip and which slice of trips carries it, (3) date its onset and line it up against every change in model, data, map and app, and (4) fix the root cause, then remove the temporary correction under the same monitoring. As SRE (site reliability engineering) lead my job is to make the investigation parallel and evidence-driven, and to keep drivers and riders protected while it runs.
Framing the problem precisely
Define signed error per trip as actual minus predicted, so positive means late. "Persistently underestimated" means the average signed error (the bias) is positive and stays positive for days. That rules out a one-off outage and points at something structural.
Also pin down which ETA (estimated time of arrival): the pickup ETA shown to the rider, the trip ETA to destination, or a delivery window promised to a customer. "Drivers miss predicted arrival windows" suggests the windows are built from our ETA, so the window logic is a suspect too.
Telemetry to gather
| Data | Why |
|---|
| Per-trip prediction log: every ETA shown, when, model version, route, feature vector (the list of numeric inputs, such as distance, time of day and road type, fed to the model for that prediction) | reconstruct exactly what we predicted |
| Per-trip ground truth: dispatch, driver accept, start moving, arrive-at-pin, rider-in-car timestamps | split the trip into legs to find where time is lost |
| Driver GPS traces, raw and map-matched (snapped to roads) | compare the route predicted vs the route driven |
| Traffic feed coverage and freshness per area | stale or missing data in part of the city |
| Map-data change log for the city: new roads, turn restrictions, speed limits, closures | map edits can shift every route |
| App and backend release history, feature flags (toggles that turn a code change on or off for some users without a full deploy), experiment assignments | find a change that coincides with onset |
| External context: construction, events, weather, school calendar | real-world shifts the model has not seen |
Statistical analyses
- Onset dating. Plot daily bias for the city with confidence bands (a shaded range around the line showing how much it could plausibly wobble from sampling noise alone, so a move outside the band is more likely a real shift) and run a changepoint test (a method that finds the date where a series' average shifts). A sharp step points to a change we made; a gradual drift points to the world changing under a stale model.
- Decompose by leg. Error = routing-time error + non-driving overhead (finding parking, building access, rider walking out). If the drive leg is accurate and all the bias is in the last 200 m, the fix is a pickup-overhead model, not the traffic data.
- Slice and look for concentration. Bias by zone, hour-of-week, trip length, road class, vehicle type, driver tenure, iOS vs Android, and model version. A city-wide average is often driven by one slice.
- Calibration curve. Bucket trips by predicted ETA and plot average actual vs average predicted. A line parallel but above the diagonal means a constant offset (every trip is off by roughly the same number of seconds, regardless of length); a steeper line means the error grows with trip length. That growth pattern points at speeds being too high: since time equals distance divided by speed, an assumed speed that is a bit too fast produces only a small timing error over a short trip but a much larger one over a long trip, so the gap between actual and predicted widens as predicted ETA increases, rather than staying constant.
- Compare to a neighbour city with similar layout: if it is fine on the same model version, the model itself is less likely the cause.
Worked example: why a city-wide multiplier is the wrong mitigation
Suppose last week's 100,000 pickups break down like this:
| Zone | Trips | Average signed error |
|---|
| Airport | 12,000 | +240 s |
| Downtown | 30,000 | +90 s |
| Suburbs | 58,000 | +10 s |
City bias:
10000012000×240+30000×90+58000×10=1000006,160,000=61.6 s
Adding 62 s to every ETA would leave airport pickups about 3 minutes late, and would make suburban ETAs about 52 s too pessimistic, causing needless cancellations there. The decomposition says the problem is mostly the airport (where a new terminal pickup lane or a changed ride-share zone would produce exactly this pattern) plus a smaller downtown effect. So the mitigation should be per slice, and the investigation should start at the airport.
Model and data checks
- Training-serving skew (the model's inputs are computed one way during training and a subtly different way when actually serving predictions, so the model sees data it was never trained on): recompute features offline for a sample of logged requests and compare to the served values. A timezone bug, a units change (km/h vs m/s) or a missing feature silently defaulted to zero can all bias one city.
- Feature freshness: is the city's traffic feed or probe aggregation (combining speed reports from our own drivers' phones, the "probes," into one speed per road segment) lagging? Stale "free-flow" speeds (the speed a road allows with no traffic on it, normally close to the speed limit, used as a default when nothing else is known) make every ETA optimistic.
- Label definition: did the "arrived" event change (a new app button, a change to the geofence radius, the invisible circle drawn around a location that triggers an event when a phone enters or leaves it)? If arrival is now recorded later, the model is being judged against a moved finish line.
- Map data: new one-way streets, a closed bridge still open in the graph, or a missing turn restriction make the predicted route shorter than any route a driver can take. Compare predicted route length to driven length per trip.
- Map-matching quality (how well raw, noisy GPS points get snapped to the actual road segment a driver is on): GPS snapped to the wrong road (an elevated highway vs the street underneath) corrupts both probe speeds and training labels.
- Training data coverage: was the model last trained before a big change (a new airport terminal, a transit strike)? Check how many training examples came from the affected zones.
- Model version and rollout: if onset matches a model release, compare old vs new on the same trips by replaying logged features.
Operational mitigations while the fix is built
- Per-slice bias correction from the trailing 7 days, applied as an additive offset per zone and hour, capped (for example at 5 minutes) and decaying automatically unless renewed, so a temporary patch cannot become permanent silently.
- Wider arrival windows in the affected zones, and pickup instructions for known trouble spots (the airport's ride-share level).
- Rollback of any model, map or feature release that coincides with onset, if the replay comparison shows it is responsible.
- Driver and rider communication: drivers are not penalized (acceptance or lateness metrics) for windows we mis-set; support gets a macro (a canned, pre-written response an agent can send with one click) explaining delays.
- Monitor the fix like a release: the bias should fall towards zero in the corrected slices without new bias appearing in uncorrected ones.
Trade-offs and pitfalls
- Treating symptoms permanently. Offsets are fast but hide the cause; give them an expiry and an owner.
- Averages lie. Always slice before acting (see the worked example).
- Survivorship in the data. Riders who saw a long ETA and cancelled produce no label; analyse cancellations by predicted ETA too.
- Assuming it is the model. Map, label and pickup-logistics causes produce exactly this signature too, and they are cheaper to check and to fix, so rule them in or out early rather than going straight to retraining.