Direct answer
A model (a program that turns input features, the columns of numbers describing each case, into predictions) is returning NaN (not-a-number, the floating-point value produced by invalid arithmetic) after a library upgrade. My report has to answer four things: what broke, what changed, the smallest way to reproduce it, and how bad it is. The method for shrinking is to freeze the inputs, then repeatedly halve the rows, keeping whichever half still fails, until one row remains, and then drop columns one at a time, keeping each drop that still fails (with only four columns, halving them is not worth the code).
Step 1: confirm and pin
- Check it is deterministic and record the old and new library versions (a comparison of the two environments' package lists shows exactly what changed).
- Freeze one input batch to a file so nothing else varies.
Step 2: shrink with a script
Below, the two predict modes stand in for the old and new library behaviour (old fills missing values with the column mean, new passes them through). In a real report you would replace them with two environments.
python
import numpy as np
rng = np.random.default_rng(7)
X = rng.normal(size=(512, 4))
X[301, 2] = np.nan # one row has a missing value in feature 2
w = np.array([0.5, -1.0, 0.25, 2.0])
def predict(batch, weights, fill_missing):
batch = batch.copy()
if fill_missing: # old behaviour: replace NaN with the column mean
means = np.nanmean(batch, axis=0)
idx = np.where(np.isnan(batch))
batch[idx] = means[idx[1]]
# batch @ weights = weighted sum per row; 1/(1+exp(-z)) squashes it to a 0-1 probability
return 1 / (1 + np.exp(-(batch @ weights)))
def fails(rows, cols, fill_missing=False):
return bool(np.isnan(predict(X[np.ix_(rows, cols)], w[cols], fill_missing)).any())
ALL = [0, 1, 2, 3]
print("old (fills) has NaN:", fails(list(range(512)), ALL, True))
print("new (no fill) has NaN:", fails(list(range(512)), ALL))
rows = list(range(512)) # shrink rows: keep whichever half still fails
while len(rows) > 1:
half = rows[: len(rows) // 2]
rows = half if fails(half, ALL) else rows[len(rows) // 2 :]
print("minimal failing row index:", rows)
cols = list(ALL) # shrink columns: drop one at a time, keep the drop if it still fails
for c in ALL:
trial = [k for k in cols if k != c]
if trial and fails(rows, trial):
cols = trial
print("columns still needed:", cols)
print("input for that row:", X[rows[0]].tolist()[:4])
print("prediction:", predict(X[np.ix_(rows, ALL)], w, False).tolist())
Output:
text
old (fills) has NaN: False
new (no fill) has NaN: True
minimal failing row index: [301]
columns still needed: [2]
input for that row: [-1.7415144478985787, -0.8905443058032078, nan, 0.8884531740936087]
prediction: [nan]
Reading it: the old behaviour gives no NaN, the new one does, halving isolates row 301, and dropping columns one at a time shows only feature 2 is needed, the one with the missing value. Halving keeps the half that still fails, which works when one thing triggers the bug. If neither half fails alone, two rows or columns interact: then remove smaller chunks instead of halves (a technique called delta debugging, rarely needed). The failing input is 1 row out of 512, so it affects 1/512 of rows (about 0.2%).
Step 3: shrink further and write it up
Cut to a standalone script that still fails without the real model or data: paste the code above with the data made inline (one row [-1.74, -0.89, NaN, 0.89], the four weights and the predict function), which is about 20 lines, and confirm it prints NaN in a fresh environment. Then check the release notes for the suspect change.
text
Title: Predictions are NaN for rows with a missing feature after upgrading <library> <old> -> <new>
Severity: Sev-2 (serious but not total). Only rows with a missing value are affected.
Impact: in the sample, 1 of 512 rows (0.2%); downstream ranking drops those users.
Timeline (illustrative): upgrade deployed 03:05 UTC; first NaN in logs 03:12 UTC.
Environment: old and new package lists attached (differences: <library> only).
Minimal repro: attached script, 20 lines, no model file needed.
Sample failing input: [-1.74, -0.89, NaN, 0.89]
Expected: a finite probability (old version imputes the column mean).
Actual: NaN.
Suspected cause (hypothesis, unverified): default handling of missing values changed.
Workaround: impute (fill in the missing value, here with the column mean) before calling predict.
If the trigger is a downstream data pipeline change instead of a library upgrade
Do the same halving on the pipeline's output before and after the change: keep the snapshot timestamps, the first bad output time, and one failing sample record. Report severity and the share of affected records in the same way.
Pitfalls
- Shrinking before freezing the input, so the failure moves.
- Filing a screenshot of the model output instead of the smallest script.
- Presenting the suspected cause as a fact.