Quality Metrics and Test Reporting Questions
Defining, computing, and communicating software quality. Covers choosing meaningful quality and test metrics (defect escape rate, defect detection effectiveness, defect density, MTTD/MTTR, pass rate, regression frequency, automation suite health and maintenance cost) versus vanity numbers; baselining, trend interpretation (real change versus normal variation), and alert thresholds; dashboards, weekly stability reports, and release-quality reports for engineering, product, and executive audiences, including composite go/no-go scores; framing unfavourable results; guarding against gamed metrics and reading what numbers such as code coverage or a high pass rate hide; investigating contradictory or shifting metrics and testing whether a quality signal really predicts customer outcomes; metric definitions, ownership, and governance; computing metrics from test-run and bug-tracker data (SQL and scripts); designing test-results reporting pipelines, storage schemas, real-time versus batch reporting, alerting, and failure fingerprinting and grouping for triage; and tying quality signals to product and business outcomes. Deciding what to automate and diagnosing individual flaky tests are covered elsewhere.
Your team wants sub-minute alerts on critical test failures and nightly analytics for trends. Would you run real-time and batch reporting separately or as one system? Describe the architecture and where it fails.
Sample Answer
Direct answer
Run one pipeline with two speeds, not two separate systems. Every test result from CI (continuous integration, the system that builds and tests each change) is published once to a durable event log (an append-only queue that keeps events so they can be read again later; Kafka is a well-known example, and cloud services offer equivalents) and stored raw. A small streaming path reads that log to raise sub-minute alerts for a short list of critical tests only. A batch path (a scheduled job over the whole stored history) computes nightly trends. Both paths use the same event schema and the same metric definitions, so an alert and the morning report can never disagree about what a failure is. I would keep them as two consumers of one source, and reserve a fully separate second system for when volume or ownership truly forces it.
Architecture
| Stage | Streaming path (alerts) | Batch path (trends) |
|---|---|---|
| Input | CI publishes one event per test result to the log (an append-only, replayable queue) | Reads the raw store filled from the same log |
| Work | Filter to critical tests, confirm, dedup, route | Aggregate, join ownership, rank, compute rates over days |
| Output | Message to the owning team | Dashboards and the nightly summary |
| Truth | Fast, approximate, may be revised (late or out-of-order events can arrive after the alert fired) | Slow, complete, authoritative |
Alert path details, in order:
- Filter: only tests tagged critical (checkout, login, billing) can page. Everything else waits for the batch report.
- Confirm to bound false positives (limit how often a false alarm pages someone): a failure alerts only if it repeats on an automatic rerun, or fails on N consecutive runs of the same commit. A flaky test (one that passes and fails on identical code) is suppressed unless its recent failure rate is low, so a sudden failure of a normally stable test stands out.
- Dedup (deduplicate: collapse repeats into one): key = test id + failure signature (a short fingerprint of the error text) + branch. Duplicate deliveries and repeated failures collapse into one open alert.
- Owner routing: look up the owning team from the code-owners file (a file in the repository, often named CODEOWNERS, mapping paths to responsible teams) at the commit under test; fall back to a shared triage channel if none matches.
- Rate limit: at most one page per (signature, owner) per window (illustratively 30 minutes), then further failures are added to the open alert and rolled into a digest. This stops a broken shared dependency from paging 200 times.
Worked example (illustrative)
At 14:02 the login test fails on main. The event reaches the stream, the rerun fails too, and no open alert exists for signature TimeoutError@auth, so one alert goes to the identity team. In the next 10 minutes 40 more login and profile tests fail with the same signature, and they attach to that alert instead of paging again. Overnight the batch job counts all 41 failures once, groups them under one cause, and the nightly summary lists that cause with a link to the alert. The two views agree because both used the same signature function and the same "counted failure" definition.
Where it fails
- Duplicate and out-of-order events: CI retries publishing, shards finish in odd order. Use an idempotent key (an identifier that makes processing the same event twice harmless, because the second copy is recognised and ignored): run id + test id + attempt, so counts do not double. Out-of-order events are events that arrive in a different order from the one in which they happened.
- Two definitions drifting apart: if streaming and batch each implement "failure", numbers diverge. Keep the definition in one shared library or view, and reconcile nightly by comparing alert counts to batch counts. Example: the stream counted 41 critical failures for a signature and the batch job counted 43; the two late events explain the gap (about 5%), and a gap much bigger than that triggers an investigation.
- Silent pipeline death: no alerts can mean no failures or a dead consumer. Alert on consumer lag (how far behind the newest event the reader is, for example more than 2 minutes) and send a periodic heartbeat test result (a known always-passing dummy test published every 5 minutes, so if none arrives for 15 minutes the pipeline itself is presumed broken).
- Alert storms and fatigue: dedup and rate limiting are the guard. Track the share of alerts that humans mark as noise and tighten confirmation when it rises.
- Stale ownership: a moved directory routes pages to the wrong team, so review the fallback channel's volume as a signal.
- Late data: batch numbers may change after an alert fired, so the report should show the revised figure and not hide the change.
What would change my call
For a small team with modest event volume, a single system running one-minute micro-batches (processing events in small groups every minute instead of one by one) meets the need and is cheaper to operate. If a separate platform team owns the alerting stack, or the alert path needs a much stricter availability target than analytics, split them physically, but still share the event schema and definitions.
Given test_results(test_id, suite, run_id, status, started_at, finished_at), write a Postgres query for pass rate per suite over the last 7 days, excluding skipped and quarantined runs, as a percentage with two decimals, lowest first. Which data-quality traps could change your numbers?
Sample Answer
Direct answer
Count rows per suite over a rolling 7-day window, keep only rows that gave a verdict (passed or failed; skipped means not run, quarantined means a known-unreliable test moved out of the blocking gate), divide passed by that total, round to two decimals and sort ascending. Use FILTER so a suite whose rows were all skipped or quarantined still appears, and NULLIF so its empty denominator gives NULL instead of a divide-by-zero error.
Setup (illustrative data, timestamps relative to now so results are reproducible)
Run the setup and then each query below in the same psql session against PostgreSQL. It uses a temporary table, so nothing persists.
CREATE TEMP TABLE test_results (
test_id text, suite text, run_id int, status text,
started_at timestamptz, finished_at timestamptz);
INSERT INTO test_results
SELECT test_id, suite, run_id, status, started_at, started_at + secs * interval '1 second'
FROM (VALUES
('t1', 'checkout', 101, 'passed', now() - interval '2 days', 4),
('t2', 'checkout', 101, 'failed', now() - interval '2 days', 10),
('t2', 'checkout', 101, 'passed', now() - interval '2 days' + interval '5 minutes', 8), -- retry
('t3', 'checkout', 101, 'passed', now() - interval '2 days', 6),
('t4', 'checkout', 101, 'skipped', now() - interval '2 days', 0),
('t1', 'checkout', 102, 'passed', now() - interval '1 day', 4),
('t2', 'checkout', 102, 'passed', now() - interval '1 day', 8),
('t3', 'checkout', 102, 'failed', now() - interval '1 day', 10),
('t4', 'checkout', 102, 'quarantined', now() - interval '1 day', 2),
('s1', 'search', 201, 'passed', now() - interval '3 days', 2),
('s2', 'search', 201, 'passed', now() - interval '3 days', 2),
('s3', 'search', 201, 'failed', now() - interval '3 days', 12),
('s4', 'search', 201, 'passed', now() - interval '3 days', 2),
('s1', 'search', 150, 'failed', now() - interval '10 days', 2), -- outside 7 days
('l1', 'legacy', 301, 'skipped', now() - interval '1 day', 0),
('l2', 'legacy', 301, 'quarantined', now() - interval '1 day', 0)
) AS v(test_id, suite, run_id, status, started_at, secs);
Plain-language guide to the SQL used
GROUP BY suitecollapses rows into one output row per suite so aggregates likecount(*)work per suite.count(*) FILTER (WHERE ...)counts only rows matching the condition, so numerator and denominator come from one pass.NULLIF(x, 0)returns NULL when x is 0, avoiding divide-by-zero.::numericconverts a value to an exact decimal type soROUND(..., 1)works.EXTRACT(EPOCH FROM interval)turns a time span into a number of seconds.WITH name AS (...)is a CTE (common table expression): a named temporary result the next query reads from.DISTINCT ON (a, b, c)keeps just the first row for each (a, b, c) group; because of theORDER BY ... started_at DESC, that first row is the latest attempt.
The query
SELECT suite,
ROUND(100.0 * count(*) FILTER (WHERE status = 'passed')
/ NULLIF(count(*) FILTER (WHERE status IN ('passed', 'failed')), 0), 2) AS pass_rate_pct
FROM test_results
WHERE started_at >= now() - interval '7 days'
GROUP BY suite
ORDER BY pass_rate_pct ASC NULLS LAST;
Output (psql command tags omitted)
suite | pass_rate_pct
----------+---------------
checkout | 71.43
search | 75.00
legacy |
(3 rows)
checkout: 5 passed of 7 verdicts = 71.43%. search: 3 of 4 = 75.00%; the 10-day-old failure is outside the window. legacy has no verdicts, so its rate is NULL and sorts last.
Data-quality traps that change the numbers
- Retries. The table has no attempt column, so
t2in run 101 appears twice (failed, then passed on retry) and the query above counts both. The retry variant below keeps only the last attempt per test per run and gives a different answer. - Denominators. Putting skipped and quarantined rows in the denominator lowers rates and punishes suites for not running. Also watch an empty denominator, as with
legacy. - Duplicates. A CI (continuous integration) system that re-sends results double-counts unless there is a unique key such as (run, test, attempt number). This table has no attempt column, so you would add one, or de-duplicate on (run, test, started_at).
- Window and time zone. A rolling 7 days is not the same as 7 calendar days, and a date cast depends on the session time zone (the time zone setting of your database connection), so
2026-07-01can be a different instant for two users. - In-flight and odd statuses. In-flight rows are tests still running, with a NULL
finished_at. Rows still running,error, or differently cased statuses (PASSED) fall outside a strictIN ('passed', 'failed'). Decide their bucket explicitly. - Small samples.
searchhas only 4 verdicts, so a single flipped test moves it by 25 points. Show the count next to the percentage. - A suite that stops reporting disappears from the output altogether, which looks like health. Left-join from a list of expected suites to catch it, for example
SELECT e.suite, r.pass_rate_pct FROM expected_suites e LEFT JOIN rates r USING (suite), where a missing suite shows up with NULL.
Retry variant: last attempt per test per run
WITH final_attempt AS (
SELECT DISTINCT ON (suite, run_id, test_id) suite, run_id, test_id, status
FROM test_results
WHERE started_at >= now() - interval '7 days'
ORDER BY suite, run_id, test_id, started_at DESC
)
SELECT suite,
ROUND(100.0 * count(*) FILTER (WHERE status = 'passed')
/ NULLIF(count(*) FILTER (WHERE status IN ('passed', 'failed')), 0), 2) AS pass_rate_pct
FROM final_attempt
GROUP BY suite
ORDER BY pass_rate_pct ASC NULLS LAST;
suite | pass_rate_pct
----------+---------------
search | 75.00
checkout | 83.33
legacy |
(3 rows)
checkout rises from 71.43% to 83.33% (5 of 6). Neither number is wrong; they answer different questions (every attempt versus final outcome), so state which one the dashboard shows.
Variant: failure rate and average duration over 30 days
SELECT suite,
ROUND(100.0 * count(*) FILTER (WHERE status = 'failed')
/ NULLIF(count(*) FILTER (WHERE status IN ('passed', 'failed')), 0), 2) AS failure_rate_pct,
ROUND(AVG(EXTRACT(EPOCH FROM finished_at - started_at))::numeric, 1) AS avg_seconds
FROM test_results
WHERE started_at >= now() - interval '30 days'
AND status IN ('passed', 'failed')
GROUP BY suite
ORDER BY failure_rate_pct DESC;
suite | failure_rate_pct | avg_seconds
----------+------------------+-------------
search | 40.00 | 4.0
checkout | 28.57 | 7.1
(2 rows)
Both rows include the old search failure because 30 days reaches back 10 days. Failures 2 of 5 = 40.00%, durations (2 + 2 + 2 + 12 + 2) / 5 = 4.0 seconds.
Complexity and edge cases
One pass with a filter, O(n) (work grows in proportion to the number of rows in the window), and an index on started_at (a lookup structure so Postgres can jump to recent rows instead of scanning all) lets it skip older data. DISTINCT ON adds a sort, O(n log n) (slightly worse than proportional). Edge cases: no rows in the window, in-flight rows with NULL finished_at (AVG silently skips them, so the average covers fewer rows than the count), and ties in started_at during retries.
Which metrics tell you whether an automated suite is healthy? How would you measure each one, what trend would worry you, and what would you do about it?
Sample Answer
Direct answer
A healthy suite is one the team trusts, that gives fast feedback, that catches real defects, and that costs a sustainable amount to maintain. No single number shows all four, so I track a small set of metrics, one group per question, and for each I know how it is measured, which trend worries me, and the one action it triggers. (CI means continuous integration, the automated system that builds and tests every change. A flaky test is one that passes and fails on the same code. Quarantine means moving an unreliable test out of the gate that blocks merges. p95 is the value 95% of observations fall below, so it describes the slow tail rather than the average, and p50 is the median. Test rot is a suite decaying as the product changes and tests stop matching it. Escaped defects are bugs found after release. Wall-clock time is elapsed real time, not CPU time. To shard and parallelise is to split the suite into pieces that run at once on several machines. A runner is the machine that executes CI jobs. A triage rota is a rotating on-call duty to look at red builds. Hard-coded waits are fixed sleeps, such as waiting 5 seconds for a page, instead of waiting for a condition. If you are starting from nothing, begin with pass rate, flake rate and duration; the cost and deflection rows are organisation-level extras.)
The metrics, grouped by the question they answer
Reading the table by question: can I trust the suite (pass rate, flake rate, quarantine size and age); is it fast (duration and queue time, time to green); does it find bugs (defect escape rate); is it affordable (maintenance cost per test, manual-test deflection).
| Metric | How I measure it | Trend that worries me | What I do |
|---|---|---|---|
| Pass rate on the main branch | passed / (passed + failed) per day, executed tests only, with skipped and quarantined counts shown beside it | Sliding for weeks, or a flat 100% while production bugs rise | Split real regressions from test rot before reacting |
| Flake rate | Share of executions on unchanged code that flip result on retry (this row only measures the rate; diagnosing why an individual test flakes is a separate investigation) | Rising, or the same tests flipping week after week | Quarantine with an owner and an expiry date, then fix or delete |
| Duration and queue time | p50 and p95 wall-clock of the full CI run, plus time waiting for a runner | p95 growing faster than test count, or people merging without waiting | Shard and parallelise, delete redundant tests, move slow ones to nightly |
| Time to green | Median time from first red build on main to the next green one | Growing, or red builds with no owner | Route failures to owners, run a triage rota |
| Defect escape rate | Escaped defects / (escaped + caught before release), per release | Flat or rising while pass rate looks good | Attribute each escape to the stage that should have caught it, add a test there |
| Quarantine size and age | Count of quarantined tests and the age of the oldest | Count only grows, entries older than the policy limit | Enforce expiry: fix or delete |
| Maintenance cost per test | Engineer hours spent fixing or updating tests per month / number of tests | Cost per test rising | Remove low-value tests, fix brittle patterns (for example, hard-coded waits) |
| Manual-test deflection | Automated cases / (automated + manual regression cases), and hours of manual regression per release | Flat while release effort stays high | Automate the manual checks repeated most often |
Targets, cadence and the organisation view
- I set targets against each team's own baseline first (for example, "no quarantined test older than a fixed number of days"), not a universal number borrowed from another company.
- The team reviews a one-page view weekly; an engineering manager reviews the trends monthly with the same definitions rolled up across teams. The organisation dashboard adds maintenance cost per test and manual-test deflection, because those are the numbers that show whether automation is paying for itself.
Worked example
A team has 1,200 tests. A quarter ago it spent 42 engineer-hours a month on test upkeep. Now it spends 60.
- Then: 42 h x 60 min / 1,200 tests = 2.1 minutes per test per month.
- Now: 60 h x 60 min / 1,200 tests = 3.0 minutes per test per month, about 43% more (60 / 42 = 1.43).
- Meanwhile pass rate stayed at 98% and quarantine grew from 10 to 45 tests. The pass rate alone says "fine"; the cost and quarantine metrics say the suite is quietly decaying. The action is to review the 45 quarantined tests, delete or fix them by their expiry date, and look at which tests take the most upkeep.
Defect escape rate traced: a release had 15 defects, 12 caught before release and 3 escaped, so 3 / (3 + 12) = 20%. If pass rate stayed 98% while that rate went from 20% to 30% over two releases, the suite is passing yet missing more.
Trade-offs and pitfalls
- Any single metric can be gamed (skipping tests lifts pass rate, deleting tests lifts everything), so I pair each with a counter-metric, a second number that exposes the cheat: pass rate with skipped and quarantined counts, duration with escape rate (a suite made fast by deleting tests will show escapes rising).
- Too many metrics means none is acted on. Each one must map to a decision; if it never changes a decision I drop it.
- Raw line coverage is deliberately absent: it says code was executed, not that behaviour was checked. I prefer coverage of critical user paths.
Design a test-reporting platform for an organisation running tens of thousands of test runs a day across many teams. What must the dashboards answer, what query latency would you commit to, how are failures grouped, and how do you control cost?
Sample Answer
Direct answer
I would build a pipeline that ingests every CI (continuous integration) system's results in one common schema, stores raw attempts cheaply (an attempt is one execution of one test), pre-computes rollups for the questions people ask every day, groups failures by a fingerprint of the error, and controls cost by tiering retention and guarding queries. The design starts from the questions the dashboards must answer, because those decide what to pre-compute.
What the dashboards must answer (by audience)
| Audience | Question |
|---|---|
| Developer | Did my change break this, or is the failure a known flake? Where is the log? |
| Team lead | Is pass rate, duration and quarantine size trending the wrong way? Which tests cost us most? |
| Engineering manager | Are escapes and time to fix improving across teams, with one definition? |
| Platform owner | How much does the platform cost, and which queries or teams drive it? |
Latency commitments (targets I would commit to only after measuring on real data)
- Freshness: a run's results visible to its developer within a couple of minutes of finishing; rollups within about 15 minutes.
- Standard dashboards read rollups: p95 (95th percentile, the slow tail) under 2 seconds.
- Per-test history drill-down: p95 under 5 seconds.
- Ad-hoc queries over raw data: best effort within a minute or so, and refused if they carry no time filter.
Terms used below. Ingest means receiving and loading data into the platform. Dedupe means dropping duplicate records. Scrub means removing sensitive text such as tokens or emails. Adapter is a small translator from one CI system's output to our schema. Queue is a buffer that decouples receiving results from processing them. Object storage is cheap file storage (such as S3) for logs and screenshots, and a URI is the address of a file there. Quarantine means a flaky test is moved out of the pass/fail gate so it stops blocking builds, while still running and being tracked. Rollup means a pre-computed summary table.
Reading the diagram. Results flow left to right: CI systems send results to the ingest API, which stores big artifacts in object storage and puts records on a queue. From the queue they land in the raw attempts table and go to the failure-grouping step, which writes a signature back onto the raw rows. Rollup jobs read raw data and fill the aggregates that dashboards read, while ad-hoc queries go straight to raw.
Architecture
flowchart LR
CI["CI systems"] --> ING["Ingest API: validate, dedupe, scrub"]
ING --> ART[("Object storage: logs, screenshots")]
ING --> Q["Queue"]
Q --> RAW[("Raw attempts, daily partitions")]
Q --> SIG["Failure grouping"]
SIG --> RAW
RAW --> ROLL["Rollup jobs"]
ROLL --> AGG[("Aggregates")]
AGG --> DASH["Dashboards"]
RAW --> ADHOC["Ad-hoc queries"]
- Ingest: adapters translate each CI system's output (many runners can emit the JUnit XML report format, a widely used file that lists each test with its name, status, duration and failure message, so one parser covers many tools) into one schema, validate it, drop duplicates by run, test and attempt, and scrub sensitive text.
- Storage: raw attempts in daily partitions, artifacts in object storage referenced by URI (tiering those artifacts is an infrastructure question I leave aside here).
- Rollups per day, per suite and per team feed the dashboards; raw data feeds investigation.
Scale and cost. Assume 50,000 runs a day with 400 results each: 20,000,000 rows a day, about 7.3 billion a year (20,000,000 x 365). At an assumed 200 bytes a row that is 4 GB a day before compression (an estimate). Over a 90-day full-attempt window that is about 360 GB (4 GB x 90), and a full year of raw rows would be about 1.5 TB (4 GB x 365), so keeping everything forever costs roughly four times the 90-day window. A daily rollup by suite and team is only thousands of rows a day, negligible next to that. Cost controls:
- Keep full attempts for a fixed window (for example 90 days), then keep failed and flaky attempts plus rollups. Failures are never sampled away.
- Keep one year of searchable metadata (run, test, status, team, commit, error fingerprint), not logs.
- Drop whole partitions for retention, cache dashboard queries, apply per-team query quotas (a cap on how much data or how many queries each team can run, so one heavy user cannot drive the bill).
Grouping failures. Group by a signature: exception type, normalised message and the top stack frame, where ids, ports, addresses and counts are replaced by placeholders so the same root cause hashes to the same value. A hash turns any text into a short fixed-length code, so identical input always gives the identical code; the top stack frame is the first line of the call trace, the file and line where the error surfaced. Reading the code: the regex lines rewrite variable parts into placeholders. Hex addresses and UUIDs are rewritten before the digit rule because otherwise 0x7f3a would be left as a jumbled 0x<n>f<n>a and the same address would fingerprint differently each time; the last rule turns every remaining number into <n>. SHA-1 truncated to 8 characters is a compact label, not a security feature.
import hashlib
import re
def signature(exc_type, message, top_frame):
"""Fingerprint a failure so the same root cause groups together."""
msg = message.lower()
msg = re.sub(r"0x[0-9a-f]+", "<hex>", msg) # memory addresses
msg = re.sub(r"[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}", "<uuid>", msg)
msg = re.sub(r"\d+", "<n>", msg) # ids, ports, counts
key = f"{exc_type}|{msg}|{top_frame}"
return hashlib.sha1(key.encode()).hexdigest()[:8]
failures = [
("TimeoutError", "waited 30012 ms for element #cart-41", "cart_page.py:88"),
("TimeoutError", "waited 30977 ms for element #cart-87", "cart_page.py:88"),
("ConnectionError", "connect to 10.0.3.17:5432 refused", "db.py:31"),
("ConnectionError", "connect to 10.0.9.44:5432 refused", "db.py:31"),
("AssertionError", "expected total 1999 got 2000", "test_totals.py:12"),
]
groups = {}
for f in failures:
groups.setdefault(signature(*f), []).append(f[1])
for sig, msgs in groups.items():
print(sig, len(msgs), msgs[0])
69cb5206 2 waited 30012 ms for element #cart-41
8e4deaa4 2 connect to 10.0.3.17:5432 refused
9bd79436 1 expected total 1999 got 2000
Five failures collapse to three groups. Timeouts and connection refusals differ only by ids and IP addresses, so they merge; the assertion stays separate. Store the signature on each row, and show each group with its first-seen commit and whether it is new.
Privacy. Error messages can contain emails, tokens or customer data. Scrub at ingest (before storage), restrict access by team, set retention on metadata, and support deletion by request.
Trade-offs and pitfalls
- Signatures can over-merge (different bugs with the same message) or under-merge (one bug, many messages). Allow manual merge and split overrides and review the largest groups.
- Low latency costs more; that is why only rollups carry the tight commitment.
- A schema too rigid for the least capable CI system causes teams to bypass it, so keep the required fields few.
Failures spike across several suites right after a shared dependency upgrade. Using your reporting data and captured artifacts, how do you quickly decide whether it is one cause or many, and how do you tell people?
Sample Answer
Direct answer
Group the failures by root-cause signature, then compare against the change. If a small number of signatures explains most failures and their first appearance lines up with the upgrade, it is one cause with a small tail of unrelated ones. If failures spread across many signatures with no shared frames, it is several causes, or noise the upgrade merely coincided with. Then confirm by running a failing test with the old dependency version, and tell people early with a first message that says what is known, what is not, and when the next update is.
Steps for the first 30 minutes
- Automated checks first (0-10 min).
- Pull failures since the last green build and fingerprint them (the fingerprint is a hash of the error type plus the innermost frames, so the same bug groups together).
- Find the time of the first failure and compare it with the upgrade merge time.
- Diff the dependency lockfile (the file that pins the exact version of every library the build uses) between the last green build and the first red one, so you know exactly which versions moved.
- Cluster (10-15 min). Count failures by signature, and count how many suites and tests each covers.
- Rule out coincidences (15-20 min). Check for an infrastructure outage, and compare against the normal daily flaky-failure count so ordinary background failures are not blamed on the upgrade.
- Confirm (20-30 min). Take one failing test from the biggest cluster and run it against the previous lockfile (a controlled A/B: the same test and code run twice with only the dependency version changed). Failing only on the new version confirms causation.
- Severity and owner. If the failing suites gate a release, treat it as high severity; assign an owner (the upgrade's author or the platform team), not "QA".
Worked example (illustrative numbers)
The upgrade merged at 14:05. 120 failures across 4 suites after the upgrade:
| Signature | Failures | Suites | First seen |
|---|---|---|---|
| AttributeError in the upgraded library's client wrapper | 96 | 4 | 14:12 |
| Timeout in the payment stub | 18 | 1 | 11:40 (before the upgrade) |
| Assertion on a date format | 6 | 1 | 14:40 |
The top signature covers 96 / 120 = 80% of failures across every suite and first appears 7 minutes after the merge, so it is the upgrade. The timeout cluster started at 11:40, over two hours before the merge, so it cannot be caused by the upgrade. The date-format cluster is the one that timing cannot settle (it appeared after the merge, but it is only 6 failures in one suite), so apply the confirm step: run that test against the previous lockfile. In this example it fails there too (the assertion depends on the machine's locale), so it is unrelated. Conclusion: one primary cause (the upgrade) plus two unrelated ones.
Deciding one cause versus many
My rule of thumb: if one signature covers most failures (say, above about 70%) and matches the time and the diff, treat it as one cause and investigate the leftover separately. This threshold is a judgement call, not a standard.
How to tell people (first message, at roughly 30 minutes)
Upgrade of the shared HTTP library merged at 14:05. Since then 120 test failures across 4 suites; 80% share one error (AttributeError in the client wrapper), confirmed by reproducing on the new version only. Impact: the release candidate (the build being prepared for release) is blocked. Owner: platform team. Options: revert the upgrade (fastest) or patch the wrapper. Next update at 15:30. Other failures (timeouts, date format) look unrelated and are tracked separately.
Include the failed-suite list link, artifacts (the logs, screenshots and reports each failed run saves), owner and next update time. Send it to the channel used by every team that has affected suites, and to the release manager directly.
Trade-offs and pitfalls
- Reverting first and investigating second is often right, because the upgrade is easy to undo and the blocked release is expensive.
- Avoid blaming the upgrade until you have reproduced it; coincidence with an infrastructure change is common.
- Do not send a message with only "tests failing"; without owner and next-update time it produces a flood of questions.
Unlock Full Question Bank
Get access to all 12 Quality Metrics and Test Reporting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.