Quality Metrics and Test Reporting Questions
Defining, computing, and communicating software quality. Covers choosing meaningful quality and test metrics (defect escape rate, defect detection effectiveness, defect density, MTTD/MTTR, pass rate, regression frequency, automation suite health and maintenance cost) versus vanity numbers; baselining, trend interpretation (real change versus normal variation), and alert thresholds; dashboards, weekly stability reports, and release-quality reports for engineering, product, and executive audiences, including composite go/no-go scores; framing unfavourable results; guarding against gamed metrics and reading what numbers such as code coverage or a high pass rate hide; investigating contradictory or shifting metrics and testing whether a quality signal really predicts customer outcomes; metric definitions, ownership, and governance; computing metrics from test-run and bug-tracker data (SQL and scripts); designing test-results reporting pipelines, storage schemas, real-time versus batch reporting, alerting, and failure fingerprinting and grouping for triage; and tying quality signals to product and business outcomes. Deciding what to automate and diagnosing individual flaky tests are covered elsewhere.
What strategies do you use to set alert thresholds on quality metrics so they are sensitive without turning into noise? Walk me through a concrete rule for weekly regression failures.
Sample Answer
Direct answer
Set the threshold from the metric's own history, not from a round number like "5 percent". Learn what a normal week looks like and how much it naturally wobbles, then alert only when a week is both unusual and big enough to matter. Use two rules, because failures show up in two shapes: a spike (one bad week) and a drift (a slow climb no single week would trip). Every alert needs an owner and a first action, otherwise it is just noise. This answer stays on quality-metric thresholds, not service-level objective (SLO) or error-budget design.
Structured elaboration
- Measure a rate, not a count. Weekly regression failures (regression means previously working behaviour that broke) = failed test executions divided by total executions. A week that ran more tests would otherwise look worse for no real reason.
- Learn the baseline (what "normal" is) from recent weeks. Use the median of the last 8 weeks and the MAD (median absolute deviation: the median distance of each week from the median). Unlike the standard deviation, one terrible week barely moves it. Multiplying MAD by 1.4826 gives a "robust standard deviation" on the same scale as the ordinary one. (Why that number: for normally distributed data, the MAD comes out at about 0.6745 of the standard deviation, and 1 / 0.6745 = 1.4826, so the multiplier converts MAD into what the standard deviation would have been.)
- Spike rule: one week above median + max(3 robust standard deviations, 1 percentage point). A percentage point is the plain difference between two percentages: going from 2.1 percent to 3.1 percent is 1 percentage point (a 48 percent relative rise, which is a different thing). The 1-point floor stops a very quiet history from producing a hair-trigger alert on a 0.3-point wobble nobody cares about.
- Drift rule: three consecutive weeks above median + 1 robust standard deviation. This catches slow rot.
- Route by severity. A spike gets triaged next working day; a drift becomes a ticket for the suite owner. A weekly metric should not page anyone.
- Keep the baseline honest. Leave out weeks with a declared outage, and re-baseline after a deliberate change (for example a big batch of new tests).
- Review the rule itself monthly. Mark each alert actionable or not; if most were not, loosen the rule.
Worked example
Save as weekly_alert.py and run it. The numbers are illustrative weekly failure rates in percent.
import statistics
# Weekly regression-suite failure rate in percent (failed executions / total executions)
weeks = [2.1, 1.8, 2.4, 2.0, 2.2, 1.9, 2.3, 2.1, 2.6, 4.6, 2.0, 2.8, 3.1, 3.4]
baseline_weeks = weeks[:8] # trailing window used to learn "normal"
med = statistics.median(baseline_weeks)
mad = statistics.median(abs(x - med) for x in baseline_weeks)
rsd = 1.4826 * mad
spike = med + max(3 * rsd, 1.0) # one week above this = alert (never closer than 1 point)
drift = med + 1 * rsd # three weeks in a row above this = alert
print("median %.2f MAD %.2f robust sd %.3f" % (med, mad, rsd))
print("spike line %.2f drift line %.2f" % (spike, drift))
streak = 0
for i, r in enumerate(weeks, start=1):
streak = streak + 1 if r > drift else 0
flags = []
if r > spike:
flags.append("SPIKE")
if streak >= 3:
flags.append("DRIFT")
print("week %2d rate %.1f %s" % (i, r, " ".join(flags)))
Output:
median 2.10 MAD 0.15 robust sd 0.222
spike line 3.10 drift line 2.32
week 1 rate 2.1
week 2 rate 1.8
week 3 rate 2.4
week 4 rate 2.0
week 5 rate 2.2
week 6 rate 1.9
week 7 rate 2.3
week 8 rate 2.1
week 9 rate 2.6
week 10 rate 4.6 SPIKE
week 11 rate 2.0
week 12 rate 2.8
week 13 rate 3.1
week 14 rate 3.4 SPIKE DRIFT
By hand: the eight baseline weeks sorted are 1.8, 1.9, 2.0, 2.1, 2.1, 2.2, 2.3, 2.4, so the median is the average of the middle two, (2.1 + 2.1) / 2 = 2.10. Each week's distance from 2.10 is 0, 0.3, 0.3, 0.1, 0.1, 0.2, 0.2, 0 in week order (2.1, 1.8, 2.4, 2.0, 2.2, 1.9, 2.3, 2.1); sorted that is 0, 0, 0.1, 0.1, 0.2, 0.2, 0.3, 0.3, and the median of those is (0.1 + 0.2) / 2 = 0.15. So weeks 1 to 8 give median 2.10 and MAD 0.15, so the robust standard deviation is 0.222. The spike line is 2.10 + max(0.67, 1.0) = 3.10 and the drift line is 2.10 + 0.222 = 2.32. Week 10 (4.6) trips the spike rule, and the next week is back at 2.0, which points to a one-off such as an environment problem. Weeks 12, 13 and 14 (2.8, 3.1, 3.4) each sit above the drift line, so the drift rule fires at week 14 even though no earlier week in that run crossed the spike line. Week 9 (2.6) is above the drift line but is only a streak of one, so it stays quiet.
Trade-offs and pitfalls
- A fixed limit ("fail above 5 percent") is either too late for a suite that normally runs at 2 percent or constant noise for one that runs at 4.
- Mean and standard deviation are dragged by the very outliers you want to detect; median and MAD are not.
- More sensitivity means more false alarms. The 1-point floor and the three-week requirement are the deliberate brakes, and both numbers are tunable conventions, not laws.
- Teams can game any threshold, for example by quarantining failing tests. Report the quarantined-test count next to the rate.
- If a suite runs very few tests per week, a rate is too jumpy. Fall back to counts or widen the window to a month.
Design the alerting for CI test failures: when does a failure notify a team channel, when does it page someone, and when is it suppressed? Choose thresholds and de-duplication windows and defend them.
Sample Answer
Direct answer
Notify the team channel for failures that need a human to look but not right now, page (wake or interrupt an on-call person through a pager or phone alert) only when the failure blocks a release or signals production risk, and suppress everything that is noise or already known. I would deduplicate by failure signature (a fingerprint of the error and location) plus branch, so one root cause produces one alert, not one per test. Every number below is a starting value to tune against a few weeks of your own alert history.
The three tiers
| Tier | Trigger | Why |
|---|---|---|
| Suppress | Failure on a feature branch or draft pull request (the author sees it in the pull request, nobody else is told); a test currently flagged quarantined (a known-unreliable test that is parked outside the blocking set but still recorded); failures during a declared CI infrastructure outage; a failure that passes on one automatic rerun | These are noise or already handled |
| Notify channel | Same signature fails on the main branch on 2 consecutive runs (after the automatic rerun); or a nightly suite's pass rate falls more than 5 percentage points below its own 14-day median (for a suite that normally sits at 97%, that means below 92%) | A real regression is likely, but a working-hours fix is fine |
| Page | A post-deploy smoke test (a small check run against the live deployment) fails twice in a row; or main is red and blocking the deploy pipeline for 60 minutes with no acknowledgement; or 30% or more of suites fail in the same run (mass failure suggests a shared cause) | Users or the release are at risk |
De-duplication windows
- Channel notification: one thread per signature and branch. Repeats inside a 4-hour window add a reply ("seen 6 more times") rather than a new message. Reason: a nightly job might legitimately fail again the next morning, but not within hours as a new event.
- Page: one open incident per signature until it is acknowledged and resolved. Repeats never re-page, they append to the incident.
Why 5 points: on a suite with a few hundred tests, a handful of flaky tests can move the pass rate by a point or two on an ordinary night, so a 5-point drop is comfortably outside normal wobble. If your own 14-day history shows wider swings, widen the margin. Tune it until at most a few channel alerts a week are false alarms.
Escalation and back-off for repeated alerts
(On-call is the person or rotation responsible for responding out of hours; primary is the first person paged, secondary is the backup. Acknowledgement means that person confirms in the paging tool that they have seen the alert and are working on it.)
- Page the primary on-call. No acknowledgement after 15 minutes: page the secondary.
- Still no acknowledgement after 30 minutes: notify the engineering manager.
- For a channel alert that is still open: reminder after 4 hours, then 24 hours, then move it into a daily digest. Back-off stops a stale alert from becoming wallpaper.
Flaky tests (pointer only)
A flaky test is one that passes and fails on the same code. Detecting and repairing flaky tests is a separate discipline. Here the rules only consume its output: a test with a quarantine flag is suppressed from paging and channel alerts but still recorded, and a quarantined test that fails while running alone still shows on the dashboard. Rerun-on-failure is a suppression aid, not a fix.
Worked example
The main-branch suite runs every hour, plus the longer nightly suite. At 02:10 the main run fails: 30 tests, 3 signatures. After auto-rerun 26 still fail and all 26 share signature A. The next hourly run fails identically at 03:10 (second consecutive run), so the channel gets one message: "Signature A, 26 tests, 2 runs, owner: payments". The smoke test after the 09:00 deploy passes, so nobody is paged. Had that smoke test failed twice, payments on-call would be paged once, with the same signature attached.
Trade-offs and pitfalls
- Lower thresholds catch bugs earlier and train people to ignore alerts. The 2-consecutive-run rule trades up to one run of delay for far fewer false alarms.
- Paging on test failures alone is usually wrong. Page on impact (deploys blocked, production signal), not on test count.
- What would change my call: a very slow suite (a run takes hours) argues for notifying on the first confirmed failure, since waiting a second run is too costly.
Your release manager wants a single go/no-go score built from several quality indicators. Would you build one? If so, what goes into it, how do you normalise and weight it, and how could it mislead or be gamed?
Sample Answer
Direct answer
Yes, but not as one blended number that decides alone. I would build a small set of hard gates (any failure blocks release) plus an advisory score shown next to the underlying indicators. A single weighted average lets a strong indicator hide a serious problem, and once a number decides a launch, people start optimising the number.
What goes in (5 to 6 indicators, each answering a different question)
- Pass rate of the release-blocking test set (are the checks that must pass passing?).
- Open high-severity defects, weighted by age (any known serious problem?).
- Coverage of the lines changed this release (did we test what we touched?).
- Share of unstable (flaky) tests (can we trust the results?).
- Escape trend from the last releases (are things reaching customers?).
Normalise and weight
Convert each indicator to 0 to 1 between a floor (worst tolerable) and a target, clamped at both ends (a value worse than the floor is cut to 0, better than the target is cut to 1, so no indicator can score below 0 or above 1):
score=target−floorvalue−floor
Plain reading: 0 means as bad as you will tolerate, 1 means on target. Weights are set by risk before the release, written down, reviewed quarterly, never tuned per release. A weighted sum then gives 0 to 1.
Worked example (illustrative)
| Indicator | Value | Floor / target | Normalised | Weight |
|---|---|---|---|---|
| Blocking-suite pass rate | 97% | 90% / 100% | 0.70 | 0.4 |
| Changed-code coverage | 72% | 50% / 80% | 0.73 | 0.3 |
| Open high-severity defects | 2 | 5 (score 0) / 0 (score 1) | 0.60 | 0.3 |
The other two indicators use the same formula. For an indicator where lower is better the floor is the bigger number, and the formula still works. Flaky share: floor 10% (score 0), target 1% (score 1), observed 4% gives (4 - 10) / (1 - 10) = 0.67. Escape trend (customer-found defects per release over the last releases): floor 8 (score 0), target 2 (score 1), observed 4 gives (4 - 8) / (2 - 8) = 0.67. A blocking pass rate of 85% against a floor of 90% would be (85 - 90) / (100 - 90) = -0.5, which clamping cuts to 0. The table above shows three indicators to keep the arithmetic short; with all five you would rebalance the weights so they still sum to 1.
Score = 0.4 x 0.70 + 0.3 x 0.733 + 0.3 x 0.60 = 0.28 + 0.22 + 0.18 = 0.68. With a go threshold of 0.75, that is a no-go. Now someone adds trivial tests that lift changed-code coverage to 80% with no drop in real risk: the score becomes 0.28 + 0.30 + 0.18 = 0.76, which is a go. Nothing about the product changed. Worse, a score of 0.76 could coexist with one open critical bug, because the average forgives it.
Where it misleads or gets gamed
- Averages compensate: excellent scores elsewhere hide one critical defect.
- Goodhart's law (a measure that becomes a target stops measuring): coverage padded with assertion-free tests, defects closed as won't fix.
- False precision: with the go threshold of 0.75 used above, 0.74 versus 0.76 is within the noise of the inputs, but it flips the decision.
- Stale weights that no longer match where the product actually breaks.
Guards: veto gates (hard pass/fail rules that block a release no matter how good the score is), the raw indicators always shown beside the score, a trend over releases, and periodic audit of how defects were closed and tests were written.
Explicit gates for a weekly-delivered SaaS product (starting points to calibrate against your own history)
- Zero open critical (P1, the highest priority: must-fix) defects.
- The release-blocking suite is green after triaged retries (a failed test is rerun automatically, and a repeat failure is looked at by a person rather than retried until it passes); unstable tests are on a visible quarantine list (tests that still run but no longer block, each with an owner and a deadline), not silently skipped.
- No open high-severity defect in an area changed this week without an approved mitigation (a documented workaround or safeguard that reduces the risk, such as a feature switched off by default).
- Changed-code coverage above an agreed floor (for example 70%), with review that new tests assert behaviour.
Exceptions: a waiver (a formal, recorded permission to ship despite a failed gate) that names the risk, approved by the release manager and the product owner of the affected area, expiring by the next weekly release, and logged. Review the count of waivers monthly; a rising count means thresholds or staffing need attention. In practice the gates are wired into the delivery pipeline as a required check that blocks the deploy step until it passes or a waiver is recorded; that wiring is engineering work outside this answer, which defines what the gates are and who may waive them.
Tell me about a time you built or improved a test-reporting dashboard or automated summary. Which metrics did you put on it, which did you leave off, and what changed as a result?
Sample Answer
Direct answer
I turned a nightly "wall of red" into a one-screen summary that showed failures grouped by root cause instead of one row per failing test. The result was that people reviewed a handful of causes each morning instead of hundreds of individual failures, and the team started fixing causes instead of re-running builds. (Below is a story skeleton; swap in your own real numbers, but keep the shape.)
Situation and task
- Our CI (continuous integration, the automated build-and-test system) ran about 2,000 automated tests overnight. The default report was a flat list of failing test names, so on a bad night 240 tests failed and nobody could tell whether that was 240 problems or one.
- I was asked to make the report usable for QA, developers and the engineering manager, who each wanted something different.
Action: metrics I put on it (each tied to a decision)
- Failures grouped by signature (a signature is a short fingerprint of a failure's error and location, so identical failures group together). Decision: which cause to fix first.
- Pass rate on the main branch only (the trunk, the shared branch releases are cut from), trended over 14 days. Decision: is the trunk healthy enough to release from?
- Flake rate per test (a flaky test is one that passes and fails on the same code; rate = runs whose result differs from the previous run divided by total runs). Decision: which tests to quarantine (temporarily stop them from blocking the build while they keep running) or repair.
- Time from first red build to first fix. Decision: is triage keeping up?
Metrics I deliberately left off
- Total test count and code coverage percentage (they go up whether or not quality does).
- Overall pass rate across every branch (draft branches drown the trunk signal).
- Per-person failure counts (it turns a data tool into a blame tool and invites people to game it).
Technical decision
I added a signature column to the results table (one row per test per run) and a second table linking a signature to its ticket. The dashboard's top panel was a simple query grouping failed rows by signature, with a count of distinct tests and builds. That schema choice, not the chart styling, is what made the rest possible. Here is that query run on a tiny sample (7 result rows, builds 41 and 42, the ticket table linking a signature to its ticket):
import sqlite3
db = sqlite3.connect(":memory:")
db.executescript("""
CREATE TABLE results (test TEXT, status TEXT, signature TEXT, build INTEGER);
CREATE TABLE signature_ticket (signature TEXT PRIMARY KEY, ticket TEXT);
""")
db.executemany("INSERT INTO results VALUES (?,?,?,?)", [
("test_login", "fail", "expired-credential", 41),
("test_profile", "fail", "expired-credential", 41),
("test_export", "fail", "expired-credential", 41),
("test_login", "fail", "expired-credential", 42),
("test_search", "fail", "timeout-search-db", 42),
("test_cart_tax", "fail", "assert-rounding", 42),
("test_billing", "pass", None, 42),
])
db.execute("INSERT INTO signature_ticket VALUES ('expired-credential', 'QA-311')")
query = """
SELECT r.signature,
COUNT(*) AS failed_results,
COUNT(DISTINCT r.test) AS distinct_tests,
COUNT(DISTINCT r.build) AS builds,
COALESCE(t.ticket, 'none yet') AS ticket
FROM results r
LEFT JOIN signature_ticket t ON t.signature = r.signature
WHERE r.status = 'fail'
GROUP BY r.signature
ORDER BY failed_results DESC, r.signature
"""
for row in db.execute(query):
print(row)
('expired-credential', 4, 3, 2, 'QA-311')
('assert-rounding', 1, 1, 1, 'none yet')
('timeout-search-db', 1, 1, 1, 'none yet')
Six failed rows became three lines: GROUP BY signature collapses rows with the same signature, and the counts show how big each cause is. The same query over the real 240 failed rows is what produced the 9 signatures below.
Result (derived from the example data, not a benchmark)
On the first night we ran it, 240 failing results collapsed into 9 signatures, and the top signature (one expired test credential) explained 150 of them, so 240 items to read became 9. Beyond that, the 15-minute weekly review shifted from arguing about which tests to blame to assigning owners to signatures.
Trade-offs and pitfalls
- Over-grouping can hide a second, unrelated bug behind a large signature, so I kept a drill-down to individual tests.
- A metric nobody acts on is decoration. I kept a metric only if I could name the decision it feeds.
- Be ready to say what you would do differently (for example, I would have interviewed the consumers before choosing metrics, not after).
You must present an unfavourable quality report to stakeholders: escapes have doubled, automation is not improving, and time to fix has grown. How do you frame the conversation to drive action rather than blame, and what do you propose?
Sample Answer
Direct answer
I would open with the shared goal and the evidence, present the three findings as one connected story about how the system let this happen (not about a person), and close with a specific, prioritised plan and a clear request for what I need from the room. I avoid blame by asking what made this the easy outcome, and by never presenting a number without its context.
Terms. A quarantined test is a flaky (unreliable) test moved out of the pass/fail gate so it stops blocking builds, but it also stops protecting you, so a long quarantine list is hidden risk. A red build is a CI run that failed. Time to green is how long a build stays red before it passes again. A triage rota is a rotating duty roster of who investigates failing builds each week. Plateaued means the number has flattened, no longer improving. A release cohort means grouping bugs by the release they shipped in. A sprint is a team's fixed work cycle, often two weeks.
Before the meeting
- Validate the data. Escapes are bugs that reached production. "Escapes doubled" could mean 4 to 8 or 40 to 80, and it matters whether releases doubled too. Example: 4 escapes across 10 releases is 0.4 per release; 8 across 10 is 0.8, a real doubling; 8 across 20 is 0.4 per release, unchanged, so the story would be volume, not quality.
- Pre-wire the owners of the affected areas (talk to them privately beforehand so nobody is surprised in public), and bring one concrete customer-impact example.
The three-number slide (illustrative numbers)
| Measure | Last quarter | This quarter | Reading |
|---|---|---|---|
| Escapes per 10 releases | 4 | 8 | Doubled, same release volume |
| Automated test count | 1,200 | 1,230 | Flat, plateaued |
| Median time to fix (report to fix) | 3 days | 5 days | 67% slower (5 / 3 is about 1.67) |
Add one line of context: 35 tests are quarantined, up from 12, and 22 of them sit on the checkout and payment paths.
Opening (say it plainly)
"Escapes are up, our automation has plateaued, and fixes take longer. I want to show how these connect, agree what to do first, and decide how we will know it is working. This is about how our process behaves, not who is at fault."
The connected story
- Escapes are up because coverage gaps and quarantined tests let defects pass the stages that should catch them.
- Automation is not improving because the team's time goes to keeping existing tests alive, so new coverage is not added.
- Time to fix has grown because failures lack clear owners and take long to triage, so bugs sit longer.
Each link is a hypothesis with evidence I bring, and I say which parts I am unsure about.
Prioritised remediation plan
| Priority | Action | Why first |
|---|---|---|
| 1 | Review every severe escape to find which stage should have caught it, and add a regression test for each fix | Stops the same bug escaping twice; a regression test is one that guards against a specific past bug returning; cheap |
| 2 | Fix or delete quarantined tests on critical paths, with an owner and a deadline for each (target: the 22 critical-path ones resolved within 60 days, total under 15) | Restores trust in the signal |
| 3 | Assign clear owners and a triage rota for red builds (target: median time to green under 2 hours) | Shortens time to fix |
| 4 | Automate the highest-risk paths that currently rely on manual checks | Raises real coverage |
What I ask the room for: protected time for priority 1 to 3 (for example a fixed share of each team's sprint), and one named sponsor for owners of red builds.
How progress is measured
- Leading indicators (visible in weeks): number and age of quarantined tests, time to green after a red build, share of severe escapes with a regression test.
- Lagging indicators (visible in months): escapes per release, counted by release cohort, and time from report to fix.
- A 30, 60 and 90 day review, reported even when the numbers are bad. Expect escapes to lag the leading measures.
Handling pushback
If someone says "QA should catch this", I respond that testing is one stage of several, and show the escape review data. If someone wants a single culprit, I redirect to what change would prevent a repeat.
Trade-offs and pitfalls
- Overpromising a quick recovery destroys credibility. Promise the leading indicators and the review cadence.
- A plan with ten priorities is a plan with none. Lead with the first two, and keep the rest visible but not funded yet.
- If leaders will not protect any capacity, say so plainly and name what quality risk they are accepting.
Unlock Full Question Bank
Get access to all Quality Metrics and Test Reporting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.