Quality Metrics and Test Reporting Questions
Defining, computing, and communicating software quality. Covers choosing meaningful quality and test metrics (defect escape rate, defect detection effectiveness, defect density, MTTD/MTTR, pass rate, regression frequency, automation suite health and maintenance cost) versus vanity numbers; baselining, trend interpretation (real change versus normal variation), and alert thresholds; dashboards, weekly stability reports, and release-quality reports for engineering, product, and executive audiences, including composite go/no-go scores; framing unfavourable results; guarding against gamed metrics and reading what numbers such as code coverage or a high pass rate hide; investigating contradictory or shifting metrics and testing whether a quality signal really predicts customer outcomes; metric definitions, ownership, and governance; computing metrics from test-run and bug-tracker data (SQL and scripts); designing test-results reporting pipelines, storage schemas, real-time versus batch reporting, alerting, and failure fingerprinting and grouping for triage; and tying quality signals to product and business outcomes. Deciding what to automate and diagnosing individual flaky tests are covered elsewhere.
Design the top-level view of a release-quality dashboard read by product, QA and engineering. Which tiles earn a place, how does each show its signal, what guardrail marks it as a concern, and what drill-downs do engineers get?
Sample Answer
Direct answer
Six tiles on one screen (read by a product manager, QA and engineers): a release verdict, critical escapes, gating-test health, recovery speed, change failure rate and open defects. Each tile shows one number, a small trend line (sparkline) and a colour, and each has a guardrail (the line that turns it amber or red) and a drill-down for engineers. Coverage and raw test counts stay out of the top level; they live one click down as context.
Structured elaboration
The tiles
| Tile | Signal shown | Calculation, source, refresh | Concern marker (starting values, tune to your baseline) | Engineer drill-down |
|---|---|---|---|---|
| Release verdict | Green / amber / red plus the list of failed criteria | Rolls up the tiles below; refreshed each pipeline run | Any red criterion makes the verdict red | Which criterion failed and since when |
| Critical escapes | Count in the last 30 days, with a 12-month trend | Severity-1 defects found in production, from the defect tracker; refreshed daily | Any escape newer than the last release goes red | List with cause, owner, regression test link |
| Gating-suite health (the tests that must pass before a release can ship) | First-attempt pass rate and flaky rate (share of tests with mixed results on the same code) | Test runner results, before retries; every run | First-attempt pass rate more than 2 percentage points below the team's trailing 8-week median, or flaky rate rising 3 weeks running | Failing tests grouped by area and by error message |
| Recovery speed | Median time to detect (MTTD, how long a defect is live before it is noticed) and time to restore (MTTR, how long from noticing to users working normally) | Incident timestamps: detected minus started, restored minus detected; daily | Median restore over 4 hours, or median detect over 1 hour (example values; set from your own baseline) | Timeline of each incident |
| Change failure rate | Share of deployments needing immediate intervention such as a hotfix (an urgent unplanned fix) or rollback, per DORA (the DevOps Research and Assessment research programme; its current guide lists five metrics: change lead time, deployment frequency, failed deployment recovery time, change fail rate and deployment rework rate, replacing the older four-key set) | Failed deployments divided by deployments over 4 weeks; weekly | Above the previous quarter's median | Which changes and which service |
| Open defects | Count by severity and age | Tracker query filtered to the release; hourly | Any open severity-1, or any severity-2 older than 5 working days (example values) | Defect list, assignee, linked commits |
DORA's current name for the recovery metric is failed deployment recovery time (its older name was MTTR); the tile above uses incident-level detect and restore times, a close cousin. Unplanned deployments caused by production incidents are counted separately by DORA, as deployment rework rate.
How each shows its signal. Number first, sparkline second, colour third, and every tile states its window ("last 30 days") and last-refresh time, so a stale tile cannot pass for current.
One alert. A severity-1 defect opened against the release candidate (the build proposed for release) posts to the release channel and pages the QA lead (sends an urgent phone alert). Everything else stays passive on the screen so the alert keeps meaning something.
What one tile looks like
+------------------------------------+
| Critical escapes (last 30 days)|
| 1 [sparkline: 0 0 1 0 0 1] |
| RED: 1 newer than last release |
| updated 09:05 | click: escape list |
+------------------------------------+
Big number on top, sparkline (a tiny trend line with no axes) beside it, one line saying why it has its colour, then window, refresh time and the drill-down link at the bottom.
Worked example
Suppose today the screen reads: 1 critical escape found after the last release (red), first-attempt pass rate inside the team's normal band (green), flaky rate steady (green), median restore time inside target (green), change failure rate 7.5 percent, which is 3 failed deployments out of 40 (green against a previous median of 10 percent), and no open severity-1 defects (green). The verdict tile is red with one line: "Critical escape has no regression test." An engineer clicks the escapes tile, sees the defect, its owner and the missing test link, and the fix and test are added; the tile then turns green on the next refresh.
Trade-offs and pitfalls
- Fewer tiles beat more: product wants a verdict, QA wants test health, engineering wants causes, and six tiles let each find theirs.
- Never show pass rate without flaky rate beside it, or a retry-heavy suite looks healthy.
- The guardrail numbers here are starting points; calibrate each to your own baseline rather than copying them.
Tell me about a time quality data changed a product or release decision. What did you measure, how did you present it, and what happened?
Sample Answer
Direct answer
Yes. I once recommended delaying a checkout redesign because our own data showed a customer-facing problem that a mostly green test report hid, and the product manager agreed to move the date by three business days. The story below is an illustrative skeleton: use your own real numbers, but keep the shape of what was measured, how the evidence was framed as a decision, and what changed afterwards.
Structured elaboration
A strong answer to this question has five parts, in this order:
- Situation and stakes (what was about to ship, and to whom).
- The specific measures you chose and why they, rather than the headline pass rate, spoke to customer harm.
- A comparison that makes the evidence hard to dismiss (a control group, a rerun, a trend).
- How you framed the decision: recommendation first, options with costs, an owner and a date.
- The result and what stuck, including what would have changed your mind.
The interviewer is listening for delivered quality results (which metric, which comparison, which decision), not for a generic "I built consensus" story.
Worked example
Situation. A checkout redesign was due on a Thursday. The regression suite (the automated tests that re-check old behaviour after each change; a failure there is a regression) was mostly green, and a canary (a release to a small share of real users first, here 5 percent of traffic) had run for a day.
Task. As the quality assurance (QA) lead I owned the quality call in the go/no-go meeting (the meeting where the team decides to ship or hold).
Action: measure. Two failing tests were in the payment path and had been labelled "known flaky" (a flaky test passes and fails on the same code with no change, so teams tend to dismiss its failures) without anyone checking. I reran both ten times on the same commit and they failed every time with the same message, the signature of a deterministic bug (one that happens every time the same steps are run) rather than flakiness. I also compared the canary's checkout error rate with a control group (users still on the old build, used as the yardstick for normal). In this illustrative case it was 1.8 percent against 0.6 percent, a 3x gap (1.8 / 0.6). At an assumed 20,000 checkouts a day, the 1.2-point gap would mean about 240 extra failed checkouts every day at full traffic (20,000 x 0.012).
Action: present. One page, decision first: "Recommend delaying three business days: payment failures are reproducible and canary users hit errors at 3x the control rate." One chart of canary versus control. The page read roughly like this (illustrative):
DECISION: Delay checkout redesign by 3 business days (owner: QA lead, decide by Wed 5pm)
WHY: 2 payment tests fail 10/10 reruns, same error (not flaky).
Canary checkout errors 1.8% vs 0.6% control (3x). Est. ~240 extra failed checkouts/day at full traffic.
OPTIONS: A ship as planned cost: ~240 failed checkouts/day, hotfix pressure
B ship, flag off cost: no launch value, 1 day work
C delay 3 days cost: marketing date moves
CHART: two lines, daily checkout error rate, canary vs control
Three options with costs: ship as planned; ship with the new checkout behind a feature flag (a switch that turns the feature on or off without a new deploy) that stays off; or delay three days. I recommended the delay and offered the flag as the fallback if the date was immovable, so the product manager had a choice instead of a veto.
Result. The team found a race (two operations that sometimes finish in the wrong order and corrupt the outcome) in refreshing the payment token, fixed it and added a regression test that failed before the fix and passed after. The canary error rate returned to the control group's level, the launch slipped three days and no customer saw the failure. We wrote the go/no-go criteria into the release checklist (no unexplained failures in the payment path; canary error rate within an agreed distance of control), so the next call needed less argument.
Trade-offs and pitfalls
- Bring evidence, not alarm. The delay was accepted because it rested on a reproducible test and a comparison group.
- Offer costed options. A bare "no" makes QA the blocker; a choice makes QA the adviser.
- Own the cost out loud. The delay cost a marketing date, and saying so built trust.
- Say what would have changed your call: if ten reruns had passed and the canary gap had been inside normal noise, I would have recommended shipping.
- Avoid the generic version: name the measure, the comparison and the decision.
You must present an unfavourable quality report to stakeholders: escapes have doubled, automation is not improving, and time to fix has grown. How do you frame the conversation to drive action rather than blame, and what do you propose?
Sample Answer
Direct answer
I would open with the shared goal and the evidence, present the three findings as one connected story about how the system let this happen (not about a person), and close with a specific, prioritised plan and a clear request for what I need from the room. I avoid blame by asking what made this the easy outcome, and by never presenting a number without its context.
Terms. A quarantined test is a flaky (unreliable) test moved out of the pass/fail gate so it stops blocking builds, but it also stops protecting you, so a long quarantine list is hidden risk. A red build is a CI run that failed. Time to green is how long a build stays red before it passes again. A triage rota is a rotating duty roster of who investigates failing builds each week. Plateaued means the number has flattened, no longer improving. A release cohort means grouping bugs by the release they shipped in. A sprint is a team's fixed work cycle, often two weeks.
Before the meeting
- Validate the data. Escapes are bugs that reached production. "Escapes doubled" could mean 4 to 8 or 40 to 80, and it matters whether releases doubled too. Example: 4 escapes across 10 releases is 0.4 per release; 8 across 10 is 0.8, a real doubling; 8 across 20 is 0.4 per release, unchanged, so the story would be volume, not quality.
- Pre-wire the owners of the affected areas (talk to them privately beforehand so nobody is surprised in public), and bring one concrete customer-impact example.
The three-number slide (illustrative numbers)
| Measure | Last quarter | This quarter | Reading |
|---|---|---|---|
| Escapes per 10 releases | 4 | 8 | Doubled, same release volume |
| Automated test count | 1,200 | 1,230 | Flat, plateaued |
| Median time to fix (report to fix) | 3 days | 5 days | 67% slower (5 / 3 is about 1.67) |
Add one line of context: 35 tests are quarantined, up from 12, and 22 of them sit on the checkout and payment paths.
Opening (say it plainly)
"Escapes are up, our automation has plateaued, and fixes take longer. I want to show how these connect, agree what to do first, and decide how we will know it is working. This is about how our process behaves, not who is at fault."
The connected story
- Escapes are up because coverage gaps and quarantined tests let defects pass the stages that should catch them.
- Automation is not improving because the team's time goes to keeping existing tests alive, so new coverage is not added.
- Time to fix has grown because failures lack clear owners and take long to triage, so bugs sit longer.
Each link is a hypothesis with evidence I bring, and I say which parts I am unsure about.
Prioritised remediation plan
| Priority | Action | Why first |
|---|---|---|
| 1 | Review every severe escape to find which stage should have caught it, and add a regression test for each fix | Stops the same bug escaping twice; a regression test is one that guards against a specific past bug returning; cheap |
| 2 | Fix or delete quarantined tests on critical paths, with an owner and a deadline for each (target: the 22 critical-path ones resolved within 60 days, total under 15) | Restores trust in the signal |
| 3 | Assign clear owners and a triage rota for red builds (target: median time to green under 2 hours) | Shortens time to fix |
| 4 | Automate the highest-risk paths that currently rely on manual checks | Raises real coverage |
What I ask the room for: protected time for priority 1 to 3 (for example a fixed share of each team's sprint), and one named sponsor for owners of red builds.
How progress is measured
- Leading indicators (visible in weeks): number and age of quarantined tests, time to green after a red build, share of severe escapes with a regression test.
- Lagging indicators (visible in months): escapes per release, counted by release cohort, and time from report to fix.
- A 30, 60 and 90 day review, reported even when the numbers are bad. Expect escapes to lag the leading measures.
Handling pushback
If someone says "QA should catch this", I respond that testing is one stage of several, and show the escape review data. If someone wants a single culprit, I redirect to what change would prevent a repeat.
Trade-offs and pitfalls
- Overpromising a quick recovery destroys credibility. Promise the leading indicators and the review cadence.
- A plan with ten priorities is a plan with none. Lead with the first two, and keep the rest visible but not funded yet.
- If leaders will not protect any capacity, say so plainly and name what quality risk they are accepting.
How would you show that your quality work moves a business outcome such as retention or conversion? What data would you join, what would you compute, and what would you say about correlation versus causation?
Sample Answer
Direct answer
Show it as a chain with a comparison at the end: quality problem, then customer-visible symptom (errors, support contacts), then behaviour (conversion, retention, churn meaning customers leaving). The cleanest evidence is a comparison between users who got the fixed build and users who did not. Be plain about what the data can and cannot prove: quality work correlates with better outcomes, and a controlled comparison makes the causal story credible.
Structured elaboration
What data to join
- Releases and deployments: release id, deploy time, flag or rollout cohort.
- Defects and incidents: introduced-in release, severity, affected feature, time to detect and restore, rollbacks.
- Product analytics events: user id, funnel step (for example checkout completed), which build the user saw.
- Support tickets and churn: contacts per 1,000 users and cancellations, per user id.
- Join at user level by build and feature exposure. Release-level joins have only a handful of rows and are dominated by other things.
A tiny sample of the joined table (the join key is user_id, plus the build the user was served):
| user_id | build | rollout_group | checkout_completed | support_contacts_30d |
|---|---|---|---|---|
| u1001 | 4.2.0 (buggy) | control | 0 | 1 |
| u1002 | 4.2.1 (fixed) | early | 1 | 0 |
| u1003 | 4.2.1 (fixed) | early | 0 | 0 |
Grouping this table by build and averaging checkout_completed gives the conversion rates compared below.
What to compute
- Symptom rates: error rate and support contacts per 1,000 users, per build.
- Outcome difference between cohorts: conversion (share of visitors who complete a purchase) or 30-day retention (share still active after 30 days) for users on each build.
- A confidence interval: the range the true difference plausibly lies in, not just the point estimate.
- Cost of quality: what escapes cost (engineer-hours on triage, hotfix and communication, plus support load) versus what prevention and testing cost. Example: 2 escapes a month at 30 engineer-hours each is 2 x 30 = 60 hours a month (an assumption to replace with your own tracked hours).
Worked example
A fix was rolled out to a randomly chosen half of users first, so half were still on the buggy checkout build and half on the fixed one. Save as rollout_compare.py:
import math
# standard error (se) = how much a measured rate would typically wobble from sampling luck alone
# Fix rolled out to a randomly chosen half of users first: half still on the buggy build, half on the fixed build
n_bug, conv_bug = 5000, 540
n_fix, conv_fix = 5000, 600
p_bug, p_fix = conv_bug / n_bug, conv_fix / n_fix
diff = p_fix - p_bug
se = math.sqrt(p_bug * (1 - p_bug) / n_bug + p_fix * (1 - p_fix) / n_fix)
print("conversion bug build %.1f%%, fixed build %.1f%%" % (100 * p_bug, 100 * p_fix))
print("difference %.2f points, 95%% CI %.2f to %.2f" % (100 * diff, 100 * (diff - 1.96 * se), 100 * (diff + 1.96 * se)))
Output:
conversion bug build 10.8%, fixed build 12.0%
difference 1.20 points, 95% CI -0.05 to 2.45
The interval is built as the difference plus or minus 1.96 standard errors. The standard error is how far a measured rate would typically wander by luck if you repeated the comparison, and 1.96 of them either side covers the middle 95 percent of those luck-driven outcomes, so the interval is the plausible range for the true difference. The fixed build converted 1.20 points better, but the 95 percent interval runs from about -0.05 to 2.45 points, so at 5,000 users per group the evidence is borderline, not conclusive. If 100,000 checkouts start a month, 1.20 points is about 1,200 orders (100,000 x 0.012), with a plausible range of roughly -50 to 2,450. The honest report is "likely helped, needs more traffic or a longer window", not "quality earned us 1,200 orders".
Trade-offs and pitfalls
Correlation versus causation
- The staged, randomized rollout above is the core method; the rest of this section explains why it beats the alternatives, and the last bullet is a fallback you will rarely need.
- Correlating releases with monthly churn is weak: releases differ in size, season and marketing. A confounder is a third factor that moves both the thing you changed and the outcome (for example a holiday sale that lifts conversion and also delays releases). Reverse causation means the arrow points the other way: for example, a team that sees engagement falling may rush out hurried fixes, so the churn drives the release activity rather than the releases driving the churn. (A busy season that pushes both releases and churn is the confounder case above, a different problem.)
- A randomized rollout (chance, not people, decides who gets the fix first) removes those confounders, because both groups face the same season and marketing.
- Without randomization, use matched cohorts (pair each exposed user with an unexposed one of similar plan, country and activity) or difference-in-differences (compare the before-and-after change in exposed users with the change in similar unexposed users) and describe results as "consistent with".
- Never expose users to a known defect just to run an experiment; use rollouts that happen anyway.
Trade-offs
- Argue the chain link by link: escapes to support contacts, contacts to incidents and rollbacks, incidents to churn. Each link is easier to defend than one leap to revenue.
- Report intervals and assumptions with every order or revenue figure.
Near a release deadline you notice a burst of defects closed as won't fix or reclassified to lower severity. How would you check whether numbers are being gamed, what evidence would you look for, and how would you respond without souring trust?
Sample Answer
Direct answer
Treat it first as a question the data can answer, quietly, and only then as a conversation. A burst of "won't fix" closures or severity downgrades before a deadline can be legitimate triage or pressure-driven relabelling; the evidence that separates them is baseline rate, timing, who is doing it, whether the reasons hold up, and what happens after release. I would gather that evidence, then talk to the owners about the pattern (not accuse anyone), and change the process so that deferring a defect is legitimate and visible instead of hidden.
How to check (evidence to collect)
- Baseline: how many closures as won't fix and downgrades happen per day in the same phase of earlier releases? Only a departure from that baseline is a signal.
- Timing: are they clustered in the last 72 hours, after hours, in a single batch?
- Actor concentration: one person or one team doing most of them, versus spread across owners.
- Distribution shift: the share of critical and high defects falling while the total stays roughly constant, without matching code changes.
- Quality of reasons: empty, copy-pasted or "per PM" notes, versus specific ones. Did the original reporter agree, or reopen it?
- Blind re-rating: pick a sample (random plus highest-impact), hide the current severity, and have a second person (QA lead plus a developer) re-rate against the written severity rubric.
- Outcome: after release, do customer-reported issues match the ones that were closed? This is the ultimate check.
Worked example (illustrative)
The usual rate is about 6 downgrades or won't-fix closures per week, roughly 0.86 per day. In the last three days there were 18, which is 6 per day, about seven times the baseline. Fifteen of the 18 were by one owner, 12 had no comment, and the critical share dropped from 20% to 8% of open defects. That is enough to investigate, not to conclude. A blind re-rate of 10 sampled items follows. If the second rater mostly agrees with the new severities, this was overdue clean-up of stale tickets; if most are re-rated back up, the data goes to the engineering manager with the sample and the rubric.
How to respond without souring trust
- Start with the owner privately: "Here is the pattern I see, walk me through your reasoning." Assume good faith until the sample says otherwise.
- Talk about the process gap, not the person: under deadline pressure the only visible options are "fix now" or "close".
- Add statuses that make risk visible: "deferred with owner and date" is separate from "won't fix", and deferred items appear in the release report so the decision is owned by the people who accept the risk.
- Require a reason for severity changes, and require the reporter or QA to acknowledge them.
- Publish the severity rubric so the argument is about criteria.
- If evidence shows deliberate relabelling, escalate to the manager with the sample, not in a public channel.
What would change my call: legitimate reasons exist (duplicate, cannot reproduce, out of scope, feature removed). A high share of these with good notes means the burst was real backlog grooming.
Pitfalls
- Announcing suspicion in a public channel before you have evidence.
- Treating every downgrade as gaming and punishing the testers who raised the defects.
- Fixing the metric (blocking closures) instead of the incentive (deadline pressure with no legitimate way to defer).
Unlock Full Question Bank
Get access to all 16 Quality Metrics and Test Reporting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.