Quality Metrics and Test Reporting Questions
Defining, computing, and communicating software quality. Covers choosing meaningful quality and test metrics (defect escape rate, defect detection effectiveness, defect density, MTTD/MTTR, pass rate, regression frequency, automation suite health and maintenance cost) versus vanity numbers; baselining, trend interpretation (real change versus normal variation), and alert thresholds; dashboards, weekly stability reports, and release-quality reports for engineering, product, and executive audiences, including composite go/no-go scores; framing unfavourable results; guarding against gamed metrics and reading what numbers such as code coverage or a high pass rate hide; investigating contradictory or shifting metrics and testing whether a quality signal really predicts customer outcomes; metric definitions, ownership, and governance; computing metrics from test-run and bug-tracker data (SQL and scripts); designing test-results reporting pipelines, storage schemas, real-time versus batch reporting, alerting, and failure fingerprinting and grouping for triage; and tying quality signals to product and business outcomes. Deciding what to automate and diagnosing individual flaky tests are covered elsewhere.
Tell me about a time you built or improved a test-reporting dashboard or automated summary. Which metrics did you put on it, which did you leave off, and what changed as a result?
Sample Answer
Direct answer
I turned a nightly "wall of red" into a one-screen summary that showed failures grouped by root cause instead of one row per failing test. The result was that people reviewed a handful of causes each morning instead of hundreds of individual failures, and the team started fixing causes instead of re-running builds. (Below is a story skeleton; swap in your own real numbers, but keep the shape.)
Situation and task
- Our CI (continuous integration, the automated build-and-test system) ran about 2,000 automated tests overnight. The default report was a flat list of failing test names, so on a bad night 240 tests failed and nobody could tell whether that was 240 problems or one.
- I was asked to make the report usable for QA, developers and the engineering manager, who each wanted something different.
Action: metrics I put on it (each tied to a decision)
- Failures grouped by signature (a signature is a short fingerprint of a failure's error and location, so identical failures group together). Decision: which cause to fix first.
- Pass rate on the main branch only (the trunk, the shared branch releases are cut from), trended over 14 days. Decision: is the trunk healthy enough to release from?
- Flake rate per test (a flaky test is one that passes and fails on the same code; rate = runs whose result differs from the previous run divided by total runs). Decision: which tests to quarantine (temporarily stop them from blocking the build while they keep running) or repair.
- Time from first red build to first fix. Decision: is triage keeping up?
Metrics I deliberately left off
- Total test count and code coverage percentage (they go up whether or not quality does).
- Overall pass rate across every branch (draft branches drown the trunk signal).
- Per-person failure counts (it turns a data tool into a blame tool and invites people to game it).
Technical decision
I added a signature column to the results table (one row per test per run) and a second table linking a signature to its ticket. The dashboard's top panel was a simple query grouping failed rows by signature, with a count of distinct tests and builds. That schema choice, not the chart styling, is what made the rest possible. Here is that query run on a tiny sample (7 result rows, builds 41 and 42, the ticket table linking a signature to its ticket):
import sqlite3
db = sqlite3.connect(":memory:")
db.executescript("""
CREATE TABLE results (test TEXT, status TEXT, signature TEXT, build INTEGER);
CREATE TABLE signature_ticket (signature TEXT PRIMARY KEY, ticket TEXT);
""")
db.executemany("INSERT INTO results VALUES (?,?,?,?)", [
("test_login", "fail", "expired-credential", 41),
("test_profile", "fail", "expired-credential", 41),
("test_export", "fail", "expired-credential", 41),
("test_login", "fail", "expired-credential", 42),
("test_search", "fail", "timeout-search-db", 42),
("test_cart_tax", "fail", "assert-rounding", 42),
("test_billing", "pass", None, 42),
])
db.execute("INSERT INTO signature_ticket VALUES ('expired-credential', 'QA-311')")
query = """
SELECT r.signature,
COUNT(*) AS failed_results,
COUNT(DISTINCT r.test) AS distinct_tests,
COUNT(DISTINCT r.build) AS builds,
COALESCE(t.ticket, 'none yet') AS ticket
FROM results r
LEFT JOIN signature_ticket t ON t.signature = r.signature
WHERE r.status = 'fail'
GROUP BY r.signature
ORDER BY failed_results DESC, r.signature
"""
for row in db.execute(query):
print(row)
('expired-credential', 4, 3, 2, 'QA-311')
('assert-rounding', 1, 1, 1, 'none yet')
('timeout-search-db', 1, 1, 1, 'none yet')
Six failed rows became three lines: GROUP BY signature collapses rows with the same signature, and the counts show how big each cause is. The same query over the real 240 failed rows is what produced the 9 signatures below.
Result (derived from the example data, not a benchmark)
On the first night we ran it, 240 failing results collapsed into 9 signatures, and the top signature (one expired test credential) explained 150 of them, so 240 items to read became 9. Beyond that, the 15-minute weekly review shifted from arguing about which tests to blame to assigning owners to signatures.
Trade-offs and pitfalls
- Over-grouping can hide a second, unrelated bug behind a large signature, so I kept a drill-down to individual tests.
- A metric nobody acts on is decoration. I kept a metric only if I could name the decision it feeds.
- Be ready to say what you would do differently (for example, I would have interviewed the consumers before choosing metrics, not after).
Several teams report quality numbers and they disagree on what pass rate or escape rate even means. How would you set up ownership, definitions and review so the numbers become trusted, and how would you settle disputes between teams?
Sample Answer
Direct answer
Numbers disagree first because the words mean different things, so I would fix definitions before debating results. I would name one accountable owner for each metric, publish a versioned definition for it, compute it centrally from raw data instead of accepting team-reported figures, review it on a fixed rhythm, and settle disputes with a written procedure that ends in a decision by a named person.
Terms. Pass rate is the share of tests that passed. Escape rate is the share of a release's defects that were found only after release (in production): for example 10 escaped out of 100 total defects is 10%. A metric dictionary is a shared document that defines each metric in one place. A data steward is the person who keeps the underlying data pipeline correct.
Show the problem with real numbers. One run: 900 tests passed, 30 failed, 50 skipped, 20 quarantined, and 25 of the 900 passes needed a retry.
- Team A (executed tests only): 900 / 930 = 96.77%.
- Team B (skipped and quarantined in the denominator): 900 / 1,000 = 90.0%.
- Team C (first-attempt passes only, so the 25 retried passes do not count): 875 / 930 = 94.09%.
Same data, three honest but different answers. Starting the meeting with this makes it about definitions, not blame.
1. Ownership. Each metric has one accountable owner (usually the SDET (software development engineer in test) or quality-platform lead) and a data steward for its pipeline. In a RACI sense (a table of who does the work, who answers for it, who gets asked first, and who is just told): the platform team is responsible for computing it, the metric owner is accountable for the definition, team leads are consulted on changes, and stakeholders are informed.
2. Definitions. Keep a metric dictionary in version control. Each entry states: name, numerator, denominator, window, what is excluded, source of truth, owner and version. Changes go through review, get a version number, and are back-computed, meaning history is recalculated under the new definition, so trends stay comparable (show old and new side by side for a cycle).
Example entry for pass rate: name pass_rate; numerator: tests whose final attempt passed (900); denominator: tests executed, so passed plus failed, excluding skipped and quarantined (930); window: one CI run; exclusions: skipped and quarantined tests, reported separately; source of truth: the central results store; owner: quality-platform lead; version: 1.0. Plugging in the numbers gives 900 / 930 = 96.77%, and every team's report must use exactly this.
Example for escape rate: 40 defects attributed to release 5.2, of which 6 were first found in production. Escape rate = 6 / 40 = 15%. Its dictionary entry states the same fields plus the measurement window (for example 30 days after release) and which severities count.
3. Central computation. The platform computes every team's numbers from raw results with the same code. Teams may add local views, but only the shared definition appears in cross-team reports.
4. Review and reliability.
- A monthly quality review looks at trends, definition changes and requests to alter alert thresholds (threshold changes need approval, so alerts are not quietly loosened).
- Monitor the metrics themselves: the share of CI runs ingested, data freshness, and the share of bugs unmapped to a release (bugs that cannot be tied to the release they belong to, so they cannot enter an escape rate). A metric with holes in its data loses trust faster than a metric with a debatable definition.
5. Settling disputes.
- Check mechanically: apply the published definition to the data. Most disputes end here.
- If the definition is ambiguous, the metric owner decides within an agreed time, records the decision, and updates the dictionary.
- If it is a real trade-off (for example whether to count a legacy suite), escalate to the engineering manager or product sponsor. Once decided, everyone follows it: disagree and commit means voicing objections openly first, then supporting the decision once it is made.
6. Onboarding a new team. Provide the results format, map suites to owners, backfill 30 days, and run a shadow period (a trial stretch where numbers are shown but not used for decisions) showing the team's own figure and the standard figure side by side, so differences are explained before they are contested.
7. Keeping the numbers honest when capacity is tight: test maintenance versus bug fixes. Definitions only stay trusted if the tests and data feeding them stay healthy, so the same governance covers how effort is split. Agree an explicit capacity split (for example a fixed share of each sprint for test health) and revisit it quarterly. Within it, first fix test problems that make the signal untrustworthy, such as quarantined tests guarding areas with recent escapes. Customer-impacting bugs take priority over cosmetic test cleanup.
Trade-offs and pitfalls
- A central team that dictates definitions without consulting teams gets compliance, not trust. Involve leads when the dictionary is drafted.
- If the underlying data is unreliable, fix ingestion first; better definitions cannot repair missing data.
- Too many metrics defeats the exercise. Standardise a few (pass rate, escape rate meaning the share of defects found only in production, and time to fix) before adding more.
You run a large automation suite whose maintenance cost keeps climbing. How would you estimate the yearly maintenance cost of an individual test, and how would you decide when to retire one?
Sample Answer
Direct answer
Estimate a test's yearly maintenance cost as the human time its failures consume (including rot: a test breaking because the product changed even though no bug exists) plus what it costs to run, and compare that with what the test protects. Retire a test only when it is expensive, has caught nothing real, and its behaviour is covered by other tests. Otherwise fix, demote (move it from the every-commit suite to a slower nightly one) or quarantine (keep running it but stop it blocking builds) it.
The estimate
Cyear=F×H×R+N×E
Where F is failures per year needing human attention (real bugs, false alarms and breakages from product changes), H is hours of triage plus repair per failure, R is the loaded hourly cost of an engineer (salary plus overhead such as benefits and equipment, divided by working hours), N is runs per year and E is the cost per run. In words: how often it bothers people, times what each interruption costs, plus compute.
Add a criticality weight W to judge whether the cost is worth paying. It is a judgement your team sets, not a measured number: 1 for a low-risk feature, 2 for medium, 4 for a business-critical flow such as payment. It scales how much upkeep you will tolerate. The base budget (here 4 hours a year for a weight-1 test) is likewise a policy choice, roughly "what a low-value test may cost before someone looks at it"; tolerable hours = base budget x W, so a payment test is allowed 4 times as much.
Worked example (illustrative rates: 80 USD per engineer-hour, 0.02 USD per run, 3 runs a day, base budget of 4 hours a year)
| Test A: legacy export screen | Test B: payment flow | |
|---|---|---|
| F | 24 failures, all false alarms or rot | 2 failures, 1 a real bug |
| H | 1.5 hours | 1.5 hours |
| Human cost | 24 x 1.5 x 80 = 2,880 | 2 x 1.5 x 80 = 240 |
| Run cost | 1,095 runs x 0.02 = 21.90 | 21.90 |
| Total per year | 2,901.90 USD | 261.90 USD |
| Weight W | 1 | 4 |
| Tolerable hours (4 x W) | 4 | 16 |
| Actual hours | 36 | 3 |
The dollar totals and the hours say the same thing (36 hours x 80 USD = 2,880 USD of the 2,901.90), so I decide in hours, which are easier to compare with a budget, and use the dollars when I need to justify the work to someone who thinks in money. Test A uses 36 hours against a tolerable 4 and has caught no real defect, so it is a retirement candidate. Test B uses 3 hours against 16 and caught a real bug, so keep it.
Deciding what to do with Test A
- Is the behaviour covered at a cheaper level (unit or API test)? If yes, retire it.
- If not, and the feature matters: rewrite it. A one-off rewrite of 8 hours (640 USD) against roughly 2,880 USD a year in upkeep pays back in about 2.7 months, if the rewrite removes essentially all of the upkeep. If it only cuts failures by 75%, the saving is 0.75 x 2,880 = 2,160 USD a year and payback is 640 / 2,160 = 0.3 years, about 3.6 months. Either way it is a strong case, but state the assumption.
- If it is flaky (passes and fails on identical code), quarantine it (run it, but do not block on it) with an owner and a deadline.
- If the feature itself is being removed, retire the test.
Where the inputs come from: continuous integration (CI) run history gives F and N, commits touching the test file and the fixes that follow give an approximate H (sample with real time tracking to calibrate), and the bug tracker shows whether the test ever caught a real defect. Treat the result as an estimate to rank tests, not an accounting figure.
Pitfalls
- Deleting a test because it is annoying rather than because it is low value. Flaky is not the same as useless.
- Ignoring hidden costs such as waiting for slow suites and lost trust in red builds.
- Leaving the criticality weight subjective and unreviewed, so it turns into "my tests are critical".
- Counting only the cost. A test that is cheap but never catches anything still costs run time and reading time.
You join a mature product that has no historical quality metrics. Lay out your first 90 days: which metrics you collect first, how you baseline them, who you involve, and what you do about the ones you cannot yet trust.
Sample Answer
Direct answer
In the first 90 days I would not build a dashboard. I would listen and inventory data in month 1, collect a small set of metrics and baseline them in month 2, and publish an honest baseline (with a trust label on each number) plus two targets in month 3. Metrics I cannot yet trust go into a "provisional" list (numbers shown with a warning label until their source is verified as complete) with a named fix, and they stay off leadership views until they pass an audit.
Days 1-30: listen and inventory
- Talk to the people who feel quality problems: engineering lead, product manager, on-call or site reliability engineers (SREs), support lead, and QA.
- Ask each: what is the last quality problem that hurt, and how did you find out?
- Inventory sources: bug tracker, CI results, incident records, support tickets, release notes. Note which are complete and which are not.
Days 31-60: collect a starting set and baseline
Start with about five metrics, each answering a different question. If time is short, start with the first two rows (escape rate and CI reliability) and the bug-ticket count; they need the least new instrumentation and the others can follow in month 3. An escape is a defect that got past pre-release testing and was found in production.
| Metric | Definition | Question |
|---|---|---|
| Defect escape rate | Production defects / (defects found before release + production defects), per release | How much does testing miss? |
| Main-branch CI pass rate and flaky-failure share | Passing runs / runs; failures that pass on rerun / failures | Is the automated signal reliable? |
| Change failure rate (called change fail rate by DORA, the DevOps Research and Assessment programme; it is one of DORA's five current delivery metrics, originally one of four, a widely used research-based set) | Deployments that cause a production problem / deployments | Are releases safe? |
| Mean time to detect and restore | Average delay before noticing; average time to recover | How quickly do we react? |
| Support tickets tagged as bugs, per week | Count from the support tool | What do customers feel? |
Baseline (the starting reference value you compare later changes against) by backfilling (reconstructing past weeks from records you already have, such as old CI runs and tickets) 8-12 weeks of history where it exists, and report the median and range, not one number. Example: the weekly flaky-failure share over 10 weeks was 4, 6, 5, 9, 5, 6, 4, 5, 7, 5 percent. Sorted, that is 4, 4, 5, 5, 5, 5, 6, 6, 7, 9; the middle two values are both 5, so the median is 5 percent and the range is 4 to 9 percent. The mean is 5.6, nudged up by the single 9 week, which is why the median is the steadier headline. Do not set targets until the baseline is stable.
Who I involve
Engineering manager (sponsor), tech leads (they will own fixes), product manager (context on severity and priorities), support lead (customer signal), and SRE (incident data). Agree the definitions with them so nobody can later dismiss the numbers as "your definitions".
Metrics I cannot trust yet
Test each with an audit. The link rate is the share of bug-tagged support tickets that link to a tracker bug. Example (illustrative): the bug tracker shows 40 production bugs last quarter, and support logged 65 bug-tagged tickets. Several tickets can describe the same bug, so 65 against 40 is not itself a gap; the real question is how many tickets have any tracker bug at all. I pick 20 tickets at random (a small sample, so the result is rough) and find 14 correctly linked to a tracker bug: 14 / 20 = a 70% link rate. About 3 in 10 customer-reported bugs are therefore missing from the tracker, so its escape rate undercounts. So: label the tracker's escape rate "provisional (about 70% coverage)", fix the process (a required field on bug creation for "found in production"), re-audit in 4 weeks, and only then publish it as trusted.
Days 61-90: publish and choose
Publish the baseline with trust labels, pick two improvement targets tied to a real pain (for example, halve the flaky-failure share; raise the link rate from 70% to 90% by making the found-in-production field required), and set a monthly review.
Trade-offs and pitfalls
- Too many metrics early produces a dashboard nobody trusts. Fewer, audited numbers beat many shaky ones.
- Metrics used to judge individuals get gamed; keep them at team or product level.
- Say what would change the plan: if a production incident occurs in week 3, start from the incident data first.
Across three releases, defect density fell steadily while customer-reported incidents rose sharply. How would you investigate the contradiction, and what might each explanation change about how you measure quality?
Sample Answer
Direct answer
Two things can both be true, so the first job is to find out which measurement is misleading. Defect density is defects per unit of code (usually per thousand lines, KLOC), so it falls whenever the code grows faster than the defect count. Customer-reported incidents measure what customers felt. I would investigate whether the numerator, the denominator, or the detection process changed, and classify the incidents by root cause before touching the testing strategy.
Worked example (illustrative numbers)
| Release | Defects | Code (KLOC) | Density | Incidents | Customers | Incidents per 1,000 customers |
|---|---|---|---|---|---|---|
| 1 | 50 | 100 | 0.50 | 10 | 4,000 | 2.5 |
| 2 | 55 | 130 | 0.42 | 18 | 5,000 | 3.6 |
| 3 | 60 | 200 | 0.30 | 30 | 6,000 | 5.0 |
Density fell 40% (0.50 to 0.30) while defects went up (50 to 60) because the code doubled. Incidents tripled and, even per customer, doubled (2.5 to 5.0). The dashboard said "better", customers said "worse", and no one is lying.
Terms used below, in plain words
- Root cause: the underlying reason an incident happened (for example a bad configuration value), as opposed to the symptom customers saw.
- Postmortem: the written review after an incident describing what happened and why.
- Happy path: the normal, everything-goes-right route through a feature; tests that only cover it miss odd inputs and failures.
- Environment parity: how closely the test environment matches production (same data shape, configuration, versions, scale).
- Telemetry gap: something goes wrong in production but no log, error report or metric records it.
- Error tracking: a tool that collects production exceptions automatically.
- Test-to-incident traceability: being able to link each incident to the test that should have caught it (or to the fact that none exists).
- Version-control churn: how much code was changed in a period, taken from commit history.
Which I would check first: the denominator and the incident classification, because both use data already in hand and take an hour or two. Detection and telemetry come next (they need comparing sources), and coverage blind spots and definition drift last (they need sampling and manual review). Rank by what the evidence shows, not by which theory is most appealing.
Hypotheses, the evidence to check, and what each would change in how I measure
| Hypothesis | How to check | What it changes |
|---|---|---|
| Denominator growth hides absolute defects | Plot raw counts and density on changed code only | Measure density on changed code, plus incidents per user |
| Incidents are not code defects (configuration, infrastructure, data, third-party) | Classify every incident by root cause from postmortems and tickets | Track incident rate by cause, not just by defect |
| Detection got worse: fewer bugs found internally | Compare share found before release versus in production per release | Track escape rate at fixed age |
| Coverage blind spots: tests run in an environment unlike production, happy paths only | Map each incident to the test that should have caught it; check test data and environment differences | Track test-to-incident traceability and environment parity |
| Telemetry gap: customers hit issues we cannot see, or support intake improved so more gets reported | Compare incident reports against error tracking and logs; check intake changes | Add production monitoring as a detection source |
| Severity or definition drift | Re-rate a sample of old defects by today's rubric | Freeze definitions, log any change |
High automation rate yet production defects rising: automation percentage says how many test cases run by machine, not how much behaviour they verify. Check whether automated tests assert meaningful outcomes, cover integrations, configuration and unusual data, run against realistic environments, and whether failing tests are ignored as noise.
Data sources: incident tickets and postmortems, support tickets, production error tracking, release and deploy history, test management and continuous integration (CI, the automated build-and-test system) results, version-control churn, customer counts.
Corrective actions (chosen from the findings): add tests where incidents actually clustered, make defect and incident classification mandatory, add production checks for the classes of failure tests cannot reach, and change the headline metric to something customers would agree with.
Pitfalls
- Fixing testing before finding whether the incidents were code defects at all.
- Trusting one metric because it is easy to compute.
- Explaining the gap with a single theory. Several can apply at once, so rank by evidence.
Unlock Full Question Bank
Get access to all 32 Quality Metrics and Test Reporting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.