Quality Metrics and Test Reporting Questions
Defining, computing, and communicating software quality. Covers choosing meaningful quality and test metrics (defect escape rate, defect detection effectiveness, defect density, MTTD/MTTR, pass rate, regression frequency, automation suite health and maintenance cost) versus vanity numbers; baselining, trend interpretation (real change versus normal variation), and alert thresholds; dashboards, weekly stability reports, and release-quality reports for engineering, product, and executive audiences, including composite go/no-go scores; framing unfavourable results; guarding against gamed metrics and reading what numbers such as code coverage or a high pass rate hide; investigating contradictory or shifting metrics and testing whether a quality signal really predicts customer outcomes; metric definitions, ownership, and governance; computing metrics from test-run and bug-tracker data (SQL and scripts); designing test-results reporting pipelines, storage schemas, real-time versus batch reporting, alerting, and failure fingerprinting and grouping for triage; and tying quality signals to product and business outcomes. Deciding what to automate and diagnosing individual flaky tests are covered elsewhere.
Design the top-level view of a release-quality dashboard read by product, QA and engineering. Which tiles earn a place, how does each show its signal, what guardrail marks it as a concern, and what drill-downs do engineers get?
Sample Answer
Direct answer
Six tiles on one screen (read by a product manager, QA and engineers): a release verdict, critical escapes, gating-test health, recovery speed, change failure rate and open defects. Each tile shows one number, a small trend line (sparkline) and a colour, and each has a guardrail (the line that turns it amber or red) and a drill-down for engineers. Coverage and raw test counts stay out of the top level; they live one click down as context.
Structured elaboration
The tiles
| Tile | Signal shown | Calculation, source, refresh | Concern marker (starting values, tune to your baseline) | Engineer drill-down |
|---|---|---|---|---|
| Release verdict | Green / amber / red plus the list of failed criteria | Rolls up the tiles below; refreshed each pipeline run | Any red criterion makes the verdict red | Which criterion failed and since when |
| Critical escapes | Count in the last 30 days, with a 12-month trend | Severity-1 defects found in production, from the defect tracker; refreshed daily | Any escape newer than the last release goes red | List with cause, owner, regression test link |
| Gating-suite health (the tests that must pass before a release can ship) | First-attempt pass rate and flaky rate (share of tests with mixed results on the same code) | Test runner results, before retries; every run | First-attempt pass rate more than 2 percentage points below the team's trailing 8-week median, or flaky rate rising 3 weeks running | Failing tests grouped by area and by error message |
| Recovery speed | Median time to detect (MTTD, how long a defect is live before it is noticed) and time to restore (MTTR, how long from noticing to users working normally) | Incident timestamps: detected minus started, restored minus detected; daily | Median restore over 4 hours, or median detect over 1 hour (example values; set from your own baseline) | Timeline of each incident |
| Change failure rate | Share of deployments needing immediate intervention such as a hotfix (an urgent unplanned fix) or rollback, per DORA (the DevOps Research and Assessment research programme; its current guide lists five metrics: change lead time, deployment frequency, failed deployment recovery time, change fail rate and deployment rework rate, replacing the older four-key set) | Failed deployments divided by deployments over 4 weeks; weekly | Above the previous quarter's median | Which changes and which service |
| Open defects | Count by severity and age | Tracker query filtered to the release; hourly | Any open severity-1, or any severity-2 older than 5 working days (example values) | Defect list, assignee, linked commits |
DORA's current name for the recovery metric is failed deployment recovery time (its older name was MTTR); the tile above uses incident-level detect and restore times, a close cousin. Unplanned deployments caused by production incidents are counted separately by DORA, as deployment rework rate.
How each shows its signal. Number first, sparkline second, colour third, and every tile states its window ("last 30 days") and last-refresh time, so a stale tile cannot pass for current.
One alert. A severity-1 defect opened against the release candidate (the build proposed for release) posts to the release channel and pages the QA lead (sends an urgent phone alert). Everything else stays passive on the screen so the alert keeps meaning something.
What one tile looks like
+------------------------------------+
| Critical escapes (last 30 days)|
| 1 [sparkline: 0 0 1 0 0 1] |
| RED: 1 newer than last release |
| updated 09:05 | click: escape list |
+------------------------------------+
Big number on top, sparkline (a tiny trend line with no axes) beside it, one line saying why it has its colour, then window, refresh time and the drill-down link at the bottom.
Worked example
Suppose today the screen reads: 1 critical escape found after the last release (red), first-attempt pass rate inside the team's normal band (green), flaky rate steady (green), median restore time inside target (green), change failure rate 7.5 percent, which is 3 failed deployments out of 40 (green against a previous median of 10 percent), and no open severity-1 defects (green). The verdict tile is red with one line: "Critical escape has no regression test." An engineer clicks the escapes tile, sees the defect, its owner and the missing test link, and the fix and test are added; the tile then turns green on the next refresh.
Trade-offs and pitfalls
- Fewer tiles beat more: product wants a verdict, QA wants test health, engineering wants causes, and six tiles let each find theirs.
- Never show pass rate without flaky rate beside it, or a retry-heavy suite looks healthy.
- The guardrail numbers here are starting points; calibrate each to your own baseline rather than copying them.
What is defect detection effectiveness (also called defect removal efficiency), how do you calculate it for a release or sprint, and why is it a poor sole basis for a release decision?
Sample Answer
Direct answer
Defect detection effectiveness (DDE), also called defect removal efficiency (DRE), is the share of all defects for a release that were caught before release. It measures how good your pre-release testing was, but it is a poor sole basis for a release decision because it looks backward, ignores severity, and can be inflated by logging trivial bugs.
Formula
DDE = defects found before release / (defects found before release + defects found after release) x 100
Intuition: of every 100 bugs the release ever had, how many did testing stop?
Worked example
A sprint's release had 45 defects found in testing. In the 30 days after release, users and monitoring found 5 more that were traced to that release.
DDE = 45 / (45 + 5) = 45 / 50 = 90%
Data sources and the lag problem
- Pre-release defects: bug tracker entries linked to the release, plus failed test runs that led to a fix.
- Post-release defects: bug tracker entries and support tickets traced back to the release, matched via commits or fix-version fields.
- Lag: late-found defects keep arriving. At 30 days the number is 90%. If 3 more surface by day 60, then 45 / (45 + 8) = 45 / 53 = about 84.9%. The score for a release falls over time, so always state the window (and compare releases only at the same age).
Why it is a poor sole basis for a release decision
- Backward-looking. It scores a release that has already happened; you cannot compute it before shipping.
- Ignores severity. Catching 45 cosmetic bugs and missing one payment bug still yields 90%.
- Depends on effort and definitions. Hunting harder for easy bugs raises it; recording duplicates or trivial findings inflates it. A release with 2 found and 0 escaped scores 100% while covering almost nothing.
- Denominator is unknowable in real time. The bugs you have not found yet are not in it.
- Says nothing about untested areas. No escapes may mean no users in that area yet.
What to use with it
Pair it with severity-weighted escapes, release-gate test status, and coverage of changed code. Use DDE as a trend across releases, not a pass/fail gate.
Which metrics tell you whether an automated suite is healthy? How would you measure each one, what trend would worry you, and what would you do about it?
Sample Answer
Direct answer
A healthy suite is one the team trusts, that gives fast feedback, that catches real defects, and that costs a sustainable amount to maintain. No single number shows all four, so I track a small set of metrics, one group per question, and for each I know how it is measured, which trend worries me, and the one action it triggers. (CI means continuous integration, the automated system that builds and tests every change. A flaky test is one that passes and fails on the same code. Quarantine means moving an unreliable test out of the gate that blocks merges. p95 is the value 95% of observations fall below, so it describes the slow tail rather than the average, and p50 is the median. Test rot is a suite decaying as the product changes and tests stop matching it. Escaped defects are bugs found after release. Wall-clock time is elapsed real time, not CPU time. To shard and parallelise is to split the suite into pieces that run at once on several machines. A runner is the machine that executes CI jobs. A triage rota is a rotating on-call duty to look at red builds. Hard-coded waits are fixed sleeps, such as waiting 5 seconds for a page, instead of waiting for a condition. If you are starting from nothing, begin with pass rate, flake rate and duration; the cost and deflection rows are organisation-level extras.)
The metrics, grouped by the question they answer
Reading the table by question: can I trust the suite (pass rate, flake rate, quarantine size and age); is it fast (duration and queue time, time to green); does it find bugs (defect escape rate); is it affordable (maintenance cost per test, manual-test deflection).
| Metric | How I measure it | Trend that worries me | What I do |
|---|---|---|---|
| Pass rate on the main branch | passed / (passed + failed) per day, executed tests only, with skipped and quarantined counts shown beside it | Sliding for weeks, or a flat 100% while production bugs rise | Split real regressions from test rot before reacting |
| Flake rate | Share of executions on unchanged code that flip result on retry (this row only measures the rate; diagnosing why an individual test flakes is a separate investigation) | Rising, or the same tests flipping week after week | Quarantine with an owner and an expiry date, then fix or delete |
| Duration and queue time | p50 and p95 wall-clock of the full CI run, plus time waiting for a runner | p95 growing faster than test count, or people merging without waiting | Shard and parallelise, delete redundant tests, move slow ones to nightly |
| Time to green | Median time from first red build on main to the next green one | Growing, or red builds with no owner | Route failures to owners, run a triage rota |
| Defect escape rate | Escaped defects / (escaped + caught before release), per release | Flat or rising while pass rate looks good | Attribute each escape to the stage that should have caught it, add a test there |
| Quarantine size and age | Count of quarantined tests and the age of the oldest | Count only grows, entries older than the policy limit | Enforce expiry: fix or delete |
| Maintenance cost per test | Engineer hours spent fixing or updating tests per month / number of tests | Cost per test rising | Remove low-value tests, fix brittle patterns (for example, hard-coded waits) |
| Manual-test deflection | Automated cases / (automated + manual regression cases), and hours of manual regression per release | Flat while release effort stays high | Automate the manual checks repeated most often |
Targets, cadence and the organisation view
- I set targets against each team's own baseline first (for example, "no quarantined test older than a fixed number of days"), not a universal number borrowed from another company.
- The team reviews a one-page view weekly; an engineering manager reviews the trends monthly with the same definitions rolled up across teams. The organisation dashboard adds maintenance cost per test and manual-test deflection, because those are the numbers that show whether automation is paying for itself.
Worked example
A team has 1,200 tests. A quarter ago it spent 42 engineer-hours a month on test upkeep. Now it spends 60.
- Then: 42 h x 60 min / 1,200 tests = 2.1 minutes per test per month.
- Now: 60 h x 60 min / 1,200 tests = 3.0 minutes per test per month, about 43% more (60 / 42 = 1.43).
- Meanwhile pass rate stayed at 98% and quarantine grew from 10 to 45 tests. The pass rate alone says "fine"; the cost and quarantine metrics say the suite is quietly decaying. The action is to review the 45 quarantined tests, delete or fix them by their expiry date, and look at which tests take the most upkeep.
Defect escape rate traced: a release had 15 defects, 12 caught before release and 3 escaped, so 3 / (3 + 12) = 20%. If pass rate stayed 98% while that rate went from 20% to 30% over two releases, the suite is passing yet missing more.
Trade-offs and pitfalls
- Any single metric can be gamed (skipping tests lifts pass rate, deleting tests lifts everything), so I pair each with a counter-metric, a second number that exposes the cheat: pass rate with skipped and quarantined counts, duration with escape rate (a suite made fast by deleting tests will show escapes rising).
- Too many metrics means none is acted on. Each one must map to a decision; if it never changes a decision I drop it.
- Raw line coverage is deliberately absent: it says code was executed, not that behaviour was checked. I prefer coverage of critical user paths.
What is defect escape rate, and how do you calculate it? Last release 120 defects were logged: 30 found by QA before release, 70 by customers and 20 by operations afterwards. Work out the rate, say how it relates to defect leakage, and what you would conclude about risk and where to focus testing.
Sample Answer
Direct answer
Defect escape rate is the share of a release's defects that got past your pre-release testing and were found afterwards. It is defects found after release divided by all defects found for that release. For this release: 70 (customers) + 20 (operations) = 90 escaped out of 120 total, so the escape rate is 90 / 120 = 75%. That is high: only a quarter of known defects were caught before release.
Calculation, step by step
- Caught before release by QA (quality assurance testing): 30
- Escaped (found after release): 70 by customers + 20 by operations = 90
- Total logged for the release: 30 + 90 = 120
- Escape rate: 90 / 120 = 0.75, or 75%
- Its mirror, the defect detection percentage (share caught before release): 30 / 120 = 25%
- As a ratio of escaped to caught: 90 / 30 = 3.0, so three defects escaped for every one caught
Escape versus leakage
Teams use the two words loosely, so state your definition. A common split: escape means the defect reached production (users or operations); leakage means a defect slipped past the stage that should have caught it (for example found in system test when a unit test should have caught it), which you can measure at every stage. Every escape is a leak, but not every leak escapes. With this release's 120 defects: the 90 escapes reached production, so the escape rate is 90 / 120 = 75%. Now suppose 10 of the 30 defects QA caught were found in system test although a unit test should have caught them (an illustrative split). Those 10 leaked past the unit-test stage, yet they did not escape. Leakage from the unit stage is therefore 10 defects (a third of the 30 caught before release), while the escape rate is unchanged at 75%. Some organisations use leakage as a synonym for escape, so confirm the local convention before comparing numbers.
What I would conclude about risk
The known risk is high, but I would qualify it before acting:
- Severity: 85 cosmetic escapes and 5 critical ones are very different stories, even though both make 90 escapes (85 + 5). Split the rate by severity.
- Who found them: 70 by customers is customer-visible pain; 20 by operations may point to configuration (settings differing between environments) or environment gaps that QA tests do not reach, so treat that group separately.
- Under-counting: customers do not report everything, so 75% is a floor for customer-visible issues.
- Window: fix the measurement window (for example the first 30 days after release) so releases compare fairly.
Where to focus testing
Do not "test more everywhere". Take the 90 escapes and group them by component (a module or part of the product, such as checkout or search), severity and root cause: no test existed, a test existed but missed it, or the environment differed. If, for instance, most severe escapes cluster in one component, add tests and review effort there first. For the 20 operations-found defects, look at environment parity (making the test environment match production in versions, data and settings), configuration testing (checking the settings that change between environments) and monitoring.
Pitfalls
The rate is gamed by not logging defects, closing them as "works as intended", or downgrading severity; cross-check with support tickets and incident counts.
The nightly pass rate fell from 98 to 92 percent after a batch of UI tests was added. How do you investigate, what evidence separates real product failures from unreliable tests or environment problems, and how do you keep confidence in the pipeline meanwhile?
Sample Answer
Direct answer
Do not assume the product got worse or that the new tests are bad. First split the drop into the old tests and the new ones, then classify every failure as a real regression, a flaky test, an environment problem or a test-data problem, using reruns and clustering as evidence. Meanwhile keep the original suite as the gate and run the new UI batch as non-blocking, so the pipeline stays trusted.
Structured elaboration
How I investigate
- Freeze the facts: which commit, which run, which runner. Rerun the same commit to see what is stable. A runner is the machine that executes the CI (continuous integration, the automated build-and-test pipeline) jobs.
- Split failures by cohort: old suite versus new UI batch, and by first-failure time.
- Rerun failures in isolation and on a clean runner: deterministic or intermittent?
- Cluster by error signature, runner and time window. An error signature is the first line of the failure message with volatile bits (ids, timestamps) removed, so identical causes group together.
- Bisect (search commits for the first bad one) if the old tests dropped.
Evidence that separates the causes
| Cause | Evidence | First action |
|---|---|---|
| New regression in the product | Fails on every rerun with the same assertion on every runner; reproduces by hand or through the API; tracks to a recent commit | File a defect, block if in a critical path |
| Unreliable (flaky: passes and fails on the same code) test | Passes on rerun with no change; varying messages (timeouts, element not found, stale element meaning the page redrew after the test grabbed a button); order-dependent; failure rate steady regardless of commits | Quarantine (park the test outside the gating run so it cannot block merges, but keep it running and visible) with an owner and expiry, then fix the test |
| Environment problem | Failures cluster on one runner, time window or dependency; one message such as connection refused or 503; disappear on another runner | Fix or replace the environment, then rerun |
| Test data problem | Shared data changed by parallel tests; results depend on order or time of day; cleared by a data reset | Isolate data per test |
Keeping confidence in the pipeline meanwhile
- Keep the original suite as the gate; its 98 percent baseline is unchanged and still means something.
- Run the new UI batch as non-blocking with a named owner and an expiry date (say two weeks), and publish both pass rates separately every day.
- Track first-attempt pass rate; silent auto-retries hide flakiness and should never be the reported number.
- Admit a new test to the gate only after a run of consecutive green nights (the length is a team convention, for example 20).
Worked example
A quick sum shows why the split matters. Say there were 500 tests with 490 passing (98 percent), and after adding 100 UI tests there are 600 with 552 passing (92 percent).
- If the old tests are unchanged (490 still pass), the new tests passed 552 - 490 = 62 of 100, a 62 percent pass rate. That points at the new tests or their environment.
- If instead all 100 new tests passed, the old ones would be at 552 - 100 = 452 of 500, which is 90.4 percent: a real regression in the old suite.
The same headline number implies opposite investigations, so the cohort split comes first.
Classifying real failures (illustrative)
| Failure message | What a rerun showed | Bucket |
|---|---|---|
Expected total 42.00 but was 40.00 | Same failure on 3 reruns, on 2 runners | Real regression |
Timeout waiting for #pay-button | Passed on rerun, different runner each time | Flaky test |
stale element reference | Passed on rerun, no code change | Flaky test |
connection refused: payments-stub:8080 | 14 failures, all on runner-3, gone on runner-1 | Environment |
duplicate key: user test@example.com | Fails only when two tests run in parallel | Test data |
Non-blocking simply means the batch still runs and reports, but a red result does not stop a merge or release.
Trade-offs and pitfalls
- Blaming the tests reflexively: some failures are real, so triage each one.
- Deleting failing tests, or quarantining with no expiry, turns the quarantine into a graveyard.
- Rerunning until green creates false confidence.
- What would change my call: if the old suite has dropped too, this is no longer a new-test problem, and I would escalate it as a product or environment issue and block the release.
Unlock Full Question Bank
Get access to all Quality Metrics and Test Reporting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.