Quality Metrics and Test Reporting Questions
Defining, computing, and communicating software quality. Covers choosing meaningful quality and test metrics (defect escape rate, defect detection effectiveness, defect density, MTTD/MTTR, pass rate, regression frequency, automation suite health and maintenance cost) versus vanity numbers; baselining, trend interpretation (real change versus normal variation), and alert thresholds; dashboards, weekly stability reports, and release-quality reports for engineering, product, and executive audiences, including composite go/no-go scores; framing unfavourable results; guarding against gamed metrics and reading what numbers such as code coverage or a high pass rate hide; investigating contradictory or shifting metrics and testing whether a quality signal really predicts customer outcomes; metric definitions, ownership, and governance; computing metrics from test-run and bug-tracker data (SQL and scripts); designing test-results reporting pipelines, storage schemas, real-time versus batch reporting, alerting, and failure fingerprinting and grouping for triage; and tying quality signals to product and business outcomes. Deciding what to automate and diagnosing individual flaky tests are covered elsewhere.
Pass rate for a large suite has drifted downward over several weeks. How would you decide whether that is a real change in quality or ordinary variation, and how would you have it flagged automatically next time?
Sample Answer
Direct answer
I would not react to the slope. First I would check that the measurement itself did not change, then compare the drop with the normal day-to-day wobble of this suite, then slice the drop to see whether it is concentrated in a few tests, one component or one date. Only then do I call it real. For next time, I turn that comparison into an automatic alert built on the suite's own history.
Step 1: did what we measure change?
- Was the set of tests different (new tests added that fail more, old ones deleted, more skipped or quarantined, meaning moved out of the gate that blocks merges)? Recompute the rate on the tests that existed at the start of the window. A falling overall rate with a flat rate on that fixed set means the mix changed, not the quality.
- Did the definition, retry policy, runner image or test data change? Infrastructure changes often show up as a spread-out drop across many unrelated suites.
Step 2: is it more than ordinary variation?
Measure the normal wobble from a stable baseline: the mean and standard deviation (sd, how far daily values typically stray from the mean) of the daily pass rate. Treat the day or the run, not the individual test, as the unit, because failures are clumped: one broken dependency can fail hundreds of tests at once, so a textbook binomial error bar (the range you would expect if every test failed independently, like coin flips) would be far too narrow. Adjust for known patterns: compare weekdays with weekdays if daily builds differ at weekends, pool sparse tests over a longer window instead of judging them daily, and use rates rather than counts when suite sizes differ.
Step 3: slice it. A genuine quality loss usually concentrates: a handful of tests, one owning team, one component, or one merge date. Then reproduce a sample of the new failures on a clean environment. Real bugs reproduce; test rot and infrastructure trouble often do not.
Step 4: automate the flag. This is a control chart approach: plot the series against limits computed from its own history and alarm when it leaves them. If you adopt only one rule first, take 8 in a row below the mean, then add the 3 sd limit for crashes. Three simple rules on the daily series, each with a different trade-off:
| Rule | Catches | Weakness |
|---|---|---|
| One day beyond mean minus 3 sd (sigma is just another name for sd; 3 is a convention that leaves very few false alarms when data is roughly bell-shaped) | Sudden drops | Slow to notice a gradual slide |
| 8 consecutive days below the mean | Sustained small shifts | Blind to a one-day crash if used alone |
| CUSUM (cumulative sum: adds up each day's shortfall and alarms when the total passes a threshold) | Small persistent drifts, early | Needs tuning of its allowance and threshold |
Change-point methods, which estimate when a shift began, are rarely needed in interviews and are useful after an alarm to date the change against releases; pairing that with the rules above is enough for most teams.
Worked example (synthetic data, seed pinned; run it unchanged)
import random
from statistics import mean, stdev
rng = random.Random(7)
# Synthetic daily pass rates (percent): 28 stable baseline days, then 21 days of slow decline.
baseline = [rng.gauss(97.0, 0.6) for _ in range(28)]
recent = [rng.gauss(97.0 - 0.06 * d, 0.6) for d in range(1, 22)]
series = baseline + recent
mu, sd = mean(baseline), stdev(baseline)
lower = mu - 3 * sd
print(f"baseline mean={mu:.2f} sd={sd:.2f} lower limit={lower:.2f}")
# Rule 1: any single day below the 3-sigma lower limit
first_limit = next((i for i, x in enumerate(series) if x < lower), None)
# Rule 2: 8 consecutive days below the baseline mean
first_run = None
streak = 0
for i, x in enumerate(series):
streak = streak + 1 if x < mu else 0
if streak == 8:
first_run = i
break
# Rule 3: one-sided CUSUM (allowance k = 0.5 sd, decision threshold h = 4 sd)
k, h, s, first_cusum = 0.5 * sd, 4 * sd, 0.0, None
for i, x in enumerate(series):
s = min(0.0, s + (x - mu) + k)
if s < -h:
first_cusum = i
break
def day(i):
return "none" if i is None else f"day {i + 1}"
print("drift starts: day 29")
print("3-sigma limit alarm:", day(first_limit))
print("8-in-a-row below mean:", day(first_run))
print("CUSUM alarm:", day(first_cusum))
baseline mean=96.98 sd=0.51 lower limit=95.46
drift starts: day 29
3-sigma limit alarm: day 44
8-in-a-row below mean: day 38
CUSUM alarm: day 38
How the CUSUM line reads: each day it adds (today's value minus the mean, plus the allowance k). Days near the mean add about zero, and the min(0.0, ...) keeps the total from drifting upward, so only shortfalls accumulate. When the running total drops below minus h, the shortfall has persisted long enough to alarm. k = 0.5 sd ignores tiny wobble; h = 4 sd is a common starting point, tuned on past data.
Days 1 to 28 are a stable baseline; from day 29 the true rate slides by 0.06 points a day. None of the rules fired during the baseline. The 8-in-a-row rule and CUSUM alarm on day 38, roughly 0.6 points below baseline, while the strict 3-sd limit needs until day 44. This is one synthetic series, so the exact days illustrate the trade-off, not a promise.
Trade-offs and pitfalls
- Tighter rules alert sooner but cry wolf more, and an alert nobody trusts is ignored. Tune on past data and review alert precision monthly (the share of alerts that turned out to be real problems; for example 3 real out of 10 alerts is 30 percent, too noisy).
- Sending the alert with the slice (which tests, which team, first bad date) makes it actionable; a bare percentage prompts a meeting.
- The method assumes the past is a fair baseline. After a deliberate change, such as a suite migration, reset the baseline.
A release dashboard shows: pass rate 96 percent, automated coverage 42 percent, flaky rate 6 percent, 2 critical escapes in the last 30 days, time to detect 4 hours, time to restore 36 hours. The product manager asks whether the release is ready. How do you read it, what else do you need, and what do you recommend?
Sample Answer
Direct answer
As shown, not ready: I would recommend a conditional no-go, meaning hold the release until three specific checks pass, then ship it gradually. The process numbers (pass rate, coverage) look tolerable, but the two numbers closest to customer harm are the weak ones: 2 critical escapes (severity-1 defects, meaning outage or data loss, found in production after release) in 30 days, and a 36-hour time to restore.
Structured elaboration
How I read each number
| Metric | What it tells me | What it hides / question to ask |
|---|---|---|
| Pass rate 96% | 4% of executed tests failed | Which tests, in which areas? A 96% made mostly of stable old tests says little about the code that changed in this release |
| Automated coverage 42% | Under half of whatever is being counted is exercised by automation | The denominator: requirements, code lines, or user flows? For example, 42% of 20,000 lines is 8,400 lines run by tests, but if the 11,600 untested lines include payment code, the number is worse than it looks; 42% of 100 user flows is 42 flows, and which 42 matters |
| Flaky rate 6% (share of tests that both pass and fail on the same code) | The suite is partly untrustworthy | Together with the pass rate, it means failures cannot be called real or fake without a rerun |
| 2 critical escapes in 30 days | The net let serious defects through recently | Root cause, and whether each fix came with a regression test |
| Time to detect 4h, time to restore 36h | Quick to notice, slow to recover. Time to detect (MTTD) is how long a defect is live before someone notices; time to restore (MTTR) is how long from noticing until users are working normally again | Here I assume the restore clock starts at detection, so each escape means about 40 hours (4 + 36) from the defect going live to customers being back to normal. If your team starts the restore clock at the defect's start instead, 36 already includes the 4. Ask which. Either way, any escape in this release is expensive |
What else I need
- The failing tests, filtered to areas this release touched.
- Open defects by severity.
- Whether a fast rollback or feature flag exists. A working rollback or flag usually restores service in minutes to an hour, so a 36-hour restore suggests fixes are being built, tested and deployed from scratch. That is an inference to confirm, not a fact.
- The trend for each number, not only the snapshot: is 2 escapes better or worse than last quarter?
- The customer impact of the change: a payment feature and a colour change carry different risk.
Recommendation for the product manager (the go/no-go). Go only if all three hold by the agreed date:
- Every failing test is classified as real, flaky or environment, with no open real failure in changed areas.
- Both critical escapes are fixed and each has a regression test (a test that fails if that same defect returns).
- The release ships behind a feature flag (a switch that turns the change on or off without a new deploy) or staged (for example 5 percent, then 25, then 100) with a rollback (return to the previous version) rehearsed and a restore target of under an hour.
Otherwise hold. I would not make the coverage number a blocker: 42 percent is a risk to monitor, not a reason to stop this release.
Worked example
Why pass rate alone misleads: on 1,000 tests a 96% pass rate is 40 failures. If 6% (60 tests) are flaky, those 40 failures could be entirely noise or entirely real, and the dashboard cannot tell you which. Rerunning the 40 failures three times each and seeing, say, 30 pass on rerun and 10 fail every time would leave 10 real suspects to investigate, which is a far smaller and more honest number than "4 percent failing".
Trade-offs and pitfalls
- I would not say "go" on a 96 percent pass rate, nor a flat "no" without conditions, which only stalls the team.
- If the failures triage as flaky or environment and rollback is proven fast, I would say go with the staged plan.
- If a single real failure is in the payment or data-loss path, it is a hard no regardless of the rest.
- Do not try to fix all flakiness before release; fix the flakiness that blocks the decision.
You report the same quality numbers to team leads, product managers and executives. How do you change what you show, the level of detail and the form of visualisation for each audience? Use one metric as a worked example.
Sample Answer
Direct answer
Keep one source of truth and one definition, and change only the zoom: team leads get detail they can act on this week, product managers get impact and trend tied to their plan, executives get one number, its direction and the decision needed. I will use defect escape rate (the share of a release's defects found only after release) as the worked metric.
Worked example data (illustrative)
Last four releases: escaped defects out of all defects found: R1 18 of 60 (30%), R2 15 of 50 (30%), R3 9 of 45 (20%), R4 6 of 40 (15%). In R4, the 6 escapes were checkout 4, search 1, profile 1.
What each audience sees
| Audience | Detail level | Form of visualisation | What it says |
|---|---|---|---|
| Team lead | Per component and per defect, with root-cause tag and link | Table of the 6 R4 escapes; bar chart by component | "Checkout has 4 of 6; three lacked a test for the refund path, so add these tests this sprint" |
| Product manager | Per feature area and severity, per release | Trend line over R1 to R4 plus a stacked bar by severity | "Escapes fell from 30% to 15%; the remaining ones sit on the checkout flow, a revenue path, so scope and dates need test time there" |
| Executive | One metric, direction, target, one sentence, one ask | Small trend line with the four points labelled, one status colour plus a text label | "Escape rate halved over four releases (30% to 15%). One area still drives most escapes. We ask for two weeks of test investment there before the next major launch" |
Principles
- Same numbers, different zoom: never change values between audiences, only how much detail sits under them.
- Every view answers "so what?": leads see what to fix, product managers see what it means for the plan, executives see what decision is needed.
- Show denominators for small samples: R4's 15% is 6 of 40; one more escape moves it by 2.5 points, so avoid over-precise claims.
- Choose forms deliberately: tables for action, line charts for trend, stacked bars for composition. Avoid pie charts and 3D, and never rely on red and green alone (add a text label so colour-blind readers get the status).
- Fewer metrics upward: an executive page with 20 metrics is not read; keep to three or four and link down for detail.
Pitfalls
Do not simplify to the point of hiding severity (15% of trivial defects versus 15% of critical ones), and do not present a single good-looking figure without the context that would let a non-technical reader challenge it.
In Python, write a function that turns a raw stack trace into a stable fingerprint so the same underlying failure groups together across runs. Say what you strip out, what you must preserve, and what over-normalising would cost you.
Sample Answer
Direct answer
A stable fingerprint hashes only the parts of a stack trace that identify the failure (the exception type plus the last few frames' file names and function names) and throws away the parts that change from run to run (line numbers, absolute paths, memory addresses, IDs and numbers in the message). Two runs of the same bug then hash to the same value, and a different bug hashes differently.
What you strip out (it varies run to run)
- Absolute path prefixes (
/home/ci/build-4711/...): keep only the file's basename. - Line numbers: they shift whenever anyone edits the file above them.
- Memory addresses (
0x7f3a9c), UUIDs, timestamps, and numeric IDs in the message. - Quoted values in the message (
'user_918'becomes'<str>').
What you must preserve (it identifies the bug)
- The exception type (
TimeoutErrorversusKeyErrorare different bugs). - Function names and file basenames of the frames, in order.
- Only the innermost few frames (the ones nearest the raise), because the outer frames are usually test-runner plumbing that is identical for every failure.
- Optionally the cleaned message, when the same code path can fail for different reasons.
What the code does, before you read it
FRAMEandEXCare regular expressions (patterns for finding text).FRAMEfinds eachFile "...", line N, in funcline and captures the path and function name;EXCfinds the last line that looks likeSomethingError: message.clean_messagerewrites the volatile parts of the message: addresses to<addr>, UUIDs to<uuid>, quoted strings to'<str>', digits to<n>.fingerprintjoins the exception type and the last 3file:functionpairs, then hashes that text with SHA-1 (a hash function that turns any text into a fixed-length code; the same text always gives the same code, and we keep the first 12 characters). Chained exceptions are theraise ... fromcase, where one error is raised while handling another, so the trace holds several tracebacks.similarityis the optional softer alternative described below.
Code (runnable as one file; output below is exactly what it prints)
import hashlib
import re
import sqlite3
from difflib import SequenceMatcher
FRAME = re.compile(r'File "(?P<path>[^"]+)", line \d+, in (?P<func>\S+)')
EXC = re.compile(r'^(?P<type>[A-Za-z_][\w.]*(?:Error|Exception|Failure|Exit|Timeout))\b:?\s*(?P<msg>.*)$')
def clean_message(msg):
msg = re.sub(r'0x[0-9a-fA-F]+', '<addr>', msg)
msg = re.sub(r'[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}', '<uuid>', msg)
msg = re.sub(r"'[^']*'", "'<str>'", msg)
msg = re.sub(r'\d+', '<n>', msg)
return msg.strip()
def frames_and_exception(trace):
frames = [(m['path'].replace('\\', '/').rsplit('/', 1)[-1], m['func'])
for m in FRAME.finditer(trace)]
exc_type, msg = 'UnknownError', ''
for line in reversed(trace.strip().splitlines()):
m = EXC.match(line.strip())
if m:
exc_type, msg = m['type'], m['msg']
break
return frames, exc_type, msg
def fingerprint(trace, top_frames=3, keep_message=False):
frames, exc_type, msg = frames_and_exception(trace)
# innermost frames are the last ones printed; keep the deepest few
parts = [exc_type] + [f'{f}:{fn}' for f, fn in frames[-top_frames:]]
if keep_message:
parts.append(clean_message(msg))
return hashlib.sha1('|'.join(parts).encode()).hexdigest()[:12]
def similarity(trace_a, trace_b):
fa, ea, _ = frames_and_exception(trace_a)
fb, eb, _ = frames_and_exception(trace_b)
seq = SequenceMatcher(None, fa, fb).ratio()
return round(0.7 * seq + 0.3 * (ea == eb), 2)
RUN1 = '''Traceback (most recent call last):
File "/home/ci/build-4711/tests/test_cart.py", line 42, in test_checkout
cart.pay(order)
File "/home/ci/build-4711/app/cart.py", line 88, in pay
gateway.charge(order.id, 0x7f3a9c)
File "/home/ci/build-4711/app/gateway.py", line 17, in charge
raise TimeoutError("charge 9001 timed out after 30s")
TimeoutError: charge 9001 timed out after 30s
'''
RUN2 = RUN1.replace('build-4711', 'build-4802').replace('line 88', 'line 91').replace('9001', '9377').replace('0x7f3a9c', '0x55d1e0')
OTHER = RUN1.replace('line 17, in charge', 'line 23, in refund').replace('charge 9001', 'refund 9001')
DIFFMSG = RUN1.replace('TimeoutError("charge 9001 timed out after 30s")', 'TimeoutError("connect to db-7 refused")').replace('TimeoutError: charge 9001 timed out after 30s', 'TimeoutError: connect to db-7 refused')
if __name__ == '__main__':
print('run1 vs run2 same:', fingerprint(RUN1) == fingerprint(RUN2))
print('run1 vs other (different function):', fingerprint(RUN1) == fingerprint(OTHER))
print('run1', fingerprint(RUN1), 'run2', fingerprint(RUN2))
print('same frames, different message, message ignored:', fingerprint(RUN1) == fingerprint(DIFFMSG))
print('same frames, different message, message kept:',
fingerprint(RUN1, keep_message=True) == fingerprint(DIFFMSG, keep_message=True))
print('message kept but cleaned, run1 vs run2:',
fingerprint(RUN1, keep_message=True) == fingerprint(RUN2, keep_message=True))
print('similarity run1/run2:', similarity(RUN1, RUN2))
print('similarity run1/other:', similarity(RUN1, OTHER))
db = sqlite3.connect(':memory:')
db.execute('create table results (test text, status text, signature text, build int)')
rows = [('test_checkout', 'fail', 'a1', 1), ('test_checkout', 'fail', 'a1', 2),
('test_refund', 'fail', 'a1', 2), ('test_login', 'fail', 'b7', 2),
('test_login', 'fail', 'b7', 3), ('test_search', 'pass', None, 3)]
db.executemany('insert into results values (?,?,?,?)', rows)
q = '''select signature, count(*) as failures, count(distinct test) as tests, count(distinct build) as builds
from results where status = 'fail' group by signature order by failures desc, signature'''
for r in db.execute(q):
print(r)
run1 vs run2 same: True
run1 vs other (different function): False
run1 407602b05015 run2 407602b05015
same frames, different message, message ignored: True
same frames, different message, message kept: False
message kept but cleaned, run1 vs run2: True
similarity run1/run2: 1.0
similarity run1/other: 0.77
('a1', 3, 2, 2)
('b7', 2, 1, 2)
Reading the output
- Run 1 and run 2 differ in build number, line number, ID and address, yet share a fingerprint: the normalising worked.
- Run 1 and the "other" trace differ in one function (
chargeversusrefund), so they get different fingerprints. - Ignoring the message merged a "charge timed out" trace with a "db refused" trace that share frames. That merge is the cost of over-normalising. Keeping the raw message would have the opposite failure: every run splits because of the ID inside it. Keeping the cleaned message is the middle path.
Two optional alternatives
- Similarity score instead of exact match.
similarity()above blends how much the frame sequences overlap (difflib's SequenceMatcher ratio) with whether the exception types match. How the 0.77 for run1 versus therefundtrace is derived: the frame lists aretest_cart:test_checkout, cart:pay, gateway:chargeversus..., gateway:refund, so 2 of 3 frames match. SequenceMatcher's ratio is 2 x matches / total items in both lists = 2 x 2 / 6 = 0.667. The exception types are equal, so the second term is 1. Score = 0.7 x 0.667 + 0.3 x 1 = 0.767, printed as 0.77. The 0.7 and 0.3 weights are a judgement call that says frame overlap matters more than the exception type alone, not a standard; tune them on labelled examples. A threshold (for example, group when 0.8 or higher) catches near-duplicates such as a frame inserted by a refactor, at the cost of being harder to explain and to index. - Same idea in SQL. Once each failed row stores its signature, the top failing causes are one grouped query (the second half of the script; its output is the last two lines above).
a1covering two tests across two builds is one cause, not two.
Complexity and edge cases
- Time is linear in the trace length (one regex scan plus one hash); no extra memory beyond the frames list.
- Edge cases: the innermost frames (the ones nearest the line that raised the error) are what we keep; a trace with no frames (fingerprint falls back to the exception type only, which over-merges, so log these separately), chained exceptions (
raise ... from), and multi-language runners (Java frames need a different regex). - Hash collisions in a truncated 12-character SHA-1 are not a practical concern at this scale; if they were, keep the full digest.
What quality and test metrics would you track for a delivery team, and for each one, what does it actually tell you and where can it mislead?
Sample Answer
Direct answer
Track a small balanced set: a few metrics about defects reaching users, a few about how fast and trustworthy the test signal is, and a couple about delivery health. For every metric, write down two things next to it: the decision it informs and the way it lies. A metric you cannot attach a decision to is decoration.
The set, what each tells you, and where it misleads
| Metric | What it tells you | Where it misleads |
|---|---|---|
| Defect escape rate (share of a release's defects found only after release) | How much the test effort is missing before users see it | Falls when customers stop reporting, or when defects get downgraded or never logged. Says nothing about severity unless you split it |
| Test pass rate | Whether the current build is healthy | Rises when tests are deleted, skipped, quarantined (parked outside the blocking set so their failures no longer count), or retried until green (rerun automatically until they happen to pass). A 100% pass rate on a weak suite is comfort, not safety |
| Flaky test rate (tests that pass and fail on the same code) | How much you can trust a red build | Hidden if retries are automatic. High flakiness makes teams ignore real failures |
| Feedback time (push to result) and per-test execution time | How long a developer waits | Averages hide the slow tail (see the worked example) |
| Change failure rate (deployments that cause a failure needing a fix, divided by all deployments) and recovery time (two of the DORA delivery metrics from the DevOps Research and Assessment programme; dora.dev now calls the recovery one failed deployment recovery time, the time to recover from a deployment that fails and needs immediate intervention, which replaced the older MTTR, mean time to restore service, and it also lists a fifth metric, deployment rework rate) | Whether fast delivery is coming at the cost of stability | Small deployments can make the rate look better while impact per failure grows. Averages hide one long outage |
| Open defects by severity and age | Release risk and neglected problems | A low count can mean nobody is looking. Severity labels are subjective |
| Customer-side signals (crash-free sessions, meaning the share of app sessions that end without a crash; support tickets per active user) | Product health as users experience it, the balance to engineering-side numbers | Lags by days or weeks and is affected by usage changes |
Code coverage (the share of code lines that tests execute) is a related signal that is deliberately left out of the table because it measures only whether code was executed, not whether any assertion checked the result (a test with no assertions still raises coverage). Use it as "where we definitely have no tests", never as proof of quality.
Quick numbers for the other rows (illustrative)
- Escape rate: a release has 40 defects logged in total, 30 found before release and 10 found by customers afterwards. Escape rate = 10 / 40 = 25%.
- Pass rate and quarantine: 1,000 tests run, 50 fail, so 950 / 1,000 = 95%. If 30 of the failures are quarantined and excluded, the report shows 950 / 970 = 97.9%. The number rose although nothing was fixed, which is why quarantined counts must be shown beside it.
- Flaky test rate: 12 of 400 tests both passed and failed on the same code this week, so 12 / 400 = 3%.
- Change failure rate: 6 of 40 deployments this month needed a fix or rollback, so 6 / 40 = 15%.
- Crash-free sessions: 9,900 of 10,000 sessions ended without a crash, so 99%. Note that 99% still means 100 crashed sessions.
Worked example: throughput and execution time under parallelism
A suite has 1,200 tests averaging 2 seconds each, so 2,400 seconds of compute. Split across 4 parallel machines (shards) the ideal wall-clock time is 2,400 / 4 = 600 seconds. In practice the shards are uneven: if the slowest takes 900 seconds, the run takes 900 seconds, not 600. Add 5 minutes (300 seconds) waiting in the queue for a free runner and a developer waits 1,200 seconds, 20 minutes, even though "average test time" is still 2 seconds. Retries also distort: if 3% of tests (36) are retried, that adds 36 x 2 = 72 seconds of compute, invisible in the average but visible as extra cost and hidden flakiness. So report feedback time end to end and its 95th percentile (the value 95% of runs come in under), plus retry count, not only average per-test time or tests per hour.
Thresholds and balance
There are no honest universal numbers. Set thresholds against your own baseline and act on direction: escape rate rising two releases in a row, flake rate creeping up, feedback time past what developers will wait for (they start skipping it). Read product and engineering metrics together: fast feedback with rising escapes means the fast tests are the wrong tests; falling escapes with falling customer usage proves nothing.
Gaming and pitfalls
Every row above can be gamed (delete tests, downgrade severity, retry to green). Prefer trends over targets, publish denominators, and never rank individuals with these numbers.
Unlock Full Question Bank
Get access to all Quality Metrics and Test Reporting interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.