Technical Debt Management and Refactoring Questions
Identifying, prioritizing, and paying down technical debt sustainably. Covers recognizing debt, making the case to invest in it, refactoring safely behind tests, and balancing debt reduction against feature velocity. Includes keeping a codebase maintainable over the long term.
Discuss the risks introduced by third-party libraries and pretrained dependencies in a codebase. Propose a governance policy covering version pinning, vulnerability scanning, licensing audits, and an upgrade path that minimizes long-term technical debt.
Sample Answer
Direct answer
Govern third-party and pretrained-model dependencies with four controls: version pinning (no floating versions in production), automated vulnerability scanning on every build, a periodic licensing audit, and a defined upgrade cadence that separates routine updates from urgent security patches, plus a specific process for tracking security-debt items through the existing defect-tracking workflow rather than a separate, easily-forgotten system.
Structured elaboration
The risk profile differs by category and both need naming before the controls make sense. Third-party libraries carry the risks the controls below target directly: known CVEs (Common Vulnerabilities and Exposures, the public ID system used to catalog known security vulnerabilities), transitive-dependency surprises, and license incompatibility, all scannable and patchable with normal tooling. Pretrained model dependencies (embeddings, base LLMs, vision backbones) carry a different, less tractable set: there is usually no CVE feed for a set of weights, so a security or bias issue baked into the training data is invisible to a standard vulnerability scanner; the license often restricts USE (for example a non-commercial or field-of-use clause) rather than redistribution, which a normal package-license scanner will miss; and a version bump can silently change behavior, since a new checkpoint is not guaranteed backward-compatible the way a patch-version library bump is, so an 'upgrade' for a pretrained dependency needs a behavioral check, not just a compatibility-test run.
- Version pinning: every dependency locked to an exact version in a lockfile, so "it worked yesterday" incidents from an unplanned transitive update don't happen; upgrades are always a deliberate, reviewed action.
- Vulnerability scanning: automated scanning (Dependabot, Snyk, or equivalent) on every build, gated by severity: critical/high vulnerabilities block merge or trigger an urgent-patch process; low-severity ones queue into the routine upgrade cadence.
- Licensing audit: a periodic (quarterly) automated scan for license compatibility, since a transitively-introduced GPL dependency in a proprietary codebase is a real, easy-to-miss risk that compounds the longer it's undetected.
- Upgrade cadence: routine dependency bumps on a scheduled cadence (weekly automated PRs for patch/minor versions, reviewed monthly for major versions) separate from the urgent path for critical CVEs, which bypasses the cadence and gets same-week attention.
- Security-debt tracking: integrate into the existing defect tracker rather than a separate spreadsheet, using dedicated fields (CVE ID or advisory link, severity, affected component, discovery date) and a severity-based SLA (critical: patch within days, medium: within the next routine cadence, low: batched quarterly), with a triage cadence (weekly for critical/high, monthly review for the rest) so items don't silently age past their SLA unnoticed.
- Pretrained-dependency-specific governance: extend the licensing audit to check model-card (a short, standardized document describing what a model was trained on, its intended use, and known limitations) and training-dataset license terms, not just package licenses, and extend the vulnerability-scanning control with a behavioral regression suite run against any new checkpoint before it replaces the pinned one in production, since a clean scan of the surrounding code says nothing about whether the checkpoint's own outputs have drifted.
Worked example
A scan flags a critical CVE in a widely-used logging library used across 15 services. Because severity is critical, it bypasses the routine monthly cadence: a ticket is filed with the CVE ID, severity, and affected services fields populated, triaged within the week (per the SLA), and an urgent coordinated upgrade PR is generated for all 15 services simultaneously using the same automated-rollout pattern used for routine dependency updates at scale, rather than 15 separate, uncoordinated fixes.
Trade-offs & pitfalls
The common failure is treating all dependency updates with the same urgency, which either creates alert fatigue (every minor version bump treated as urgent) or, worse, causes a genuinely critical CVE to get lost in a queue of low-priority noise; severity-based triage with a real SLA is what prevents both failure modes.
Design a decision framework for when maintaining an obsolete internal library is no longer worth the effort of learning or continuing to update. Include a cost-benefit analysis, migration-effort estimates, a risk assessment covering incidents and security, the learning cost for engineers, and a phased deprecation strategy with rollback plans.
Sample Answer
Direct answer
Decide to stop maintaining an obsolete internal library when the ongoing learning and migration cost for the team clearly exceeds the cost of a phased deprecation, weighing incident and security risk heavily, since an unmaintained library's risk compounds silently over time even without any new code being written against it.
Structured elaboration
- Cost-benefit analysis: compare the ongoing cost of engineers learning and working around the obsolete library's quirks (real, if hard to measure precisely, in onboarding time and workaround effort) against the one-time cost of migrating off it.
- Migration-effort estimate: count actual call sites and usages across the codebase, and estimate effort per call site based on complexity (a simple 1:1 API swap versus a call site requiring real logic changes), rather than assuming uniform migration cost across all usages.
- Risk assessment: incidents historically traced to this library, and security posture, specifically whether it's still receiving security patches from any source, since an abandoned library with no patches is a compounding, silent risk even if the team never touches that code again.
- Learning cost for engineers: how much onboarding or debugging time is spent specifically on this library's quirks, a real cost even when no incidents have resulted yet.
- Phased deprecation strategy: stop new usage immediately (a lint rule or code-review policy blocking new call sites), migrate existing call sites in priority order (highest-risk or highest-change-frequency first), and set a hard sunset date with a rollback plan for each migrated call site in case the replacement introduces a regression.
Worked example
An internal HTTP client library, unmaintained for 3 years, has 340 call sites across 25 services. A security audit finds it hasn't received a patch for a moderate-severity vulnerability disclosed 18 months ago (real, compounding risk). Migration effort estimate: roughly 200 call sites are simple 1:1 swaps to the standard library (a few hours each with tooling assistance), 100 require moderate rework (custom retry logic built on top of the old client), and 40 are complex enough to need individual review. Given the vulnerability exposure and the growing onboarding cost engineers report, the decision is to deprecate: block new usage immediately, migrate the 200 simple call sites first via a semi-automated codemod within a month, then the 100 moderate ones over the following quarter, and the 40 complex ones on a slower, individually-scoped timeline, with a 9-month full sunset target.
Trade-offs & pitfalls
The risk in phased deprecation is the long tail (the 40 complex call sites here) stalling indefinitely once the easy wins are done and momentum fades; setting a hard sunset date up front, with visible tracking of remaining call sites, is what prevents the deprecation from silently stalling at 85% complete forever.
You must decide whether to rewrite a legacy module or incrementally refactor it. Using a structured decision framework, list the criteria you would evaluate, how you would estimate each, and the decision thresholds you would use to choose one approach over the other.
Sample Answer
Direct answer
Decide with four criteria evaluated explicitly, not by gut feel: measurable risk of the change, customer impact if something goes wrong, effect on ongoing developer productivity, and time to deliver either path. Default to incremental refactor; only choose a full rewrite when the criteria clearly and jointly favor it, since rewrites systematically underdeliver against their promised timeline.
Structured elaboration
| Criterion | Favors incremental refactor | Favors full rewrite |
|---|---|---|
| Risk | Existing behavior partially understood, tests exist or can be added incrementally | Existing behavior is essentially undocumented AND untestable in place |
| Customer impact | Feature delivery must continue in parallel | The module is isolated enough that a parallel rewrite doesn't block other work |
| Developer productivity | The pain is localized (one team, one service) | The pain is systemic and actively blocking multiple teams |
| Time to deliver | Any bounded time horizon | Only when there's genuine slack (a rewrite with a hard deadline is a red flag, not a plan) |
A useful decision threshold: choose full rewrite only if you can characterize the CURRENT system's behavior well enough to know what "done" means for the rewrite (via characterization tests or a clear spec), and if you can ship the rewrite behind a flag with a real rollback path. If either of those isn't true, the risk of a rewrite is understated no matter how bad the existing code looks. Estimating each criterion concretely, not just qualitatively, is part of the framework too: risk from the incident count traced to the module over the last 6 months (more incidents, higher risk); customer impact from the percentage of traffic or revenue that touches the affected code path; developer productivity from the engineer-hours per month currently spent on workarounds in that module; and time to deliver from an engineer-weeks estimate for each path (an incremental fix is estimated directly from the known change, while a rewrite is estimated from re-implementing every currently-used behavior, which is why rewrite estimates are systematically less reliable).
Worked example
A legacy pricing module: undocumented but has decent test coverage (72%), is used by three other services, and the team has two sprints of slack before the next major deadline. Risk is moderate (tests exist), customer impact of an incremental approach is low (small, reviewable changes), developer productivity impact is currently moderate (the module is annoying but not blocking), and the time-to-deliver criterion clearly favors incremental (a rewrite would need to reproduce three services' worth of edge cases from scratch). Two sprints might look like the "genuine slack" the table's threshold calls for, but it isn't enough here: reproducing three services' worth of undocumented edge cases from scratch is realistically a multi-month effort, not a two-sprint one, so the slack criterion is not actually satisfied for THIS rewrite even though slack exists in the abstract; genuine slack means enough runway for the scope of rewrite under consideration, not just any nonzero buffer. Verdict: incremental refactor, prioritizing tests where coverage gaps exist first.
Trade-offs & pitfalls
The classic mistake is choosing rewrite because the existing code is unpleasant to work in, which is a developer-experience signal, not a risk/customer-impact/time signal. A second-order trap: a rewrite that starts as "just the pricing logic" and scope-creeps into touching everything the original module touched, at which point its risk profile has silently become worse than the incremental path it was chosen over.
Define test debt as a category of technical debt, and describe two practical tactics to reduce it quickly, such as targeted automation of the highest-risk paths and enforcing a test-pyramid shape. Explain how you would decide which tests to automate first.
Sample Answer
Direct answer
Test debt is the accumulated gap between the tests a codebase needs and the tests it has: missing coverage, brittle or flaky tests, and tests that pass without meaningfully verifying behavior. Two fast, practical tactics to reduce it: automate the highest-risk paths first rather than chasing a blanket coverage number, and enforce a test-pyramid shape (many fast unit tests, fewer integration tests, very few slow end-to-end tests) so the suite stays fast enough that people don't start skipping it.
Structured elaboration
- Targeted automation of critical paths: rank untested code by a combination of change frequency and business criticality (a checkout flow beats an admin settings page), and automate tests there first; this delivers risk reduction per engineer-hour far faster than chasing overall coverage percentage.
- Enforce the test pyramid: audit the current suite's shape (count of unit vs integration vs end-to-end tests and their respective run times); if end-to-end tests dominate, they're slow and flaky by nature, which erodes trust in the whole suite and is itself a source of NEW test debt as people start skipping or ignoring red builds.
Deciding what to automate first: score candidate areas on (a) how often the code changes (higher change frequency = higher risk of a regression slipping through) and (b) business criticality (revenue-path or compliance-relevant code outranks internal tooling), and start with the intersection of both being high.
Worked example
A codebase's untested areas: an admin-only reporting page (rarely changes, low criticality) and the checkout discount-code logic (changes almost every sprint, directly revenue-relevant). Despite the admin page having zero coverage and the checkout logic having partial coverage, the checkout logic is the higher-priority automation target, since its change frequency and criticality both point to it as the place most likely to regress silently and cause real damage.
Trade-offs & pitfalls
Chasing a blanket coverage percentage target ("get to 80% everywhere") often produces the OPPOSITE of risk reduction: teams write low-value tests against low-risk code to move the number, while the highest-risk, highest-change-frequency code stays under-tested because it's harder to test well. Targeted, risk-based automation is slower to show a clean percentage but faster to reduce real incident risk.
Identify the common sources of technical debt in a software organization. For each source, give a practical indicator or metric you would monitor to detect it early, and a simple detection method you could run against an existing codebase.
Sample Answer
Direct answer
Technical debt in a real organization comes from six recurring sources: rushed delivery under deadline pressure, deferred upgrades, missing tests, duplicated solutions, knowledge loss from turnover, and shortcuts taken without a repayment plan. Each has its own early-warning signal, so detecting them requires more than one metric.
Structured elaboration
| Source | Early indicator | Detection method |
|---|---|---|
| Rushed delivery | Spike in hotfixes shortly after a release | Correlate deploy dates with incident timestamps |
| Deferred upgrades | Dependencies several major versions behind | Automated dependency-audit scan (e.g. npm outdated, pip list --outdated) run on a schedule |
| Missing tests | Coverage trending down on recently-changed files | Coverage-delta check per PR, not just a global number |
| Duplicated solutions | The same logic appears with small variations across files/services | Static duplication detection (e.g. jscpd, PMD CPD) |
| Knowledge loss | One person authors the majority of commits to a module, then leaves | Git blame / commit-author concentration analysis per module |
| Undisciplined shortcuts | "TODO" / "FIXME" / "HACK" comment density rising | Grep-based comment audit, tracked over time |
Worked example
A quarterly audit of a mid-size codebase finds: 40 unresolved TODO/FIXME comments concentrated in the billing module (undisciplined-shortcut signal), the billing module's primary contributor left the company two months prior with no documented handoff (knowledge-loss signal), and npm outdated shows the payment SDK is three major versions behind (deferred-upgrade signal). All three point at the same module, which is exactly the kind of correlated signal that should escalate a module from "routine backlog item" to "priority review," even though no single metric alone would have triggered that.
Trade-offs & pitfalls
A single detection method catches only its own blind spot: static duplication detection misses semantically duplicated logic written differently, and commit-author concentration can flag a module that's simply owned by a specialist by design, not one that's actually at knowledge risk. Combine at least two independent signals before escalating a module, rather than acting on any single metric in isolation.
Unlock Full Question Bank
Get access to all 15 Technical Debt Management and Refactoring interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.