Technical Debt Management and Refactoring Questions
Identifying, prioritizing, and paying down technical debt sustainably. Covers recognizing debt, making the case to invest in it, refactoring safely behind tests, and balancing debt reduction against feature velocity. Includes keeping a codebase maintainable over the long term.
Design a decision framework for when maintaining an obsolete internal library is no longer worth the effort of learning or continuing to update. Include a cost-benefit analysis, migration-effort estimates, a risk assessment covering incidents and security, the learning cost for engineers, and a phased deprecation strategy with rollback plans.
Sample Answer
Direct answer
Decide to stop maintaining an obsolete internal library when the ongoing learning and migration cost for the team clearly exceeds the cost of a phased deprecation, weighing incident and security risk heavily, since an unmaintained library's risk compounds silently over time even without any new code being written against it.
Structured elaboration
- Cost-benefit analysis: compare the ongoing cost of engineers learning and working around the obsolete library's quirks (real, if hard to measure precisely, in onboarding time and workaround effort) against the one-time cost of migrating off it.
- Migration-effort estimate: count actual call sites and usages across the codebase, and estimate effort per call site based on complexity (a simple 1:1 API swap versus a call site requiring real logic changes), rather than assuming uniform migration cost across all usages.
- Risk assessment: incidents historically traced to this library, and security posture, specifically whether it's still receiving security patches from any source, since an abandoned library with no patches is a compounding, silent risk even if the team never touches that code again.
- Learning cost for engineers: how much onboarding or debugging time is spent specifically on this library's quirks, a real cost even when no incidents have resulted yet.
- Phased deprecation strategy: stop new usage immediately (a lint rule or code-review policy blocking new call sites), migrate existing call sites in priority order (highest-risk or highest-change-frequency first), and set a hard sunset date with a rollback plan for each migrated call site in case the replacement introduces a regression.
Worked example
An internal HTTP client library, unmaintained for 3 years, has 340 call sites across 25 services. A security audit finds it hasn't received a patch for a moderate-severity vulnerability disclosed 18 months ago (real, compounding risk). Migration effort estimate: roughly 200 call sites are simple 1:1 swaps to the standard library (a few hours each with tooling assistance), 100 require moderate rework (custom retry logic built on top of the old client), and 40 are complex enough to need individual review. Given the vulnerability exposure and the growing onboarding cost engineers report, the decision is to deprecate: block new usage immediately, migrate the 200 simple call sites first via a semi-automated codemod within a month, then the 100 moderate ones over the following quarter, and the 40 complex ones on a slower, individually-scoped timeline, with a 9-month full sunset target.
Trade-offs & pitfalls
The risk in phased deprecation is the long tail (the 40 complex call sites here) stalling indefinitely once the easy wins are done and momentum fades; setting a hard sunset date up front, with visible tracking of remaining call sites, is what prevents the deprecation from silently stalling at 85% complete forever.
How would you redesign team incentives and OKRs so engineering, product, and customer-success teams are jointly accountable for reducing technical debt and improving product quality over the next year? Provide example first-quarter OKRs that reflect this alignment and describe how you would track progress.
Sample Answer
Direct answer
Redesign incentives so engineering, product, and customer success share OKRs tied to shared outcomes (not separate function-specific goals that happen to touch quality), with each function's individual OKRs explicitly linked to the shared one, so debt reduction stops being "engineering's problem" that others merely tolerate.
Structured elaboration
- Shared outcome OKR: a single, org-level objective all three functions contribute to, for example "reduce customer-impacting quality issues while maintaining feature velocity," with key results spanning all three functions' natural metrics (engineering: incident rate and cycle time; product: feature-adoption-adjusted for quality issues; customer success: quality-related ticket volume).
- Function-level OKRs link back explicitly: each function's own quarterly OKRs include at least one key result that ties directly to the shared objective, so it's visible in each function's own planning, not just in a separate cross-functional document nobody revisits.
- Joint accountability mechanism: a shared quarterly review where all three functions report against the shared objective together, not three separate updates in three separate meetings, so trade-offs and mutual dependencies are visible in the room.
Worked example, first-quarter OKRs:
- Shared objective: "Reduce customer-impacting quality issues by 30% while holding feature velocity flat."
- Engineering KR: reduce incident rate in the top 3 highest-incident services by 30%, and hold cycle time within 10% of current baseline.
- Product KR: for every major feature shipped this quarter, a documented quality/debt review occurs before commitment (process KR, since a direct quality-outcome KR for product alone would be too indirect).
- Customer success KR: quality-related ticket volume down 25%, tracked and reported alongside the engineering incident-rate KR in the same joint review, not separately.
Tracking progress: a shared quarterly dashboard showing all three KRs together, reviewed jointly rather than function-by-function, so a shortfall in one area (say, product's process KR slipping) is visible as a shared problem in the same conversation as engineering's incident-rate progress, not siloed into a separate report nobody cross-references.
Trade-offs & pitfalls
The common failure in cross-functional OKR design is each function getting a token, disconnected KR that technically checks the "joint accountability" box without creating any real shared incentive; the design above avoids that by requiring each function's KR to be reported and reviewed IN THE SAME meeting as the others, which is what actually creates the shared visibility and mutual pressure that makes joint accountability real rather than nominal.
Briefly summarize what legacy modernization means for a product platform. List three early warning signals that indicate modernization may be necessary, and propose one low-cost proof-of-concept experiment to validate the need before committing to a full modernization effort.
Sample Answer
Direct answer
Legacy modernization means deliberately investing to bring an aging system's architecture, tooling, or platform up to a standard that supports the business going forward, distinct from ongoing maintenance; three early warning signals it may be needed: recurring hacks/workarounds to keep the system functioning, a vendor or platform the system depends on approaching end of life, and rising onboarding time for new engineers joining the team.
Structured elaboration
- Repeated hacks: if the team keeps reaching for workarounds instead of proper fixes for the same class of problem, that's evidence the underlying architecture can no longer absorb normal change gracefully, a stronger signal than any single hack in isolation.
- Vendor or platform end-of-life: a dependency (a database version, a language runtime, a third-party platform) approaching its support end date converts "nice to modernize eventually" into a forcing function with a real deadline.
- Growing onboarding time: if new engineers take measurably longer to become productive than they did a year or two ago, that's a leading indicator the system has become harder to reason about, often before it shows up in any other metric.
A low-cost validation experiment before committing to full modernization: pick the single most painful, most-frequently-hit workaround (or, if a hard vendor/platform end-of-life date is the driving signal instead, the smallest isolated module that depends on it), and prototype (not ship) a proper fix or migration for just that one case, timeboxed to one to two weeks, to get real signal on whether the underlying architecture genuinely needs restructuring or whether targeted fixes can address the pain without a full modernization program.
Worked example
A payments platform running on a database version reaching end-of-vendor-support in 8 months (a hard forcing function), combined with three separate recent incidents traced to the same undocumented workaround pattern (repeated hacks), and a new-hire ramp-up time that's grown from 3 weeks to 6 weeks over the past year (onboarding signal). All three signals independently point toward modernization being genuinely needed, not just aesthetically desirable; a two-week proof-of-concept migrating a single, low-risk module to the target database version validates the migration approach's feasibility before committing the full team to an 8-month program.
Trade-offs & pitfalls
The risk in over-applying this framing is treating any old-feeling codebase as needing modernization; the signals above are deliberately concrete and measurable (an end-of-life date, a recurring incident pattern, a ramp-up-time trend) specifically to avoid triggering a large investment purely on the vague, common feeling that "this code feels old."
You observe developer velocity declining, bug rates rising, and build times increasing. Develop a model to quantify the annualized cost of this technical debt to the business, and forecast its impact on feature delivery over the next 12 months if left unaddressed. State the inputs and assumptions your model needs and show the equations you would use.
Sample Answer
Direct answer
Model the annualized cost of debt as the sum of three components: extra developer time spent working around the debt (velocity tax), the cost of incidents attributable to it, and the opportunity cost of delayed feature delivery, each estimated from trend data and projected forward if nothing changes.
Structured elaboration
Inputs needed:
- Baseline versus current cycle time, to estimate the velocity tax in engineer-hours per week.
- Fully-loaded engineer cost per hour (salary plus overhead, typically available from finance).
- Incident count and average cost per incident (support time, and, if available, an estimate of revenue impact) attributable to the affected area.
- Estimated feature-delivery delay, in weeks, and the delayed feature's revenue rate (value per week or per year once shipped), so the delay cost is the value rate times the duration of the delay, not the feature's entire value.
A simple equation:
Annualized cost=52×(Δcycle-time-hours-per-week×hourly-cost)+(incidents-per-year×cost-per-incident)+(52delayed-feature-annual-value×weeks-delayed)
The opportunity-cost term is a rate multiplied by a duration, the same way the velocity tax is; treating a delayed feature's full annual (or total) value as though it were entirely lost overstates the cost, since the feature still ships, just later.
Worked example
A team's cycle time has risen from 8 to 14 hours (a 6-hour weekly tax per engineer, on 10 engineers = 60 engineer-hours/week), at a fully-loaded cost of $85/hour: 52 * 60 * 85 = $265,200/year in velocity tax alone. Incidents attributable to the affected module run at roughly 1 per month at an estimated $6,000 per incident: 12 * 6000 = $72,000/year. If the roadmap has one feature worth an estimated $200,000/year in incremental revenue once shipped, and it is delayed by an estimated 4 weeks due to the velocity tax, the opportunity cost of that delay is the weekly value rate times the delay, not the full feature value: (200,000 / 52) * 4 = $3,846.15 * 4 ≈ $15,385, a one-time cost this year (not annualized, since it is a one-time delay, not a recurring cost). Total estimated annual cost: 265,200 + 72,000 + 15,385 ≈ $352,585, with the velocity tax as the dominant, recurring term.
Forecasting forward 12 months: if the cycle-time trend continues linearly (a real risk, since debt tends to compound rather than plateau on its own), the velocity tax alone could grow another 20-30% by year end absent intervention, an estimate stated as a range rather than false precision, since compounding rates are inherently uncertain.
Trade-offs & pitfalls
The velocity-tax term dominates in most real cases and is also the hardest to defend, since it depends on attributing a cycle-time change specifically to debt rather than to other causes (control for confounders before asserting this attribution). State the assumption explicitly ("assuming the full cycle-time increase is attributable to this debt") rather than presenting the total as an unqualified fact; a defensible range beats a false-precision point estimate that collapses under a single pointed follow-up question. Also watch the opportunity-cost term specifically: it is tempting to add a delayed feature's entire value as a one-time cost, but that conflates the feature's total value with the cost of the delay itself. A 4-week slip on a $200,000/year feature costs roughly $15,400 in delayed value, not $200,000; only a permanently cancelled feature would justify using the full value.
What signals and telemetry would you monitor to detect accumulating technical debt across multiple teams and repositories? Name at least five, explain why each is useful, and what tooling you would use to collect it. Also explain how you would guard against teams gaming these metrics once they know they are being tracked.
Sample Answer
Direct answer
Monitor at least five signals: the trend (not just the snapshot) of cyclomatic complexity, the trend of test coverage, mean build and deploy time, PR size or review latency, and incident frequency. Each one is useful because it is a leading or lagging proxy for the thing you actually care about (developer velocity and system reliability), and together they cover code, process, and operational health rather than just one dimension.
Structured elaboration
| Signal | What it approximates | Why it's useful | Tooling |
|---|---|---|---|
| Cyclomatic complexity trend | How hard code is to reason about and test | Rising complexity predicts rising bug rate and slower changes | Static analysis (SonarQube, CodeClimate, radon) |
| Test coverage trend | How much of the system is safety-netted | Falling coverage predicts riskier changes and slower reviews | Coverage tools wired into CI (coverage.py, jacoco, istanbul) |
| Mean build and deploy time | Developer feedback loop speed | A slowing loop is often the earliest visible symptom of debt, before bugs show up | CI dashboards (GitHub Actions insights, Jenkins) |
| PR size or review latency | How safely work can be decomposed and reviewed | Growing PR size or review time signals coupling and unclear boundaries | Git provider analytics (GitHub Insights, LinearB) |
| Incident frequency | Real-world reliability cost of the debt | Ties abstract code metrics to concrete business impact | Incident tracker (PagerDuty, Opsgenie) |
For ML/AI-flavored teams, add training-time-per-epoch and retrain frequency as domain-specific analogues of build time and deploy frequency.
Worked example
A team's dashboard shows cyclomatic complexity flat for six months, test coverage dropping from 78% to 61% over the same period, and mean build time climbing from 6 to 14 minutes. None of these individually triggers an incident, but together they predict that the team's next quarter will show slower PR cycle time and a rising bug rate, roughly two months before that actually shows up in the incident tracker. That lead time is the entire point of tracking trends instead of point-in-time snapshots.
Trade-offs & pitfalls
The moment a metric becomes a target, people optimize the metric rather than the underlying thing it measures (Goodhart's law). Coverage percentage is the classic case: teams write low-value tests (asserting a getter returns what was just set) purely to move the number, without reducing real risk. Guard against this by pairing coverage with a second signal, such as mutation-testing score (deliberately inserting small bugs into the code and checking what fraction the test suite actually catches) or bug-escape rate, that is harder to game, and by treating any single metric as a conversation starter rather than a scorecard.
Unlock Full Question Bank
Get access to all Technical Debt Management and Refactoring interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.