Incident Response and Management Questions
The end-to-end operational lifecycle of a production incident: detection, triage, mitigation, resolution, and handoff. Covers response coordination, ownership of an active incident, and the mechanics of restoring service quickly. This is the generalist backbone topic that most operational roles are drilled on.
Describe a moment where you had to choose between a quick workaround to restore service and a longer-term architectural fix. What factors did you weigh (risk, cost, customer impact, how much runway you had), and what did you actually decide?
Sample Answer
Direct answer
A strong answer names the specific factors weighed (how much risk the workaround carries, what it costs to maintain, how much customer impact continuing to be broken causes, and how much runway there actually was before a proper fix could ship) and is honest about what was actually decided, including if the workaround turned out to be the wrong call in hindsight.
Structured elaboration
- Risk. Does the quick workaround introduce new risk of its own (a hacky patch that could fail in an unexpected way) versus being a safe, well-understood stopgap?
- Cost. What does it cost to build and maintain the workaround, especially if 'temporary' fixes have a track record of becoming permanent at your organization?
- Customer impact of waiting. How bad is the status quo for customers while the long-term fix is being built, and does that urgency justify accepting the workaround's downsides?
- Runway. How much time is realistically available before the long-term fix ships, and is the workaround actually needed to bridge that gap or is the long-term fix closer than it first appears?
- A genuinely good answer names a real trade-off that was made, including a downside that was accepted, rather than presenting the decision as costless; a decision with no acknowledged downside usually means the story is being told too cleanly.
Worked example
An illustrative skeleton: 'We had a service that was intermittently timing out under load, and the real fix (redesigning how it handled a slow downstream dependency) was going to take a couple of weeks of actual engineering work. In the meantime, I put in a quick workaround: an aggressive timeout with a fallback response, which meant some requests got a slightly degraded but fast response instead of a slow failure. I weighed the risk (the fallback response was less precise, which had a real, measurable cost for a subset of users) against the customer impact of continuing to time out entirely, and decided the workaround was worth it given a two-week runway, but I made sure the fallback usage was itself monitored so we'd know exactly how often it was being hit and wouldn't lose track of it once the real fix shipped.'
Trade-offs and pitfalls
The main weakness is presenting the decision as obviously correct with no real cost, when a genuine trade-off almost always has an accepted downside worth naming honestly. A second pitfall is describing a 'temporary' fix that quietly became permanent without acknowledging that as a real outcome worth reflecting on; a candidate who's aware of that risk and actively guarded against it (through monitoring, a tracked follow-up, an explicit deadline) demonstrates more judgment than one who simply moved on. This decision-making pattern, weighing risk, cost, customer impact, and available runway, applies whether you're making the call directly as the engineer or explaining the same trade-off to someone else who has to decide, like a customer evaluating their own quick-fix-versus-rebuild choice.
You are paged for a sudden spike in errors on a critical production service. Walk through what you do in the first 15 to 30 minutes: what you check first, how you decide whether to page anyone else, and what you would and would not do in that opening window.
Sample Answer
Direct answer
In the first 15 to 30 minutes: acknowledge the page immediately so others know someone is on it, check your two or three primary signal sources (error-rate dashboard, recent deploys, recent config changes) to get a rough sense of scope and cause, apply an obviously safe temporary mitigation if one exists, post a first status update even if it just says 'investigating,' and decide whether you need more people. What you don't do: silently investigate alone for the whole window, or start deep root-causing before you know how big the problem is.
Structured elaboration
- Acknowledge and orient (minute 0-2). Confirm the page is real, not a flapping alert, and check the primary dashboard to see current error rate, latency, and traffic. This tells you if you're dealing with a full outage or a partial degradation.
- Check for an obvious cause (minute 2-8). Look at what changed recently: deploys, config pushes, feature flag flips, infrastructure changes. Most production incidents correlate with a recent change, and this is the highest-value early check.
- Decide on scope and escalation (minute 5-10). Based on what you've seen, decide if this affects one service or many, one region or all, and whether you need to page anyone else. Getting help early is cheap; struggling alone for 25 minutes before asking is expensive.
- Apply a safe temporary action if available (minute 10-20). If a recent deploy correlates with the problem, rolling it back is usually the safest first move, since it's typically reversible. If there's no obvious cause, avoid guessing with an irreversible action.
- Post a first status update (by minute 15-ish). Even 'we're aware, investigating, will update in 15 minutes' is far better than silence, both for stakeholders and for your own discipline of staying on a cadence.
- What not to do: don't start a deep root-cause investigation before you understand the blast radius; don't make an irreversible change (a manual database edit, a permanent config change) under time pressure without a second person's input; don't go quiet.
Worked example
Page fires at 14:02 for elevated error rate on the checkout service. 14:02-14:03: engineer acknowledges, opens the dashboard, sees error rate at 8% (up from a normal 0.1%), affecting checkout only, not the whole site. 14:03-14:06: checks recent deploys, finds one shipped at 13:55, seven minutes before the alert. 14:06-14:08: decides this is likely deploy-related, pages a second engineer to help verify while they prepare a rollback, and opens an incident ticket marked SEV2 (partial, revenue-affecting feature). 14:08-14:11: rolls back the deploy. 14:11-14:14: watches error rate drop back to 0.1% and holds there. 14:15: posts 'checkout errors have returned to normal following a rollback of a recent deploy; monitoring to confirm and will follow up with a summary.'
Trade-offs and pitfalls
Spending the first 15 minutes purely gathering information without acting risks letting an easily-mitigated problem run longer than necessary; acting immediately without any scoping risks mitigating the wrong thing entirely, or worse, taking an action whose blast radius you don't understand. The 'lone hero' pattern, where a responder tries to fully resolve the incident single-handedly before looping anyone in, is a recurring failure mode: it delays getting a second set of eyes on a decision made under pressure, and it delays stakeholder communication that people are actively waiting on.
A CPU spike is causing service timeouts for a subset of users. Distinguish containment, mitigation, and recovery as distinct phases of your response, and give one concrete action for each: something that limits how far the problem can spread, something that reduces the impact customers feel, and something that restores full functionality. Explain the reasoning and any safety checks behind each action.
Sample Answer
Direct answer
For a CPU spike causing timeouts: containment is isolating the affected host or throttling the traffic causing the spike so the problem stops spreading to healthy instances; mitigation is shedding non-critical load or scaling out so timeouts stop for most users even before you know the root cause; recovery is restoring the service to its normal capacity and confirming latency has genuinely returned to baseline, not just dropped once.
Structured elaboration
- Containment answers 'how do I stop this from getting worse.' For the CPU spike: pull the overloaded instance out of the load balancer pool so it stops receiving new traffic and can't drag healthy instances down with it (for example through a shared connection pool or retry storm). Safety check: confirm the load balancer has enough remaining healthy capacity before pulling the instance, or you've just made the outage worse for everyone else.
- Mitigation answers 'how do I reduce what customers feel right now,' independent of understanding the root cause yet. For the CPU spike: shed non-critical requests (defer background jobs, disable an expensive feature) or scale out horizontally to spread load. Safety check: shedding load has to be reversible and clearly logged, so nobody forgets a feature is disabled after the incident ends.
- Recovery answers 'is the service actually back to normal.' For the CPU spike: bring capacity back to its normal level, watch CPU and latency hold at baseline for a sustained window, and re-enable anything that was shed. Safety check: recovery isn't declared on a single good data point, since a transient dip can look like recovery for a minute and then relapse.
Worked example
Suppose CPU on the affected instance is pinned at 95% and p99 latency has risen from a normal 120ms to 4 seconds, causing client-side timeouts. Containment: the on-call engineer removes the instance from the load balancer pool, redistributing its traffic across the remaining healthy instances. Mitigation: because the remaining instances are now each carrying more load, they scale out from 4 to 6 instances and temporarily disable a CPU-heavy recommendation feature; CPU across the fleet drops to a more sustainable 60% and p99 latency returns to around 300ms, well below the timeout threshold even if not yet fully back to baseline. Recovery: once the root cause (an inefficient query introduced by a recent deploy) is identified and fixed, the team restores the original instance count, re-enables the recommendation feature, and watches p99 latency hold near 120ms for 30 minutes before declaring the incident resolved.
Trade-offs and pitfalls
Containment that's too aggressive (pulling too many instances, or isolating a component that other services still depend on) can turn a partial degradation into a full outage; containment that's too narrow (missing that the real blast radius includes a shared resource like a connection pool or cache) leaves the spread unaddressed. Mitigation can also mask the symptom in a way that delays real diagnosis: shedding load makes CPU look healthy again, but if nobody tracks that the shed feature is still disabled, the team can lose track of what 'fully recovered' actually means. The general pattern (same containment options: read-only mode, feature toggles, traffic shaping, or temporary scaling) applies just as well to a database write-outage as it does to a CPU spike; the phase you're in, not the specific technology, determines what action is appropriate.
A high-severity incident has caused a six-hour outage affecting customers. As the on-call engineer or service owner, describe your immediate response: how you contain and mitigate the impact, how you decide what to communicate and to whom while you are still investigating, and how you validate that the service is genuinely healthy again before standing down.
Sample Answer
Direct answer
My immediate response has three simultaneous threads: contain and mitigate the technical impact, communicate what I actually know (not what I'm guessing) on a steady cadence, and validate that the fix genuinely held before standing down. None of these wait for the others; a responder who only does the technical work and goes quiet for six hours has failed the incident just as much as one who only communicates and never fixes anything.
Structured elaboration
- Contain and mitigate. First isolate whatever is failing so it can't spread (pull a bad instance from rotation, disable a feature flag, freeze further changes), then apply the fastest safe action that reduces customer pain, which is very often a rollback to the last known-good state if a recent change correlates with the onset.
- Communicate on a cadence, not just when there's news. Post a first update as soon as you have a rough sense of scope, even if it's just 'investigating, will update in 15 minutes.' Keep to that cadence even when the honest update is 'still investigating,' because silence reads as worse than an uneventful update, and stakeholders making their own decisions (support teams, other engineering teams, leadership) need to know what you know, not what you hope.
- Validate real recovery before standing down. Don't declare resolved on the first green data point; watch the key metrics hold at baseline for a sustained window, and get explicit confirmation from anyone whose data, transactions, or workflows were affected before formally closing.
- Handle degraded evidence gracefully. Sometimes your own tooling is part of the problem: if the logging or monitoring pipeline itself is down, you're triaging with less information than normal. In that case fall back to direct, lower-level checks (host-level metrics, spot-checking a handful of real requests, canary user reports) and be more conservative about declaring anything resolved, since you're flying with reduced visibility.
- Some failures don't have a clean single-instance fix. A storage array with a degraded RAID rebuilding for 12 hours, or a network partition that's caused two database replicas to diverge (a split-brain), both require deliberately not rushing: forcing a failover mid-rebuild can worsen data loss, and blindly failing over during a split-brain risks writing to the wrong side of a diverged pair. The instinct to 'just fail over' has to be checked against whether failover itself is currently safe.
Worked example
A six-hour outage: a database failover during peak traffic silently sends a subset of write traffic to a replica that hadn't fully caught up, so some writes appear to succeed but are later lost on failback. Immediate response: the on-call engineer confirms the scope (which customers, which write types affected) within the first 20 minutes and posts a first update. Containment: they freeze automated failover so the system can't flip again mid-investigation. Mitigation: they route new writes exclusively to the confirmed-primary replica, restoring correct behavior for new activity even though the earlier lost writes aren't yet recovered. Communication: updates go out every 30 minutes to internal stakeholders and, once the customer-facing impact is confirmed, to affected customers, describing what's known and what's still being investigated, not speculating about root cause before it's confirmed. Recovery: after the fix, the team monitors write consistency across replicas for two hours (not two minutes) before declaring the incident resolved, given how deceptive an initially-healthy-looking replica had already proven to be.
Trade-offs and pitfalls
The most senior-discriminating mistake here is premature stand-down: declaring victory on the first healthy-looking check when the underlying problem (like replica divergence) can silently resurface. A second common mistake is communicating guesses as facts under pressure ('this was caused by X') before root cause is actually confirmed, which erodes trust when the real cause turns out to be different. A third is treating every outage the same way regardless of its actual shape: a clean single-service outage tolerates a fast rollback, while a data-divergence or hardware-degradation scenario often calls for patience and extra verification precisely when the instinct is to act fastest.
A dependency you do not control (a vendor or a third-party provider) starts failing intermittently, causing real customer impact. Decide between putting in a temporary mitigation yourself versus waiting for the vendor to fix it, and explain the criteria and risks behind that choice.
Sample Answer
Direct answer
My default is to build a temporary mitigation myself rather than wait, unless the vendor's ETA is both short and credible, because customer experience shouldn't depend on a timeline I don't control and can't verify. The key criteria are how reliable the vendor's own status communication has historically been, how cheap and low-risk a mitigation is to build, and whether building a mitigation might itself mask a vendor problem that needs to stay escalated.
Structured elaboration
- Assess the vendor's ETA credibility. Vendors are often optimistic about their own timelines; check their status page's track record, whether they've given a specific, committed time or a vague 'we're looking into it,' and weigh that against your own tolerance for continued impact.
- Assess mitigation cost and risk. A cheap, well-understood mitigation (a circuit breaker to fail fast instead of hanging, a cached fallback response, a degraded-but-functional mode) is usually worth building even for a short outage; a mitigation that requires new, untested code under time pressure carries its own risk and might not be worth it for a vendor issue expected to resolve in minutes.
- Don't let your mitigation hide a problem that still needs escalating. If you build a good enough fallback, make sure someone is still tracking and escalating the underlying vendor issue; a well-built mitigation can quietly make an important problem invisible to leadership or to the vendor relationship owner.
- Verify the vendor issue is what you think it is, not just assumed: check the vendor's own status page or reach out directly rather than purely inferring from your own symptoms, since 'looks like a vendor is down' and 'confirmed the vendor is down' warrant different confidence levels.
Worked example
A third-party payment processor starts intermittently failing, causing checkout errors for a subset of users. The vendor's status page shows 'investigating' with no ETA, and their historical incidents have often run longer than initially communicated. The team decides not to wait: they add a circuit breaker so failing calls fail fast rather than hanging and degrading the whole checkout flow, and enable a secondary, lower-priority payment path they'd already built for exactly this kind of situation. They keep the primary vendor issue actively tracked and continue monitoring the vendor's status page, so once the vendor recovers they can cleanly revert to the primary path, rather than letting the fallback silently become permanent.
Trade-offs and pitfalls
Building a mitigation under pressure risks shipping throwaway code that never gets properly cleaned up and becomes unplanned permanent technical debt; it's worth explicitly flagging a mitigation as temporary and following up after the incident. Waiting on a vendor whose communicated ETA turns out to be optimistic (vendors very often are) means your customers experience longer impact than necessary, purely because you trusted someone else's timeline you had no way to verify. The same underlying judgment applies whether the failing dependency is a payment processor, a cloud provider's specific service, or a network transit provider throttling traffic to a region: the calibration is always about ETA credibility versus mitigation cost, not about the specific vendor.
Unlock Full Question Bank
Get access to all 11 Incident Response and Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.