Incident Response and Management Questions
The end-to-end operational lifecycle of a production incident: detection, triage, mitigation, resolution, and handoff. Covers response coordination, ownership of an active incident, and the mechanics of restoring service quickly. This is the generalist backbone topic that most operational roles are drilled on.
Tell me about a specific production incident you triaged hands-on. Walk through the actual monitoring signals, tools, and commands you used to narrow down the problem, the immediate fix you applied, and what you changed afterward to prevent a repeat.
Sample Answer
Direct answer
This question is looking for hands-on technical fluency under pressure, not just process: the specific signals you checked, the specific tools and commands you actually ran, and how those concrete steps led to the fix, told as a real, technically credible narrative rather than a high-level summary.
Structured elaboration
What separates a strong answer here from a more general first-responder story is the technical texture:
- Monitoring and observability signals. Name the actual dashboards, metrics, or logs you looked at first and why those were the right starting point given the symptom.
- Tools and commands. Be specific about what you actually ran (checking pod status and recent events in a container orchestrator, tailing logs for a specific service, checking resource utilization on a host) and what each step told you, not just that you 'checked the logs.'
- The diagnostic chain. Walk through how one check led to the next: what you ruled out, what pointed you toward the real cause, and where you might have gone down a wrong path before correcting.
- Immediate remediation. What specific action restored service, and why that was the right call given what you'd found (a restart, a rollback, a manual failover, a config change), including any manual verification you did to confirm it actually worked rather than just assuming.
- Long-term change. A concrete artifact that came out of it: a new alert, a runbook entry, an automated check, something that reduces reliance on someone remembering the right diagnostic steps next time.
Worked example
An illustrative skeleton: 'I got paged for a service that had stopped responding to health checks. I checked our container orchestrator's dashboard and saw several pods in a crash-loop state, so I described one of the pods to see recent events and found repeated out-of-memory kills. I tailed the pod's logs right before each restart and saw a specific request pattern that correlated with memory spikes, which pointed at a recent code change rather than an infrastructure problem. I rolled back that deploy, watched the pods stabilize and stop crash-looping over the next several minutes, and manually verified a few real requests were succeeding before considering it resolved. Afterward, I added a memory-usage alert tied to that service specifically, since the existing alerting hadn't caught the gradual climb before the crash loop started.'
Trade-offs and pitfalls
The main weakness in answers here is staying at too high a level ('I looked into the logs and found the issue') without the concrete tool-level detail that actually demonstrates hands-on competence; an interviewer asking this specific version of the question is usually trying to distinguish someone who directs others during an incident from someone who can personally do the diagnostic work. A second pitfall is describing a diagnostic path that sounds suspiciously clean and linear; a credible answer often includes at least one wrong turn or ruled-out hypothesis, since that's how real diagnosis actually goes, and its total absence can read as rehearsed rather than genuine.
Two people disagree on how to respond to a bad deployment: roll it back, which loses a day of writes, or patch forward, which risks the underlying problem spreading further. Walk through how you would decide, what information you would gather first, and how you would explain the decision to the people affected by whichever data loss or risk you accept.
Sample Answer
Direct answer
The real question isn't which option is faster, it's which one is cheaper to be wrong about given what you currently know. I'd gather a quick read on how quantifiable and recoverable the rollback's data loss actually is, and how bounded or open-ended the patch's propagation risk is, before deciding, and I'd document the decision and get a second person's sign-off given the stakes.
Structured elaboration
- Quantify the rollback's cost. Is the lost day of data truly unrecoverable, or can some of it be reconstructed from logs, upstream systems, or replay? A rollback that loses genuinely unrecoverable customer data is a much bigger deal than one that loses data you can mostly reconstruct.
- Bound the patch's risk. Is 'risking propagation' a vague fear or a specific, boundable failure mode? If you can identify exactly what could go wrong and put a guard around it (a feature flag, a canary rollout of the patch itself), the risk becomes much more manageable than an open-ended 'we're not sure how bad this could get.'
- Consider reversibility of each option, not just its immediate outcome. A rollback is usually fast to reverse if it turns out to be wrong (roll forward again); a patch that goes wrong under time pressure is often harder to cleanly undo, especially if it's already started propagating.
- Decide who needs to sign off. For a decision with real data loss or real risk of making things worse, this shouldn't be a unilateral call under pressure; loop in whoever owns the data (for the rollback) or whoever understands the propagation risk best (for the patch), even if briefly.
- Document the decision and the reasoning, not just the action taken, so it can be explained afterward and revisited in the postmortem.
- Explain the decision honestly to the people who bear its cost. If you accept the rollback's data loss, tell the specific team or customers whose day of writes is gone what was lost and why, rather than a vague status update; if you accept the patch's propagation risk, tell whoever owns the systems it could spread to what you're watching for. People affected by a real cost should hear the reasoning, not just the outcome.
- Watch for a subtlety: your 'safe' default might not be safe. Sometimes the rollback mechanism itself is the risky part (a rollback tool that has its own failure modes, or that could trigger a different kind of cascading failure), so the instinct to always default to rollback as 'the safe choice' needs the same scrutiny as the riskier-looking option.
Worked example
A bad deployment has two disagreeing camps: roll back (losing a day of writes) or patch forward (risking the underlying bug spreading to more data). First, the team checks whether the day of writes can be reconstructed from an upstream event log; it turns out about 80% of it can be replayed after rollback, meaningfully reducing the real cost of that option. Second, they check whether the patch's propagation risk can be bounded; the bug only affects a specific, identifiable code path, so a scoped patch with a feature flag around just that path is possible rather than a full risky redeploy. Given both, the team chooses the scoped patch behind a flag, since the true data loss from rollback (even reduced by replay) is now higher-cost than a well-bounded patch, and they get a second engineer to review the patch's blast radius before shipping it, given the stakes.
Trade-offs and pitfalls
A common mistake is defaulting to rollback simply because it feels psychologically safer ('undo the thing we just did'), without actually quantifying whether the data loss is worse than the alternative risk; rollback is not automatically the conservative choice. Another mistake, at the opposite extreme, is defaulting to a quick patch because it avoids acknowledging any data loss, without honestly bounding how far the underlying problem could actually spread. A related edge case: sometimes the rollback path itself carries risk (a partial rollback failing midway, or a rollback tool triggering a separate cascading failure under load), so 'roll back, it's the safe option' deserves the same scrutiny as the alternative, not an automatic pass.
You are paged for a sudden spike in errors on a critical production service. Walk through what you do in the first 15 to 30 minutes: what you check first, how you decide whether to page anyone else, and what you would and would not do in that opening window.
Sample Answer
Direct answer
In the first 15 to 30 minutes: acknowledge the page immediately so others know someone is on it, check your two or three primary signal sources (error-rate dashboard, recent deploys, recent config changes) to get a rough sense of scope and cause, apply an obviously safe temporary mitigation if one exists, post a first status update even if it just says 'investigating,' and decide whether you need more people. What you don't do: silently investigate alone for the whole window, or start deep root-causing before you know how big the problem is.
Structured elaboration
- Acknowledge and orient (minute 0-2). Confirm the page is real, not a flapping alert, and check the primary dashboard to see current error rate, latency, and traffic. This tells you if you're dealing with a full outage or a partial degradation.
- Check for an obvious cause (minute 2-8). Look at what changed recently: deploys, config pushes, feature flag flips, infrastructure changes. Most production incidents correlate with a recent change, and this is the highest-value early check.
- Decide on scope and escalation (minute 5-10). Based on what you've seen, decide if this affects one service or many, one region or all, and whether you need to page anyone else. Getting help early is cheap; struggling alone for 25 minutes before asking is expensive.
- Apply a safe temporary action if available (minute 10-20). If a recent deploy correlates with the problem, rolling it back is usually the safest first move, since it's typically reversible. If there's no obvious cause, avoid guessing with an irreversible action.
- Post a first status update (by minute 15-ish). Even 'we're aware, investigating, will update in 15 minutes' is far better than silence, both for stakeholders and for your own discipline of staying on a cadence.
- What not to do: don't start a deep root-cause investigation before you understand the blast radius; don't make an irreversible change (a manual database edit, a permanent config change) under time pressure without a second person's input; don't go quiet.
Worked example
Page fires at 14:02 for elevated error rate on the checkout service. 14:02-14:03: engineer acknowledges, opens the dashboard, sees error rate at 8% (up from a normal 0.1%), affecting checkout only, not the whole site. 14:03-14:06: checks recent deploys, finds one shipped at 13:55, seven minutes before the alert. 14:06-14:08: decides this is likely deploy-related, pages a second engineer to help verify while they prepare a rollback, and opens an incident ticket marked SEV2 (partial, revenue-affecting feature). 14:08-14:11: rolls back the deploy. 14:11-14:14: watches error rate drop back to 0.1% and holds there. 14:15: posts 'checkout errors have returned to normal following a rollback of a recent deploy; monitoring to confirm and will follow up with a summary.'
Trade-offs and pitfalls
Spending the first 15 minutes purely gathering information without acting risks letting an easily-mitigated problem run longer than necessary; acting immediately without any scoping risks mitigating the wrong thing entirely, or worse, taking an action whose blast radius you don't understand. The 'lone hero' pattern, where a responder tries to fully resolve the incident single-handedly before looping anyone in, is a recurring failure mode: it delays getting a second set of eyes on a decision made under pressure, and it delays stakeholder communication that people are actively waiting on.
During a live incident, the root cause turns out to live in a shared service owned by a different team than yours. Describe how you would work with that team while the incident is still active: how you get the right people engaged quickly, and how you keep the response moving without waiting on a formal handoff.
Sample Answer
Direct answer
When the root cause lives in a service another team owns, my first move is getting the right person from that team engaged directly and fast, usually by paging their on-call rather than routing through a manager, and then working in parallel rather than blocking: I keep making progress on whatever I can control while they investigate their side.
Structured elaboration
- Get the right person, not just any person. Page the owning team's on-call directly if your investigation clearly points at their service, rather than escalating through management layers that add delay without adding expertise.
- Be specific about what you need from them. Rather than a vague 'something's wrong with your service,' share exactly what you've observed and why you believe the root cause is there, which lets them start from your findings instead of re-deriving them from scratch.
- Work in parallel, not sequentially. While the owning team investigates their side, continue anything you can independently do on your side (further mitigation, additional monitoring, keeping stakeholders updated), rather than sitting idle waiting for their update.
- Don't take over their system without context. Even if you technically have the access to poke at their service directly, doing so without their domain knowledge risks causing a second problem; the better move is close collaboration, not unilateral action on a system you don't own.
- Don't silently wait either. If you've reached out and haven't heard back within a reasonable window given the severity, escalate again rather than assuming they're already on it.
Worked example
An incident's root cause traces to a shared authentication service owned by a different team. Rather than waiting for a formal handoff process, the responder directly pages that team's on-call with specific findings ('auth requests from our service are timing out starting at 14:02, correlating with your deploy at 13:58'), which lets the other team's engineer start investigating their deploy immediately rather than starting from scratch. While waiting, the original responder adds a client-side retry with backoff on their own service as a partial mitigation, something within their own control, rather than being fully blocked on the other team's fix.
Trade-offs and pitfalls
The most common failure here is silently waiting on the other team without actively escalating, which can leave an incident stalled far longer than necessary if that team is slow to notice or prioritize it. The opposite failure, someone outside the owning team taking matters into their own hands and directly modifying a system they don't fully understand, risks introducing a second, unrelated incident on top of the first. The right balance is proactive, specific engagement paired with continuing to make progress on what you do control, rather than either extreme.
Describe a moment where you had to choose between a quick workaround to restore service and a longer-term architectural fix. What factors did you weigh (risk, cost, customer impact, how much runway you had), and what did you actually decide?
Sample Answer
Direct answer
A strong answer names the specific factors weighed (how much risk the workaround carries, what it costs to maintain, how much customer impact continuing to be broken causes, and how much runway there actually was before a proper fix could ship) and is honest about what was actually decided, including if the workaround turned out to be the wrong call in hindsight.
Structured elaboration
- Risk. Does the quick workaround introduce new risk of its own (a hacky patch that could fail in an unexpected way) versus being a safe, well-understood stopgap?
- Cost. What does it cost to build and maintain the workaround, especially if 'temporary' fixes have a track record of becoming permanent at your organization?
- Customer impact of waiting. How bad is the status quo for customers while the long-term fix is being built, and does that urgency justify accepting the workaround's downsides?
- Runway. How much time is realistically available before the long-term fix ships, and is the workaround actually needed to bridge that gap or is the long-term fix closer than it first appears?
- A genuinely good answer names a real trade-off that was made, including a downside that was accepted, rather than presenting the decision as costless; a decision with no acknowledged downside usually means the story is being told too cleanly.
Worked example
An illustrative skeleton: 'We had a service that was intermittently timing out under load, and the real fix (redesigning how it handled a slow downstream dependency) was going to take a couple of weeks of actual engineering work. In the meantime, I put in a quick workaround: an aggressive timeout with a fallback response, which meant some requests got a slightly degraded but fast response instead of a slow failure. I weighed the risk (the fallback response was less precise, which had a real, measurable cost for a subset of users) against the customer impact of continuing to time out entirely, and decided the workaround was worth it given a two-week runway, but I made sure the fallback usage was itself monitored so we'd know exactly how often it was being hit and wouldn't lose track of it once the real fix shipped.'
Trade-offs and pitfalls
The main weakness is presenting the decision as obviously correct with no real cost, when a genuine trade-off almost always has an accepted downside worth naming honestly. A second pitfall is describing a 'temporary' fix that quietly became permanent without acknowledging that as a real outcome worth reflecting on; a candidate who's aware of that risk and actively guarded against it (through monitoring, a tracked follow-up, an explicit deadline) demonstrates more judgment than one who simply moved on. This decision-making pattern, weighing risk, cost, customer impact, and available runway, applies whether you're making the call directly as the engineer or explaining the same trade-off to someone else who has to decide, like a customer evaluating their own quick-fix-versus-rebuild choice.
Unlock Full Question Bank
Get access to all 14 Incident Response and Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.