Incident Response and Management Questions
The end-to-end operational lifecycle of a production incident: detection, triage, mitigation, resolution, and handoff. Covers response coordination, ownership of an active incident, and the mechanics of restoring service quickly. This is the generalist backbone topic that most operational roles are drilled on.
Ransomware has started encrypting file shares across two business units and it's still spreading. As incident lead, walk me through your first hour, including who beyond engineering you bring in and how you decide whether your backups can be trusted.
Sample Answer
Direct answer
First priority is stopping the spread, not restoring anything: isolate affected segments and, if the entry point is unknown, disable shared credentials the ransomware could use to jump further. In parallel I loop in legal and executive leadership, since ransomware is a business and legal event as much as a technical one. I don't trust backups until I've verified they predate the compromise.
Structured elaboration
Following a NIST/SANS-style response: contain, then assess, then decide on recovery.
Containment (first 15-30 min): isolate infected hosts and shares, reset credentials the ransomware could use to spread laterally, and preserve one infected machine uncontained for forensics instead of wiping everything.
Who's looped in: legal (breach notification duties, and ransom-payment legality varies by jurisdiction), leadership (business impact and ransom authority), and often external incident response or law enforcement.
Trusting backups: check the timestamp against your best estimate of the initial compromise, not just when encryption started, since attackers often sit inside a network for days first. Restore into an isolated environment and scan for the same indicators of compromise before promoting it to production.
Worked example (story skeleton)
Shares in two departments start encrypting minutes apart. The lead isolates that segment, confirms the others are clean, and works with IT to find the last known-good backup timestamp while legal is notified.
Trade-offs and pitfalls
Restoring fast from an untested backup risks reintroducing the malware. Waiting too long extends the outage. A common mistake is scoping only to systems visibly encrypting, when the attacker may already have persistence elsewhere that hasn't triggered yet.
What the interviewer probes next
How you'd decide whether to consider paying the ransom and who holds that authority, and how your plan changes if the backups are also compromised.
Walk through the lifecycle of a production incident end to end, from before anything goes wrong through the post-incident review. For each phase (preparation, detection, triage, containment, mitigation, recovery, and post-incident review), name the key activity, one artifact you would expect to see (a dashboard, a ticket, a timeline), and who is typically involved. Use a concrete example action at one phase to ground your answer.
Sample Answer
Direct answer
A production incident moves through seven phases: preparation, detection, triage, containment, mitigation, recovery, and post-incident review. Preparation happens before anything breaks (runbooks written, on-call staffed, alerts wired up); detection is the moment you learn something is wrong; triage scopes and prioritizes it; containment stops it from getting worse; mitigation reduces the pain customers feel; recovery restores full normal service; and the post-incident review turns the experience into a lasting fix. Each phase has a different owner, a different artifact, and a different question it answers.
Structured elaboration
- Preparation. Activity: writing and testing runbooks, defining on-call rotations, wiring alerts to real signals. Artifact: the runbook itself and the on-call schedule. Owner: the team that runs the service, done continuously, not reactively.
- Detection. Activity: an alert fires or a human notices something is off. Artifact: the alert or the first ticket. Owner: whoever is paged, or whoever notices first.
- Triage. Activity: scoping how bad it is and who needs to know. Artifact: an incident ticket with severity, scope, and a first status note. Owner: the first responder, sometimes handed to an incident commander for anything large.
- Containment. Activity: stopping the blast radius from growing (isolating a host, throttling a bad client, disabling a feature flag). Artifact: a decision log entry noting what was done and why. Owner: whoever is closest to the failing component.
- Mitigation. Activity: making the customer-visible symptom smaller even before the root cause is fixed (failing over, serving cached data, degrading gracefully). Artifact: an updated status note describing customer impact before and after. Owner: the responder or incident commander.
- Recovery. Activity: restoring full functionality and confirming it holds, not just that one metric blipped green. Artifact: a recovery validation checklist and the all-clear message. Owner: the responder, with sign-off from anyone whose data or workflow was affected.
- Post-incident review. Activity: reconstructing the timeline, finding the root cause, and turning it into owned action items. Artifact: the postmortem document. Owner: usually the incident lead, with input from everyone involved.
Worked example
An API starts returning 5xx errors to 15% of traffic. Preparation already exists: there's a runbook for 'elevated 5xx rate' and an on-call SRE. Detection: a synthetic check pages the on-call engineer. Triage: the engineer opens a ticket, sees the error rate and which endpoints are affected, and judges this a SEV2. Containment: they notice the errors correlate with a recent deploy and freeze further deploys to that service so nothing else changes mid-investigation. Mitigation: they roll back the deploy, which drops the error rate from 15% to under 1% within two minutes. Recovery: they watch the error rate and latency stay at baseline for 20 minutes before declaring the incident resolved, since a single good data point after a rollback isn't proof the fix held. Post-incident review: a review two days later finds the deploy introduced a null-pointer bug in an edge case, and the action items are a missing test case plus a canary step that would have caught it before full rollout.
Trade-offs and pitfalls
The most common mistake is skipping straight from detection to mitigation without a real triage step, which means responders end up mitigating the wrong thing or missing that three separate alerts are actually one incident. The second common mistake is calling recovery too early: one healthy-looking dashboard refresh is not the same as a service that has held steady long enough to trust. A third, subtler pitfall is treating containment and mitigation as the same step; containment is about preventing the problem from spreading (a freeze, an isolation), while mitigation is about reducing what customers currently feel (a rollback, a failover) - conflating them means teams sometimes stop at containment and believe the incident is handled when customers are still seeing errors.
During initial triage, what signs would make you suspect you are looking at a security incident rather than a purely operational one, and what changes once you suspect that?
Sample Answer
Direct answer
Signs pointing toward a security incident rather than a purely operational one include unexplained privilege or permission changes, authentication failures or account lockouts clustering in an unusual pattern, traffic or data-access patterns that look like exfiltration rather than normal load, and any sign of unauthorized file or configuration changes that nobody on the team made. Once you suspect any of these, the biggest change is that you stop trying to 'just fix it': you preserve evidence instead of immediately remediating, and you loop in a security responder rather than continuing to triage it as a routine outage.
Structured elaboration
- Signals that lean operational: the timing correlates with a known deploy or infrastructure change, the failure pattern matches a resource exhaustion or a known dependency issue, and the behavior is explainable by something the team did on purpose.
- Signals that lean security: access or configuration changes nobody recognizes, authentication anomalies (a spike in failed logins, logins from unusual locations, tokens being used in ways that don't match normal patterns), data being read or moved in volumes or patterns that don't match normal usage, or any indicator resembling a known attack pattern (credential stuffing, privilege escalation, lateral movement).
- What changes once you suspect it. You stop applying your normal 'fix it fast' instincts on the affected system, because touching it (restarting a process, wiping a disk, rotating credentials without first documenting state) can destroy evidence a security investigation needs. You loop in whoever owns security response, and from that point the deep investigation, containment technique, and evidence-handling discipline live with that team rather than being improvised by whoever happened to be on call.
- Who to involve: the security on-call or incident response function, as early as suspicion arises, not after you've already tried to resolve it yourself.
Worked example
A service starts throwing errors and the on-call engineer initially assumes it's a bad deploy, since that's the most common cause. But checking recent deploys shows nothing changed, and instead they notice a spike in failed authentication attempts against an admin endpoint in the minutes before the errors started, followed by a permissions change on a service account that nobody on the team made. That combination (no correlated deploy, authentication anomaly, unexplained permission change) is the tell that this isn't a routine outage; the engineer stops attempting further remediation, preserves the current state (avoids restarting the affected service, which could wipe useful logs), and escalates to the security team rather than continuing to debug it as an availability problem.
Trade-offs and pitfalls
The main risk under pressure is dismissing security signals too quickly because restoring service feels more urgent, which can mean actively destroying evidence (restarting a compromised host, deleting suspicious files 'to clean up') before anyone with security expertise has looked at it. The opposite risk is over-escalating every anomaly as a security incident, which burns the security team's time and can create alert fatigue that makes real security incidents harder to distinguish from noise; the right calibration is a small, well-understood set of signals (like the ones above) rather than a vague sense that 'something feels off.'
You've just confirmed an employee's laptop is compromised and may be exfiltrating data. Walk me through how you preserve evidence while you contain the threat, and why chain of custody matters here.
Sample Answer
Direct answer
I'd document what I observe first, then use forensically sound methods, a write-blocked disk image and a memory capture before shutdown, instead of deleting files or reimaging on the spot. Chain of custody means logging who handled the evidence and what they did to it, so if this becomes a legal or HR matter, nobody can claim it was tampered with.
Structured elaboration
Preserving evidence: isolate the laptop from the network rather than powering it off, since RAM holds active attacker processes you'd otherwise lose, capture memory then a disk image with a write blocker, and log timestamps and every action from the moment it's flagged.
Chain of custody: a written log of who accessed the evidence and when, and storing the images in an access-controlled location separate from the systems under investigation.
Worked example
The analyst isolates the laptop, images it with a write blocker, and hands the image to a forensic examiner who signs for receipt.
Trade-offs and pitfalls
Rushing to reimage or wipe the device before capturing evidence destroys artifacts you cannot recreate. A common mistake is treating chain of custody as only needed "if it goes to court," when you rarely know that up front.
What the interviewer probes next
What changes if the compromised device is a production server you cannot take offline, and how you balance forensic rigor against pressure to restore service fast.
In incident response, what's the difference between containment and eradication, and why might a security team deliberately hold off on eradicating a threat even after they've contained it?
Sample Answer
Direct answer
Containment stops a threat from spreading right now, like isolating a host. Eradication removes the root cause, like deleting the malware or closing the exploited hole. Teams often delay eradication after containing, because acting too fast can destroy evidence needed to find every other place the attacker got in.
Structured elaboration
NIST's incident lifecycle treats these as separate phases: containment buys time and limits damage (isolate a host, force password resets); eradication removes what let the attacker in or persist, so it can't simply return.
Worked example
An analyst isolates a beaconing workstation from the network but leaves the malware in place for a few hours while the team hunts for other infected machines, instead of wiping it immediately.
Trade-offs and pitfalls
Eradicate too soon and you can tip off the attacker and lose forensic evidence before you know the full scope. Wait too long and some level of attacker access stays live.
What the interviewer probes next
How you'd verify containment is actually holding, and what would push you to eradicate immediately instead of waiting.
Unlock Full Question Bank
Get access to all 10 Incident Response and Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.