Incident Response and Containment Questions
Managing security incidents from detection through recovery. Covers incident response process and playbooks, containment and remediation, data-breach investigation methodology, data-exfiltration detection and analysis, root-cause and post-incident analysis, and fraud and complex-attack investigation. The operational 'a compromise is happening, now what' discipline, distinct from broader production-outage incident management.
You need to scope and respond to a suspected large-scale data exfiltration event (for example uploads to a personal or external cloud account, or unusual database export activity). Describe how you would rapidly identify all potentially affected systems, confirm what data left and when, produce an evidentiary summary of the scope, and contain the exfiltration channel while minimizing further leakage.
Sample Answer
Direct answer
Use asset inventories and access logs to rapidly enumerate every system the suspicious activity could have touched, confirm what actually left by correlating access and transfer logs against known-legitimate patterns, and contain the exfiltration channel itself while you finish scoping rather than waiting for a complete picture first.
Structured elaboration
Rapidly identifying affected systems. Start from the specific account, host, or service where the exfiltration was first observed, then use asset inventories, service dependency maps, and access logs to trace outward: what else did this account or host have access to, and does activity on those systems show the same anomalous pattern. Automated queries against your asset inventory (which services does this account have access to, which hosts share this network segment) are far faster than manually working through a list, especially at any meaningful scale.
Confirming what left and when. Correlate access logs (who or what accessed the data) against transfer or egress logs (what actually moved, and to where) to build a defensible summary distinguishing data the attacker could reach from data that was actually copied out; access alone doesn't prove exfiltration, but combined with unusual outbound transfer volume or destination, it becomes a credible evidentiary chain.
Producing an evidentiary summary. Document what you found, from what sources, with timestamps, in a form that can support both the technical remediation decision and any downstream legal or regulatory reporting obligation, since this summary may need to stand up to scrutiny well after the immediate incident is resolved.
Containing while still scoping. Don't wait for a complete picture before acting on the exfiltration channel itself; block or revoke the specific access path in use as soon as you have reasonable confidence, and continue scoping in parallel, since every additional minute of uncontained access is potential additional harm regardless of how complete your current understanding is.
For a customer-facing API caught exfiltrating data, this plays out on a compressed timeline: in the first 60 minutes, contain the specific vulnerable endpoint or credential and begin evidence collection; in the first 24 hours, complete the scoping (what data, how much, which customers affected) well enough to make an informed decision on customer notification and regulatory reporting obligations. Where the data involved is personal or otherwise sensitive, balance preserving evidence against minimizing further exposure: don't leave a vulnerable path open longer than necessary just to gather more forensic detail, since the harm from continued exposure typically outweighs the incremental investigative value once the channel is well understood.
Worked example
A service account is found to have downloaded an unusually large volume of records overnight. Using the asset inventory, the team quickly confirms this account only has access to two specific databases, narrowing the scope immediately rather than needing to check every system in the environment. Correlating the account's access logs against network egress logs confirms not just that records were queried, but that a matching volume of data left via an outbound transfer to an unfamiliar destination in the same window, distinguishing "could have accessed" from "actually exfiltrated." The vulnerable credential is revoked within the first hour, well before the full scope (exactly which record types and roughly how many) is finalized over the following several hours, since delaying containment to finish the count first would have left the channel open longer than necessary.
Trade-offs and pitfalls
Waiting until scoping is fully complete before containing the channel is the most common and costly mistake, trading a small amount of investigative completeness for a real, ongoing exposure window. The opposite mistake, treating mere access as proof of exfiltration without corroborating actual data transfer, risks overstating the incident's severity and triggering unnecessary notification obligations on data that was reachable but never actually left.
Design an end-to-end incident-response architecture for a large-scale AI/LLM inference platform (on the order of 100 million inferences per day). Requirements: fast detection of quality or safety degradation, automated mitigations (rollback or fallback models), forensic data capture (prompts, retrievals, outputs) with cross-region replication, and an immutable audit trail sufficient for regulatory review.
Sample Answer
Direct answer
At this scale, fast detection needs automated quality and safety monitoring rather than relying on user reports, mitigation needs pre-built rollback and fallback paths that can trigger without human intervention for the fastest-moving failure modes, and the audit trail needs to capture enough of the model's actual inputs and outputs to support a real regulatory review, not just aggregate metrics.
Structured elaboration
Fast detection. At 100 million inferences a day, waiting for user complaints is far too slow; automated monitoring needs to track output-quality signals (confidence distributions, refusal rates, safety-classifier flags) in near real time and alert on statistically significant deviations from baseline, ideally within minutes rather than hours.
Automated mitigations. Pre-built rollback to a previous known-good model version, or fallback to a simpler, more conservative model or a rule-based response for the specific failure pattern detected, should be triggerable automatically for well-understood failure signatures, with human review for anything novel or ambiguous, mirroring the same auto-versus-human-gate logic used in traditional security automation: reversible, well-understood actions can run automatically, while anything with broader impact or uncertainty escalates to a human.
Forensic data capture with cross-region replication. Capturing prompts, retrieval context, and outputs at this volume requires careful engineering (sampling strategies for routine traffic, full capture triggered around detected anomalies) balanced against storage cost and privacy considerations; cross-region replication ensures this data survives a regional outage or failover event, which matters both for continuity of investigation and because a serious incident and an infrastructure failure can plausibly coincide.
Immutable audit trail for regulatory review. The audit trail needs to demonstrate, after the fact, exactly what the model was given, what it produced, and what automated or human decisions were made in response, stored in a way that can't be quietly altered after the fact (append-only storage, cryptographic integrity checks), since regulatory review of an AI system's behavior increasingly expects this level of evidentiary rigor rather than aggregate dashboards alone.
Key trade-offs. Full capture of every prompt and output at this scale is expensive and raises its own privacy considerations, so most designs sample routine traffic while capturing fully around detected anomalies; automated rollback reduces response latency dramatically but requires very high confidence in the detection signal to avoid false-positive rollbacks disrupting service unnecessarily; and cross-region replication of potentially sensitive prompt data needs to respect data-residency requirements, which can conflict with a simple "replicate everywhere" approach.
flowchart TD
A[Inference requests] --> B[Model serving, multi-region]
B --> C[Safety and quality classifiers]
C --> D{Anomaly detected?}
D -->|no| E[Sampled prompt/output logging]
D -->|yes, known pattern| F[Automatic rollback or fallback model]
D -->|yes, novel pattern| G[Human review]
E --> H[Immutable audit store, cross-region replicated]
F --> H
G --> H
F --> I[Serving resumes on prior known-good model]
Worked example
A large-scale LLM-serving platform detects a sudden spike in a safety classifier's flag rate for one specific category of harmful output, concentrated on requests routed through one particular model version deployed an hour earlier. Automated detection catches the deviation within minutes based on the classifier-flag baseline. Given this matches a pre-defined, well-understood failure signature (a specific classifier threshold breach tied to a recent deployment), the system automatically rolls back that model version to the prior one without waiting for human approval, while flagging the incident for review. Prompt and output logs around the anomaly window are captured in full (rather than the standard sampled rate) and replicated across regions, providing the detailed evidentiary basis needed for the post-incident review and any subsequent regulatory inquiry into what specifically went wrong.
Trade-offs and pitfalls
Relying on user reports or manual dashboard review as the primary detection mechanism at this scale means real harm accumulates for far longer than automated monitoring would allow; the investment in automated, near-real-time quality and safety signals is what actually makes fast detection possible at 100 million inferences a day. Over-broad automatic rollback, triggered on a detection signal that isn't actually reliable enough, is the main risk on the mitigation side, and requires the same discipline around confidence thresholds and blast-radius awareness used in any automated containment system.
Define a severity classification scheme for machine-learning-system security incidents (for example Sev1 to Sev4) that combines business impact, personal-data exposure, and technical impact. Give concrete thresholds for common ML failure modes (model unavailability, accuracy collapse, PII leakage, regulator-impacting errors) and describe how you would coordinate the first hours of a SEV1 ML incident with engineering, legal, and executive leadership.
Sample Answer
Direct answer
Combine business impact, personal-data exposure, and technical severity into a single Sev1-to-Sev4 scale with concrete thresholds for common ML failure modes, and treat a confirmed Sev1 (for example, PII leakage or regulator-impacting errors at scale) with the same first-hours coordination discipline as any other major security incident: engineering, legal, and executives in the room immediately.
Structured elaboration
A practical scheme, roughly:
- Sev1: confirmed personal-data exposure at scale, or model errors directly causing regulator-impacting harm (for example, systematically incorrect loan denials affecting a protected class), or full model/service unavailability for a business-critical use case.
- Sev2: significant accuracy collapse on a meaningful user segment without confirmed PII exposure, or a smaller-scale, contained data exposure affecting a limited number of users.
- Sev3: localized or edge-case model degradation with limited business impact, no confirmed data exposure, affecting a small population or a non-critical feature.
- Sev4: cosmetic or negligible-impact anomalies requiring tracking but not urgent response.
Give each threshold concrete, measurable criteria rather than leaving it to judgment call each time: for example, define "accuracy collapse" as a specific percentage-point drop against a rolling baseline, define "PII exposure" thresholds by estimated affected-user count, and define "regulator-impacting" by whether the affected decision category (lending, employment, healthcare) is one your legal team has already flagged as high-risk.
Coordinating the first hours of a Sev1. Convene engineering, legal, and executive leadership together, not sequentially, since a confirmed PII leak or regulator-impacting error needs all three perspectives simultaneously: engineering to contain and assess technical scope, legal to assess notification obligations and regulatory exposure, and executives to make resourcing and external-communication calls quickly. Assign a single incident commander to own the coordination rather than letting three parallel, uncoordinated workstreams emerge.
Worked example
An automated loan-approval model is found to be systematically denying applications from a specific demographic at a disproportionate rate due to a data-pipeline bug introducing a biased feature. Given the regulator-impacting nature of lending decisions, this is classified Sev1 even though there's no direct PII leakage and the model itself hasn't crashed. Within the first hour, engineering, legal, and an executive sponsor are convened together: engineering begins tracing the specific feature causing the bias and prepares an immediate mitigation (reverting to the prior model version), legal assesses fair-lending regulatory exposure and notification obligations, and the executive sponsor authorizes the immediate rollback despite the operational cost of reverting a recently-improved model, given the regulatory severity classification.
Trade-offs and pitfalls
Defining severity thresholds vaguely ("high impact" without a measurable definition) leads to inconsistent classification between incidents and slows down the decision of how urgently to respond; concrete, pre-agreed thresholds remove that ambiguity exactly when speed matters. A second common mistake is classifying by technical severity alone (did the model crash) without weighing business and regulatory impact, which can under-classify an incident like the lending-bias example above, where the model technically "worked" but caused serious harm.
Tell me about a time you organized, led, or participated in a tabletop exercise or incident drill. Describe your role, a key decision point, what the exercise revealed, and one concrete operational change that resulted from it.
Sample Answer
Direct answer
The most useful version of this story names a specific, concrete gap the exercise revealed, describes the decision point where it became visible, and ends with a real change made afterward, not just "it went well" or "we learned a lot."
Structured elaboration
Structure the answer around: your specific role (participant, scribe, technical lead, incident commander), one genuine decision point during the exercise where the team had to actually choose between options rather than just discuss abstractly, what that decision revealed (a gap in the playbook, a missing owner for some step, a tool that didn't behave as expected), and the concrete operational change that came out of it afterward (a specific playbook update, a new pre-authorized action, a training gap addressed).
If you haven't personally led an exercise, describing one you contributed to is entirely legitimate: focus honestly on your specific contribution and what you personally observed or learned, rather than inflating your role or speaking for the whole exercise's outcome.
Worked example
"I participated in a tabletop simulating a ransomware outbreak, playing the on-call security analyst role. Partway through, the facilitator injected new information that our most recent backup was found to be incomplete due to an unrelated failure. As the analyst, I realized in the moment that our actual playbook had no explicit step requiring backup-currency verification before a restore decision, meaning in a real incident we might have proceeded with a stale or incomplete restore without anyone catching it. I raised this during the debrief, and the exercise lead added it as a tracked action item. Two weeks later, the playbook was updated to include an explicit backup-verification checkpoint with a named owner, and the following quarter's tabletop specifically tested whether that step was now being followed."
Trade-offs and pitfalls
A vague answer ("it was a good learning experience, we identified some gaps") gives an interviewer nothing to evaluate and reads as though you didn't actually engage deeply with the exercise. Overstating your role, claiming to have led an exercise you only participated in, is a common temptation but is usually easy for an experienced interviewer to probe past with a follow-up question about a specific decision you personally made.
You are notified that a live production database primary shows signs of compromise (suspicious administrative queries, an unexpected new privileged account, or noisy unauthorized writes). Design a short-term containment plan that minimizes downtime while preserving forensic evidence: options include isolating or failing over the primary, taking read-only mode, snapshotting for forensics, and communicating with dependent application owners. State the assumptions you make.
Sample Answer
Direct answer
Contain by isolating or failing over the compromised primary while preserving forensic evidence, favoring a read replica promotion or read-only mode over an abrupt hard shutdown, and communicate with dependent application owners before and during the change since a database failover has broad ripple effects.
Structured elaboration
Options, roughly from least to most disruptive:
- Read-only mode. If the concern is unauthorized writes specifically, switching the primary to read-only stops further tampering while keeping reads (and therefore much of the application) functioning, buying time to investigate without a full outage.
- Promote a replica, isolate the primary. If a healthy read replica exists, promoting it to primary and isolating the original (rather than powering it off) preserves the compromised instance for forensic snapshotting while minimizing downtime, since application traffic can be redirected to the newly promoted primary.
- Snapshot before any destructive action. Whichever path you take, snapshot the compromised primary's current state (disk and, if feasible, memory) before doing anything that would alter it further, since this may be the only forensic evidence of exactly what the suspicious administrative queries or the new privileged account actually did.
- Full isolation as a last resort. If no healthy replica exists and read-only mode isn't sufficient (for example, the compromise itself is at the infrastructure layer, not just the application-data layer), full isolation with accepted downtime may be necessary, but this should be the option of last resort given its business cost.
Communicating with dependent application owners. A database failover or read-only switch has broad ripple effects across every service reading or writing to it; notify dependent teams before executing wherever the timeline allows a few minutes' warning, and immediately after if speed required acting first, so they aren't debugging a mystery outage in parallel with your investigation.
Assumptions to state explicitly, since this scenario is genuinely ambiguous without more detail: whether a healthy, up-to-date replica exists and how far behind it might be (a stale replica means promoting it could lose recent legitimate writes); whether the suspicious activity is ongoing or already stopped (ongoing activity favors faster, more disruptive containment); and whether the new privileged account itself is still active (if so, containing the account, not just the database, is equally urgent).
Worked example
A production database primary shows a new, unrecognized privileged account and unusual administrative queries querying and exporting large volumes of customer data. Assuming a healthy, near-real-time replica exists (state this assumption explicitly): the team promotes the replica to primary, redirects application traffic to it, and isolates the original primary at the network layer rather than shutting it down, preserving it for a forensic snapshot. In parallel, the newly discovered privileged account is immediately disabled everywhere it might have reach, not just on this one database. Dependent application owners are notified of the brief failover window and asked to watch for any application-level errors during the cutover. The isolated original primary is later imaged for forensic analysis to determine the full scope of the suspicious queries before it's decommissioned.
Trade-offs and pitfalls
Failing over to a replica that's more than a few seconds or minutes stale risks silently losing legitimate recent writes, a real cost that needs to be weighed against the benefit of a fast, low-downtime containment; teams sometimes assume replication lag is negligible without actually checking it in the moment. Powering off the primary rather than isolating and snapshotting it is a common instinct under pressure but destroys volatile evidence (an active malicious session's state, in-memory query cache) that a network-level isolation would have preserved.
Unlock Full Question Bank
Get access to all Incident Response and Containment interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.