Operational Risk Management Questions
Identifying, assessing, and reducing operational risk before it becomes an incident. Covers operational risk categories (process, people, supplier, technology, execution), surfacing the risks in a large program such as a cloud migration, risk registers and ownership cadence, likelihood and impact scoring (heat maps, qualitative vs quantitative, expected loss, ranges and Monte Carlo, estimating with little history), scenario analysis and structured failure-mode review before a risky change, key risk indicators and early-warning signals, risk response strategies (avoid, reduce, transfer, accept) including contracts and insurance for supplier exposure, prioritizing mitigations by expected loss and cost per unit of risk reduced, residual risk reporting and escalation to leadership, key-person risk and single points of failure, risk appetite and risk-versus-speed trade-offs, systemic and recurring risk including human error, the three lines model (formerly three lines of defense), and building organizational resilience (resilience metrics, resilience testing programs, culture). Proactive risk reduction, not reactive incident handling. Disaster recovery and continuity planning, incident command, vendor due diligence, security and privacy risk, and project schedule risk are covered elsewhere.
What is a key risk indicator, and how does it differ from a KPI or an SLO? Give three examples for a major feature rollout and explain what makes a trigger threshold actionable instead of noise.
Sample Answer
Direct answer. A key risk indicator (KRI) is a leading signal that exposure to a specific risk is rising. A KPI (key performance indicator) reports how well you are doing against a goal. An SLO (service-level objective) is a target for a service-level measurement, such as availability. The same number can play different roles, but a KRI always points forward to a risk and carries a threshold that triggers an action.
| KRI | KPI | SLO | |
|---|---|---|---|
| Question | Is a risk getting more likely or bigger? | Are we hitting the business goal? | Is the service reliable enough? |
| Looks | Forward | Mostly backward | Over a window against a target |
| Example | Support contacts per 1,000 active users doubling | Weekly active users | 99.9% of requests succeed in 30 days |
| Triggers | A planned response | A performance review | An error-budget policy (the pre-agreed rule for what to do as the budget runs down, such as freezing releases) |
Error budget in one line. An SLO of 99.9% over 30 days allows 0.1% failure, which is 43.2 minutes of downtime. That allowance is the error budget; burn rate is how fast it is being spent.
Three KRIs for a major feature rollout. Illustrative thresholds.
- Error budget burn rate for the feature's SLO. A rate of 1 uses the budget exactly over the period; 2 would exhaust a 30-day budget in 15 days. Amber at 2, red at 5 (budget gone in 6 days).
- Support contacts per 1,000 active users of the feature, against the pre-launch baseline. Illustrative baseline: 4 contacts per 1,000 users per week, so amber at 1.5x is 6 and red at 2x is 8.
- Days since the last successful rollback rehearsal. Amber above 30, red above 60. It warns that your escape route may no longer work.
What makes a threshold actionable instead of noise
- It comes from a baseline or a limit, not a round number.
- It has one owner and a stated action at each level.
- It fires early enough to act and rarely enough to be believed.
- It uses a window (such as 15 minutes) so one spike does not page anyone.
- You review it after every miss and every false alarm.
Contingency triggers. A KRI that crosses its red line can start a response that was agreed in advance. Typical examples for the same rollout, with the numbers (4 hours, 30 minutes, 2 days) set per contract and per risk. SLA below means the service-level agreement, the contractual availability promise from a partner:
| Trigger | Immediate action | Owner | Success criterion |
|---|---|---|---|
| Partner misses its SLA for 4 hours | Route traffic to the fallback provider | On-call lead | Fallback serves the traffic within 30 minutes |
| Regulatory notice received | Pause the feature in the affected jurisdiction | Compliance lead | Written response filed before the regulator's deadline |
| Payment failures above baseline for 2 days in a row | Enable backup processor and retries | Payments lead | Payment success back to baseline within a day |
Pitfall. A "KRI" that is only a KPI renamed, with no threshold, owner or action.
You are about to run a schema migration that touches many services. Before you approve it, how do you run the risk review: what could go wrong, what blast radius you will accept, what rollback and progressive rollout you would require, and which triggers and signals decide go or no-go?
Sample Answer
Direct answer
I approve only a plan where every step before the final cleanup is reversible, the first exposure is small, and the go or no-go decision comes from signals agreed beforehand. The method is expand, migrate, contract (a parallel change): add the new structure so old and new code both work, move data and traffic gradually, and delete the old structure only after a hold period.
Example: renaming column customer_name to full_name
- Expand: add the
full_namecolumn (nullable, so existing code is unaffected). - Migrate: deploy code that writes both columns, backfill (copy existing rows in small batches) old values into
full_name, then switch reads to the new column behind a feature flag (a switch that turns new behaviour on or off without a deploy). - Contract: after the hold period, drop
customer_name. Only this step cannot be undone, so it gets its own approval. Renaming in one step instead would break every service still reading the old name.
What could go wrong (six failure modes)
- Old and new code versions disagree during rollout, since services deploy at different times.
- A lock or table rewrite blocks writes (some changes make the database hold the table exclusively or copy it, so writes queue up).
- Replication lag (how far read copies of the database trail the primary) or load spikes hurt other work.
- A backfill silently corrupts or drops data.
- Unknown consumers (batch jobs, analytics, partners) break.
- An irreversible step (dropping a column) has no way back.
Blast radius I will accept
The blast radius is how much can break if this goes wrong. Start with one low-risk service and a small slice (illustrative: 1% of traffic, one shard, which is one horizontal partition of the data, or one tenant group). I accept losing a slice of requests briefly; I do not accept data loss or corruption. Preconditions: rehearsal on production-sized data, a restore test of backups, an abort command in the runbook, and a dependency inventory built from query logs, not memory.
Rollback and progressive rollout (illustrative stages)
| Stage | Exposure | Hold |
|---|---|---|
| 1 | 1% | 30 minutes |
| 2 | 5% | 1 hour |
| 3 | 25% | 2 hours |
| 4 | 100% | 24 hours before the contract step |
The first three holds total 3.5 hours. Until the contract step, rollback is a feature flag flip or stopping the backfill. Contract is the point of no return: a separate approval, a verified restore, and no Friday-afternoon runs (little time and staff to react) or freeze-window runs (periods when changes are banned, such as peak season).
Go or no-go triggers (thresholds illustrative, set from baselines)
| Signal | No-go or halt |
|---|---|
| Error rate | Above 2x baseline for 5 minutes |
| 99th percentile latency (the time 99 in 100 requests beat) | Above 1.5x baseline for 5 minutes |
| Replication lag | Above the limit derived from the recovery point objective (RPO: maximum tolerable data loss) |
| Blocked queries or lock waits | Beyond the agreed limit |
| Dual-write comparison (checking that old and new columns hold the same data) | Any mismatch |
| People | Owners not on call |
Measure each trigger on the exposed slice against the unexposed remainder (a control group), not against the whole fleet. At 1% exposure a fleet-wide "2x baseline" trigger would only fire if the slice's own error rate were about 100 times normal (0.99 x e + 0.01 x s = 2e gives s = 101e), so a badly broken slice would pass. Also set a minimum request count per window so a quiet 5 minutes at 1% cannot look healthy by having almost no traffic.
Who signs what
Each affected service owner signs compatibility; the SRE or database owner signs the rollout plan; one named migration lead can halt at any stage without a committee; an engineering manager or executive accepts the residual risk of the contract step.
What would change my call
If the system is small and a maintenance window is acceptable, a simple window is cheaper. If the consumer inventory is incomplete, I delay until it is.
You discovered repeated human error was causing high-severity incidents. Design a plan combining process, tooling, and training to eliminate the top 3 human-error causes. Include measures to detect recurrence and KPIs to demonstrate improvement.
Sample Answer
Direct answer
Treat "human error" as a symptom, not a cause: ask why the system made the mistake easy. Rank the causes, then eliminate or guard against each one in the order the hierarchy of controls prefers. The hierarchy ranks fixes by how little they depend on people remembering something, so the strongest comes first. Using "someone runs a destructive command in production" as the example: (1) remove the hazard, so nobody has standing production write access; (2) build guardrails into tooling, so the command asks you to type the environment name before it runs; (3) change process, so a second person reviews risky changes; (4) training last, because a reminder fades while a guardrail keeps working. Detect recurrence with tags and audit signals, and prove improvement with rates, not counts.
Baseline and the top three causes (illustrative data)
Suppose a review of 20 severity-1 and severity-2 incidents (the most serious two levels on a scale where 1 is worst) over two quarters finds 16 caused by three patterns: manual edits to production configuration (7), commands run in the wrong environment (5), missed runbook steps during deploys (4). Together 7 + 5 + 4 = 16 of 20, or 80%.
| Cause | Process | Tooling (guardrail) | Training | Detect recurrence |
|---|---|---|---|---|
| Manual prod config edits | Config changes only through reviewed pipeline | Schema validation plus automatic canary (release to a small slice of traffic first and roll back if it looks bad); break-glass access (emergency access that is allowed but logged and reviewed) with audit | Short walkthrough of the pipeline | Alert on any direct edit; weekly count should be zero |
| Wrong-environment commands | Separate credentials per environment | Distinct prompts and colours, dry-run by default (the command only shows what it would do until you add a flag to execute), typed-environment confirmation for prod. Before: the same terminal prompt for staging and production, so a delete aimed at staging can hit prod. After: the prompt turns red in prod and the command refuses until you type "prod" | Onboarding scenario | Audit log: prod commands from non-prod sessions |
| Missed runbook steps | Turn runbooks into checklists | Pre-flight checks in the deploy pipeline block when a step is missing | Game-day practice | Incident tag "missed step"; skipped-check count |
KPIs
- Outcome: human-error incidents per 100 changes. Baseline: if the same period had 1,600 production changes, 16 / 1,600 x 100 = 1.0 per 100. Illustrative target: 0.5 within two quarters, which is 8 incidents at the same volume.
- Leading: share of production changes through the pipeline (target 98% or more), break-glass uses per month, checks that blocked a bad change.
- Counter-measure: lead time for changes must not get worse, or people will bypass the guardrails.
- Not a KPI: training completion, which shows attendance, not safety.
Process for learning
Blameless postmortems (reviews that ask what let the failure happen, not who to blame) record contributing factors with a consistent tag, reviewed monthly, so the top three are re-ranked as the old ones disappear.
Pitfalls
Retraining the people involved and calling it fixed; adding approvals that slow everyone; measuring raw incident counts when change volume moves.
Pick an operational risk such as cascading failures during deployments. Outline a mitigation plan that covers detection signals, automated mitigation (e.g., circuit breakers), human runbooks, and post-incident monitoring. Explain how you would present this plan differently to SREs versus product managers.
Sample Answer
Risk chosen: a bad deployment that triggers a cascading failure (a failure that spreads from one service to the services depending on it). Typical path: a new version slows one service, callers time out, callers retry, and the retries overload the slow service further.
Why cascades grow. Take a chain of four services A calls B, B calls C, C calls D, and each of the three callers (A, B, C) makes up to 3 retries on its call (4 attempts). One user request can then become 4 x 4 x 4 = 64 calls arriving at the bottom service D. (With only three services in the chain there are two retrying hops, which gives 4 x 4 = 16; each extra retrying hop multiplies the load again.) Mitigation has to cap that amplification.
1. Detection signals
- Canary (releasing to a small share of traffic first) compared to baseline on error rate, p99 latency (the response time that 99% of requests beat, so it shows the slow tail) and saturation (how close a resource such as CPU, threads or connections is to its limit).
- Retry rate and queue depth rising on callers of the changed service.
- SLO burn-rate alerts (an SLO is a reliability target such as 99.9% of requests succeed; burn rate is how fast the allowed failures are being used up) on the customer-facing path.
- Deploy markers on every graph so a change is correlated in seconds.
2. Automated mitigation. Start with the controls that stop a bad release and stop the amplification; the later items add protection once those are in place.
- Progressive rollout (for example 1%, 10%, 50%, 100%) with automatic rollback when canary metrics breach thresholds.
- Timeouts shorter than the caller's own deadline.
- Circuit breaker: after failures cross a threshold (illustrative: 50% of the last 20 calls), stop calling for a cool-off (illustrative: 30 seconds), then let a few trial calls through (half-open) before closing again.
- Retries done safely: exponential backoff (wait longer after each failure, for example 1 s, 2 s, 4 s), jitter (a random offset so thousands of callers do not retry at the same instant) and a retry budget (a cap such as retries may add at most 10% extra traffic).
- Bulkheads (separate pools per dependency, like watertight compartments in a ship) and load shedding (deliberately rejecting low-priority requests when overloaded so critical ones still succeed), so one slow dependency cannot take all threads.
- Feature flags as a kill switch: a setting that turns a feature off instantly without a new deploy.
3. Human runbook (a short step-by-step guide for the on-call engineer)
- Check the deploy marker; if a change landed in the last 30 minutes, roll it back first, diagnose second.
- If rollback does not help, open breakers manually or shed non-critical traffic.
- Declare an incident when customer impact is confirmed, assign an incident commander (the one person who coordinates the response and decides, rather than debugging), post status updates.
- Record the timeline for review.
4. Post-incident monitoring
- Extended bake time (the watch period after a change at full traffic before calling it safe, for example 24 hours).
- New alert or test for the exact failure signature, and a game day (a planned failure drill) to prove the breaker trips.
- Track change failure rate (share of deployments that cause an incident or rollback) and time to restore (how long from impact starting to service restored).
Presenting differently
- SREs: thresholds, breaker settings, retry budgets, alert routing, runbook steps, what pages whom. They want mechanism and failure modes. Sample slide line: "Breaker opens at 50% errors over 20 calls, 30 s cool-off; canary auto-rolls back if p99 is 30% above baseline for 5 minutes; retries capped at 10% extra load."
- Product managers: customer impact in plain terms ("a bad release affects at most 1% of users for minutes, not everyone for an hour"), the cost in release speed, what trade-off we ask them to accept (slower full rollout), and the decision they own. Sample slide line: "Rollouts now take about two days instead of one hour. In return, a bad release reaches 1% of customers for a few minutes. Decision for you: accept the slower rollout, or fund the faster automation."
Pitfall. Circuit breakers with untested thresholds either never trip or trip constantly; rehearse them.
You are asked to build a resilience testing program for a platform organization that today only does ad hoc failure tests. Design it as risk reduction, not incident response: what you test and how often, how you measure whether resilience is improving, and how findings turn into funded fixes.
Sample Answer
Direct answer
A game day is a scheduled exercise where a team deliberately breaks something (or simulates it) and practises responding. Fault injection is causing a failure on purpose with a tool. Tier-1 services are the ones whose failure hurts revenue or safety most; tier-2 are the next most important. A risk register is the running list of known risks with owners. Treat every test as an experiment against a named risk: "if X fails, we believe Y still works." Tier the tests (automated spot tests, team game days, cross-team drills, one full-scale rehearsal a year), judge improvement by outcomes of those experiments, and route every finding into the risk register with an owner, a due date and reserved capacity so it becomes a funded fix. Example of one experiment: "If the primary orders database fails, we believe checkout still works from the replica within 5 minutes." Run it, record whether it held, and file the gap if it did not. This is risk reduction before customers are hurt, so disaster-recovery planning and incident command stay out of scope.
What to test (six areas; the listed order is a default, not a ranking)
- Dependency failure (a database, queue or supplier slows or dies).
- Overload and capacity (traffic spike, a noisy neighbour: another workload sharing the same hardware that hogs resources).
- Bad change (faulty deploy or config, and whether rollback works).
- Data (a restore from backup actually meets its recovery point objective, RPO: the most data you accept losing).
- People and process (on-call response when the usual expert is unavailable).
- Zone or region loss.
The order above is only a starting default. Re-rank the six areas from your own risk register: score each by how likely it is to hurt tier-1 services and how much a test would reduce that risk. That is why the year-one sequence below starts with bad-change rollback even though it is listed third.
Start with tier-1 services (the revenue or safety path). Run in staging first, then production behind guardrails: a defined blast radius (the most users or systems the test can affect, for example one zone or 1% of traffic), an abort condition (the signal that ends the test, such as error rate above a set limit) and a named stop authority (the one person who can halt it without asking anyone). In year one, a sensible sequence is automated spot tests on tier-1 dependencies and bad-change rollback, then team game days, then the broader areas.
Annual calendar (illustrative sizing)
| Level | Frequency | Participants | Objective | Person-days |
|---|---|---|---|---|
| Automated spot tests | Weekly, tier-1 and tier-2 | Run by tooling, owners review | Catch regressions in known weaknesses | tooling cost |
| Team game days | 12 tier-1 services, twice a year = 24 | About 6 people each | Validate runbooks, alerts, failover | 24 x 6 = 144 |
| Cross-team drills | Quarterly = 4 | About 20 people | Test shared dependencies and handoffs | 4 x 20 = 80 |
| Full-scale rehearsal | Once a year | About 40 people for 3 days | Multi-service failure under realistic pressure | 40 x 3 = 120 |
Basis: each team game day and cross-team drill is assumed to take one day for every participant (the rehearsal is stated as 3 days), and preparation and follow-up time are not counted, so real effort is higher. Participation totals 344 person-days, about 1.6 full-time people at 220 working days a year (344 / 220 = 1.56). The budget request is therefore: that time (mostly existing staff), a fault-injection tool, a production-like test environment, and one or two program owners.
Is resilience improving?
- Hypothesis hold rate: share of experiments where the system behaved as predicted. It may fall at first as tests get harder, which is progress.
- Time to detect and recover from injected faults versus targets.
- Count and age of single points of failure.
- Coverage: share of tier-1 services tested in the last six months (target 100%).
- Recurrence of the same finding class.
Findings to funded fixes
Each finding gets a severity, an entry in the register, one owner and a due date by severity (illustrative: critical in 30 days, high in 90). Teams reserve a standing share of capacity for such fixes (a policy choice, for example 15%), cross-team items go to a monthly leadership review, and the case is made in expected-loss terms (probability of the loss times its size per year). Illustrative: a game day finds that failover to the replica was never tested and would take hours. Estimated 10% a year chance of a $300,000 outage, so about $30,000 a year of expected loss; the fix takes 20 person-days, about $16,000 at an assumed $800 a day, and removes most of the exposure (say 90%, or $27,000 a year). The fix is funded because $27,000 a year avoided exceeds the $16,000 cost. Put another way, the fix pays back in about 7 months (16,000 / 27,000 x 12 = 7.1), and because it compares a one-time cost with a yearly benefit it only holds if the fix keeps working for at least that long. A declined fix is formally accepted by an executive, not silently dropped.
Pitfalls
Tests that only confirm what is already known; no stop authority; findings that age without owners.
Unlock Full Question Bank
Get access to all 11 Operational Risk Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.