Operational Risk Management Questions
Identifying, assessing, and reducing operational risk before it becomes an incident. Covers operational risk categories (process, people, supplier, technology, execution), surfacing the risks in a large program such as a cloud migration, risk registers and ownership cadence, likelihood and impact scoring (heat maps, qualitative vs quantitative, expected loss, ranges and Monte Carlo, estimating with little history), scenario analysis and structured failure-mode review before a risky change, key risk indicators and early-warning signals, risk response strategies (avoid, reduce, transfer, accept) including contracts and insurance for supplier exposure, prioritizing mitigations by expected loss and cost per unit of risk reduced, residual risk reporting and escalation to leadership, key-person risk and single points of failure, risk appetite and risk-versus-speed trade-offs, systemic and recurring risk including human error, the three lines model (formerly three lines of defense), and building organizational resilience (resilience metrics, resilience testing programs, culture). Proactive risk reduction, not reactive incident handling. Disaster recovery and continuity planning, incident command, vendor due diligence, security and privacy risk, and project schedule risk are covered elsewhere.
A critical third-party API your service relies on is intermittently failing. You can build a local caching/fallback layer (weeks) or press the vendor for SLA improvements and dedicated support (uncertain timeline). As a senior SRE, describe your decision process, immediate risk mitigations, long-term strategy, and vendor management considerations.
Sample Answer
Direct answer. I would not choose between the two options: I would do both in sequence. This week I add immediate protections on my side (they need no vendor cooperation), measure the real impact, start a scoped fallback for the highest-value paths, and escalate to the vendor with evidence in parallel. I would not stake the service on an uncertain vendor timeline.
Decision process
- Measure. Failure rate, which calls fail, revenue or user impact, and whether failures cluster in time. Example: 1,000,000 calls a month at 2% failure is 20,000 failed calls.
- Classify the traffic. Reads that can be served stale (cacheable) versus writes that cannot (payments, orders). A cache only helps the first class.
- Put a price on it. Compare the cost of the fallback with the loss it removes. Illustrative: if each failed call costs about $1 in lost margin and support, the 20,000 failed calls cost about $20,000 a month. A fallback that answers 70% of them saves about $14,000 a month; if it takes 6 engineer-weeks at about $4,000 a week ($24,000), it pays back in under 2 months. If the failing share were 0.2%, the loss would be $2,000 a month and the same build would take over 12 months to pay back, so I would only add timeouts and retries and press the vendor. Also weigh the vendor option as a probability, not a promise: a fix that might land in 3 months saves nothing in the meantime, and dedicated support does not remove failures that cluster in bursts.
- Decide. If a large share of failures hit cacheable paths and impact is real, build the fallback now. If the impacted calls are non-cacheable, caching is the wrong tool: queue-and-retry or a second provider is.
Immediate mitigations (days, no weeks of build). Timeouts, retries with backoff, a circuit breaker and a cache or degraded mode do most of the work for most services; the later items matter only when the specific situation applies.
- Timeouts and a retry budget (a cap on how many retries you allow) with exponential backoff (each wait longer than the last, such as 1s, 2s, 4s) and jitter (randomised delay so clients do not retry in lockstep), honouring any Retry-After header (the vendor telling you how long to wait).
- A circuit breaker: after repeated failures, stop calling for a period and fail fast.
- Idempotency keys: a unique ID sent with each request so the vendor recognises a repeat of the same request and does the action only once, so a retry does not double-charge or double-create.
- Alerting on error rate, and a status-page subscription.
- Degrade gracefully: hide the feature or show cached data.
Illustration: suppose 70% of the failing calls are reads for which we hold a saved copy, possibly slightly stale. Those can be answered from the cache, so only the other 30% still fail for users: 20,000 x 0.3 = 6,000. One independent retry would cut a 2% failure rate to 0.04% (0.02 squared), but intermittent failures usually cluster in bursts, so retries help much less than that independence figure.
Long-term strategy. A thin adapter in front of the vendor so the code depends on an interface, not the vendor. Add a secondary provider on warm standby (already deployed and taking a trickle of traffic, so it can take over fast) for critical paths, and a cache with explicit staleness limits.
Vendor management
- Build an evidence packet (timestamps, request IDs, error codes) and use the named escalation contact, not the general support queue.
- Read the SLA: a 99.9% monthly SLA allows 43.2 minutes of downtime per 30 days (43,200 minutes x 0.001), and credits rarely cover lost revenue. Our own availability is bounded by the vendor's, since both must work.
- At renewal ask for notice periods on limit changes, termination for chronic failure, data export rights, and a better SLA.
Rate-limit change with no notice. Same toolkit: client-side throttling (a token bucket: tokens refill at a fixed rate and each request spends one, so short bursts are allowed up to the bucket size while the long-run rate stays capped), backoff on 429 responses (the HTTP status meaning 'too many requests'), request coalescing (merging identical simultaneous requests into one call to the vendor), shedding low-priority calls first (dropping the least important work to protect the important), then a cross-functional response: product reduces calls, finance prices a higher tier, procurement renegotiates.
Single-source authentication or search vendor. Short term: cache sessions or tokens so signed-in users continue, keep a break-glass path (an emergency, audited bypass that lets staff in when the vendor is down). Long term: abstraction layer, a second vendor tested periodically, and contract changes (exit rights, SLA credits, change notice).
Pitfall. Building for weeks without data on what is failing.
Your product depends on one payment processor, and an outage would cost roughly $200,000 per hour. You are weighing paying for a second provider versus accepting the dependency. How would you build an expected-loss model for this supplier failure to decide, and which inputs would you find hardest to estimate?
Sample Answer
Direct answer
Model the annual loss with and without a second provider, count only the outage hours the second provider can actually bypass, and compare the saving with its yearly cost. With the illustrative inputs below the second provider pays for itself in the middle case and loses money in the low case, so the answer depends on outage frequency, which I would measure before committing. The hardest inputs are the share of outages failover can truly bypass, the quality of failover, and the tail of outage duration.
The model
Variables: failover means switching traffic to the backup provider when the main one fails. The covered share is the portion of outages that are the processor's own fault and so can be bypassed. The recovered share is the portion of lost sales customers retry later. The failover gap is the minutes of loss while the switch happens.
loss now = outages per year x mean hours x $200,000 x (1 - recovered share)
loss with 2nd = loss now x (1 - covered share) + failover gap loss
failover gap = outages x covered share x failover hours x $200,000 x (1 - recovered share)
Outages from your own bugs or shared networks stay in the loss. Example of estimating covered share from data: if 24 months of logs show 5 outages (2.5 a year, between the mid and high cases) and 4 were the processor's own fault, covered share is 4 / 5 = 0.8. Better still, weight by outage hours, because the model multiplies hours: if the processor-caused outages were the short ones, the count share overstates what failover saves.
Run it (inputs illustrative: $90,000 a year to run plus $250,000 to build spread over 3 years)
cost_per_hour = 200_000
def annual_loss(outages, mean_hours, recovered):
# recovered: share of lost sales that customers retry later and complete
return outages * mean_hours * cost_per_hour * (1 - recovered)
def loss_with_second_provider(outages, hours, recovered, covered, failover_minutes):
# covered: share of outages the second provider can bypass
base = annual_loss(outages, hours, recovered)
gap = outages * covered * (failover_minutes / 60) * cost_per_hour * (1 - recovered)
return base, base * (1 - covered) + gap
second_cost = 90_000 + 250_000 / 3 # yearly run cost + build cost spread over 3 years
print(f"second provider annual cost ${second_cost:,.0f}")
cases = {
"low": (1.0, 1.0, 0.40, 0.70, 15),
"mid": (2.0, 1.5, 0.30, 0.80, 15),
"high": (3.0, 2.5, 0.15, 0.90, 30),
}
for label, args in cases.items():
base, after = loss_with_second_provider(*args)
saving = base - after
print(f"{label:5s} loss now ${base:>9,.0f} after ${after:>9,.0f} saving ${saving:>9,.0f} net ${saving - second_cost:>10,.0f}")
second provider annual cost $173,333
low loss now $ 120,000 after $ 57,000 saving $ 63,000 net $ -110,333
mid loss now $ 420,000 after $ 140,000 saving $ 280,000 net $ 106,667
high loss now $1,275,000 after $ 357,000 saving $ 918,000 net $ 744,667
Reading the result
Mid case: loss now is 2 x 1.5 x $200,000 x 0.7 = $420,000; the second provider removes $336,000 of it but failover time costs $56,000 of new loss, so the saving is $280,000 and the net is about $107,000 a year. In the low case the saving ($63,000) does not cover the cost ($173,333). In the high case the net is roughly $745,000.
Hardest inputs to estimate
- Covered share: most outages are partly caused by things a second processor cannot fix (your own code, a shared bank or card network issue). Review incident history for cause.
- Failover quality: code that has never run in production often fails. A second provider's approval rates (share of card payments it accepts), routing rules (which transactions go to which provider) and reconciliation (matching payments to the books) also matter.
- Frequency and duration tail: a few years of status-page history is thin, and costly outages are long ones.
- Recovered share: depends on checkout behaviour.
Decision
Gather 24 months of the processor's incident history and our own logs first, then reassess. If the mid case holds, I would fund the second provider, and rehearse failover quarterly so the covered share is a tested figure. Also check what the contract's service credits (refunds the provider owes for downtime) pay: they rarely cover lost sales. Accepting the dependency is rational only if low-case inputs are confirmed.
Pitfalls
Counting every outage hour as saved; ignoring the second provider's own failures and complexity cost.
You are about to run a schema migration that touches many services. Before you approve it, how do you run the risk review: what could go wrong, what blast radius you will accept, what rollback and progressive rollout you would require, and which triggers and signals decide go or no-go?
Sample Answer
Direct answer
I approve only a plan where every step before the final cleanup is reversible, the first exposure is small, and the go or no-go decision comes from signals agreed beforehand. The method is expand, migrate, contract (a parallel change): add the new structure so old and new code both work, move data and traffic gradually, and delete the old structure only after a hold period.
Example: renaming column customer_name to full_name
- Expand: add the
full_namecolumn (nullable, so existing code is unaffected). - Migrate: deploy code that writes both columns, backfill (copy existing rows in small batches) old values into
full_name, then switch reads to the new column behind a feature flag (a switch that turns new behaviour on or off without a deploy). - Contract: after the hold period, drop
customer_name. Only this step cannot be undone, so it gets its own approval. Renaming in one step instead would break every service still reading the old name.
What could go wrong (six failure modes)
- Old and new code versions disagree during rollout, since services deploy at different times.
- A lock or table rewrite blocks writes (some changes make the database hold the table exclusively or copy it, so writes queue up).
- Replication lag (how far read copies of the database trail the primary) or load spikes hurt other work.
- A backfill silently corrupts or drops data.
- Unknown consumers (batch jobs, analytics, partners) break.
- An irreversible step (dropping a column) has no way back.
Blast radius I will accept
The blast radius is how much can break if this goes wrong. Start with one low-risk service and a small slice (illustrative: 1% of traffic, one shard, which is one horizontal partition of the data, or one tenant group). I accept losing a slice of requests briefly; I do not accept data loss or corruption. Preconditions: rehearsal on production-sized data, a restore test of backups, an abort command in the runbook, and a dependency inventory built from query logs, not memory.
Rollback and progressive rollout (illustrative stages)
| Stage | Exposure | Hold |
|---|---|---|
| 1 | 1% | 30 minutes |
| 2 | 5% | 1 hour |
| 3 | 25% | 2 hours |
| 4 | 100% | 24 hours before the contract step |
The first three holds total 3.5 hours. Until the contract step, rollback is a feature flag flip or stopping the backfill. Contract is the point of no return: a separate approval, a verified restore, and no Friday-afternoon runs (little time and staff to react) or freeze-window runs (periods when changes are banned, such as peak season).
Go or no-go triggers (thresholds illustrative, set from baselines)
| Signal | No-go or halt |
|---|---|
| Error rate | Above 2x baseline for 5 minutes |
| 99th percentile latency (the time 99 in 100 requests beat) | Above 1.5x baseline for 5 minutes |
| Replication lag | Above the limit derived from the recovery point objective (RPO: maximum tolerable data loss) |
| Blocked queries or lock waits | Beyond the agreed limit |
| Dual-write comparison (checking that old and new columns hold the same data) | Any mismatch |
| People | Owners not on call |
Measure each trigger on the exposed slice against the unexposed remainder (a control group), not against the whole fleet. At 1% exposure a fleet-wide "2x baseline" trigger would only fire if the slice's own error rate were about 100 times normal (0.99 x e + 0.01 x s = 2e gives s = 101e), so a badly broken slice would pass. Also set a minimum request count per window so a quiet 5 minutes at 1% cannot look healthy by having almost no traffic.
Who signs what
Each affected service owner signs compatibility; the SRE or database owner signs the rollout plan; one named migration lead can halt at any stage without a committee; an engineering manager or executive accepts the residual risk of the contract step.
What would change my call
If the system is small and a maintenance window is acceptable, a simple window is cheaper. If the consumer inventory is incomplete, I delay until it is.
Your team has a roadmap full of features and a flat cost target, and reliability work keeps getting pushed back. How do you decide when to invest in resilience versus defer it, and what rubric would you show leadership?
Sample Answer
Direct answer
Stop arguing resilience versus features in the abstract and give leadership a rubric that turns reliability work into a decision with evidence and a price. My proposal: a trigger from the service-level objective (SLO: the reliability target, for example 99.9% of requests succeed over 30 days), a payback test for sizing, and a protected capacity slice. A flat cost target means each resilience investment must either pay back soon or be funded by something stopped.
Rubric to show leadership
| Criterion | Rule (starting proposal) | Outcome |
|---|---|---|
| Error budget | The error budget is the amount of unreliability the SLO permits (0.1% of a 30-day window, computed below). Trigger: budget fully spent | Reliability work preempts features until restored |
| Budget warning | 70% of the budget spent before the window ends | Tell leadership and slow risky releases; the trigger has not fired yet |
| Repeat cause | Same root cause caused two incidents in a quarter | Schedule the fix this quarter |
| Payback | One-time cost divided by yearly expected loss avoided is 12 months or less | Do now |
| Single point of failure (one component or person whose loss alone stops a service) on the revenue path | One failure stops billing or checkout | Do now regardless of payback |
| Everything else | Over 12 months and none of the above | Defer, with a named revisit date |
The thresholds are proposals the leadership team tunes; the value is that they are written before the next argument.
Worked example (illustrative)
A 99.9% SLO over 30 days allows 30 x 24 x 60 x 0.001 = 43.2 minutes of downtime: that 43.2 minutes is the error budget. If 70% is used, 30.2 minutes are gone and 13.0 remain. That crosses the warning line but not the trigger (the trigger needs all 43.2 minutes spent).
- Option 1, automated failover (a standby copy takes over automatically when the primary fails): costs $90k. Incidents: 3 a year at 2 hours each and $12k per hour = $72k a year expected loss (the same $72k base is used for Option 2). The fix avoids 75%, so $54k a year. Payback 90 / 54 = 1.67 years, about 20 months. Defer (over the 12-month bar) unless it is also a single point of failure or the budget is spent.
- Option 2, a $30k partial fix avoiding 50%: $36k a year avoided, payback 30 / 36 = 0.83 years, about 10 months. Do now.
Consistency check on the incident sizes: a full 2-hour outage is 120 minutes, which is 2.8 times the 43.2-minute budget (120 / 43.2), so the trigger would fire after a single one in a 30-day window. For the Defer verdict on Option 1 to hold, read each of the 3 yearly incidents as a partial degradation that uses about 30 budget minutes (for example 2 hours with a quarter of requests failing, 69% of the budget, just under the 70% warning line), falling in different windows, with the $12k an hour as the loss during degradation. If the incidents are full outages, the budget is spent by the first one, the rubric says reliability work preempts features, and Option 1 moves to Do now.
Handling the flat cost target
Fund resilience by trade, not by hope: swap a lower-value feature out, remove redundant tooling, or choose cost-neutral changes (runbooks: written step-by-step procedures; alert tuning; game days: rehearsals where the team deliberately breaks something to practice the response). Reserve a capped slice of capacity (say 15 to 20%, my proposal) so reliability work is a standing line item.
Pitfalls: using only incident counts, ignoring that the avoided-loss estimate is uncertain (show a low and high case), and a rubric that nobody signs off.
You are scoping a nine-month migration of a monolith to the cloud for a mid-sized client. Before the plan is signed, how would you systematically surface the technical and operational risks, and what would you hand the team afterward so those risks drive decisions?
Sample Answer
Direct answer. Before signing, I would run a short, evidence-led risk discovery: map what the monolith really does, hold one structured workshop, prove the scariest unknowns with small spikes (time-boxed technical experiments), then price and gate the plan on what is found. Afterwards I would hand the team a ranked risk register with owners and triggers, a wave plan ordered by risk, go/no-go gates, rollback plans and a contingency reserve sized from the top risks.
Terms used here: cutover is the moment live traffic switches from the old system to the new; a wave is one batch of services migrated together; a go/no-go gate is a checkpoint where named people decide to proceed or stop based on a checklist; a rollback plan is the tested way back to the old system; a contingency reserve is spare time or budget held for expected slips; a trigger is the signal that says a risk is materialising.
1. Gather evidence first.
- Dependency and data mapping: batch jobs, file drops, scheduled tasks and integrations nobody documented (the usual source of surprises).
- Traffic and data profile: peak load, database size, cutover window the business can tolerate.
- Constraints: licences tied to hardware, compliance and data-residency rules, team skills, client decision speed.
- History: what went wrong on the client's earlier projects.
2. Run a 180-minute workshop. Participants: client engineering lead, operations or SRE (site reliability engineering), database owner, security, product owner, finance or procurement, our delivery lead and architect.
| Segment | Minutes |
|---|---|
| Scope and ground rules | 15 |
| Pre-mortem: everyone silently writes why the migration failed in month nine | 30 |
| Walk categories (data, integration, performance, people, supplier, regulatory) | 45 |
| Score likelihood and impact | 30 |
| Assign owners and responses | 40 |
| Agree what must be resolved before signature | 20 |
That is six segments summing to 180 minutes.
Scales used in the scoring segment (illustrative): likelihood is Low under 20%, Medium 20% to 50%, High over 50% within the nine months. Impact 1 is under a week of slip, 2 is 1 to 2 weeks, 3 is 3 to 5 weeks, 4 is 6 to 10 weeks or a customer-visible outage, 5 is a lost deadline or a data or security event. In a real session the probabilities would come from owners giving a low, likely and high estimate, checked against what happened on the client's earlier projects; the reserve table below uses illustrative values, not measurements.
3. Prove the top unknowns. Rehearse a data migration on a copy, load-test one extracted service, and confirm the hardest integration works end to end. A result changes the price or scope; a guess does not.
4. Hand-off package.
- Register: cause, event, consequence, score, owner, trigger date or signal, response. One row (illustrative): cause: undocumented nightly batch jobs read the monolith's tables; event: a job breaks after cutover; consequence: invoices stop for days; likelihood 40% (Medium), impact 4; owner: client database lead; trigger: any new dependency found in the discovery map; response: rehearse the job on the migrated copy before wave 3.
- Assumptions log with who validates each and by when.
- Wave plan: low-risk services first to learn; the riskiest cutover after rehearsal, each wave with a rollback plan and a go/no-go checklist. Illustrative: wave 1 is the internal reporting service (rollback: switch DNS back within an hour); wave 2 is the customer portal; wave 3 is the billing database cutover, allowed to proceed only if the rehearsal finished inside the agreed cutover window and the parallel run showed matching totals.
- Contingency reserve from expected slip (illustrative probabilities):
| Risk | Probability | Slip if it happens (weeks) | Expected (weeks) |
|---|---|---|---|
| Hidden integrations found late | 40% | 4 | 1.6 |
| Data volume breaks cutover window | 30% | 6 | 1.8 |
| Key engineer unavailable | 25% | 3 | 0.75 |
| Licence or vendor delay | 20% | 5 | 1.0 |
| Client approvals slow | 15% | 8 | 1.2 |
The five expected values sum to 6.35 weeks, about 16% of a 39-week plan (nine months is roughly 39 weeks). Summing assumes the risks are independent, so I would present the figure as a basis for a reserve, not a promise. A reserve equal to the expected value is also only a coin-flip level of protection: simulating these five risks as independent events gives about a 57% chance that the total slip is at or below 6.35 weeks, while a reserve of about 10 weeks covers roughly 80% of outcomes and about 13 weeks roughly 90% (the worst case, all five landing, is 26 weeks). So I would agree the confidence level with the client (for example 80%) and size the reserve from that percentile rather than from the average.
Inheriting a project with no register. In the first 48 hours: day one, read the plan and interview five people, asking what worries them and what they are assuming (for example, "who would be paged if the nightly job failed?"); draft a starter register by the end of the day. Day two, validate it with the sponsor, assign owners, set a review cadence and book the first spikes.
Market entry or alliance. The same cycle applies: swap the evidence (local legal and regulatory review, partner operations) and the participants (partner and channel owners), and run the pre-mortem on why the entry or alliance failed. For example, for entering a new country, a pre-mortem note might read "we could not sign local customers because data had to stay in-country", which becomes a register row with a legal owner and a spike to confirm hosting options.
Pitfalls. A register nobody links to the schedule. Risks without owners. Treating a signed contract as the end of risk discovery.
Unlock Full Question Bank
Get access to all 12 Operational Risk Management interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.