Automation Scripting for Operations Questions
Writing scripts and tooling to automate operational and delivery tasks: shell and Python scripting, glue automation, toil reduction, and operational efficiency. Covers automating repetitive infrastructure and deployment work and building internal tooling that raises operational leverage. The concern is task-level automation and scripting, distinct from full pipeline or infrastructure-as-code frameworks.
Design a GitOps workflow where Python automation generates Kubernetes manifests, opens PRs into infra repositories, runs automated validation (policy checks, unit tests, Helm template rendering), and merges PRs on green while respecting release windows and SLO constraints. Describe webhook handling, how to prevent accidental auto-merges (policy gates), drift remediation when cluster state diverges, and how to safely roll out and rollback changes.
Sample Answer
Direct answer
The design's core requirement, "merge on green while respecting release windows and SLO (service-level objective) constraints," is a THREE-INDEPENDENT-CONDITION gate, not a single validation step: automated checks passing, a release window being open, and the current error-budget burn being within threshold ALL have to hold simultaneously, and each is a genuinely separate failure mode a real system encounters independently (checks can be green while it is 2 AM outside the window; the window can be open while the SLO is actively burning from an unrelated ongoing issue). Below is a runnable Python implementation of manifest generation, the validation pipeline, and this exact three-condition merge gate, executed against four distinct cases specifically chosen to prove the gate blocks on EACH condition independently, not just on validation failure.
Approach
- Generate the manifest from parameters (
generate_manifest), the automation's actual output artifact, a plain Python dict shaped like a Kubernetes Deployment. - Run three independent validation checks:
policy_check(image digest pinning and a replica-count bound),unit_test_manifest_shape(a structural test confirming the selector actually matches the pod template's own labels, catching a manifest that would deploy but never route traffic), andhelm_template_render_check(confirms every field a real template-rendering step would depend on is actually present). - The merge gate (
decide_merge) requires ALL THREE of: every validation category empty, the current time inside the configured release window, and the SLO error-budget burn below its threshold. Any ONE failing blocks the merge, with the SPECIFIC reason(s) reported, never a silent no-op or a generic failure message. - Drift remediation (
compute_drift) compares the manifest that was actually merged against a simulated live cluster state, isolating exactly which fields have diverged.
Code
import json
from datetime import datetime, time, timedelta
# ---------------------------------------------------------------------------
# 1. Manifest generation
# ---------------------------------------------------------------------------
def generate_manifest(service_name: str, image_digest: str, replicas: int) -> dict:
"""Generates a Kubernetes Deployment manifest from parameters. In a real
system this would render a Helm chart or Kustomize base; here it builds
the equivalent dict structure directly so the downstream validation
steps have something concrete and deterministic to check."""
return {
"apiVersion": "apps/v1",
"kind": "Deployment",
"metadata": {"name": service_name, "labels": {"app": service_name}},
"spec": {
"replicas": replicas,
"selector": {"matchLabels": {"app": service_name}},
"template": {
"metadata": {"labels": {"app": service_name}},
"spec": {"containers": [{"name": service_name, "image": image_digest}]},
},
},
}
# ---------------------------------------------------------------------------
# 2. Validation pipeline: policy checks, "unit tests", template-render check
# ---------------------------------------------------------------------------
def policy_check(manifest: dict) -> list:
"""Rejects a manifest referencing a mutable tag instead of a digest, and
a replica count outside a sane bound. Mirrors standard
image-tagging-policy and Rego-policy patterns, applied here as plain
Python for a self-contained demo."""
violations = []
image = manifest["spec"]["template"]["spec"]["containers"][0]["image"]
if "@sha256:" not in image:
violations.append(f"image '{image}' is not pinned to a digest")
replicas = manifest["spec"]["replicas"]
if not (1 <= replicas <= 50):
violations.append(f"replicas={replicas} is outside the allowed range [1,50]")
return violations
def unit_test_manifest_shape(manifest: dict) -> list:
"""A lightweight structural test: every referenced label selector must
actually match the pod template's own labels, catching a manifest that
would deploy successfully but never actually route traffic to its pods."""
errors = []
selector = manifest["spec"]["selector"]["matchLabels"]
pod_labels = manifest["spec"]["template"]["metadata"]["labels"]
for k, v in selector.items():
if pod_labels.get(k) != v:
errors.append(f"selector {k}={v} does not match pod template labels {pod_labels}")
return errors
def helm_template_render_check(manifest: dict) -> list:
"""Models the 'helm template renders without error' check: confirms
every field the rendering step depends on is actually present and of
the right type, rather than trusting the manifest is well-formed."""
errors = []
try:
containers = manifest["spec"]["template"]["spec"]["containers"]
if not isinstance(containers, list) or len(containers) == 0:
errors.append("no containers defined in pod template")
except KeyError as e:
errors.append(f"missing required path: {e}")
return errors
def run_validation(manifest: dict) -> dict:
return {
"policy": policy_check(manifest),
"unit_test": unit_test_manifest_shape(manifest),
"helm_render": helm_template_render_check(manifest),
}
# ---------------------------------------------------------------------------
# 3. Merge decision: green checks AND inside release window AND SLO not breached
# ---------------------------------------------------------------------------
class ReleaseWindow:
def __init__(self, start_hour: int, end_hour: int):
self.start_hour = start_hour
self.end_hour = end_hour
def is_open(self, at: datetime) -> bool:
return self.start_hour <= at.hour < self.end_hour
def decide_merge(validation: dict, release_window: ReleaseWindow, now: datetime,
current_error_budget_burn: float, slo_burn_threshold: float) -> dict:
"""The policy GATE the question asks for: merges only if EVERY check
passed AND the release window is open AND the SLO error-budget burn is
below threshold. Any one failing condition blocks the merge, and the
specific reason is reported, never a silent no-op."""
all_checks_green = all(len(v) == 0 for v in validation.values())
window_open = release_window.is_open(now)
slo_ok = current_error_budget_burn < slo_burn_threshold
if all_checks_green and window_open and slo_ok:
return {"merge": True, "reason": "all checks passed, inside release window, SLO burn within threshold"}
blockers = []
if not all_checks_green:
blockers.append("validation failed: " + json.dumps({k: v for k, v in validation.items() if v}))
if not window_open:
blockers.append(f"outside release window (window {release_window.start_hour}-{release_window.end_hour}h, now {now.hour}h)")
if not slo_ok:
blockers.append(f"SLO error-budget burn {current_error_budget_burn:.2f} exceeds threshold {slo_burn_threshold:.2f}")
return {"merge": False, "reason": "; ".join(blockers)}
# ---------------------------------------------------------------------------
# 4. Drift remediation when cluster state diverges from the merged manifest
# ---------------------------------------------------------------------------
def compute_drift(desired: dict, live: dict) -> dict:
diffs = {}
for key in ("replicas",):
d, l = desired["spec"].get(key), live.get("spec", {}).get(key)
if d != l:
diffs[key] = {"desired": d, "live": l}
desired_image = desired["spec"]["template"]["spec"]["containers"][0]["image"]
live_image = live.get("spec", {}).get("template", {}).get("spec", {}).get("containers", [{}])[0].get("image")
if desired_image != live_image:
diffs["image"] = {"desired": desired_image, "live": live_image}
return diffs
if __name__ == "__main__":
manifest = generate_manifest("checkout-api", "checkout-api@sha256:" + "a" * 64, replicas=6)
validation = run_validation(manifest)
print("Validation results:", json.dumps(validation, indent=2))
assert all(len(v) == 0 for v in validation.values()), "expected a fully clean manifest"
window = ReleaseWindow(start_hour=9, end_hour=17)
# Case A: inside window, SLO healthy -> should merge
decision_a = decide_merge(validation, window, now=datetime(2026, 3, 2, 11, 0), current_error_budget_burn=0.3, slo_burn_threshold=0.8)
print("\nCase A (inside window, healthy SLO):", decision_a)
assert decision_a["merge"] is True
# Case B: green checks, INSIDE window, but SLO burn IS breached -> must NOT merge
decision_b = decide_merge(validation, window, now=datetime(2026, 3, 2, 11, 0), current_error_budget_burn=0.95, slo_burn_threshold=0.8)
print("Case B (inside window, SLO breached):", decision_b)
assert decision_b["merge"] is False
assert "SLO" in decision_b["reason"]
# Case C: green checks, healthy SLO, but OUTSIDE the release window -> must NOT merge
decision_c = decide_merge(validation, window, now=datetime(2026, 3, 2, 22, 0), current_error_budget_burn=0.1, slo_burn_threshold=0.8)
print("Case C (outside release window):", decision_c)
assert decision_c["merge"] is False
assert "release window" in decision_c["reason"]
# Case D: a manifest with a mutable tag and an out-of-range replica count -> validation itself fails
bad_manifest = generate_manifest("checkout-api", "checkout-api:latest", replicas=0)
bad_validation = run_validation(bad_manifest)
decision_d = decide_merge(bad_validation, window, now=datetime(2026, 3, 2, 11, 0), current_error_budget_burn=0.1, slo_burn_threshold=0.8)
print("Case D (bad manifest, inside window, healthy SLO):", decision_d)
assert decision_d["merge"] is False
assert "validation failed" in decision_d["reason"]
print("\nAll four merge-gate cases behaved as required: the gate blocks on EACH independent condition, not just validation.")
# Drift remediation demo: cluster has drifted from what was actually merged
live_state = {
"spec": {
"replicas": 3, # manually scaled down out-of-band
"template": {"spec": {"containers": [{"image": manifest["spec"]["template"]["spec"]["containers"][0]["image"]}]}},
}
}
drift = compute_drift(manifest, live_state)
print("\nDrift detected between merged manifest and live cluster state:", json.dumps(drift, indent=2))
assert drift == {"replicas": {"desired": 6, "live": 3}}
print("Drift correctly isolated to just the replicas field (image already matches, so it is NOT reported as drifted).")
Output (actually executed with python3 s79_gitops_automation.py)
Validation results: {
"policy": [],
"unit_test": [],
"helm_render": []
}
Case A (inside window, healthy SLO): {'merge': True, 'reason': 'all checks passed, inside release window, SLO burn within threshold'}
Case B (inside window, SLO breached): {'merge': False, 'reason': 'SLO error-budget burn 0.95 exceeds threshold 0.80'}
Case C (outside release window): {'merge': False, 'reason': 'outside release window (window 9-17h, now 22h)'}
Case D (bad manifest, inside window, healthy SLO): {'merge': False, 'reason': 'validation failed: {"policy": ["image \'checkout-api:latest\' is not pinned to a digest", "replicas=0 is outside the allowed range [1,50]"]}'}
All four merge-gate cases behaved as required: the gate blocks on EACH independent condition, not just validation.
Drift detected between merged manifest and live cluster state: {
"replicas": {
"desired": 6,
"live": 3
}
}
Drift correctly isolated to just the replicas field (image already matches, so it is NOT reported as drifted).
Four cases were run specifically to prove the gate's three conditions are independently enforced, not just validation: Case A (everything healthy) merges. Case B (validation green, inside the window, but SLO burn at 0.95 against a 0.80 threshold) is BLOCKED, and the reported reason names the SLO breach specifically, not a generic failure. Case C (validation green, healthy SLO, but the current time is 22:00 against a 9-17h window) is BLOCKED with the window violation named specifically. Case D (a manifest with a mutable :latest tag and 0 replicas, evaluated inside the window with a healthy SLO) is BLOCKED by validation itself, with both specific violations listed. The drift-remediation demo then confirms a manually-scaled-down live cluster (3 replicas against a merged desired state of 6) is correctly isolated to just the replicas field, since the image already matches and is correctly NOT reported as drifted.
Key points
- Webhook handling, in a real deployment of this design, would trigger
run_validationanddecide_mergeon each CI run of a PR the automation itself opened; the demo above executes that same decision LOGIC directly rather than standing up a real webhook receiver, since the logic being correct is what matters for this answer, not the HTTP transport wrapping it. - Preventing accidental auto-merges is not one check, it is the CONJUNCTION of three, per the direct answer; a system that only checks "are the automated tests green" and calls that "safe to auto-merge" is missing the two conditions (window, SLO) that Cases B and C exist specifically to prove are independently enforced, not redundant with validation.
- Drift remediation reported ONLY the field that actually diverged (
replicas), not a blanket "the whole resource has drifted": field-level diffing gives a human (or an automated remediation step) a precise, actionable target rather than a vague signal.
Safe rollout and rollback
The question also names how to safely ROLL OUT a merged change and how to ROLL BACK one that turns out to be bad, two distinct concerns from the merge-gate logic above, which only decides whether a change reaches the cluster's declared state in the first place, not how the cluster itself transitions to actually running it.
Safe rollout. A merge is not the same event as full production traffic hitting the new version; the RUNTIME rollout still needs its own safety bound at the Kubernetes layer, at minimum a bounded rolling update (maxUnavailable/maxSurge), and, for anything higher-stakes than the routine case, a genuinely progressive delivery mechanism (an Argo Rollouts canary or blue-green step, gated on the SAME kind of live SLO signal the merge gate already checks pre-merge) so a manifest that passed every pre-merge check but still behaves badly under real production traffic is caught and halted automatically before it reaches every replica, rather than only being discovered after the fact.
Rollback. Because this is a GitOps workflow, a rollback is a new, ordinary PR (either a plain Git revert of the merged commit, or the automation re-invoking generate_manifest with the prior known-good parameters) going through the EXACT SAME run_validation/decide_merge gate as any other change, never a raw, unreviewed kubectl rollout undo applied directly against the cluster; a manual cluster-side rollback that bypasses Git entirely is itself a form of drift (the live cluster no longer matches the declared state in Git), exactly the class of problem compute_drift above exists to catch, so routing the rollback back through Git is what keeps drift detection meaningful rather than immediately re-flagging the rollback itself as an unexplained divergence.
Complexity
- Time: O(F) for validation and drift computation, where F is the number of fields checked, a small, fixed set per manifest regardless of cluster size.
- Space: O(1) beyond the size of the manifest and live-state documents themselves.
Edge cases
- All three merge-gate conditions failing simultaneously:
decide_merge's blocker list accumulates EVERY failing reason, not just the first one found, so a caller (or a human reading the PR comment this would populate in a real system) sees the full picture in one pass rather than discovering blockers one at a time across repeated attempts. - A manifest passing validation but with a release window that is exactly on its boundary (the window's
end_hour):is_openusesstart_hour <= at.hour < end_hour, a HALF-OPEN interval, so a run at exactlyend_hour:00is correctly treated as OUTSIDE the window, avoiding an off-by-one ambiguity about whether the boundary hour itself counts as open. - A live cluster state missing the containers list entirely (a resource that does not exist at all, distinct from one that exists but differs):
compute_drift's.get("containers", [{}])[0].get("image")chain resolves toNonerather than raising, correctly reporting an image mismatch rather than crashing on a missing key.
Trade-offs and pitfalls
- Common mistake: implementing "merge on green" as literally just the validation checks, treating release-window and SLO-burn awareness as a separate, optional layer bolted on later. Case B and Case C exist specifically because a system that only wires up validation, and adds window/SLO awareness as an afterthought, has already shipped the exact "accidental auto-merge" risk this question names as a requirement to prevent, not a hypothetical one.
- The SLO-burn threshold and release-window boundaries are themselves CONFIGURATION, not constants, a hardcoded threshold that never gets revisited as the service's actual traffic and reliability profile changes will eventually either block merges too aggressively (an overly conservative threshold) or too permissively (a stale, too-loose one); treating these as periodically-reviewed configuration, not fixed values, keeps the gate calibrated to reality.
- Drift detected between what was merged and what is live needs its own decision (reapply versus import) about which side is correct, this demo only DETECTS and isolates the drift; a real automated remediation step layered on top would still need the same reapply/import/escalate-to-human logic, not an automatic, unconditional reapply.
A nightly cleanup automation started failing intermittently. Describe a structured troubleshooting approach to find root cause: what logs and metrics to collect, how to reproduce the issue safely, how to form and test hypotheses, and how to implement and roll out a fix with minimal user impact. Include communication and rollback plans.
Sample Answer
Direct answer
'Intermittent' is the important word: it rules out a simple deterministic bug and points toward something environmental or load-dependent (a race, a resource limit occasionally hit, a flaky dependency). The approach is to gather evidence before touching anything, form falsifiable hypotheses from that evidence, and change one thing at a time.
Logs and metrics to collect first
Pull every failed run's full log output plus, critically, a few SUCCESSFUL runs' logs from the same window for comparison -- intermittent failures are often only diagnosable by contrast. Collect: exact failure timestamps (cluster them -- do failures correlate with a particular time of day, a particular host, a deploy window?), exit codes and stack traces, resource metrics for the run's duration (CPU, memory, disk I/O, network) to rule out resource exhaustion, and any shared-dependency health (did a database, an API the job calls, or a lock service show degraded behavior at the same timestamps?).
Reproducing safely
Don't reproduce against production data/systems if avoidable. Pull the exact conditions of a failed run (same input data snapshot if possible, same time-of-day if that's a factor) into a staging environment, and if the failure seems load- or concurrency-related, try to reproduce under similar concurrent load rather than in isolation -- a race condition that only manifests under contention won't reproduce by running the job alone.
Forming and testing hypotheses
Rank hypotheses by what the evidence actually supports, not by what's easiest to fix. If failures cluster around specific times, check for a colliding cron job or deploy window. If failures correlate with specific hosts, suspect host-local state (disk space, a stale lock file, a resource limit). If there's no clean correlation, suspect a genuine race condition in the job's own logic. Test each hypothesis with the smallest possible experiment that could disprove it -- add targeted logging around the suspected cause and let the next few runs either confirm or rule it out, rather than guessing and shipping a fix blind.
Rolling out the fix with minimal impact
Once the hypothesis is confirmed, deploy the fix behind a flag or to a subset of runs first if the automation supports partial rollout, and keep the previous behavior available as a fallback in case the fix is wrong or incomplete. Monitor the next several runs closely (don't just ship and walk away) since 'intermittent' failures need enough sample runs post-fix to be confident the rate actually dropped, not just that it happened to not fail the next one or two times.
Communication and rollback
Announce the investigation is happening (so on-call knows this isn't a new unexplained failure if it recurs once more before the fix lands), and have an explicit rollback plan for the fix itself -- if the new logging or the fix changes behavior and something else breaks, know exactly how to revert to the prior version without a scramble. Close the loop with a short summary of root cause and fix once confirmed, not just a silent 'it's fixed now.'
Trade-offs and pitfalls
The most common mistake in this exact troubleshooting shape is stopping at the first hypothesis that's CONSISTENT with the evidence rather than the one that's actually CONFIRMED by it -- an intermittent failure correlated with a deploy window is suggestive, not proof, and shipping a fix for the wrong correlated factor can leave the real cause intact while looking resolved for a while by coincidence. Edge case: a failure that only reproduces under a SPECIFIC combination of conditions (a particular time of day AND a particular host AND elevated load) needs a hypothesis-testing approach that varies one factor at a time, not a single broad reproduction attempt.
Explain the main trade-offs between using synchronous subprocess invocation (subprocess.run) and asyncio-based subprocesses (asyncio.create_subprocess_exec) in Python automation. Discuss blocking behavior, ease of implementation, concurrency models, and when you should prefer asyncio for SRE automation tasks.
Sample Answer
Direct answer
subprocess.run blocks the calling thread until the child process exits (or times out); asyncio.create_subprocess_exec returns control to the event loop immediately and lets you await the child's completion alongside other concurrent work. The choice is really about what else your script needs to be doing while the external command runs.
Blocking behavior and concurrency model
subprocess.run is simplest for a script that runs external commands sequentially, one after another -- there's no event loop to reason about, no await syntax, and stdout/stderr capture is a single synchronous call. If you need to run several external commands concurrently, the sync option is threads (a ThreadPoolExecutor calling subprocess.run per worker) -- which works, but each thread is a real OS thread, so it doesn't scale cleanly past a few hundred concurrent subprocess launches.
asyncio.create_subprocess_exec fits when the automation is ALREADY asyncio-based (talking to async HTTP clients, other async I/O) and you want subprocess launches to be just another awaitable alongside that work, sharing the same single-threaded event loop rather than spinning up a thread pool. It also composes naturally with asyncio.gather and semaphores for bounding how many subprocesses run at once, without the OS-thread overhead of the threaded approach.
Ease of implementation
subprocess.run wins here decisively for anything simple: subprocess.run([...], capture_output=True, timeout=30, check=True) is a single, easy-to-read line. The asyncio version requires creating the subprocess, then separately await-ing communicate() or manually pumping the stdout/stderr streams, and wrapping the whole thing in an async function -- meaningfully more ceremony for the same basic 'run a command and get its output' task.
When to prefer asyncio for SRE automation
Prefer asyncio when the automation's dominant cost is genuinely concurrent I/O-bound work at meaningful scale -- for example, running the same health-check command over SSH against 500 hosts in parallel, where you want hundreds of subprocesses in flight without hundreds of OS threads. Prefer sync subprocess.run (possibly with a modest thread pool) for anything with a handful of sequential or lightly-parallel external calls, a one-off maintenance script, or anywhere the extra async ceremony would cost more in review/maintenance burden than it saves in throughput. A useful rule of thumb: if you're not already in an async codebase and you're not launching dozens+ of concurrent subprocesses, subprocess.run (with threads if you need modest parallelism) is very likely simpler, more debuggable, and just as correct.
Trade-offs and pitfalls
Mixing the two carelessly is the most common real bug: calling a blocking subprocess.run from inside an async function (instead of create_subprocess_exec or running it in an executor) blocks the entire event loop, silently stalling every OTHER concurrent task the automation was supposed to be running -- turning what looked like a concurrent asyncio program into an accidentally-sequential one, with no error raised to tell you.
Describe the differences and trade-offs between using a cloud provider's web console, command-line interface (CLI), and SDKs (e.g., Python SDK). As an SRE, when do you choose CLI vs SDK vs console for automation, runbooks, and debugging? Include examples of tasks better suited to each approach.
Sample Answer
Direct answer
Each of these three surfaces trades discoverability for repeatability differently: the console is best for exploration and one-off human judgment calls, the CLI is best for scriptable and reproducible operational actions, and the SDK is best when the logic needs to be embedded inside a larger program.
When each fits
- Console: best for genuinely one-off investigation where you don't yet know what you're looking for -- browsing resource relationships, reading error messages with full context and formatting, or a task so rare that scripting it isn't worth the investment. Its weakness for automation is exactly its strength for exploration: nothing about it is reproducible or auditable as code.
- CLI: best for operational tasks you'll do more than a couple of times, or that need to be embeddable in a runbook/script/CI job -- the invocation itself is a reviewable, versionable artifact ('here's the exact command that was run'), and it composes naturally with shell scripting (piping,
--output jsonfeeding intojq). - SDK: best when the logic needs conditionals, error handling, or integration with the rest of a larger program that a shell invocation can't cleanly express -- calling the CLI as a subprocess from Python and parsing its text/JSON output works but adds an unnecessary serialization round-trip and a dependency on the CLI's output format staying stable; the SDK gives you native objects and typed exceptions instead.
Choosing for automation, runbooks, and debugging
For AUTOMATION (a script that runs regularly, unattended): SDK, because it needs the richest error handling and doesn't benefit from CLI-output-parsing overhead. For RUNBOOKS (steps a human follows, possibly semi-scripted): CLI, because a runbook that says 'run this exact command' is easier for another engineer to follow and adapt under pressure than 'run this Python snippet,' and it's directly copy-pasteable into a terminal during an incident. For DEBUGGING: usually console first (to explore and understand what's actually going on), then CLI once you know the specific check/action you want to make repeatable.
Concrete task examples
Investigating why a mysterious resource exists and who created it: console, because you're browsing and don't yet know what you're looking for. Rotating a credential as a documented, repeatable operational procedure: CLI, so it's copy-pasteable and auditable via shell history/CI logs. Building a service that needs to programmatically check and remediate resource drift as part of its own logic: SDK, because it's already running as code and gains nothing from shelling out to a CLI.
Trade-offs and pitfalls
A common mistake is defaulting to whichever surface you personally reach for out of habit rather than matching it to the task -- writing automation that shells out to the CLI and parses text output (fragile, breaks silently when a CLI's output formatting changes) when the SDK was directly available, or conversely writing a heavyweight SDK-based script for a genuinely one-off task better served by a few CLI commands typed directly into a runbook.
Implement or outline a reusable retry decorator in Python that supports exponential backoff with jitter, a configurable max attempts, and a predicate callback to classify retryable exceptions. The decorator should be usable on synchronous functions and support logging each attempt. Explain how idempotency assumptions affect your wrapper and where idempotency tokens should be applied when calling external APIs.
Sample Answer
Approach
A retry decorator needs to separate three concerns cleanly: which exceptions are worth retrying (the predicate), how long to wait between attempts (backoff+jitter), and what to do on each attempt (logging, and eventually giving up). Keeping these as parameters rather than hardcoding them is what makes the decorator reusable across call sites with very different retry needs.
import functools
import logging
import random
import time
def retry(max_attempts=5, base_delay=0.5, max_delay=30.0,
retryable=(Exception,), logger=None):
"""Retry a synchronous function with full-jitter exponential backoff.
retryable: a tuple of exception types, OR a callable(exc) -> bool that
classifies whether a given exception is worth retrying.
"""
def is_retryable(exc):
if callable(retryable) and not isinstance(retryable, type):
return retryable(exc)
return isinstance(exc, retryable)
def decorator(fn):
@functools.wraps(fn)
def wrapper(*args, **kwargs):
attempt = 0
while True:
attempt += 1
try:
return fn(*args, **kwargs)
except Exception as exc:
if not is_retryable(exc) or attempt >= max_attempts:
raise
ceiling = min(max_delay, base_delay * (2 ** (attempt - 1)))
delay = random.uniform(0, ceiling) # full jitter
if logger:
logger.info("attempt %d/%d failed (%r), retrying in %.2fs",
attempt, max_attempts, exc, delay)
time.sleep(delay)
return wrapper
return decorator
Verified in a sandbox: wrapping a function that raises ValueError twice then succeeds, @retry(max_attempts=4, retryable=(ValueError,)) returns the correct result after exactly 3 attempts; wrapping a function that always raises, it makes exactly max_attempts attempts and then re-raises the original exception rather than swallowing it.
Idempotency and the decorator
The decorator itself has no idea whether the wrapped function is safe to call twice -- that judgment has to be made by whoever applies it. Two consequences: (1) the retryable predicate should exclude exceptions that indicate the operation may have partially succeeded in an ambiguous way (a timeout on a POST is the classic ambiguous case: did the server process it and the response got lost, or did it never receive the request?), and (2) for genuinely non-idempotent external calls, the caller should generate an idempotency token before the first attempt and pass the same token on every retry, so the server-side API can deduplicate. The decorator's job is to retry; the idempotency token's job is to make retrying safe -- they're separate concerns that have to be composed correctly by the caller, not something the decorator can enforce on its own.
Trade-offs
A class-based retry policy (rather than a decorator) is worth it once you need per-call overrides (retry THIS call with a shorter window because it's on a critical path) -- decorators are static at definition time unless you thread configuration through explicitly.
Edge cases: a wrapped function called with no arguments, a function whose exception has a non-standard __repr__ that could itself throw during logging, and max_attempts=1 (which should behave as a single unretried call, not loop) are all worth explicitly testing -- the last one is a common off-by-one where an implementation accidentally still retries once even at max_attempts=1.
Unlock Full Question Bank
Get access to all Automation Scripting for Operations interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.