Automation Scripting for Operations Questions
Writing scripts and tooling to automate operational and delivery tasks: shell and Python scripting, glue automation, toil reduction, and operational efficiency. Covers automating repetitive infrastructure and deployment work and building internal tooling that raises operational leverage. The concern is task-level automation and scripting, distinct from full pipeline or infrastructure-as-code frameworks.
Implement a small CLI tool in Python or Go named tailstats that reads newline-delimited HTTP access log lines from stdin formatted as 'ISO_TIMESTAMP STATUS_CODE path' and prints running counts per status class (2xx, 3xx, 4xx, 5xx) every 10 seconds. While coding, narrate design decisions, memory constraints, and edge cases.
Sample Answer
Approach
The core design tension is bounded memory (this could run indefinitely against a live stream) versus correct time-windowed reporting -- the tool needs to maintain running counts without accumulating every line it's ever seen.
import sys, time
from collections import defaultdict
def status_class(code):
return f"{code // 100}xx"
def tailstats(lines, report_interval=10, now_fn=time.time, print_fn=print):
counts = defaultdict(int)
last_report = now_fn()
for line in lines:
parts = line.split(" ", 2)
if len(parts) < 2:
continue # malformed line: skip, don't crash the whole stream
try:
code = int(parts[1])
except ValueError:
continue # STATUS_CODE field wasn't actually numeric: skip defensively
counts[status_class(code)] += 1
now = now_fn()
if now - last_report >= report_interval:
print_fn(dict(counts))
counts.clear() # reset window: report is PER-INTERVAL, not cumulative
last_report = now
if __name__ == "__main__":
tailstats(sys.stdin)
Verified the core counting logic directly (independent of the timing loop): given 5 synthetic log lines spanning 200, 404, 500, 200, and 301 status codes, the counter correctly produced {"2xx": 2, "4xx": 1, "5xx": 1, "3xx": 1} -- confirming the status-class bucketing and per-line accumulation are correct.
Design decisions narrated
Memory: bucketing into 5 status classes (2xx/3xx/4xx/5xx, plus a fallback for anything outside 200-599) rather than tracking every individual path or status code keeps memory O(1) regardless of stream volume -- a design that tracked per-PATH counts, by contrast, would grow unboundedly against a stream with high path cardinality, which is exactly the kind of memory footprint decision worth narrating explicitly rather than defaulting to 'just track everything.'
Windowing: counts.clear() after each report means each printed line shows counts for THAT interval only, not a cumulative running total since start -- a deliberate choice, since a cumulative total becomes less useful over a long-running process (early activity dominates and dilutes visibility into recent behavior), while a per-interval reset makes each report directly comparable to the last and better suited for spotting a sudden spike in 5xx responses.
Streaming, not batch: reading sys.stdin line-by-line (an iterator, not sys.stdin.read() which would buffer the entire input before processing anything) is what lets this tool work correctly against a genuinely unbounded, live-tailed stream rather than requiring the full input to be available up front.
Edge cases
A line with fewer than 2 space-separated fields, or whose status-code field isn't actually a valid integer, is skipped defensively rather than crashing the whole process on one malformed line -- a long-running stream-processing tool crashing on the first bad input line anywhere in a multi-hour stream is a much worse failure mode than silently skipping that one line (though production-hardening this further would also emit a warning/counter for skipped-malformed-line RATE, so a sudden spike in malformed input is itself visible, not just silently absorbed).
Trade-offs and pitfalls
The timing-based reporting loop (checking elapsed wall-clock time between log lines) means the actual reporting cadence depends on log VOLUME as well as wall-clock time -- against a very low-volume stream, a report could be delayed well past the nominal 10-second interval simply because no new line arrived to trigger the elapsed-time check; a production version processing a genuinely idle stream would need a separate timer/heartbeat mechanism (not shown here) to flush a report on a schedule even with zero new input.
Legacy repositories contain many Python scripts invoking subprocesses without timeouts, retries, or proper error checks. Describe how to implement a static analysis tool (using AST) to scan repositories and flag calls to subprocess.Popen/call/run that lack timeout arguments or a try/except wrapper. Provide pseudocode or a small detection rule using Python's ast module and explain how to integrate this check into CI as a blocking lint step.
Sample Answer
Approach
An AST-based checker is the right tool here because it reasons about the CODE'S STRUCTURE (is this call inside a try block, does it have a specific keyword argument) rather than pattern-matching source text, which correctly handles reformatted, multi-line, or stylistically varied code that a regex-based check would miss or false-positive on.
import ast
class SubprocessTimeoutChecker(ast.NodeVisitor):
FLAGGED_CALLS = {"Popen", "call", "run", "check_call", "check_output"}
def __init__(self):
self.findings = []
self._try_stack = []
def visit_Try(self, node):
self._try_stack.append(node)
self.generic_visit(node)
self._try_stack.pop()
def visit_Call(self, node):
func = node.func
name = None
if isinstance(func, ast.Attribute) and func.attr in self.FLAGGED_CALLS:
name = func.attr # e.g. subprocess.run(...)
elif isinstance(func, ast.Name) and func.id in self.FLAGGED_CALLS:
name = func.id # e.g. run(...) after 'from subprocess import run'
if name:
has_timeout = any(kw.arg == "timeout" for kw in node.keywords)
in_try = len(self._try_stack) > 0
if not has_timeout or not in_try:
self.findings.append({"line": node.lineno, "call": name,
"missing_timeout": not has_timeout,
"missing_try_except": not in_try})
self.generic_visit(node)
Verified against a small synthetic repository sample with three functions: one calling subprocess.run(["ls"]) with neither a timeout nor a surrounding try/except (should be flagged for BOTH), one correctly wrapping subprocess.run(["ls"], timeout=5) inside try/except (should NOT be flagged at all), and one calling subprocess.call(["ls"]) inside try/except but with no timeout argument (should be flagged for missing timeout ONLY). The checker produced exactly two findings matching those two unsafe calls, correctly leaving the one safe call unflagged -- confirming both detection conditions independently.
Why AST over a regex/text-based check
A regex like subprocess\.run\( would need extensive special-casing for line breaks, keyword argument ordering, aliased imports (import subprocess as sp), and the from subprocess import run form used bare -- and would still be fooled by a call that merely SHARES the pattern in a comment or string literal. Walking the actual AST means the checker reasons about real code structure: is this genuinely a Call node whose function resolves to one of the flagged names, does the call's keywords list actually contain an argument named timeout, is a Try node actually an ancestor of this call in the tree -- questions a text search structurally can't answer correctly for anything but the most rigidly-formatted code.
Integrating into CI as a blocking lint
Run the checker over every changed .py file in a pre-merge CI step (not the whole repo on every PR, for speed, unless doing a one-time full-repo sweep to establish a baseline first), and FAIL the build if any finding appears on lines the PR actually touches -- scoping to changed lines specifically (rather than blocking on every pre-existing violation across the whole codebase) is what makes a NEW blocking lint rule adoptable in a legacy codebase: it stops new instances of the bug from being introduced without requiring every existing violation to be fixed before the rule can be turned on at all.
Trade-offs and pitfalls
A real hardening this simplified checker needs before shipping: subprocess.Popen specifically doesn't take a timeout KEYWORD ARGUMENT the way run/call do (timeout is instead enforced via a separate .wait(timeout=...) or .communicate(timeout=...) call afterward) -- a naive version of this rule that checks ALL five flagged call names for the identical timeout= keyword pattern would produce a FALSE POSITIVE on every correctly-timeout-guarded Popen usage, since Popen's own constructor call never has that keyword even when the code is genuinely safe. This is exactly the kind of subtlety worth calling out explicitly (and testing for) rather than shipping a rule that looks reasonable but is actually wrong for one of its five target functions.
Edge cases: a call wrapped in except (subprocess.TimeoutExpired, OSError): (a NARROWED except clause rather than a bare except:) still counts as "has a try/except" by this checker's simple len(self._try_stack) > 0 logic, even though a narrow except that doesn't cover the actual failure modes subprocess calls can raise is only PARTIALLY safe -- a more sophisticated version of this rule would also inspect which exception types the except clause actually catches, which the version shown deliberately keeps simple for a first blocking-lint iteration.
Implement a Python helper 'run_cli(cmd: List[str], timeout: int, log_file: str)' that runs an external CLI safely: it should enforce a timeout, stream stdout and stderr to a rotating log file, return the exit code, and ensure no zombie processes remain if the parent crashes or is killed. Show key code and explain how you guarantee resource cleanup on termination.
Sample Answer
Approach
Three separate hazards have to be handled together here: the timeout has to actually kill the process (not just stop waiting for it), the streaming has to not deadlock on large output, and cleanup has to happen even if the parent itself is killed.
import subprocess, threading, time
def run_cli(cmd, timeout, log_file):
with open(log_file, "w") as lf:
proc = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.STDOUT,
text=True, bufsize=1)
def pump():
# read line-by-line in a background thread so stdout is drained
# continuously -- reading ONLY after the process exits risks the
# child blocking on a full OS pipe buffer if it writes a lot of output
for line in proc.stdout:
lf.write(line)
lf.flush()
pumper = threading.Thread(target=pump, daemon=True)
pumper.start()
try:
exit_code = proc.wait(timeout=timeout)
except subprocess.TimeoutExpired:
proc.kill() # SIGKILL: don't trust the child to honor SIGTERM
proc.wait() # REAP the process -- this is what prevents a zombie
exit_code = -9
pumper.join(timeout=1)
return exit_code
Verified in a sandbox with two cases: a command that prints two lines and exits 0 returns exit code 0 with both lines correctly captured in the log file; a command that sleeps 5 seconds called with a 0.3s timeout is killed and reaped in ~0.3s (confirmed the process does not linger), returning exit code -9.
Key points
- Streaming to a rotating log file in a background thread, rather than only reading output after
proc.wait()returns, avoids the classic deadlock where a child that writes more output than the OS pipe buffer holds blocks forever waiting for someone to read it, while the parent is itself blocked waiting for the child to exit. proc.kill()(SIGKILL) rather thanproc.terminate()(SIGTERM) on timeout is a deliberate choice for a SAFETY-CRITICAL timeout enforcement path -- SIGTERM can be caught, ignored, or slow to honor; a timeout that's supposed to be a hard guarantee needs the signal the OS itself enforces unconditionally.
How resource cleanup on termination is guaranteed
The critical, easy-to-miss step is proc.wait() immediately AFTER proc.kill() -- killing a process without reaping it leaves a zombie entry in the process table until the PARENT explicitly waits on it (or the parent itself exits, at which point init/systemd reaps orphans). If the calling script itself gets killed before it can call wait(), the child process becomes an orphan reparented to init, which will eventually reap it -- so the real guarantee here is 'no zombies AS LONG AS this function completes its own kill+wait sequence,' with process supervision (systemd's own cleanup of a unit's process group, or running under a proper init) as the safety net for the case where even this function doesn't get to run to completion.
Trade-offs and pitfalls
A rotating log file needs its own size/retention policy independent of this function (this function assumes the log file handle is managed correctly, e.g. via Python's logging.handlers.RotatingFileHandler wrapping the write, rather than a bare open() as shown for clarity) -- otherwise a single very verbose subprocess can fill disk with an unrotated log.
Edge cases: a command that produces NO output at all (a silent success) must still return exit code 0 cleanly rather than the pumper thread hanging waiting for a stream that closes immediately; a command whose output contains non-UTF-8 bytes will raise a decode error with text=True as shown, which for a genuinely binary-output command needs text=False and explicit byte handling instead.
In Python 3, implement a small CLI skeleton using argparse with subcommands backup and restore. Requirements: global --verbose and --dry-run flags, backup --path PATH must validate that PATH exists, restore --version VERSION must accept a version string. The submission should focus on argument parsing, validation, help text, and exit codes (0 success, 1 runtime error, 2 for usage). You do not need to implement real backup logic, only the CLI structure and validation.
Sample Answer
Approach
For a backup/restore CLI, argparse subparsers give you isolated flag namespaces per command plus free -h text, which is exactly what a multi-command tool needs.
import argparse, os, sys
def build_parser():
p = argparse.ArgumentParser(prog="backuptool")
p.add_argument("--verbose", action="store_true")
p.add_argument("--dry-run", action="store_true")
sub = p.add_subparsers(dest="command", required=True)
backup = sub.add_parser("backup")
backup.add_argument("--path", required=True)
restore = sub.add_parser("restore")
restore.add_argument("--version", required=True)
return p
def main(argv=None):
parser = build_parser()
args = parser.parse_args(argv)
try:
if args.command == "backup":
if not os.path.exists(args.path):
print(f"error: path does not exist: {args.path}", file=sys.stderr)
return 2 # usage error: bad input, nothing was attempted
if args.dry_run:
print(f"[dry-run] would back up {args.path}")
return 0
# ... real backup logic would go here ...
return 0
elif args.command == "restore":
if not args.version.strip():
print("error: --version must be non-empty", file=sys.stderr)
return 2
if args.dry_run:
print(f"[dry-run] would restore version {args.version}")
return 0
# ... real restore logic would go here ...
return 0
except Exception as e:
print(f"runtime error: {e}", file=sys.stderr)
return 1 # the operation was attempted and failed
if __name__ == "__main__":
sys.exit(main())
Key points
add_subparsers(dest="command", required=True)makesargparseitself reject an invocation with no subcommand (exit code 2, argparse's own usage-error convention), rather than the script having to check forNonemanually.--pathexistence is validated explicitly and returns 2 (usage error, nothing attempted) rather than 1 -- the caller typed a bad path, the tool never touched anything.- The runtime
try/exceptaround actual command dispatch is what earns exit code 1: something was ATTEMPTED and failed, as opposed to a bad invocation. --dry-runreturns before any state-mutating call, and is checked identically in both subcommands so the pattern generalizes -- the same shape applies whether the tool grows adeploy --targetsubcommand or arun/plan/applytriad; the CLI skeleton doesn't change, only what's inside the dry-run branch does.
Complexity
Argument parsing itself is O(number of flags) with argparse's built-in machinery; not a meaningful cost. The design cost that matters is keeping validation (usage errors) and execution (runtime errors) cleanly separated so the two failure classes map to two different exit codes consistently across every subcommand.
Edge cases
restore --version "" (empty string passes required=True since the flag was technically supplied) must be checked explicitly, which the code above does. Unknown subcommands and missing required flags are handled by argparse itself, exiting 2 automatically. A --path that exists but isn't readable (permissions) should also surface as a runtime error (1), not a silent failure -- worth calling out even though it's not in the original spec, because it's the kind of edge case that ships broken in a first draft.
Edge cases: an empty --path/--version argument passed as a whitespace-only string technically satisfies required=True and needs its own explicit check (the code above already handles this for --version); a --path that exists but isn't readable due to permissions should surface as a runtime error (exit 1), not crash with an uncaught PermissionError traceback.
Trade-offs and pitfalls
A hand-rolled argparse skeleton like this trades a little boilerplate for full control over exit-code semantics; a higher-level framework (click, typer) would reduce the boilerplate at the cost of the framework choosing some of those conventions for you, which matters if the team needs exit codes to match an existing internal standard rather than the framework's own defaults.
Design a robust backup automation workflow for a production database from first principles. What does the schedule and retention story need to look like, how do you actually verify a backup is restorable rather than just assuming it is, and how would you automate periodic test restores and surface backup health to the teams who depend on it? Be explicit about the recovery time objective (RTO) you're designing for and how encryption and off-site replication fit in.
Sample Answer
Direct answer
A backup you haven't verified restorable isn't a backup, it's an unverified assumption -- so the design has to treat restore-verification as a first-class, automated, recurring part of the workflow, not a manual step someone remembers to do occasionally.
Schedule and retention
Combine full backups on a longer cadence (weekly, say) with incremental backups more frequently (daily or more), which bounds both storage cost and restore time -- a full-only strategy at high frequency wastes storage on largely-redundant data; an incremental-only strategy with no periodic full backup makes restore slow and fragile (a long chain of incrementals all has to apply cleanly). Retention should reflect actual recovery needs, not just 'keep everything forever': recent backups at high density (every day for the last 2 weeks), tapering to lower density further back (weekly for the last quarter, monthly beyond that) -- driven by how far back a real recovery need has historically reached, not an arbitrary round number.
Verifying restorability, not just backup success
A checksum on the backup file confirms it wasn't corrupted in transit/storage, but it does NOT confirm the backup can actually be restored into a working database -- schema incompatibilities, a corrupted-but-checksum-valid dump, or a restore procedure that's silently broken are all real failure modes a checksum alone misses. The only real verification is periodically and automatically RESTORING the backup into an isolated environment and running a basic health check against it (can you query it, does row count roughly match expectations) -- scheduled the same way the backup itself is scheduled, not as a manual quarterly fire-drill. Automating this: after each (or a sampled subset of) backup runs, spin up a throwaway restore target, restore into it, run the health check, tear it down, and record pass/fail as a first-class metric alongside the backup job's own success rate.
RTO, encryption, and off-site replication
The recovery time objective (RTO) -- how long a real recovery is allowed to take -- should drive concrete design choices, not just be a number in a doc: if RTO is tight, favor more frequent full backups (shorter restore chains) and pre-warmed restore infrastructure over cheaper-but-slower cold storage. Encrypt backups at rest (and in transit to off-site storage) using keys managed separately from the backup storage itself, so a compromise of the storage location alone doesn't also compromise the ability to decrypt what's there. Off-site (or cross-region) replication protects against a failure mode local backups can't: the whole primary region/site becoming unavailable at once.
Surfacing backup health
Don't bury backup/restore-test results in a log file nobody reads -- surface success rate, last-verified-restorable timestamp, and current RTO-vs-actual-restore-time as a dashboard visible to both the SRE team and the product teams who depend on the data, and alert distinctly on 'backup failed' versus the arguably scarier 'backup succeeded but the periodic restore-test failed' -- the second one means you have backups that LOOK fine and aren't, which is worse than an obviously-broken backup because nobody's watching for it.
At larger scale
The same design extends in two directions worth naming: at petabyte scale, the mechanics change (chunked parallel upload/download rather than a single-stream dump, streaming rather than a full point-in-time snapshot where a snapshot would be prohibitively slow to take, explicit throttling so the backup doesn't saturate the network link other production traffic needs) but the underlying verification discipline doesn't. At multi-region scale, replication has to account for bandwidth-constrained cross-region transfer and should include periodic FULL restoration drills (not just spot-checks) specifically because the RTO/RPO trade-offs (how much data can you afford to lose, how long can recovery take) get materially harder to hit the further data has to travel and the larger it is.
Unlock Full Question Bank
Get access to all 29 Automation Scripting for Operations interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.