Shell Scripting and Automation Questions
Shell-language craft for operations work: Bash and POSIX sh scripting, pipes and redirection, quoting and word splitting, exit codes and set -euo pipefail (including failures inside pipelines), traps and cleanup, argument parsing with getopts, functions and arrays, parameter expansion, here-documents, and text processing with grep, sed, awk and jq, including streaming pipelines over very large log, CSV and JSON Lines files. Also covers writing robust, idempotent, portable scripts (locking, atomic file updates, retries with backoff, safe temp files, GNU versus BSD differences), background jobs and signals, bounded parallelism and SSH fan-out, cron-driven jobs such as log rotation, backups, atomic deploys and health-check watchdogs, secure scripting (eval injection, secrets, path traversal), and debugging, testing and linting shell with ShellCheck and bats. Boundary: general automation design and Python or Go tooling, Linux host administration tasks, observability and alerting design, security detection engineering, and generic algorithmic coding problems are covered elsewhere.
Design a Bash script that runs an incremental load from cloud object storage: list new files since the last run, prevent concurrent runs, download and transform them, verify checksums, record progress safely and recover cleanly from a mid-run failure.
Sample Answer
Direct answer
Keep a ledger of finished object keys, and define "new" as the keys in the bucket that are not in the ledger. Take an exclusive lock with flock so only one run executes, do all work in a per-run temp directory, verify each object against a checksum before transforming it, publish the result with an atomic rename to a name derived from the key, and append the key to the ledger last. Every step is then safe to repeat, so a crash anywhere leaves either a ledger entry with a complete output or no entry, and the next run simply redoes the key.
Why a ledger and not a "last timestamp"
The ledger is an append-only text file with one finished key per line. A watermark (the last key or LastModified seen) breaks when objects arrive late, are re-uploaded, or sort before the watermark; those files are lost forever without an error. S3's listing is in lexicographic key order (plain dictionary order of the text, so dt=10 sorts before dt=9), which only helps if keys are strictly increasing, and --start-after is a documented parameter of list-objects-v2 for general purpose buckets. The ledger approach costs a full listing of the prefix each run (restrict it to recent date partitions if the prefix is huge) and buys correctness.
The script
The bucket is a local directory here (list_keys and FETCH_CMD are the two seams). For S3 you would replace list_keys with aws s3api list-objects-v2 --bucket B --prefix P --query 'Contents[].Key' --output json | jq -r '.[]?' (not --output text: the AWS CLI documentation says that a text query filtered to a single field prints one line of tab-separated values, so a key containing a space or tab cannot be told apart from a separator, and text output is paginated before the query runs; JSON output is parsed as one structure, and jq -r '.[]?' prints one key per line and nothing at all for null, which is what the query returns when the prefix is empty) and FETCH_CMD with aws s3 cp --only-show-errors. Those AWS calls were not run against real S3 here; the jq half was run locally on ["a b","c"] (one key per line, the space preserved) and on null (no output). Objects are gzipped CSV files with a producer-written KEY.sha256 sidecar.
#!/usr/bin/env bash
# loader.sh : incremental load. Environment: BUCKET_DIR STATE_DIR OUT_DIR [FETCH_CMD]
set -euo pipefail
export LC_ALL=C # sort and comm must agree on ordering
: "${BUCKET_DIR:?}" "${STATE_DIR:?}" "${OUT_DIR:?}"
FETCH_CMD=${FETCH_CMD:-cp} # S3 example: aws s3 cp --only-show-errors
ledger=$STATE_DIR/processed.keys # append-only: one finished key per line
mkdir -p "$STATE_DIR" "$OUT_DIR"; touch "$ledger"
# 1. One run at a time. The kernel drops the flock if the process dies, so no stale lock.
exec 9> "$STATE_DIR/loader.lock"
flock -n 9 || { echo "another run holds the lock, exiting"; exit 0; }
# 2. Crash recovery. We hold the lock, so any run.* directory is debris from a dead run.
for d in "$STATE_DIR"/run.*; do
[[ -d $d ]] || continue
echo "removing debris from an earlier crashed run: ${d##*/}"; rm -rf -- "$d"
done
rm -f -- "$OUT_DIR"/.publish.* # half-written outputs from a dead run
work=$(mktemp -d "$STATE_DIR/run.XXXXXX"); trap 'rm -rf -- "$work"' EXIT
# 3. What is new? Every key in the bucket minus every key in the ledger. This does not
# depend on key order or timestamps, so a late-arriving old-looking key is still found.
list_keys() { # stand-in for: aws s3api list-objects-v2 ... --output json | jq -r '.[]?'
find "$BUCKET_DIR" -type f -name '*.csv.gz' -printf '%P\n'
}
list_keys | sort -u > "$work/all.keys"
sort -u "$ledger" > "$work/done.keys"
comm -23 "$work/all.keys" "$work/done.keys" > "$work/todo.keys"
echo "new keys: $(wc -l < "$work/todo.keys")"
ok=0; bad=0
while IFS= read -r key; do
out=${key//\//__}; out=${out%.gz} # dt=2026-10-01/a.csv.gz -> dt=2026-10-01__a.csv
obj=$work/object.gz; sum=$work/object.sha256
# 4. Download object and its producer-written checksum, retrying transient failures.
fetched=0
for try in 1 2 3; do
if $FETCH_CMD "$BUCKET_DIR/$key" "$obj" && $FETCH_CMD "$BUCKET_DIR/$key.sha256" "$sum"; then
fetched=1; break
fi
sleep "$try"
done
(( fetched )) || { echo "FAIL download: $key"; bad=$((bad + 1)); continue; }
# 5. Verify BEFORE transforming. A mismatch leaves the key out of the ledger.
read -r want _ < "$sum"
read -r have _ < <(sha256sum -- "$obj")
if [[ $have != "$want" ]]; then
echo "FAIL checksum: $key"; bad=$((bad + 1)); continue
fi
# 6. Transform into a temp file INSIDE OUT_DIR, then publish with an atomic rename.
# rename(2) is atomic only within one filesystem, and OUT_DIR may not share one with
# STATE_DIR, so the temp file lives next to its final name.
tmp=$(mktemp "$OUT_DIR/.publish.XXXXXX")
gzip -dc -- "$obj" | awk -F, -v k="$key" 'NR > 1 { print $0 "," k }' > "$tmp"
chmod 0644 -- "$tmp" # mktemp creates 0600
mv -- "$tmp" "$OUT_DIR/$out"
# 7. Record progress last. A crash between 6 and 7 re-processes the key next run and
# overwrites the same output name, which is harmless.
printf '%s\n' "$key" >> "$ledger"; sync -- "$ledger"
echo "ok: $key"; ok=$((ok + 1))
done < "$work/todo.keys"
echo "done: ok=$ok failed=$bad"
(( bad == 0 ))
Reading the shell lines
comm -23.commcompares two files that are both sorted and prints three columns: lines only in file 1, lines only in file 2, and lines in both.-2hides the second column and-3hides the third, so what remains is the lines only in file 1: the keys in the bucket that are not in the ledger. If the inputs are not sorted the same way,commreports that they are out of order, which is why every list goes throughsortwithLC_ALL=C.exec 9> "$STATE_DIR/loader.lock"andflock -n 9. A file descriptor is the number a process uses for an open file.exec 9> FILEopens the lock file for writing as descriptor 9 (creating it) and keeps it open for the whole script.flock -n 9asks the kernel for an exclusive lock on that open file;-nmeans do not wait, fail at once if someone else holds it. The lock belongs to the open descriptor, so it disappears when the process exits or is killed, and no stale lock file can block later runs.${d##*/}. Removes the longest leading part of$dthat matches*/, leaving only the last path component, so/state/run.31a86lbecomesrun.31a86l.${key//\//__}. Replaces every/in$keywith__, so a key with directories becomes one flat file name.${out%.gz}then removes a trailing.gz.- Other terms. Idempotent means repeating a step gives the same end state. A sidecar is a small companion file stored next to the object (here
KEY.sha256). An atomic rename (mvwithin one filesystem) is a single step, so a reader sees the old file or the new one and never half of it. Across two filesystemsmvcannot rename: it copies into the final name and then deletes the source, so a crash mid-copy leaves a truncated file under the real name (checked withstrace:renameat2 ... = -1 EXDEV). That is why the script writes its temp file insideOUT_DIR, next to its final name, and not in the state directory.syncforces buffered writes to disk. The ETag is the identifier S3 reports for an object, which is why the note below explains why it is not a trustworthy checksum.
Run on their own (Ubuntu 24.04 container):
export LC_ALL=C
printf '%s\n' a b c > one; printf '%s\n' b c d > two
comm one two | cat -A | sed 's/\^I/<TAB>/g'
echo '--- comm -23'; comm -23 one two
d=/state/run.31a86l; echo "\${d##*/} -> ${d##*/}"
key='dt=2026-10-01/a.csv.gz'; out=${key//\//__}; echo "\${key//\//__} -> $out"; echo "then \${out%.gz} -> ${out%.gz}"
exec 9> lock
flock -n 9 && echo "first holder: got the lock"
( exec 8> lock; flock -n 8 && echo got || echo "second holder: lock busy, status $?" )
a$
<TAB><TAB>b$
<TAB><TAB>c$
<TAB>d$
--- comm -23
a
${d##*/} -> run.31a86l
${key//\//__} -> dt=2026-10-01__a.csv.gz
then ${out%.gz} -> dt=2026-10-01__a.csv
first holder: got the lock
second holder: lock busy, status 1
Reading it: with no flags comm puts a in column 1 (no tabs), d in column 2 (one tab) and b, c in column 3 (two tabs); -23 leaves only a. The two expansions give the run directory name and the flat output name. A second descriptor opened on the same lock file cannot take the lock while the first still holds it.
How each requirement is met
| Requirement | Mechanism |
|---|---|
| List new files since the last run | comm -23 of the sorted bucket listing against the sorted ledger |
| No concurrent runs | flock -n on a lock file; the kernel releases it if the process dies |
| Download and transform | Fetch with 3 retries into a run directory, `gzip -dc |
| Verify checksums | Compare sha256sum of the downloaded bytes with the sidecar before transforming; a mismatch skips the key and does not enter the ledger |
| Record progress safely | Append to the ledger last, then sync that file |
| Recover from a mid-run failure | Debris removal at start (we hold the lock, so any run.* directory or .publish.* file is dead), idempotent output names, ledger written last |
Why ETag is not used as the checksum: for multipart uploads and some encryption modes it is not an MD5 of the content, so a sidecar checksum written by the producer, or the provider's additional checksum feature, is the reliable comparison.
Worked example
Four objects exist: a, b, c d (a space in the key), and e with a corrupted body. Run 1 uses a fetch command that kills the script with SIGKILL (a power cut) while fetching the third object. Run 2 recovers. Then the producer re-uploads e, and a new object with an older-looking key (dt=2026-09-30/late.csv.gz) appears. A concurrent start is also tried. The random suffix in the debris directory name differs per run.
set -u
export BUCKET_DIR=/bucket STATE_DIR=/state OUT_DIR=/out
put() { # put KEY CONTENT [corrupt]: store a gzipped CSV and its sidecar checksum
mkdir -p "$(dirname "$BUCKET_DIR/$1")"
printf 'id,v\n%s\n' "$2" | gzip -n > "$BUCKET_DIR/$1"
sha256sum < "$BUCKET_DIR/$1" | awk '{print $1 " -"}' > "$BUCKET_DIR/$1.sha256"
[[ ${3:-} ]] && echo junk >> "$BUCKET_DIR/$1"
return 0
}
put dt=2026-10-01/a.csv.gz '1,x'
put dt=2026-10-01/b.csv.gz '2,y'
put "dt=2026-10-02/c d.csv.gz" '3,z'
put dt=2026-10-02/e.csv.gz '4,w' corrupt
cat > /usr/local/bin/dying-cp <<'EOT'
#!/bin/bash
# simulates a power cut while fetching the third object
[[ $2 == *object.gz && $1 == *"c d.csv.gz" ]] && kill -9 "$PPID"
exec cp -- "$@"
EOT
chmod +x /usr/local/bin/dying-cp
echo '--- run 1 (dies while fetching the third object)'
FETCH_CMD=dying-cp bash loader.sh; echo "exit=$?"
echo '--- state after the crash'; cat /state/processed.keys; ls -d /state/run.* | sed 's/run\..*/run.<random>/'
echo '--- run 2'
bash loader.sh; echo "exit=$?"
echo '--- producer re-uploads e, and a late object with an OLDER key name appears'
put dt=2026-10-02/e.csv.gz '4,w'; put dt=2026-09-30/late.csv.gz '0,l'
echo '--- run 3'
bash loader.sh; echo "exit=$?"
echo '--- run 4 (nothing new)'
bash loader.sh; echo "exit=$?"
echo '--- concurrent start'
( exec 9> /state/loader.lock; flock 9; sleep 3 ) & sleep 1
bash loader.sh; echo "exit=$?"; wait
echo '--- outputs'; ls /out; cat "/out/dt=2026-10-02__c d.csv"
Output:
--- run 1 (dies while fetching the third object)
new keys: 4
ok: dt=2026-10-01/a.csv.gz
ok: dt=2026-10-01/b.csv.gz
exit=137
demo.sh: line 22: 33 Killed FETCH_CMD=dying-cp bash loader.sh
--- state after the crash
dt=2026-10-01/a.csv.gz
dt=2026-10-01/b.csv.gz
/state/run.<random>
--- run 2
removing debris from an earlier crashed run: run.31a86l
new keys: 2
ok: dt=2026-10-02/c d.csv.gz
FAIL checksum: dt=2026-10-02/e.csv.gz
done: ok=1 failed=1
exit=1
--- producer re-uploads e, and a late object with an OLDER key name appears
--- run 3
new keys: 2
ok: dt=2026-09-30/late.csv.gz
ok: dt=2026-10-02/e.csv.gz
done: ok=2 failed=0
exit=0
--- run 4 (nothing new)
new keys: 0
done: ok=0 failed=0
exit=0
--- concurrent start
another run holds the lock, exiting
exit=0
--- outputs
dt=2026-09-30__late.csv
dt=2026-10-01__a.csv
dt=2026-10-01__b.csv
dt=2026-10-02__c d.csv
dt=2026-10-02__e.csv
3,z,dt=2026-10-02/c d.csv.gz
Reading it: run 1 died after two objects, leaving two ledger lines and a debris directory. Run 2 removed the debris, loaded c d, and flagged e as a checksum failure (exit 1, and e is not in the ledger, so it is retried). Run 3 picked up both the repaired e and the late 2026-09-30 object, which a "last key seen" watermark would have missed. Run 4 found nothing new. A concurrent start exited at the lock. The Killed line is printed by the demo's own shell, not by loader.sh, and its position relative to exit=137 can shift between runs.
Staging from an on-premises host through a bastion
When the objects first have to be pulled from a host reachable only through a bastion (jump host) instead of a bucket, the same rule applies to that hop: verify before the next step. Stage the files with rsync and check them before the cloud upload step:
rsync -a --partial --checksum -e "ssh -J ops@bastion.example.com" \
appsrv:/export/daily/ /stage/daily/
-J tells ssh to connect to the target through the jump host (OpenSSH manual: it first connects to the jump host and then forwards a TCP connection to the final destination). --partial keeps a half-transferred file so a retry resumes. Then prove the stage equals the source before uploading, with a checksum dry run, which prints nothing when every file matches. This was run locally between two directories:
mkdir -p /src /stage; echo one > /src/a; echo two > /src/b
rsync -a --partial /src/ /stage/
rsync -rc --dry-run --itemize-changes /src/ /stage/ # prints nothing: identical
touch -r /src/b /stage/b; printf 'twx\n' > /stage/b; touch -r /src/b /stage/b
rsync -rc --dry-run --itemize-changes /src/ /stage/ # same size, same mtime, different bytes
>fc.T...... b
Reading >fc.T...... b (the --itemize-changes code, one character per property, then the file name): > means the file would be transferred to the local side, f means it is a regular file, c means the checksum differs, and T means the modification time would be set to the transfer time. Each . is a property that does not differ (size, permissions, owner, group and so on). The c marks a checksum difference, so the corrupted copy is caught even though size and modification time match, which is exactly what rsync's default quick check would miss. Only after the dry run is silent does the cloud upload step start, and the same ledger and idempotent naming apply to it.
Trade-offs and pitfalls
- Exit codes under cron: exiting 0 on a lock conflict hides a run that is stuck for days. Alert on "no successful run in N hours", not on the exit code alone.
- A ledger is state. Back it up, and keep it on durable storage. Losing it means reprocessing everything, which is safe only because outputs are idempotent.
- Poison objects (a permanently bad file) fail every run. Add a failure counter and a quarantine prefix so they stop blocking the alert.
- Output idempotency matters more than the lock. If the downstream write is not overwrite-safe (an append to a table), the ledger cannot save you: load into a staging table keyed by object key and merge.
Read many gzipped JSON Lines files, extract id, timestamp and value, write one CSV with a single header, and remove duplicate ids. Show the pipeline and explain how it stays streaming.
Sample Answer
Direct answer
Decompress each file with gzip -dc as a stream, parse each line with jq into a CSV row, and remove duplicate ids with awk '!seen[$1]++', which keeps the first row seen for each id. Print the header once, outside the pipeline. All stages read and write one line at a time, so nothing buffers a whole file. The one exception is the de-duplication table, which holds every distinct id; that is the part to swap for sort -u when ids no longer fit in memory.
The pipeline
(A quick decode of the less obvious pieces. awk '!seen[$1]++' uses an array seen keyed by the first field; seen[$1]++ returns the old count and then adds 1, so it is 0 the first time an id appears and 1, 2, ... after; ! turns 0 into true, so awk's default action, print the line, runs only on the first sight of each id. For ids 1, 2, 1, 3, 2 the expression gives 1, 1, 0, 1, 0, and the output is ids 1, 2, 3. In the sed stage, $a\ means "at the last line, append nothing"; GNU sed appends a newline only if the last line lacks one, so every file ends cleanly. In the script it is written "\$a\\" because it sits inside double quotes, where the shell would otherwise read $a as a variable and eat one backslash. xargs -0 -r -n1 reads NUL-separated names, does nothing on empty input and runs the command once per name.)
#!/usr/bin/env bash
# usage: merge_dedupe.sh DIR > merged.csv (keeps the first record seen for each id)
set -euo pipefail
dir=${1:?usage: merge_dedupe.sh DIR}
printf 'id,timestamp,value\n'
find "$dir" -name '*.jsonl.gz' -print0 | sort -z |
xargs -0 -r -n1 bash -o pipefail -c 'gzip -dc -- "$1" | sed -e "\$a\\"' _ |
jq -Rr 'fromjson? | select(.id != null) | [.id, .ts, .value] | @csv' |
awk -F, '!seen[$1]++'
Test data: four files, one of them without a trailing newline on its last line. Four cases are in there: id 2 appears in two files (the first has value null), id 1 appears twice, one record has no id, and one record has no value.
rm -rf in; mkdir in
printf '%s\n' '{"id":1,"ts":"2026-10-01T00:00:00Z","value":10}' '{"id":2,"ts":"2026-10-01T00:00:05Z","value":null}' | gzip > in/a.jsonl.gz
printf '%s\n' '{"id":2,"ts":"2026-10-01T00:01:00Z","value":7}' '{"id":3,"ts":"2026-10-01T00:01:05Z","value":3.5}' '{"id":1,"ts":"2026-10-01T00:01:09Z","value":11}' | gzip > in/b.jsonl.gz
printf '%s\n' '{"ts":"2026-10-01T00:02:00Z","value":1}' | gzip > in/c.jsonl.gz
printf '%s' '{"id":4,"ts":"2026-10-01T00:02:05Z"}' | gzip > in/d.jsonl.gz # no trailing newline
bash merge_dedupe.sh in; echo "exit=$?"
echo "--- corrupt file:"
head -c 20 in/b.jsonl.gz > in/e.jsonl.gz
bash merge_dedupe.sh in > out.csv || echo "exit=$?"
Output:
id,timestamp,value
1,"2026-10-01T00:00:00Z",10
2,"2026-10-01T00:00:05Z",
3,"2026-10-01T00:01:05Z",3.5
4,"2026-10-01T00:02:05Z",
exit=0
--- corrupt file:
gzip: in/e.jsonl.gz: unexpected end of file
exit=123
Reading the result:
- One header line, then ids 1, 2, 3, 4 once each. The record with no id is dropped by
select(.id != null). - Id 1 keeps
value10 froma.jsonl.gzand the later 11 is discarded; id 2 keeps the earlier row, whosevalueis empty (null). "First wins" is a policy, and here it kept a worse row. If the latesttsshould win, the dedupe must compare timestamps, which needs a sort byidthentsdescending (see trade-offs). - Id 4 appears even though its file has no final newline: each file goes through
sed -e '$a\'(GNU sed), which adds a newline only when one is missing. Without it,gzip -dc a b cglues the last line of one file onto the first line of the next and silently corrupts two records. - The last block is a deliberately truncated archive:
gzipreportsunexpected end of file(its message starts with a newline, which is the blank line in the transcript, not output from the script), andxargsexits 123, the value xargs uses for "a command it ran exited with status 1 to 125", so the script fails instead of publishing a short CSV as if it were complete. The innerbash -o pipefail -c(pipefailmakes a pipeline fail if any stage fails, not only the last) is what carriesgzip's failure through the| sed; without it the pipe would report onlysed's success.
Why it stays streaming
find lists names; xargs -n1 runs one decompress per file in sorted order (sort -z keeps NUL separators and a deterministic order, which matters because "first wins" depends on file order); jq -R reads raw lines and @csv emits a row immediately; awk prints a row the moment it is judged new. At any instant memory holds a few pipe buffers plus awk's seen table. fromjson? skips a malformed line; add a counter if silent skipping is not acceptable.
Trade-offs and pitfalls
- Bounded memory dedupe: for hundreds of millions of ids, replace the
awkstage withsort -s -t, -k1,1 -u -S 1G -T /big/tmp, an external merge sort (sorts chunks in memory, writes them to temporary files, merges them) that spills to disk. The flags:-t,uses a comma as the field separator,-k1,1compares only field 1 (the id),-ukeeps one row per equal key,-smakes the sort stable so the first of equal ids in input order is the one kept,-S 1Gcaps the memory buffer at 1 GB, and-T /big/tmpnames the directory for temporary files. With-sand-uGNU sort keeps the first of each run of equal ids in input order (checked on a four-line sample), and the output is ordered by id, not arrival. - Latest wins instead of first wins: emit
tsfirst, sort by id thentsdescending and take the first row per id. - Assumption: ids contain no commas or quotes, so the first CSV field is the whole id. A string id arrives quoted (
"abc"); that still works as a dedupe key. If ids can contain commas, dedupe inside jq before formatting. - One thread: the pipeline is limited by the slowest stage, usually
jq. For hundreds of files usexargs -Pfor the decompress-and-parse step and dedupe once at the end, because dedupe has to see all rows. - Cost of
-n1: onebashstart per file, fine for thousands of files.
Write a Bash function that runs a command and retries it on failure with exponential backoff and jitter, up to a configurable number of attempts, preserving the command's last exit status and logging each attempt. Where would you apply it, and where would retrying be wrong?
Sample Answer
Direct answer
A retry function takes the attempt limit, a base delay and a delay cap, then the command. It runs the command, returns immediately on success, and on failure logs the attempt, sleeps a random time between 0 and the current delay (exponential backoff with "full jitter"), doubles that delay up to the cap, and returns the command's own last exit status once the attempts are used up. Retrying is right for transient, repeatable failures (a flaky network, a throttled API, a service that is restarting) and wrong for failures that will not change (bad credentials, a typo, a 404) and for operations that are not safe to repeat (creating an order, charging a card) unless the receiver deduplicates them.
The function
# retry.sh: source this file. retry MAX_ATTEMPTS BASE_SECONDS CAP_SECONDS COMMAND [ARGS...]
# Also: RETRY_FATAL="76 77" lists exit statuses that must NOT be retried.
retry() {
local max=$1 base=$2 cap=$3
shift 3
local attempt=1 rc delay=$base window wait_for
while :; do
if "$@"; then
(( attempt > 1 )) && echo "retry: succeeded on attempt $attempt" >&2
return 0
else
rc=$?
fi
echo "retry: attempt $attempt/$max failed with status $rc: $*" >&2
if [[ " ${RETRY_FATAL:-} " == *" $rc "* ]]; then
echo "retry: status $rc is marked fatal, not retrying" >&2
return "$rc"
fi
(( attempt >= max )) && return "$rc"
window=$(( delay > cap ? cap : delay )) # 1, 2, 4, 8 ... for base 1, never above cap
wait_for=$(( RANDOM % (window + 1) )) # "full jitter": uniform 0..window
echo "retry: sleeping ${wait_for}s (window 0-${window}s)" >&2
sleep "$wait_for"
delay=$(( window * 2 )) # doubles the capped window, so it cannot overflow
attempt=$(( attempt + 1 ))
done
}
How each requirement is met:
- Exponential backoff: the window starts at
baseand doubles after every failed attempt, never exceeding the cap, so withbase=1and a large cap the windows are 1, 2, 4, 8 seconds and so on. The next window is computed by doubling the previous capped window rather than asbase * 2 ** (attempt - 1): with that closed form, a 64-bit shell overflows at attempt 64 (executed withretry 70 1 8 false: the window printed as0--9223372036854775808s, one run drew a sleep of 25,020 seconds (random, so the value differs per run), and every later attempt used a 0 second window, so the backoff vanished exactly when the outage was longest). - Jitter: the actual sleep is uniform between 0 and the window. If a hundred jobs fail at the same moment (a shared dependency restarts), pure exponential backoff makes all hundred retry together at 1 s, 2 s, 4 s; randomizing spreads them out so the dependency is not hit by a wave each time.
- Last exit status preserved: the status is captured in the
elsebranch ofif "$@", where$?is still the command's. It is returned unchanged on the final failure. - Safe under
set -e: because the command runs as the condition of anif, a failure never aborts a script that usesset -e, and the caller decides what to do with the returned status. - Logging each attempt: to stderr, so the command's stdout can still be piped. In production log only the command name (
$1) instead of$*, because arguments can hold tokens and passwords. - Fatal statuses:
RETRY_FATALlists exit statuses that mean "do not try again", which matters for the HTTP case below.
Reading the shell syntax
shift 3and"$@": the first three arguments are the settings, andshift 3drops them so that"$@"holds only the command and its arguments."$@"expands to every remaining argument as a separate word with its internal spaces kept, soretry 3 1 4 bash -c 'echo trying; exit 7'hands bash the script string intact. Unquoted$@would split on spaces and break it (shown below).$?: the exit status of the most recent command, where 0 means success and anything else means failure.RANDOM: a bash variable that gives a new pseudo-random whole number from 0 to 32767 each time it is read.RANDOM % (window + 1)is the remainder after dividing bywindow + 1, which always lands between 0 andwindow. AssigningRANDOM=7seeds the generator, so the same numbers come out in the same order on the same bash version, in that shell process; a subshell (( ... )or$( ... )) starts from a fresh seed in bash 5.2.window=$(( delay > cap ? cap : delay )): inside(( ))and$(( )),a ? b : cpicksbwhen the testais true andcotherwise, so the window isdelayunless that exceedscap.delay=$(( window * 2 ))then doubles it for the next failure (1, 2, 4, 8 for base 1).RANDOMtops out at 32767, so a window above 32767 seconds is never fully used: the sleep stays within 0 to 32767 seconds.- The padded test
[[ " ${RETRY_FATAL:-} " == *" $rc "* ]]: the list and the status are each wrapped in spaces so that only a whole number matches. Without the padding, status 6 would match inside "76".${RETRY_FATAL:-}means "the variable, or empty if unset", which keeps the function safe underset -u. - Exit statuses 75 and 76: an arbitrary convention chosen by this script (any two unused numbers work): 75 means "transient, try again" and 76 means "permanent, stop".
code=$(curl ...) && rc=0 || rc=$?: read it in three stages.$(curl ...)runs curl and captures what it prints, here only the HTTP status code (-w '%{http_code}'prints it,-o /dev/nulldiscards the body,-sShides the progress meter but still shows errors,-m 5gives up after 5 seconds). An assignment's own exit status is the status of the command inside it, so if curl succeeded,&& rc=0runs; if curl failed (connection refused, timeout),|| rc=$?records curl's failure status. Afterwardsrcis non-zero only when no HTTP response arrived, andcodeholds the HTTP status otherwise.
A small run shows the pieces. The sleeps in the test below come from the pinned RANDOM values, so you can redo them by hand: the first window is 0-1 s, and 19344 % 2 = 0 gives a 0 s sleep; the second window is 0-2 s, and 26956 % 3 = 1 gives 1 s. The same stream continues into the second test case (7409 % 2 = 1 s, then 20442 % 3 = 0 s).
RANDOM=7
for w in 1 2 1 2; do r=$RANDOM; echo "RANDOM=$r window=0-$w sleep=$(( r % (w + 1) ))"; done
show() { printf '[%s] ' "$@"; echo; }
set -- 'one two' three # fake two script arguments: "one two" and "three"
echo "== quoted \"\$@\" keeps words"; show "$@"
echo "== unquoted \$@ re-splits"; show $@
RETRY_FATAL="76 77"; rc=6
[[ " $RETRY_FATAL " == *" $rc "* ]] && echo "padded: 6 matches" || echo "padded: 6 does not match"
[[ "$RETRY_FATAL" == *"$rc"* ]] && echo "unpadded: 6 matches (wrong)" || echo "unpadded: no match"
RANDOM=19344 window=0-1 sleep=0
RANDOM=26956 window=0-2 sleep=1
RANDOM=7409 window=0-1 sleep=1
RANDOM=20442 window=0-2 sleep=0
== quoted "$@" keeps words
[one two] [three]
== unquoted $@ re-splits
[one] [two] [three]
padded: 6 does not match
unpadded: 6 matches (wrong)
Test run: the function on its own
The jitter sequence is pinned with RANDOM=7 so the same sleeps appear every time on bash 5.2. The set -e case runs in a subshell, and bash 5.2 gives a subshell its own fresh random seed (executed: twelve subshells after RANDOM=7 printed twelve different number pairs, while the main shell printed 19344 26956 every time), so that case seeds again inside the subshell.
demo.sh:
#!/usr/bin/env bash
cd /w
source ./retry.sh; source ./post.sh
RANDOM=7 # pins the jitter sequence
n=0
flaky() { n=$(( n + 1 )); (( n >= 3 )); } # fails on calls 1 and 2, succeeds on call 3
echo "== command fails twice, then succeeds"
retry 5 1 8 flaky; echo "status: $?"
echo "== command never succeeds: last status is preserved"
retry 3 1 4 bash -c 'echo trying; exit 7'; echo "status: $?"
echo "== runs under set -e without killing the caller"
( set -e; RANDOM=7; retry 2 1 1 false || echo "caller still alive, rc=$?"; echo after ) # a subshell gets a fresh seed, so seed again inside it
== command fails twice, then succeeds
retry: attempt 1/5 failed with status 1: flaky
retry: sleeping 0s (window 0-1s)
retry: attempt 2/5 failed with status 1: flaky
retry: sleeping 1s (window 0-2s)
retry: succeeded on attempt 3
status: 0
== command never succeeds: last status is preserved
trying
retry: attempt 1/3 failed with status 7: bash -c echo trying; exit 7
retry: sleeping 1s (window 0-1s)
trying
retry: attempt 2/3 failed with status 7: bash -c echo trying; exit 7
retry: sleeping 0s (window 0-2s)
trying
retry: attempt 3/3 failed with status 7: bash -c echo trying; exit 7
status: 7
== runs under set -e without killing the caller
retry: attempt 1/2 failed with status 1: false
retry: sleeping 0s (window 0-1s)
retry: attempt 2/2 failed with status 1: false
caller still alive, rc=1
after
The first case fails twice and succeeds on the third call, after sleeps of 0 s and 1 s drawn from windows of 0-1 s and 0-2 s. The second shows the exit status 7 coming back out unchanged after three attempts. The third shows the caller surviving under set -e.
The curl POST case: retry only what is retryable, with a cap
For HTTP the command's exit status is not enough, because curl exits 0 for any response it receives, a 404 or a 503 included (only the -f flag turns an HTTP error into exit 22), so the status code has to be read separately. The wrapper reads the status code, maps it to a small set of exit codes, and retry treats one of them as fatal. Network errors, 408 (request timeout), 429 (too many requests) and 5xx responses are transient (exit 75). Other 4xx responses are permanent (exit 76), because sending the same bad request again cannot succeed. The delay cap bounds the worst wait.
post.sh:
# post.sh: POST JSON, classify the outcome so retry() can tell transient from permanent.
# 0 = success, 75 = transient (network error, 408, 429, 5xx), 76 = permanent (other 4xx).
post_json() {
local url=$1 body=$2 code rc
code=$(curl -sS -m 5 -o /dev/null -w '%{http_code}' -X POST \
-H 'Content-Type: application/json' -H "Idempotency-Key: $IDEM_KEY" \
-d "$body" "$url") && rc=0 || rc=$?
if (( rc != 0 )); then echo "post: curl error $rc" >&2; return 75; fi
case $code in
2??) return 0 ;;
408|429|5??) echo "post: HTTP $code (transient)" >&2; return 75 ;;
*) echo "post: HTTP $code (permanent)" >&2; return 76 ;;
esac
}
The test server answers /a with 503, 503, 200 and /b with 400, then 200, and records the Idempotency-Key header it sees:
flaky.py:
import http.server,sys
seq={"/a":[503,503,200],"/b":[400,200],"/c":[200]}
class H(http.server.BaseHTTPRequestHandler):
def log_message(self,*a): pass
def do_POST(self):
n=int(self.headers.get("Content-Length",0)); self.rfile.read(n)
s=seq[self.path]; code=s.pop(0) if len(s)>1 else s[0]
sys.stdout.write(f"server saw {self.path} key={self.headers.get('Idempotency-Key')} -> {code}\n"); sys.stdout.flush()
self.send_response(code); self.end_headers()
http.server.HTTPServer(("127.0.0.1",7010),H).serve_forever()
demo2.sh:
#!/usr/bin/env bash
cd /w
source ./retry.sh; source ./post.sh
python3 flaky.py & sleep 1
RANDOM=7; export IDEM_KEY=order-1001
echo "== 503, 503, 200"
RETRY_FATAL="76" retry 5 1 4 post_json http://127.0.0.1:7010/a '{"x":1}'; echo "status: $?"
echo "== 400 is permanent"
RETRY_FATAL="76" retry 5 1 4 post_json http://127.0.0.1:7010/b '{"x":1}'; echo "status: $?"
echo "== connection refused is transient"
RETRY_FATAL="76" retry 2 1 1 post_json http://127.0.0.1:7999/z '{"x":1}'; echo "status: $?"
== 503, 503, 200
server saw /a key=order-1001 -> 503
post: HTTP 503 (transient)
retry: attempt 1/5 failed with status 75: post_json http://127.0.0.1:7010/a {"x":1}
retry: sleeping 0s (window 0-1s)
server saw /a key=order-1001 -> 503
post: HTTP 503 (transient)
retry: attempt 2/5 failed with status 75: post_json http://127.0.0.1:7010/a {"x":1}
retry: sleeping 1s (window 0-2s)
server saw /a key=order-1001 -> 200
retry: succeeded on attempt 3
status: 0
== 400 is permanent
server saw /b key=order-1001 -> 400
post: HTTP 400 (permanent)
retry: attempt 1/5 failed with status 76: post_json http://127.0.0.1:7010/b {"x":1}
retry: status 76 is marked fatal, not retrying
status: 76
== connection refused is transient
curl: (7) Failed to connect to 127.0.0.1 port 7999 after 0 ms: Couldn't connect to server
post: curl error 7
retry: attempt 1/2 failed with status 75: post_json http://127.0.0.1:7999/z {"x":1}
retry: sleeping 1s (window 0-1s)
curl: (7) Failed to connect to 127.0.0.1 port 7999 after 0 ms: Couldn't connect to server
post: curl error 7
retry: attempt 2/2 failed with status 75: post_json http://127.0.0.1:7999/z {"x":1}
status: 75
The server saw the same Idempotency-Key on every attempt, the 503s were retried until the 200, the 400 stopped immediately (status 76, one request), and a refused connection was treated as transient and retried. A Retry-After header on a 429 should override the computed sleep in a production version; this one does not read it.
Where retrying is wrong
- Operations that are not idempotent (safe to repeat with the same effect). If a POST times out, the server may have processed it. Retrying can create a second order or charge a card twice. Retry only with an idempotency key the server honors (shown above), or after checking whether the first attempt took effect.
- Failures that will not change: authentication errors, validation errors, "not found", a syntax error in the command. Retrying only delays the failure by the sum of the sleeps. Classify the outcome and mark those statuses fatal.
- Stacked retries. If the script retries 3 times, the client library inside it retries 3 times and a proxy retries 3 times, one logical request becomes up to 27 attempts against a dependency that is already struggling. Retry at one layer.
- Steps that must fail fast, such as a pre-deployment health gate, or any command under a deadline: bound the total time as well as the attempts.
- Commands that leave partial effects (a half-copied file, a half-applied migration) unless the command is written to be re-run safely.
Where it fits: package and image downloads, kubectl or cloud API calls during a control-plane restart, connecting to a database that is still starting, and calls to rate-limited APIs.
How do you parse JSON reliably in a Bash script? Show how to extract fields, handle missing keys and arrays, and what you would do if the usual tool is not installed.
Sample Answer
Direct answer
Use jq, a JSON-aware command-line processor, and never grep, sed or awk on JSON (they cannot tell a key from text inside a string and break on whitespace, nesting and escapes). jq -r extracts fields as plain text, // default supplies a value for null or missing keys, []? iterates arrays that may be absent, and jq -e sets the exit status so scripts can test for presence. If jq is not installed, install it (package manager or the static release binary) or fall back to python3 -c 'import json...', which is on most hosts.
The filter language in brief, so the script reads easily: .user.name walks into nested objects; .items[] emits every element of an array as a separate output; | feeds one stage's output into the next; // means "or this default" when the left side is null, false or missing; a trailing ? swallows an error for that one step. Process substitution < <(command) makes a command's output look like a file, so mapfile -t ids < <(jq ...) reads jq's lines into the array ids (mapfile -t strips the newlines) without running the loop in a subshell.
Worked script
#!/usr/bin/env bash
set -euo pipefail
cat > resp.json <<'JSON'
{"user": {"name": "Ada", "email": null},
"items": [{"id": 1, "tags": ["a", "b"]}, {"id": 2}, {"id": 3, "tags": []}],
"next": null}
JSON
echo "1) plain extraction: $(jq -r '.user.name' resp.json)"
echo "2) null prints null: $(jq -r '.user.email' resp.json)"
echo "3) default for null: $(jq -r '.user.email // "none"' resp.json)"
echo "4) missing key: $(jq -r '.user.phone // "none"' resp.json)"
echo "5) has() vs null: email=$(jq '.user | has("email")' resp.json) phone=$(jq '.user | has("phone")' resp.json)"
echo "6) array of ids:"
mapfile -t ids < <(jq -r '.items[].id' resp.json)
printf ' id=%s\n' "${ids[@]}"
echo "7) tags per item (missing tags handled):"
jq -r '.items[] | "\(.id): \((.tags // []) | join(","))"' resp.json
echo "8) .tags[] on an item with no tags:"
jq -r '.items[].tags[]' resp.json 2>&1 || echo " jq exit=$?"
echo " with ?: $(jq -r '.items[].tags[]?' resp.json | paste -sd' ')"
echo "9) -e turns null/false into exit 1, no output into exit 4:"
jq -e '.next' resp.json >/dev/null && echo " next present" || echo " next null, exit=$?"
jq -e '.items[] | select(.id == 99)' resp.json >/dev/null && echo " found" || echo " no match, exit=$?"
echo "10) shell variable in, safely:"
want=2
jq --argjson want "$want" '.items[] | select(.id == $want) | .id' resp.json
echo "11) bad JSON stops the script:"
echo '{"a": ' | jq . 2>&1 || echo " jq exit=$?"
echo "12) fallback without jq (python3):"
python3 -c 'import json,sys; print(json.load(sys.stdin)["user"]["name"])' < resp.json
echo "13) fail fast on a required field:"
if ! phone=$(jq -er '.user.phone' resp.json); then echo " required field user.phone is missing"; fi
Output (jq 1.7 and bash 5.2 on Ubuntu 24.04):
1) plain extraction: Ada
2) null prints null: null
3) default for null: none
4) missing key: none
5) has() vs null: email=true phone=false
6) array of ids:
id=1
id=2
id=3
7) tags per item (missing tags handled):
1: a,b
2:
3:
8) .tags[] on an item with no tags:
jq: error (at resp.json:3): Cannot iterate over null (null)
a
b
jq exit=5
with ?: a b
9) -e turns null/false into exit 1, no output into exit 4:
next null, exit=1
no match, exit=4
10) shell variable in, safely:
2
11) bad JSON stops the script:
jq: parse error: Unfinished JSON term at EOF at line 2, column 0
jq exit=5
12) fallback without jq (python3):
Ada
13) fail fast on a required field:
required field user.phone is missing
What each case teaches
- Extraction (1) and
-r.-rprints a string without JSON quotes. Without it you get"Ada"with the quotes. - Null versus missing (2 to 5).
-rprints a literalnullfor a null value, and a script that stores that gets the four-character stringnull.// "none"replaces bothnulland absent keys (and alsofalse, which matters for booleans). When the difference between "key present with null" and "key absent" matters, usehas("key"). - Arrays (6, 7).
.items[].idemits one value per line, somapfile -t ids < <(...)(bash 4+) loads them into an array without word-splitting surprises. Item 7 shows(.tags // []) | join(",")handling an item with notags. - Missing array (8).
.items[].tags[]works on item 1, then aborts on item 2 withCannot iterate over null(exit 5 in this run) after output has already been produced. In the transcript the error line is printed beforeaandbeven though jq producedaandbfirst: with2>&1into a pipe or file, jq's stdout is buffered and flushed at exit, while stderr is written immediately. Run separately (>o.txt 2>e.txt),o.txtholdsaandbande.txtholds the error, andstdbuf -oL jq ...puts them in true order. Add?(.tags[]?) to suppress the error for absent or null arrays. - Exit status (9). With
-e, jq exits 1 if the last output isfalseornull, and 4 if there was no output at all, soif jq -e '.next' f >/dev/nullworks as a presence test. - Passing shell values safely (10).
--arg name value(always a string) and--argjson name value(typed JSON) put values in as data. Never splice$varinto the filter text, which breaks on quotes and allows filter injection. For example, withwant='2 or true', the spliced filter.items[] | select(.id == 2 or true)selects every item (the run printed1 2 3), while--argjsonrefuses the same text withinvalid JSON text passed to --argjson, and--argwould keep it as an inert string. - Bad input (11). A parse error is a non-zero exit with a message on stderr, so under
set -eit stops the script, which is usually what you want. - No
jq(12).python3withjson.loadis a safe fallback for simple extractions. - Required fields (13). Capture into a variable inside
if ! var=$(jq -er ...). A bareecho "$(jq ...)"hides jq's failure, becauseechosucceeds.
Pitfalls
set -eand command substitution.x=$(jq ...)aborts the script on failure;echo "x=$(jq ...)"does not. Capture first, then use.- Do not
evaljq output. Treat extracted text as data, quote every expansion, and prefermapfileorreadto unquotedfor x in $(jq ...). - Big numbers. In jq 1.7,
{"id":12345678901234567890}passes through.idunchanged, but.id + 1prints12345678901234567000: arithmetic uses IEEE doubles, the standard 64-bit floating-point format, which stores whole numbers exactly only up to 2^53 (about 9 quadrillion). Do not do arithmetic on ids, and treat them as strings when you can (older jq versions may round even a pass-through, so check yours). - One document per call. For large inputs,
jq -con JSON Lines streams a record at a time;--streamhandles giant single documents.
Each line of a JSON Lines file looks like {"user":{"id":123,"name":"a"},"events":[...]}. Produce CSV rows with user id, name and event count in one shell pipeline, and explain how it handles nulls and missing fields.
Sample Answer
Direct answer
Read the file as raw text and parse each line separately (jq -R with fromjson?), take the nested fields with a default for missing parents, turn events into a count only if it is actually an array, and finish with @csv, which quotes and escapes correctly. Nulls and missing fields become empty CSV cells, a missing or null events counts as 0, and a line that is not valid JSON is skipped (and counted) instead of aborting the run.
The pipeline
jq -Rr '
fromjson? # skip lines that are not valid JSON
| (.user // {}) as $u
| [ $u.id, # null becomes an empty CSV field
$u.name,
(.events | if type == "array" then length else 0 end) ]
| @csv
' users.jsonl | { echo 'user_id,name,event_count'; cat; }
echo "--- bad lines dropped:"
echo $(( $(wc -l < users.jsonl) - $(jq -Rc 'fromjson?' users.jsonl | wc -l) ))
echo "--- same pipeline without -R/fromjson? stops at the bad line:"
jq -r '[.user.id, .user.name, (.events | length)] | @csv' users.jsonl; echo "exit=$?"
Input users.jsonl used for the run (each row exercises one case):
{"user":{"id":123,"name":"a"},"events":[{"t":1},{"t":2},{"t":3}]}
{"user":{"id":124,"name":null},"events":[]}
{"user":{"id":125},"events":null}
{"user":{"id":126,"name":"Lee, \"Jo\""}}
{"user":null,"events":[{"t":1}]}
{"user":{"id":127,"name":""},"events":"oops"}
{this is not json}
{"user":{"id":128,"name":"z"},"events":[{"t":9}]}
Output:
user_id,name,event_count
123,"a",3
124,,0
125,,0
126,"Lee, ""Jo""",0
,,1
127,"",0
128,"z",1
--- bad lines dropped:
1
--- same pipeline without -R/fromjson? stops at the bad line:
jq: parse error: Invalid literal at line 7, column 6
123,"a",3
124,,0
125,,0
126,"Lee, ""Jo""",0
,,1
127,"",4
exit=5
Reading the filter piece by piece. -R (raw input) makes jq hand each input line over as a plain string instead of parsing it as JSON, and -r prints strings without JSON quotes. fromjson parses that string as JSON, and the ? after it means "if parsing fails, produce nothing for this line and carry on", which is how a bad line is skipped. (.user // {}) as $u computes .user, replaces a null or missing value with an empty object {}, and stores the result in a variable named $u, so $u.id and $u.name are then null rather than an error when user is absent. In .events | if type == "array" then length else 0 end, the pipe passes .events to an if that counts elements only when the value is an array. The three values go into a list [ ... ], and @csv turns that list into one correctly quoted CSV line. The chained { echo ...; cat; } prints the header first and then copies jq's rows through.
How each row is handled
| Input row | Output | Why |
|---|---|---|
| 123 with three events | 123,"a",3 | the normal case |
name is null | 124,,0 | @csv renders null as an empty cell |
events is null | 125,,0 | not an array, so the count is 0 |
events key missing | 126,"Lee, ""Jo""",0 | .events is null when absent; the comma and embedded quotes in the name are escaped for you |
user is null | ,,1 | (.user // {}) turns it into an empty object so .id and .name are null |
events is the string "oops" | 127,"",0 | name is an empty string and prints ""; events is not an array, so 0 |
{this is not json} | (dropped) | fromjson? swallows the parse error |
Two details worth noticing. An empty-string name prints as "" while a null name prints as nothing, so the two stay distinguishable if the consumer cares. And the type == "array" guard matters: the simpler .events | length reports the string length of "oops" (4) and counts as 0 only for null.
The last two commands in the script show the two safeguards. The first prints the number of dropped lines (1): wc -l counts 8 lines in the file, jq -Rc 'fromjson?' | wc -l counts the 7 lines that parsed, and 8 - 7 = 1. Without -R and fromjson?, jq stops at the bad line (jq: parse error, exit 5) after emitting the rows before it, so the output is silently truncated, and the plain length version prints 127,"",4.
Why it stays a streaming one-liner
jq -R reads one line at a time, and @csv writes it immediately, so memory is one record. The header is added by { echo ...; cat; } so jq never needs to know it. Time is O(n) in the number of lines, meaning the work grows in direct proportion to the line count.
Pitfalls
- Silent drops are a decision.
fromjson?hides corruption. Always count the dropped lines (as shown) and fail the job above a threshold you choose. - Field order is by position, so the header and the array in the jq filter must be edited together.
- Embedded newlines in a string value are legal in CSV after
@csvquoting, but they break line-oriented consumers further down the pipe. @tsvis the alternative when the consumer cannot parse CSV quoting; it escapes tabs and newlines as\tand\ninstead.- If the same user can appear in many lines, aggregate later; the pipeline above reshapes one line at a time.
That is every published Shell Scripting and Automation question for Cloud Engineer so far. Browse the other topics in this category, or practice this one interactively.