Shell Scripting and Automation Questions
Shell-language craft for operations work: Bash and POSIX sh scripting, pipes and redirection, quoting and word splitting, exit codes and set -euo pipefail (including failures inside pipelines), traps and cleanup, argument parsing with getopts, functions and arrays, parameter expansion, here-documents, and text processing with grep, sed, awk and jq, including streaming pipelines over very large log, CSV and JSON Lines files. Also covers writing robust, idempotent, portable scripts (locking, atomic file updates, retries with backoff, safe temp files, GNU versus BSD differences), background jobs and signals, bounded parallelism and SSH fan-out, cron-driven jobs such as log rotation, backups, atomic deploys and health-check watchdogs, secure scripting (eval injection, secrets, path traversal), and debugging, testing and linting shell with ShellCheck and bats. Boundary: general automation design and Python or Go tooling, Linux host administration tasks, observability and alerting design, security detection engineering, and generic algorithmic coding problems are covered elsewhere.
How do you make a Bash script clean up after itself when it exits normally, fails part-way, is interrupted with Ctrl-C, or is terminated by another process? Explain what you hook, what happens to the exit status, and what behaves differently in subshells and functions.
Sample Answer
Direct answer. Register one cleanup function with trap cleanup EXIT. The EXIT pseudo-signal fires on every way the shell ends: normal completion, an exit call, a failure under set -e and, in Bash, death from SIGINT (Ctrl-C), SIGTERM (the default signal of kill) or SIGHUP, even when you have set no signal traps (executed on bash 5.2.21: an EXIT-only script that sent itself each of those signals printed its cleanup line and left with status 130, 143 and 129). Add trap 'exit 130' INT and trap 'exit 143' TERM anyway: they turn the signal into an ordinary exit with a documented status, so a parent such as Python's subprocess sees return code 143 instead of -15, and they keep the script correct under shells that do not run the EXIT trap on a signal (executed: dash printed nothing on SIGTERM). The one thing you cannot clean up after is SIGKILL (kill -9), which the kernel delivers without giving the process a chance to run anything.
Terms used in the examples
- Signal. A small numbered message the operating system delivers to a process. SIGINT (2) is what Ctrl-C sends, SIGTERM (15) is what a plain
kill PIDsends, and SIGKILL (9) cannot be caught or ignored. - Pseudo-signal.
EXITandERRare names Bash accepts intrapthat are not real operating-system signals:EXITmeans "this shell is about to end" andERRmeans "a command just failed". $$. The process ID of the running script itself, sokill -INT $$makes the script send itself the signal Ctrl-C would send.- Subshell. A child copy of the shell, created by
( ... ), a pipeline stage or$( ... ). It has its own copy of the variables and its own exit. - Process group. The set of processes (the script and the programs it started) that the terminal signals together when you press Ctrl-C.
set -E(errtrace). Makes theERRtrap apply inside functions and subshells too.
Four hooks appear in the worked example, but only one of them does the cleaning: trap cleanup EXIT. The INT and TERM lines turn those signals into an ordinary exit with a recognisable status (Bash would run the EXIT trap for them anyway, as the direct answer shows), and the ERR line only logs which command failed. A script with just the EXIT line already cleans up after normal ends, exit calls, set -e failures and, in Bash, SIGINT, SIGTERM and SIGHUP.
Which hook for which event
| Event | What you hook | Status the script leaves with |
|---|---|---|
Normal end or exit N | trap cleanup EXIT | The status it was already leaving with |
Command fails under set -e | EXIT (and optionally ERR to log the line) | The failing command's status |
| Ctrl-C | EXIT runs (in Bash even without an INT trap); trap 'exit 130' INT makes it a normal exit | 130 (128 plus signal 2) |
kill PID | EXIT runs (in Bash even without a TERM trap); trap 'exit 143' TERM makes it a normal exit | 143 (128 plus signal 15) |
kill -9 | Nothing can be hooked | 137, no cleanup runs |
The exit status inside and after the trap
At the moment the EXIT trap starts, $? still holds the status the script was exiting with, so read it first (local rc=$?). The trap's own commands then run, and when the trap finishes the script exits with the original status, unless the trap itself calls exit N, which replaces it. The safe pattern is to capture rc, clean up, and exit "$rc". Clear the traps inside the handler (trap - EXIT INT TERM) so a second signal during cleanup does not re-enter it.
Worked example
#!/usr/bin/env bash
set -Eeuo pipefail
tmpdir=$(mktemp -d)
cleanup() {
local rc=$? # status the script was exiting with
trap - EXIT INT TERM # avoid re-entry while cleaning up
rm -rf -- "$tmpdir"
echo "cleanup: removed tmpdir, original status $rc"
exit "$rc" # preserve it
}
trap cleanup EXIT
trap 'exit 130' INT # 128 + SIGINT(2); EXIT trap then runs
trap 'exit 143' TERM # 128 + SIGTERM(15)
trap 'echo "ERR trap: line $LINENO: $BASH_COMMAND" >&2' ERR
case ${1:-ok} in
ok) echo "work done" ;;
fail) false; echo "not reached" ;;
int) kill -INT $$; sleep 1; echo "not reached" ;;
term) kill -TERM $$; sleep 1; echo "not reached" ;;
sub) ( exit 5 ) || echo "subshell left with $?, parent cleanup has not run yet"
( trap 'echo "subshell own EXIT trap"' EXIT; exit 6 ) || echo "subshell left with $?" ;;
func) f() { false; }; f; echo "not reached" ;;
esac
Each mode, run separately (kill -INT $$ sends the signal to the script itself, which is what Ctrl-C does to a foreground script):
$ bash cleanup.sh ok
work done
cleanup: removed tmpdir, original status 0
exit=0
$ bash cleanup.sh fail
ERR trap: line 19: false
cleanup: removed tmpdir, original status 1
exit=1
$ bash cleanup.sh int
cleanup: removed tmpdir, original status 130
exit=130
$ bash cleanup.sh term
cleanup: removed tmpdir, original status 143
exit=143
$ bash cleanup.sh sub
subshell left with 5, parent cleanup has not run yet
subshell own EXIT trap
subshell left with 6
cleanup: removed tmpdir, original status 0
exit=0
$ bash cleanup.sh func
ERR trap: line 24: false
cleanup: removed tmpdir, original status 1
exit=1
Note that with set -E the ERR trap also fires inside functions (func case, line 24), and that the cleanup ran exactly once per script even though three traps were registered.
Reading the sub case: ( exit 5 ) ends only the subshell, so $? becomes 5 while the script itself keeps running, and the script's cleanup has not fired, because it fires when the parent shell ends. This matters for ownership: the temp directory belongs to the parent, so the trap that removes it must be registered by the parent. A trap registered inside a subshell runs when that subshell ends, which would delete the directory while the parent still needs it. In the second line the subshell registers its own EXIT trap, which fires at its own exit (subshell own EXIT trap).
Subshells and functions
- Subshells. A subshell (
( ... ), a pipeline stage,$( )) resets caught traps to the default. The parent's EXIT trap does not run when the subshell exits: in thesubcase,( exit 5 )returned and the parent's cleanup had not yet run. A subshell may set its own trap, as the second line shows. Put the cleanup in the process that owns the resource. - Functions. Traps are process-wide, not function-scoped. An EXIT trap set inside a function replaces the script's one and fires at script exit, not when the function returns. (Checked: a trap set in
fwas still present, and ran, afterfreturned.) - ERR trap and
set -E. Without-E(errtrace), the ERR trap is not inherited by shell functions, command substitutions or subshells. With the trap set andset -ebut no-E, a failingfalseinside a function exits the script with status 1 and the ERR trap never prints. Adding-Emakes it printERR trap fired: falsefirst (both variants run, exit status 1 in each).
Timing traced with numbers
SECONDS is a Bash variable holding the whole seconds since the script started. Two scripts, each sent SIGTERM from outside one second after they start (SIGTERM is used so the demo can run in the background; any trapped signal follows the same rule):
# foreground.sh # background-wait.sh
trap 'echo "TERM trap ran at ${SECONDS}s"; exit 143' TERM
sleep 3 sleep 3 & wait $!
-- foreground.sh
TERM trap ran at 3s
-- background-wait.sh
TERM trap ran at 1s
The signal arrived at second 1 in both runs. With sleep 3 in the foreground, Bash held the trap until that command finished, at second 3, a two-second delay. With sleep 3 & followed by wait $!, the wait builtin is interrupted by the signal, the trap ran at second 1, and the script exited at once. A 10-minute child would delay the trap by up to 10 minutes in the first form.
Timing, trade-offs and pitfalls
- Bash runs a trap only after the current foreground command finishes. If the script is inside
sleep 3and receives SIGINT after 1 second, the INT trap runs at 3 seconds (measured), whereas thewaitbuiltin returns at once. For long children, run them in the background andwait, or forward the signal to the child's process group. - Ctrl-C goes to the whole foreground process group, so the child also gets SIGINT; your handler cleans the script's resources, not the child's.
- The same trap on EXIT for a temp directory is safer than putting
rm -rfat the end of the script: usemktemp -dandrm -rf -- "$tmpdir"with the variable quoted and set before the trap is registered. - SIGKILL and power loss leave debris. For resources that must not leak (locks, mounts, temp databases), also use a startup sweep, or a lock that the kernel releases when the process dies (
flock). - Give the INT and TERM traps an explicit status (
exit 130,exit 143) so callers can tell an interrupted run from a failed one.
What is the difference between a shell variable and an environment variable? Explain why VAR=1 bash -c 'echo $VAR' prints nothing, how to make a value visible to child processes, and how you would set one for a single command versus every login session.
Sample Answer
Direct answer
A shell variable lives only inside the shell process that set it. An environment variable is a variable the shell has marked for export, so it is copied into the block of NAME=value strings handed to every program the shell starts. One correction to the question first: VAR=1 bash -c 'echo $VAR' does print 1. A prefix assignment on a command line puts the variable into that command's environment, and the single quotes make the child shell, not the parent, expand $VAR. The commands that print nothing are VAR=1 on its own line followed by bash -c 'echo $VAR', and VAR=1 echo $VAR, which are explained below.
What each case does
unset VAR
echo '1) VAR=1 bash -c '"'"'echo $VAR'"'"
VAR=1 bash -c 'echo "child sees: [$VAR]"'
echo '2) VAR=1 on its own line, then the child'
VAR=1
bash -c 'echo "child sees: [$VAR]"'
echo "parent sees: [$VAR]"
echo '3) VAR=1 echo $VAR (the parent expands $VAR before the assignment applies)'
unset VAR
VAR=1 echo "echo printed: [$VAR]"
echo '4) export, then child changes its copy'
export VAR=1
bash -c 'echo "child sees: [$VAR]"; VAR=99; echo "child now: [$VAR]"'
echo "parent still: [$VAR]"
echo '5) is it exported?'
declare -p VAR
env | grep '^VAR='
echo '6) env VAR=2 for one command, env -u to remove it'
env VAR=2 bash -c 'echo "child sees: [$VAR]"'
env -u VAR bash -c 'echo "child sees: [${VAR-unset}]"'
echo '7) login shell reads ~/.profile'
export HOME=$PWD/fakehome; mkdir -p "$HOME"
echo 'export EDITOR=vim' > "$HOME/.profile"
unset EDITOR
bash -c 'echo "non-login: [${EDITOR-unset}]"'
bash -lc 'echo "login: [${EDITOR-unset}]"'
Output:
1) VAR=1 bash -c 'echo $VAR'
child sees: [1]
2) VAR=1 on its own line, then the child
child sees: []
parent sees: [1]
3) VAR=1 echo $VAR (the parent expands $VAR before the assignment applies)
echo printed: []
4) export, then child changes its copy
child sees: [1]
child now: [99]
parent still: [1]
5) is it exported?
declare -x VAR="1"
VAR=1
6) env VAR=2 for one command, env -u to remove it
child sees: [2]
child sees: [unset]
7) login shell reads ~/.profile
non-login: [unset]
login: [vim]
- Prefix assignment.
VAR=1 bash -c ...exportsVARto that one child. The child sees1. - Plain assignment.
VAR=1on its own line creates a shell variable that is not exported. The child gets no copy, so it prints an empty value, while the parent still has1. This is the case that prints nothing. - Prefix assignment with
echo $VAR. The parent shell expands$VARbefore the assignment takes effect for the command, and it was unset, soechoreceives an empty argument. Expansion happens in the parent, not in the command it launches. export. Afterexport VAR=1the child sees1. The child changed its own copy to 99, and the parent still holds 1: the environment is copied downward atexecve(the system call that starts a program) and never flows back up.- Checking.
declare -p VARshows the-xflag (exported), andenvlists only exported variables. - One command, no shell syntax.
env VAR=2 cmdsets a variable forcmd, andenv -u VAR cmdremoves one. These work with any command and in any shell.
Reading the demo
- Login versus non-login shell. A login shell is the first shell of a session, started when you log in (a console or
sshlogin,su -, orbash -l). Bash marks it as a login shell and reads the profile files (/etc/profile, then~/.profileand friends). A non-login interactive shell is one you open after you are already logged in (a new terminal tab, or typingbash); it reads~/.bashrc. A non-interactive shell (bash script.sh,bash -c '...') reads neither. Bash can tell you which one you are in:bash -c 'shopt login_shell'prints a line withlogin_shellandoff, andbash -lc 'shopt login_shell'prints the same line ending inon(executed;shoptpads the two columns with whitespace). - The quoting in test 1.
echo '1) VAR=1 bash -c '"'"'echo $VAR'"'"glues four pieces together: the single-quoted text1) VAR=1 bash -c, a single quote written as"'", the single-quoted textecho $VAR, and another"'". It is only a way to print a label that itself contains single quotes (1) VAR=1 bash -c 'echo $VAR'). The real test is the line after it. - The
HOMEtrick in test 7. Bash finds~/.profilethrough theHOMEvariable. PointingHOMEat a scratch directory (fakehome) lets the demo create a.profilewithout touching the real one. execve. The Linux system call that replaces a process with a new program and hands it its arguments and its environment. Exported variables are copied into that environment at this moment, which is why later changes never flow back.declare -p VAR. Prints how Bash holds the variable;declare -x VAR="1"has the-xflag, meaning exported./proc/PID/environ. Linux shows each running process as a folder under/proc; the fileenvironholds the environment the process started with, separated by NUL bytes. Executed:VAR=1 sleep 5 &thentr '\0' '\n' < /proc/$!/environ | grep '^VAR='printsVAR=1.env_reset. A setting ofsudo(on by default) that starts the command with a cleaned-up environment instead of yours.
Making a value visible to children
Use export VAR=value, or VAR=value; export VAR for older POSIX shells. In Bash, declare -x does the same. Exporting is a one-way flag for that shell and its descendants. It does not affect sibling shells or the parent, and it does not survive logout.
One command versus every login session
Choose by who starts the process; the list is not a sequence to work through. For your own interactive sessions, one export line in ~/.profile is the usual answer. The other entries cover other starters.
- One command:
VAR=value commandorenv VAR=value command. Nothing remains afterwards. - Every login session of one user: put
export VAR=valuein~/.profile. Bash login shells read/etc/profileand then the first of~/.bash_profile,~/.bash_login,~/.profilethat exists. Test 7 above shows the effect: a plainbash -cdid not seeEDITOR, andbash -lc(a login shell) did. Interactive non-login shells, such as a new terminal tab on many desktops, read~/.bashrcinstead, so people often source~/.bashrcfrom the profile or the reverse. - Every user:
/etc/profile.d/*.shfor login shells, or/etc/environmenton Debian and Ubuntu, which is read by the PAM modulepam_env(Pluggable Authentication Modules) and takes plainNAME=valuelines with no shell syntax. - Services and cron jobs: they do not read profile files. For a systemd unit use
Environment=orEnvironmentFile=. A cron job starts with a minimal environment, so set what it needs in the crontab or script.
Pitfalls
- Environment variables are readable by the same user and by root through
/proc/PID/environ, and are inherited by every child, so a secret exported in a login profile is handed to every program that user starts. Prefer passing secrets to the one command that needs them. sudoresets most variables by default (itsenv_resetsetting), so an exported variable can disappear acrosssudounless the policy keeps it.- Variables set inside a
( subshell ), a pipeline stage or a$(...)do not return to the parent.
When would you reach for grep, sed or awk? Give a short example of each: finding lines that match an IP address, rewriting a token in a stream, and extracting and summing a column.
Sample Answer
Direct answer
Reach for grep when the question is "which lines?", for sed when it is "change this text as it streams past", and for awk when the data has columns and you need to select by field, compute, or reformat. They compose: grep narrows, sed rewrites, awk aggregates.
When each one fits
grep(global regular expression print): filter lines by pattern, count them (-c), print only the matching part (-o), invert (-v). It never changes data.sed(stream editor): line-oriented edits, chiefly substitutions/old/new/flags, plus deleting or printing line ranges.-iedits a file in place, and-i.bakkeeps a backup copy first. The attached-suffix form-i.bakworks in both GNU and BSD/macOS sed (checked on macOS), but a bare-iwith no suffix is GNU only: BSD/macOS sed treats the next argument as the suffix, so the portable no-backup spelling there is-i ''(without it macOS sed fails with an error such ascommand c expects \ followed by textorinvalid command code f, depending on the first letter of the file name it mistakes for a script).awk: a small programming language that splits each line into fields ($1,$2, ... withNFas the field count) and has variables, arrays,BEGINandENDblocks (BEGIN { ... }runs once before the first line is read,END { ... }once after the last). An awk program is a list ofpattern { action }pairs: for each line, if the pattern is true, the action runs, and a pair with no pattern runs on every line. Pick it when the logic needs a condition on one column and an action on another, or a running total.
Worked example
One sample file with five columns (client address, method, path, status, bytes), then the three tasks asked for.
In the script, cat > access.log <<'EOF' ... EOF is a here-document: it writes the lines between the two EOF markers into the file access.log (the quotes around 'EOF' stop the shell from expanding anything inside).
cat > access.log <<'EOF'
10.0.0.5 GET /index.html 200 512
10.0.0.7 GET /api/items 500 128
999.1.1.1 GET /odd 404 64
-- rotated --
10.0.0.5 POST /api/items 201 256
EOF
echo '--- grep: lines with an IPv4-looking address'
grep -E '([0-9]{1,3}\.){3}[0-9]{1,3}' access.log
echo '--- grep: strict octets, print only the address'
oct='(25[0-5]|2[0-4][0-9]|1[0-9]{2}|[1-9]?[0-9])'
grep -oE "\b($oct\.){3}$oct\b" access.log
echo '--- sed: rewrite a token in a stream'
sed 's/POST/PUT/' access.log | head -n 5
echo '--- sed: only on lines 1-2, edit in place with backup'
cp access.log copy.log && sed -i.bak '1,2s/GET/HEAD/' copy.log && head -n 2 copy.log && ls copy.log*
echo '--- awk: sum the bytes column (field 5), count rows with 5 fields'
awk 'NF==5 {sum += $5; n++} END {print n " rows, " sum " bytes"}' access.log
echo '--- awk: bytes per client'
awk 'NF==5 {b[$1]+=$5} END {for (ip in b) print ip, b[ip]}' access.log | sort
Output:
--- grep: lines with an IPv4-looking address
10.0.0.5 GET /index.html 200 512
10.0.0.7 GET /api/items 500 128
999.1.1.1 GET /odd 404 64
10.0.0.5 POST /api/items 201 256
--- grep: strict octets, print only the address
10.0.0.5
10.0.0.7
10.0.0.5
--- sed: rewrite a token in a stream
10.0.0.5 GET /index.html 200 512
10.0.0.7 GET /api/items 500 128
999.1.1.1 GET /odd 404 64
-- rotated --
10.0.0.5 PUT /api/items 201 256
--- sed: only on lines 1-2, edit in place with backup
10.0.0.5 HEAD /index.html 200 512
10.0.0.7 HEAD /api/items 500 128
copy.log
copy.log.bak
--- awk: sum the bytes column (field 5), count rows with 5 fields
4 rows, 960 bytes
--- awk: bytes per client
10.0.0.5 768
10.0.0.7 128
999.1.1.1 64
- Finding IP-shaped lines: the loose pattern
([0-9]{1,3}\.){3}[0-9]{1,3}also accepts999.1.1.1. The strict version spells out each octet range (0 to 255) and uses-oto print only the address, and it correctly skips999.1.1.1. - The strict address pattern, piece by piece. An octet is one of the four numbers in a dotted address, and each must be 0 to 255. The variable
octlists four alternatives, any of which may match:25[0-5]is 250 to 255,2[0-4][0-9]is 200 to 249,1[0-9]{2}is 100 to 199 ({2}means exactly two more digits), and[1-9]?[0-9]is 0 to 99 (an optional first digit 1 to 9, then one digit).($oct\.){3}means an octet and a dot, three times, and the final$octis the fourth octet.\bis a word boundary, which stops a match starting or ending in the middle of a longer run of digits. Tested one octet at a time, the pattern accepted 0, 7, 42, 99, 100, 199, 200, 249, 250 and 255 and rejected 256 and 300. - Rewriting a token:
sed 's/POST/PUT/'changes the first match on each line; addgfor every match. The1,2s/GET/HEAD/form limits the edit to lines 1 to 2. - Extract and sum:
awk 'NF==5 {sum += $5; n++} END {...}'reads as patternNF==5(the line has exactly five fields) with action{sum += $5; n++}(add field 5 tosumand count the row), and theENDblock prints the totals once. It skips the malformed-- rotated --line, which has only three fields. Its four rows give 512 + 128 + 64 + 256 = 960 bytes. - Per client:
b[$1] += $5uses an associative array, a variable indexed by text instead of by number, sob["10.0.0.5"]holds the running total for that address (512 + 256 = 768).for (ip in b) print ip, b[ip]visits each key once, in no particular order, which is why the output is piped tosort.
Trade-offs and pitfalls
sedandawkrewrite text but know nothing about quoting: for CSV with quoted commas, use a real CSV parser.- A quick
grepregex for IP addresses validates shape, not meaning. If you need correctness (for example to reject999.1.1.1), use the strict octet pattern or a tool such as Python'sipaddressmodule. sed -iworks by writing a temporary file and renaming it over the original, so readers see either the old or the new content, but the result is a new file (a new inode): hard links are broken, and a process that already has the old file open, such as a logger appending to it, keeps writing to the replaced copy and those lines are lost. Do not runsed -ion a live log; edit a copy and swap it in during a quiet moment.- Use
awkrather thancutwhen fields are separated by runs of spaces, becausecuttreats each single space as a delimiter.
How would you collect the files found by find into a Bash array and then loop over them, even when names contain odd characters? What are the memory and portability trade-offs?
Sample Answer
Direct answer
Have find print NUL-terminated names with -print0 and read them with mapfile -d '' -t (Bash 4.4 or newer) or, on older Bash, a while IFS= read -r -d '' loop. Both read from < <( ... ), a process substitution: Bash runs the command inside <( ... ) and hands its output over as if it were a file, so < <(cmd) feeds that output to the reader while the reader still runs in your current shell. NUL is the one byte that cannot appear in a Unix file name, so spaces, newlines, tabs, leading dashes and glob characters all survive. Then quote every expansion: "${files[@]}". The costs: the whole list sits in memory, passing the whole array to an external command can fail with "Argument list too long", and mapfile -d is not available in the Bash 3.2 that macOS ships.
The collection
The demo directory holds six awkward names: one with a space, one with an embedded newline, one with a tab (in a subdirectory), *.txt (a literal asterisk), -rf.txt (looks like options) and a plain one.
Save this block as setup.sh. Every demo below starts with . ./setup.sh: the dot command runs that file's commands in the current shell, so the six files exist (and the cd into the demo directory sticks) before the demo's own lines run.
# setup.sh
rm -rf /tmp/tree; mkdir -p /tmp/tree/sub; cd /tmp/tree || exit 1
touch -- 'plain.txt' 'two words.txt' $'line\nbreak.txt' '*.txt' '-rf.txt' 'sub/tab'$'\t''bed.txt'
. ./setup.sh
mapfile -d '' -t files < <(find . -type f -name '*.txt' -print0 | sort -z)
printf '%d files\n' "${#files[@]}"
for f in "${files[@]}"; do
printf '[%q]\n' "$f"
done
# find's own exit status is lost inside <( ), so collect it explicitly
mapfile -d '' -t none < <(find /nonexistent -print0)
wait "$!"; echo "find status: $?"
# the naive version: word splitting and globbing mangle the names
naive=($(find . -type f -name '*.txt'))
printf 'naive array: %d elements\n' "${#naive[@]}"
6 files
[./\*.txt]
[./-rf.txt]
[$'./line\nbreak.txt']
[./plain.txt]
[$'./sub/tab\tbed.txt']
[./two\ words.txt]
find: '/nonexistent': No such file or directory
find status: 1
naive array: 13 elements
The quotes around the path are what GNU findutils 4.9 and newer print; BusyBox find prints the same message without them.
The %q format prints each name in a form that shows the odd characters ($'./line\nbreak.txt' is the name with a real newline). The array has all 6 files, each as exactly one element. Three details matter:
-tstrips the trailing delimiter from each element;-d ''sets the delimiter to NUL.sort -z(GNU) orders the NUL-separated list, becausefindreturns names in directory order, not sorted order.- Inside
< <( ... ), the exit status offindis invisible toset -e. Bash 4.4 or newer lets youwait "$!"for the process substitution, which returnedfind's status 1 for the missing directory. Without that check, a failedfindsilently gives an empty array and the loop does nothing.
The naive form naive=($(find . -name '*.txt')) produced 13 elements from the same 6 files: the unquoted expansion splits on every space, tab and newline, then runs glob expansion on each piece, so *.txt expands into whatever matches in the current directory. Never build file lists that way.
Tracing the 13, name by name (counts measured by splitting each name the same unquoted way):
5 ./\*.txt the glob ./*.txt matches 5 files in the current directory
1 ./-rf.txt
2 $'./line\nbreak.txt' split at the newline
1 ./plain.txt
2 $'./sub/tab\tbed.txt' split at the tab
2 ./two\ words.txt split at the space
total 13
The shell expands the glob ./*.txt into its matches and never splits those matches again, so each of the 5 matched names (including the ones with a space or a newline) counts as one element. 5 + 1 + 2 + 1 + 2 + 2 = 13.
Reading the pieces of the collection line:
< <( ... )versus a pipe.find ... | while readruns the loop in a subshell (a separate child copy of the shell), and a child's variable changes disappear when it ends. Executed: counting two NUL-separated items gavecount=0after the pipe loop andcount=2after the same loop fed with< <(printf 'a\0b\0').echo <(true)prints/dev/fd/63, which shows that<( ... )really is replaced by a file-like name.%qis aprintfformat that prints its argument in a form the shell could read back, so$'a\nb'shows a newline as\nandc dprints asc\ d. That is why the odd names are visible in the transcript.sort -zis the NUL-aware sort:-zmakes it treat NUL, not newline, as the end of each record.wait "$!":$!holds the process id of the most recently started background process, and a process substitution counts as one in Bash 4.4 or newer.wait PIDpauses until that process ends and returns its exit status, which is howfind's failure (status 1 for the missing directory) becomes visible.
Looping is then just:
for f in "${files[@]}"; do
printf 'processing %s\n' "$f" # every use of $f quoted; use -- before it for commands that take options
done
Portability
mapfile arrived in Bash 4.0, but the -d option only in 4.4. Executed against official images:
probe.sh: line 1: mapfile: -d: invalid option
mapfile: usage: mapfile [-n count] [-O origin] [-s count] [-t] [-u fd] [-C callback] [-c quantum] [array]
mapfile -d works on 4.4.23(1)-release
probe.sh: line 1: mapfile: command not found
The first line is Bash 4.3 (option rejected), the second Bash 4.4 (works), the third Bash 3.2, the version macOS ships (mapfile does not exist at all). The fallback works on all three versions and gave 6 files each time:
. ./setup.sh
files=()
while IFS= read -r -d '' f; do
files+=("$f")
done < <(find . -type f -name '*.txt' -print0)
printf '%s: %d files\n' "$BASH_VERSION" "${#files[@]}"
3.2.57(1)-release: 6 files
4.3.48(1)-release: 6 files
5.3.20(1)-release: 6 files
Other options and their limits:
| Approach | Handles odd names | Memory | Portability |
|---|---|---|---|
mapfile -d '' -t a < <(find ... -print0) | yes | whole list | Bash 4.4+, find -print0 (GNU, BSD, BusyBox) |
while IFS= read -r -d '' loop, optionally appending to an array | yes | whole list if you append, constant if you act inside the loop | Bash 3.2+ |
find ... -exec cmd {} + | yes | constant (find batches the calls) | POSIX |
find ... -print0 | xargs -0 cmd | yes | constant (xargs batches) | -print0 and -0 were long GNU/BSD/BusyBox extensions and were added to POSIX in the 2024 edition |
files=(**/*.txt) with shopt -s globstar nullglob | yes | whole list | Bash 4.0+, no find tests like -mtime |
Memory trade-off
The array stores every path in the shell's own memory: roughly the sum of all path lengths plus a small overhead per element. For thousands of files that is nothing; for tens of millions it is real. If you only need to act on each file once, do it inside the read -d '' loop, which holds one name at a time.
The second cost is the argument limit. Expanding "${files[@]}" into an external command puts every name on one command line, and the kernel caps that total (ARG_MAX, which on Linux scales with the stack limit). This run pins the stack limit to make it reproducible: 3000 names of 200 characters is about 600 KB, over the 256 KiB that a 1 MiB stack allows.
ulimit -s 1024 # pins ARG_MAX at a quarter of the stack limit, 256 KiB
mkdir -p /tmp/many && cd /tmp/many || exit 1
long=$(printf 'x%.0s' {1..200}) # 200-character names, 3000 of them: about 600 KB in total
for i in $(seq 3000); do : > "$long$i"; done
files=(*)
echo "array holds ${#files[@]} names"
/bin/echo "${files[@]}" > /dev/null; echo "external command, all names as arguments: status $?"
n=0; for f in "${files[@]}"; do n=$((n + 1)); done
echo "builtin loop over the same array: $n iterations"
array holds 3000 names
argmax.sh: line 7: /bin/echo: Argument list too long
external command, all names as arguments: status 126
builtin loop over the same array: 3000 iterations
The external /bin/echo fails with "Argument list too long" (status 126), but the Bash for loop over the same array works, because loops and builtins never go through exec. The fixes are to batch (xargs -0 splits the list into several invocations on its own, find -exec ... + does too) or to loop in the shell.
Pitfalls
- Reading a pipeline into a loop (
find ... | while read) runs the loop in a subshell, so variables set inside vanish; the process substitution form keeps the loop in the main shell. readwithoutIFS=strips leading and trailing whitespace; without-rit eats backslashes.- A file created or removed while
findruns can appear in the list and be gone by the time you use it; handle the "no such file" case at the point of use.
Schedule a nightly ETL job to run at 02:30 with cron. Provide the crontab line, explain its fields, and describe why a script that works in your terminal can fail under cron and how you would make failures visible and prevent overlapping runs.
Sample Answer
Direct answer. The crontab line is 30 2 * * * /opt/etl/run_etl.sh >> /var/log/etl/nightly.log 2>&1. The five fields are minute, hour, day of month, month and day of week, so 30 2 * * * means minute 30 of hour 2 (02:30, 24-hour clock) every day. A script that works in your terminal can fail under cron (the daemon that runs scheduled jobs) because cron starts it with a different environment: a minimal PATH, no profile files, a different working directory and /bin/sh as the shell. Make failures visible by logging with timestamps, exiting non-zero on failure, and alerting when the job did not succeed. Prevent overlapping runs with a lock, flock -n (flock takes a lock on an open file, -n means fail at once instead of waiting), so a run that takes longer than expected is not started twice.
The fields
| Field | Allowed | In 30 2 * * * |
|---|---|---|
| Minute | 0 to 59 | 30 |
| Hour | 0 to 23 | 2 |
| Day of month | 1 to 31 | * (every day) |
| Month | 1 to 12 | * |
| Day of week | 0 to 7 (0 and 7 are Sunday) | * |
After the five time fields comes the command, run by /bin/sh unless the crontab sets SHELL=. ETL (extract, transform, load) is the pipeline that copies data from sources into a warehouse. Cron times use the system clock's time zone: run servers in UTC, because on days when clocks change the 02:30 slot may not exist or may occur twice, and cron implementations differ in what they do about it.
Why it works in a terminal but not under cron
In a clean Ubuntu 24.04 container with cron installed, a probe job scheduled every minute logged the environment cron really provides:
SHELL=/bin/sh PATH=/usr/bin:/bin HOME=/root PWD=/root
- PATH has no
/usr/local/binor/sbin, so tools installed there (and anything added by your profile,nvm, a virtualenv or~/.local/bin) are "command not found" (status 127). - No profile.
.bashrcand.profileare not read, so exported credentials and variables are absent. - Working directory is the home directory, so relative paths break. Use absolute paths or
cdfirst. - Shell is
/bin/sh, so bashisms ([[ ]], arrays,pipefail) fail unless the script has a#!/usr/bin/env bashline (when the script is run directly rather than as an argument tosh) or the crontab setsSHELL=/bin/bash. - Percent signs are special in a crontab. An unescaped
%ends the command and the rest becomes standard input. Writingecho "stamp $(date +%F)"in a crontab produced no log file at all in that test, because the shell received a truncated, invalid line.$(date +\%F)worked and loggedstamp 2026-10-06. Put the command in a script and call that. - No terminal and no interactive prompts. Output goes to mail (if configured) or is lost, so redirect it.
A wrapper that fixes the environment, avoids overlap and reports failure
#!/usr/bin/env bash
# /opt/etl/run_etl.sh : wrapper cron calls. Makes the environment explicit.
set -euo pipefail
export PATH=/usr/local/bin:/usr/bin:/bin # do not rely on cron's PATH
cd /opt/etl # do not rely on cron's cwd
log() { printf '%s nightly-etl: %s\n' "$(date '+%F %T')" "$*"; }
# One run at a time. Descriptor 9 holds the lock until this process exits.
exec 9> /var/lock/nightly-etl.lock
if ! flock -n 9; then
log "previous run still holds the lock, skipping this one"
exit 75 # EX_TEMPFAIL: not an ETL failure
fi
on_error() { log "FAILED at line $1 (exit $2)"; exit "$2"; }
trap 'on_error "$LINENO" "$?"' ERR
log "start"
etl-tool --extract # a command that lives in /usr/local/bin
sleep "${ETL_SECONDS:-1}" # stands in for the real transform and load
log "done"
touch /var/lib/etl-last-success
Reading the wrapper line by line:
-
exec 9> /var/lock/nightly-etl.lockopens the lock file for writing and keeps it open as file descriptor 9. A file descriptor is just a small number a process uses to refer to a file it has open (0, 1 and 2 are standard input, output and error already, so 9 is a free number).execwith only a redirection changes the current shell instead of starting a program, so the file stays open until the script exits. Opening the file does not lock it; the next line does. -
flock -n 9asks the kernel for an exclusive lock on whatever descriptor 9 points to. If another process already holds that lock the command fails immediately, and theif !branch logs and exits. Measured in a container (readlink /proc/$$/fd/9shows what descriptor 9 points to):/tmp/demo.lock first: got the lock second: busy (flock exit 1) third: got the lockThe first shell opened the file on descriptor 9 and locked it, a second process that tried to lock the same file was refused, and after descriptor 9 was closed (
exec 9>&-) a third process got the lock. When a script exits, the kernel closes its descriptors, and the lock is released when the last open copy of descriptor 9 is closed. Children the wrapper starts inherit descriptor 9, so a child that outlives a killed wrapper keeps the lock held (executed: afterkill -9of a shell holding the lock, a still-runningsleepchild madeflock -non the same file fail, while a child started with9>&-did not). That is usually what you want, because the job is still running; start any helper that must not hold the lock with9>&-. -
exit 75is the status for a skipped run. 75 isEX_TEMPFAIL("temporary failure, try later") from the BSDsysexits.hlist of conventional exit codes, so a monitor can tell a skip from a real failure (any other non-zero status). -
trap 'on_error "$LINENO" "$?"' ERRregisters a handler for commands that fail whileset -eis on. The single quotes matter:$LINENO(the line of the failing command) and$?(that command's exit status) are only filled in when the trap fires, not when it is registered. Measured on this toy script (saved as a file, with the shebang as line 1), which registers the same trap and then runsfalse, a command that always fails:bash#!/usr/bin/env bash set -euo pipefail on_error() { echo "FAILED at line $1 (exit $2)"; exit "$2"; } trap 'on_error "$LINENO" "$?"' ERR echo start false echo not reachedstart FAILED at line 6 (exit 1) wrapper exit status: 1(the line number counts from the top of that toy file, so it differs from the 20 reported for the real wrapper below.)
-
${ETL_SECONDS:-1}means "the value ofETL_SECONDSif it is set and not empty, otherwise1":sleepwaits 1 second normally, and the lock test below sets it to 75.
The crontab entry calls the wrapper and appends both streams to a log:
30 2 * * * /opt/etl/run_etl.sh >> /var/log/etl/nightly.log 2>&1
To check the lock, the wrapper was run from real cron every minute with the job taking 75 seconds (ETL_SECONDS=75):
2026-10-06 06:06:01 nightly-etl: start
extracted
2026-10-06 06:07:01 nightly-etl: previous run still holds the lock, skipping this one
2026-10-06 06:07:16 nightly-etl: done
The 06:07 run found the lock held and skipped itself instead of starting a second copy. The kernel releases the lock once no process holds the descriptor open any more, even if the holder was killed, so there is no stale lock file to clean up. With the extractor replaced by a stand-in that prints an error and exits 3, the wrapper logged FAILED at line 20 (exit 3), exited with status 3 and did not create the success marker.
Making failures visible
- Log with timestamps and keep stderr in the log (
2>&1), as above. - Exit non-zero on failure so whatever supervises the job can see it. The wrapper uses
set -euo pipefailand an ERR trap. - Alert on absence as well as failure. If cron itself is down or the host is off, no error is ever logged. Record a success marker (the wrapper touches
/var/lib/etl-last-success) and run a separate check that alerts when it is older than your threshold (for example 26 hours for a daily job), or ping an external heartbeat service at the end of a successful run. - Mail:
MAILTO=ops@example.comat the top of the crontab sends any output the job produces to that address, which needs a working mail transfer agent (MTA, the program on the host that actually delivers email, such as Postfix) on the host.
Trade-offs and pitfalls
- Skipping versus queuing.
flock -nskips the overlapping run. If the data must be processed, run the next job withflockwithout-nso it waits, but then a stuck run blocks every night after it. A timeout (flock -w 600) is a middle ground. - A skipped run is not a failure, so the wrapper exits 75 (the BSD
EX_TEMPFAILconvention) to keep it distinguishable from a real error. - Idempotency (running twice gives the same result as running once). Cron can run the job twice (manual rerun, host time change), so the ETL should be safe to repeat, for example by loading by date partition.
- systemd timers are an alternative: a service unit is not started again while it is still running, output goes to the journal, and
OnFailure=can trigger an alert. Cron is fine for one nightly job when it has the wrapper above. - Cron does not retry. A failure at 02:30 is not retried until the next day unless you build it in.
Unlock Full Question Bank
Get access to all 47 Shell Scripting and Automation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.