Shell Scripting and Automation Questions
Shell-language craft for operations work: Bash and POSIX sh scripting, pipes and redirection, quoting and word splitting, exit codes and set -euo pipefail (including failures inside pipelines), traps and cleanup, argument parsing with getopts, functions and arrays, parameter expansion, here-documents, and text processing with grep, sed, awk and jq, including streaming pipelines over very large log, CSV and JSON Lines files. Also covers writing robust, idempotent, portable scripts (locking, atomic file updates, retries with backoff, safe temp files, GNU versus BSD differences), background jobs and signals, bounded parallelism and SSH fan-out, cron-driven jobs such as log rotation, backups, atomic deploys and health-check watchdogs, secure scripting (eval injection, secrets, path traversal), and debugging, testing and linting shell with ShellCheck and bats. Boundary: general automation design and Python or Go tooling, Linux host administration tasks, observability and alerting design, security detection engineering, and generic algorithmic coding problems are covered elsewhere.
Write a script that waits for a local service to accept connections on a port, retrying with growing delays and a per-attempt timeout, and fails with a non-zero status if it never comes up. How would you use it in a container entrypoint or CI step?
Sample Answer
Direct answer
Open a TCP connection in a loop using bash's built-in /dev/tcp, wrap each attempt in timeout so a connect that never gets answered cannot hang the loop, sleep 1, 2, 4, 8 seconds between attempts (capped at 8), and stop when a total deadline passes. Exit 0 as soon as one connection succeeds and exit 1 if the deadline arrives first, so a container entrypoint or CI step that runs under set -e fails fast with a clear status. Open port is not the same as ready, so the usual pairing is this script followed by an HTTP check on the service's health URL.
The script
#!/usr/bin/env bash
# wait-for-port.sh HOST PORT [TOTAL_SECONDS] [PER_ATTEMPT_SECONDS]
# Exit 0 when a TCP connection succeeds, 1 if TOTAL_SECONDS pass first, 2 on bad usage.
set -u
host=${1:-} port=${2:-}
total=${3:-60} per_attempt=${4:-3}
usage() { echo "usage: ${0##*/} HOST PORT [TOTAL_SECONDS] [PER_ATTEMPT_SECONDS]" >&2; exit 2; }
# At most 5 digits for the port and 6 for the budget (no overflow), and a per-attempt limit of at least 1.
[[ -n $host && $port =~ ^[0-9]{1,5}$ && $total =~ ^[0-9]{1,6}$ && $per_attempt =~ ^[1-9][0-9]{0,3}$ ]] || usage
# 10# forces base 10: a value such as 08 would otherwise be read as a bad octal number.
port=$((10#$port)) total=$((10#$total))
(( port >= 1 && port <= 65535 )) || usage
# One connection attempt in its own bash, killed by timeout if the SYN is never answered.
# The single quotes are deliberate: $0 and $1 are that inner bash's arguments.
try_connect() {
# shellcheck disable=SC2016
timeout "$1" bash -c 'exec 3<>"/dev/tcp/$0/$1"' "$host" "$port" 2>/dev/null
}
start=$SECONDS delay=1 attempt=0
while :; do
remaining=$(( total - (SECONDS - start) ))
if (( attempt > 0 && remaining <= 0 )); then
echo "$host:$port still down after $(( SECONDS - start ))s ($attempt attempts)" >&2
exit 1
fi
attempt=$(( attempt + 1 ))
limit=$per_attempt
(( remaining < limit )) && limit=$(( remaining > 0 ? remaining : 1 )) # never run past the budget
if try_connect "$limit"; then
echo "$host:$port is up after $(( SECONDS - start ))s (attempt $attempt)"
exit 0
fi
remaining=$(( total - (SECONDS - start) ))
(( delay > remaining )) && delay=$remaining
if (( delay > 0 )); then
echo "attempt $attempt failed, retrying in ${delay}s" >&2
sleep "$delay"
fi
delay=$(( delay * 2 > 8 ? 8 : delay * 2 ))
done
Why it is built this way:
exec 3<>"/dev/tcp/$0/$1"is a bash feature: opening that path makes bash callconnect(); if the connection is refused, the redirection fails and the shell exits non-zero. Noncorcurlhas to exist in the image, only bash. (It does not work inshas implemented by dash or Alpine's ash, so the container needs bash.)- Each attempt runs in its own
bash -cchild undertimeout. A refused connection fails in milliseconds, but a host that silently drops packets (a firewall, a not-yet-attached network) leavesconnect()waiting far longer than a second, which is why the per-attempt limit exists. The child exits right after connecting, which closes the socket again. - The delay doubles but is capped at 8 seconds, and is shortened so the last sleep never runs past the deadline. The per-attempt limit is also shortened to the time remaining, so the whole script respects
TOTAL_SECONDSinstead of overshooting it by one attempt. - Exit statuses are distinct: 0 up, 1 never came up, 2 bad arguments.
set -escripts and CI runners stop on 1 and 2 without special handling. The arguments are checked before any attempt, because the obvious check (^[0-9]+$alone) lets bad values through: port 99999 or 0 would retry for the whole budget and then report "still down" with status 1 instead of 2; a budget such as08makes bash stop withvalue too great for base(it reads a leading zero as octal); and a 20-digit budget overflows the arithmetic, so the loop never reaches its deadline. Hence the digit limits, the10#conversion and the 1 to 65535 range check. - Name resolution failures (a service name that does not exist yet) look like any other failed attempt and are retried, which is what you want for a container that is still being created.
The less obvious pieces, explained
- SYN and the handshake. Before two programs exchange data, TCP sets up a connection: the client sends a SYN ("synchronise") packet, the server answers, the client confirms. If the server's machine is up but nothing listens on the port, it answers at once with a refusal, so
connect()fails in milliseconds (executed: refused in about 0.001 s). If the SYN is simply dropped, nothing comes back, and the client keeps waiting and re-sending, soconnect()can block for a long time. Only a time limit protects the loop from that case. - File descriptor,
exec 3<>. A file descriptor is the small number a process uses for an open file or connection (0 standard input, 1 standard output, 2 standard error).exec 3<>"/dev/tcp/HOST/PORT"opens the connection as descriptor 3 for reading and writing (<>). With only a redirection after it,execchanges the current shell's descriptors instead of starting a program. ${0##*/}. Removes everything up to the last/from$0, the script's own path, so the usage message showswait-for-port.shand not the full path.- The
bash -c '...' "$host" "$port"trick. The words after the command string become the inner shell's$0,$1and so on. Because the string is in single quotes, the outer shell does not expand$0and$1; the inner bash does. Executed:bash -c 'echo "zero=[$0] one=[$1]"' db.example 5432printszero=[db.example] one=[5432]. - Entrypoint
set -euo pipefail.-eends the script at the first command that fails, so ifwait-for-port.shexits 1 the application is never started (afalseunderset -eended a test script with status 1 before its next line ran).-ustops on unset variables andpipefailmakes a pipeline fail if any stage fails. exec "$@"and PID 1."$@"is the list of arguments the container was started with (theCMD).execreplaces the shell with that command instead of running it as a child, so the application keeps the shell's process ID: a test printedpid 17before and after anexec. The first process in a container has process ID 1 (PID 1), and the container runtime delivers the stop signal to it. Withexecthat process is the application; without it, a shell would sit in front and receive the signal in the application's place.- Backlog and the accept queue in
hang.py. A listening socket has a queue of connections that the kernel has completed but the program has not yet taken withaccept(). The number given tolisten()is the backlog, the queue's size.hang.pynever callsaccept(), so the queue fills. Linux lets a backlog of 0 hold one waiting connection, whichss -ltnshowed asRecv-Q 1; the three connection attempts thathang.pymakes fill it. A further SYN finds the queue full and is dropped without a reply, which is the silent case. Executed against that listener:
$ ss -ltn 'sport = :7003'
State Recv-Q Send-Q Local Address:Port Peer Address:PortProcess
LISTEN 1 0 127.0.0.1:7003 0.0.0.0:*
$ timeout 1 bash -c 'exec 3<>"/dev/tcp/127.0.0.1/7003"'; echo "status=$?"
status=124
124 is the status timeout uses when it had to stop the command, and the attempt took 1.002 s, exactly the limit.
Using it in a container entrypoint and in CI
An entrypoint that waits for its dependency, then replaces itself with the real process so the application becomes the container's first process (PID 1) and receives stop signals directly:
#!/usr/bin/env bash
set -euo pipefail
./wait-for-port.sh "${DB_HOST:-db}" "${DB_PORT:-5432}" "${WAIT_TOTAL:-60}"
exec "$@"
In a Dockerfile this becomes WORKDIR /app, COPY wait-for-port.sh entrypoint.sh /app/, ENTRYPOINT ["/app/entrypoint.sh"] and CMD ["python", "app.py"]. In a CI job it is one step before the tests, with a diagnostic if it fails: ./wait-for-port.sh localhost 8080 60 || { docker compose logs web; exit 1; }.
Test run
The driver starts a web server after 4 seconds, then exercises a late start, a service that never appears, bad arguments, a listener that never answers connections, and the entrypoint in both outcomes. hang.py creates the unanswered listener: it listens with a backlog of zero and opens three connections it never accepts, so further connection attempts stall instead of being refused.
hang.py:
import socket,time
s=socket.socket(); s.bind(("127.0.0.1",7003)); s.listen(0)
c=[]
for _ in range(3):
k=socket.socket(); k.setblocking(False)
try: k.connect(("127.0.0.1",7003))
except BlockingIOError: pass
c.append(k)
time.sleep(30)
run.sh:
cd /w
shellcheck -s bash wait-for-port.sh entrypoint.sh && echo "shellcheck clean"
echo "== service starts after 4s"
( sleep 4; python3 -m http.server 7001 --bind 127.0.0.1 >/dev/null 2>&1 ) &
./wait-for-port.sh 127.0.0.1 7001 20; echo "exit: $?"
echo "== never comes up, 5s budget"
./wait-for-port.sh 127.0.0.1 7002 5; echo "exit: $?"
echo "== bad usage"
./wait-for-port.sh 127.0.0.1 abc; echo "exit: $?"
echo "== listener whose accept queue is full: SYN dropped, per-attempt timeout fires"
python3 hang.py & sleep 1
./wait-for-port.sh 127.0.0.1 7003 4 1; echo "exit: $?"
echo "== as an entrypoint"
DB_HOST=127.0.0.1 DB_PORT=7001 ./entrypoint.sh echo "app started with args: a b"
DB_HOST=127.0.0.1 DB_PORT=7002 WAIT_TOTAL=2 ./entrypoint.sh echo "app started"; echo "entrypoint exit: $?"
Result on bash 5.2 (Ubuntu 24.04 container):
shellcheck clean
== service starts after 4s
attempt 1 failed, retrying in 1s
attempt 2 failed, retrying in 2s
attempt 3 failed, retrying in 4s
127.0.0.1:7001 is up after 7s (attempt 4)
exit: 0
== never comes up, 5s budget
attempt 1 failed, retrying in 1s
attempt 2 failed, retrying in 2s
attempt 3 failed, retrying in 2s
127.0.0.1:7002 still down after 5s (3 attempts)
exit: 1
== bad usage
usage: wait-for-port.sh HOST PORT [TOTAL_SECONDS] [PER_ATTEMPT_SECONDS]
exit: 2
== listener whose accept queue is full: SYN dropped, per-attempt timeout fires
attempt 1 failed, retrying in 1s
attempt 2 failed, retrying in 1s
127.0.0.1:7003 still down after 4s (2 attempts)
exit: 1
== as an entrypoint
127.0.0.1:7001 is up after 0s (attempt 1)
app started with args: a b
attempt 1 failed, retrying in 1s
attempt 2 failed, retrying in 1s
127.0.0.1:7002 still down after 2s (2 attempts)
entrypoint exit: 1
The late start shows the cost of growing delays: the service came up at about 4 seconds, but attempts ran at 0, 1, 3 and 7 seconds, so the script noticed at 7. The hanging listener shows the timeout doing its job: each attempt consumed its 1 second limit, and the script still returned at 4 seconds, its total budget, with status 1.
Trade-offs and pitfalls
- Port open is not ready. Many servers listen before migrations finish or caches warm. Follow the wait with
curl -fsS http://host:8080/health(using the retry pattern with a time limit) when the port alone is not enough. - A wait script hides a design problem. The durable fix is for the application to retry its own dependency connection with backoff, because it will meet the same outage again in production. The wait script is the right tool for CI, for one-shot jobs and for images you cannot change.
- Orchestrators already model this. In Docker Compose use a
healthcheckwithdepends_on: condition: service_healthy; in Kubernetes use readiness probes and init containers. Use the script where those are not available. - Pick the budget deliberately. 60 seconds is generous for a local database and far too short for a cold cloud database; an entrypoint that gives up too early causes a restart loop, one that waits too long hides an outage.
Write a script that makes sure a particular cron job line exists in the current user's crontab without creating duplicates, and that is safe if two instances run at once. Explain how it behaves on repeated runs.
Sample Answer
Direct answer
Read the current crontab, add the line only if an identical line is not already there, and install the result with crontab file, all while holding an exclusive lock (flock). The grep (grep -qxF, whole line, fixed string) makes repeated runs idempotent: the first run prints "added", every later run prints "already present" and changes nothing. The lock is needed because crontab has no merge: read, modify and write are three separate steps, and two instances that interleave them will silently lose one of the edits.
The script
#!/usr/bin/env bash
# Usage: ensure-cron.sh '<complete crontab line>'
set -euo pipefail
line=${1:?usage: ensure-cron.sh '<crontab line>'}
lock=/tmp/ensure-cron-$(id -u).lock
exec 9> "$lock"
flock 9 # serialise the whole read-modify-write
current=$(mktemp)
trap 'rm -f "$current"' EXIT
if err=$(crontab -l 2>&1 > "$current"); then
: # an existing crontab was read into $current
elif [[ $err == "no crontab for "* ]]; then
: > "$current" # a first run for this user: start from empty
else
echo "cannot read the crontab, refusing to overwrite it: $err" >&2
exit 1
fi
if grep -qxF -- "$line" "$current"; then
echo "already present"
exit 0
fi
cp -- "$current" "$HOME/.crontab.bak"
if [[ -s $current && -n $(tail -c 1 "$current") ]]; then echo >> "$current"; fi
printf '%s\n' "$line" >> "$current"
crontab "$current" # installs the file in one step; rejects bad syntax
echo "added"
Points that are easy to get wrong:
- "No crontab yet" is not an error, but every other failure is.
crontab -lexits 1 and printsno crontab for <user>when the user has none. A script that doescrontab -l 2>/dev/null | ...treats a permission error the same way and then installs a crontab containing only the new line, wiping the real one. This script captures stderr, starts from empty only for that specific message, and aborts on anything else. grep -qxF -- "$line":-xmatches the whole line (so a commented-out copy or a longer line does not count),-Ftreats the text as a fixed string (the*and.in a schedule are not regex),--protects a line that starts with a dash.- Backup before write: the previous contents go to
~/.crontab.bak. - A trailing-newline guard so the new line does not get glued to a last line that had none.
- Lock:
exec 9> lockfileopens a descriptor on the lock file andflock 9blocks until it holds the exclusive lock, released automatically when the script exits.flockis part of util-linux (Linux); on systems without it,mkdir lockdirworks as a portable lock.
Reading the trickiest lines
The schedule fields. A crontab line is five time fields followed by the command: minute, hour, day of month, month, day of week (0 or 7 is Sunday). A * means every value, and */5 means every fifth value. So */5 * * * * runs every fifth minute (at :00, :05, :10 and so on), 0 3 * * * runs at 03:00 every day, and 15 4 * * 0 runs at 04:15 every Sunday.
err=$(crontab -l 2>&1 > "$current"). Redirections are applied left to right, and inside $( ... ) the starting stdout is the pipe that feeds the capture. 2>&1 first points stderr at that pipe, then > "$current" moves stdout to the file. The capture therefore receives only the error text, and the file receives only the real crontab. Swapping the two redirections would put both streams in the file and leave err empty. Executed with a stand-in command that prints one line to each stream:
captured in err: [no crontab for root]
written to file: [the listing]
--- the reverse order
captured in err: []
written to file: [the listing|no crontab for root|]
[[ -s $current && -n $(tail -c 1 "$current") ]]. -s is true when the file exists and is not empty. tail -c 1 prints the file's last byte, and $( ... ) strips trailing newlines from what it captures, so a file that ends with a newline gives an empty string (-n is false: nothing to fix) while a file whose last line has no newline gives that line's last character (-n is true: add a newline first). Executed on two small files:
no-nl.txt: last byte is [b], so add a newline
with-nl.txt: last byte is a newline (the substitution strips it, leaving empty), nothing to add
Without the guard, appending to the file that lacks a final newline glues the new line onto the old one (bNEW).
trap 'rm -f "$current"' EXIT runs that command whenever the script ends for any reason (normal exit, exit 1, or set -e stopping it), so the temporary copy of the crontab is never left behind.
Executed demonstration
The container is Debian with the cron package installed; the user is root. The first two blocks use a naive one-liner (unsafe.sh) that sleeps between reading and writing to make the race deterministic:
#!/usr/bin/env bash
# naive version: read, pause (the other instance runs here), then write back
{ crontab -l 2>/dev/null; sleep 1; echo "$1"; } | crontab -
exec 2>&1
set -u
chmod +x ensure-cron.sh unsafe.sh
A='*/5 * * * * /usr/local/bin/backup.sh >> /var/log/backup.log 2>&1'
B='0 3 * * * /usr/local/bin/rotate.sh'
C='15 4 * * 0 /usr/local/bin/report.sh'
echo '--- naive append, same line twice'
./unsafe.sh "$A"; ./unsafe.sh "$A"; crontab -l
crontab -r
echo '--- naive append, two different lines at once'
./unsafe.sh "$A" & ./unsafe.sh "$B" & wait
crontab -l
crontab -r
echo '--- ensure-cron.sh on a user with no crontab, then again'
./ensure-cron.sh "$A"; ./ensure-cron.sh "$A"; crontab -l
echo '--- ten instances at once, the same new line'
for i in 1 2 3 4 5 6 7 8 9 10; do ./ensure-cron.sh "$B" & done | sort | uniq -c
wait
crontab -l
echo '--- a different line and a repeat, at once'
{ ./ensure-cron.sh "$C" & ./ensure-cron.sh "$B" & wait; } | sort | uniq -c
crontab -l | wc -l | sed 's/^/lines in crontab: /'
echo '--- cron only loosely checks the line it installs'
./ensure-cron.sh '61 * * * * /bin/true'; echo "exit $?"
--- naive append, same line twice
*/5 * * * * /usr/local/bin/backup.sh >> /var/log/backup.log 2>&1
*/5 * * * * /usr/local/bin/backup.sh >> /var/log/backup.log 2>&1
--- naive append, two different lines at once
0 3 * * * /usr/local/bin/rotate.sh
--- ensure-cron.sh on a user with no crontab, then again
added
already present
*/5 * * * * /usr/local/bin/backup.sh >> /var/log/backup.log 2>&1
--- ten instances at once, the same new line
1 added
9 already present
*/5 * * * * /usr/local/bin/backup.sh >> /var/log/backup.log 2>&1
0 3 * * * /usr/local/bin/rotate.sh
--- a different line and a repeat, at once
1 added
1 already present
lines in crontab: 3
--- cron only loosely checks the line it installs
added
exit 0
What this shows:
- The naive append is not idempotent: running it twice gave two identical lines.
- Two naive instances adding different lines at once: only
rotate.shsurvived. Both read the same starting crontab, and the second write replaced the first. That is the lost-update race. ensure-cron.sh: "added" then "already present"; the file holds one line.- Ten instances at once with the same new line: exactly one printed "added" and nine "already present", and the crontab has one copy of each line.
- A different line racing a repeat: one added, one already present, three lines in total (
backup,rotate,report).
Repeated runs are therefore safe to put in a provisioning script, a cloud-init step or a config-management hook that executes on every boot.
Limits and pitfalls
- Exact-line matching means "same schedule and command". If you later change the schedule, the script adds a second job instead of replacing the old one. For managed jobs, tag the line with a comment marker (
# managed:backup), drop lines carrying that marker, and then append the new one. - Cron's own syntax check is loose. The last block shows Debian's
crontabaccepted61 * * * * /bin/truewithout complaint (exit 0), so validate schedules yourself before installing if the input is not trusted. - Percent signs in a crontab command mean "newline" unless escaped as
\%, which matters fordate +%Fin a command. - Lock files in
/tmpare per-host; if the crontab is managed from several machines (a shared home directory), a host-local lock does not serialise them. - For system-wide jobs prefer a file in
/etc/cron.d/written with the temp-file-and-rename pattern: no read-modify-write of a shared file at all.
Design the skeleton of a Bash command-line tool with start, stop and status subcommands and a global verbose flag. It should print helpful usage on invalid input and return exit codes that callers can script against.
Sample Answer
Direct answer
Structure the tool as: constants for exit codes at the top, small helpers (usage, die, usage_error), one function per subcommand (cmd_start, cmd_stop, cmd_status), a loop that parses global options (before the subcommand), and a case that dispatches. Invalid input prints a one-line error plus the usage text to stderr and exits 2; --help prints usage to stdout and exits 0. Exit codes are a contract callers script against, so define them once and document them: 0 success, 1 operational failure, 2 usage error, 3 "not running" from status (the convention used by Linux init-script status actions, so monitoring tools that expect it understand it).
The skeleton
#!/usr/bin/env bash
# svc: start, stop and status for one background process.
# exit codes: 0 ok | 1 operational failure | 2 usage error | 3 status: not running
set -u
readonly PROG=${0##*/}
readonly EX_OK=0 EX_FAIL=1 EX_USAGE=2 EX_NOT_RUNNING=3
STATE_DIR=${SVC_STATE_DIR:-/tmp/svc}
PIDFILE=$STATE_DIR/svc.pid
STOP_TIMEOUT=${SVC_STOP_TIMEOUT:-5} # seconds to wait after SIGTERM before SIGKILL
read -r -a SVC_CMD <<< "${SVC_CMD:-sleep 300}" # the managed command, split once into an array
verbose=0
usage() {
cat <<USAGE
Usage: $PROG [-v|--verbose] <command>
Commands:
start launch the service (exit 0 if it is already running)
stop stop the service (exit 0 if it is already stopped)
status exit 0 if running, 3 if not
Exit codes: 0 ok, 1 failure, 2 usage error, 3 not running (status only)
USAGE
}
log() { printf '%s\n' "$*"; }
vlog() { if ((verbose)); then printf '%s: %s\n' "$PROG" "$*" >&2; fi; return 0; }
die() { printf '%s: %s\n' "$PROG" "$1" >&2; exit "$EX_FAIL"; }
usage_error() { printf '%s: %s\n' "$PROG" "$1" >&2; usage >&2; exit "$EX_USAGE"; }
running_pid() { # print the pid and return 0 only if that process is alive
local pid
[[ -r $PIDFILE ]] || return 1
read -r pid < "$PIDFILE" || return 1
[[ $pid =~ ^[0-9]+$ ]] && kill -0 "$pid" 2>/dev/null || return 1
printf '%s\n' "$pid"
}
cmd_start() {
(($# == 0)) || usage_error "start takes no arguments"
local pid
if pid=$(running_pid); then log "already running"; vlog "pid $pid"; return "$EX_OK"; fi
mkdir -p "$STATE_DIR" || die "cannot create $STATE_DIR"
vlog "launching: ${SVC_CMD[*]}"
nohup "${SVC_CMD[@]}" > "$STATE_DIR/svc.log" 2>&1 &
pid=$!
printf '%s\n' "$pid" > "$PIDFILE"
sleep 0.2
kill -0 "$pid" 2>/dev/null || { rm -f "$PIDFILE"; die "exited immediately, see $STATE_DIR/svc.log"; }
log "started"; vlog "pid $pid"
}
cmd_stop() {
(($# == 0)) || usage_error "stop takes no arguments"
local pid i
if ! pid=$(running_pid); then rm -f "$PIDFILE"; log "not running"; return "$EX_OK"; fi
kill "$pid" 2>/dev/null || die "cannot signal pid $pid"
for ((i = 0; i < STOP_TIMEOUT * 10; i++)); do
kill -0 "$pid" 2>/dev/null || break
sleep 0.1
done
if kill -0 "$pid" 2>/dev/null; then
vlog "SIGTERM ignored for ${STOP_TIMEOUT}s, sending SIGKILL"
kill -KILL "$pid" 2>/dev/null
sleep 0.1
kill -0 "$pid" 2>/dev/null && die "pid $pid will not die"
fi
rm -f "$PIDFILE"
log "stopped"
}
cmd_status() {
(($# == 0)) || usage_error "status takes no arguments"
if running_pid > /dev/null; then log "running"; return "$EX_OK"; fi
log "not running"; return "$EX_NOT_RUNNING"
}
while (($#)); do # global options come before the subcommand
case $1 in
-v|--verbose) verbose=1; shift ;;
-h|--help) usage; exit "$EX_OK" ;;
--) shift; break ;;
-*) usage_error "unknown option: $1" ;;
*) break ;;
esac
done
(($#)) || usage_error "missing command"
command=$1; shift
case $command in
start) cmd_start "$@" ;;
stop) cmd_stop "$@" ;;
status) cmd_status "$@" ;;
*) usage_error "unknown command: $command" ;;
esac
exit $?
Words used in the script
- pid file: a small file holding the process id (the number the operating system gives a running process), so a later
svc stoporsvc statusknows which process to check or signal. nohup: runs the command so it ignores the hangup signal sent when the terminal that started it closes; the trailing&puts it in the background.- SIGTERM and SIGKILL: signals are messages the kernel delivers to a process.
kill PIDsends SIGTERM, a polite "please terminate" that a program may handle or ignore;kill -KILL PIDsends SIGKILL, which cannot be caught or ignored and ends the process immediately. kill -0: sends no signal at all, but reports through its exit status whether the process exists, which makes it a liveness test./proc/PID/comm: the Linux virtual file that holds the short name of the command running as that pid (executed: for the demo service it printssleep), so a script can check that a pid still belongs to the expected program.readonly: marks a variable as constant for the rest of the script, soEX_USAGEcannot be reassigned by accident.${0##*/}:$0is the path the script was started as;##*/removes the longest prefix ending in a slash, leaving the bare file name. Run as./svcor/usr/local/bin/svc,PROGbecomessvceither way.
Design choices worth defending:
- Global flag before the subcommand (
svc -v start): the option loop runs first and stops at the first non-option word, which is the subcommand. Subcommand-specific flags would be parsed inside eachcmd_*function. Each subcommand rejects stray arguments (start extrais a usage error), so typos fail loudly. - Idempotent
startandstop: starting a running service and stopping a stopped one both exit 0, because the desired state already holds. A caller that loops "stop, then start" in a deploy script should not break on the second half of a retry.statusis the command that reports state through its exit code (0 or 3), so scripts can writeif svc status; then .... - Verbose goes to stderr (
vlog) so it never pollutes output a caller captures from stdout. - Liveness is checked with
kill -0, which sends no signal but reports whether the process exists. A pid file that points at a dead process counts as "not running", and stop cleans it up. (Pid reuse after a long outage can make an old pid file point at an unrelated process; for production use, also compare the process name from/proc/PID/comm, or let systemd own the process.) - Stop is TERM, wait, then KILL: polite first, bounded wait (
SVC_STOP_TIMEOUT), force only if ignored. If even SIGKILL fails, exit 1. - The managed command is split into an array once (
read -r -a) and launched as"${SVC_CMD[@]}", with noeval.
Executed run
export SVC_STATE_DIR=/tmp/svc-demo
rm -rf "$SVC_STATE_DIR"
t() {
printf '$ svc'; printf ' %s' "$@"; printf '\n'
out=$(./svc "$@" 2>&1); rc=$?
printf '%s\n' "$out" | head -n 2 | sed 's/^/ /'
echo " (exit $rc)"
}
t status
t -v start
t start
t status
t stop
t status
t stop
t frobnicate
t --colour start
t
t start extra
echo '--- a service that ignores SIGTERM'
export SVC_CMD=./stubborn.sh SVC_STOP_TIMEOUT=1
t start
t -v stop
t status
stubborn.sh stands in for a service that ignores the polite signal: trap "" TERM tells the shell to ignore SIGTERM, and exec sleep 300 replaces the shell with sleep, which keeps that ignore setting.
#!/bin/sh
trap "" TERM
exec sleep 300
$ svc status
not running
(exit 3)
$ svc -v start
svc: launching: sleep 300
started
(exit 0)
$ svc start
already running
(exit 0)
$ svc status
running
(exit 0)
$ svc stop
stopped
(exit 0)
$ svc status
not running
(exit 3)
$ svc stop
not running
(exit 0)
$ svc frobnicate
svc: unknown command: frobnicate
Usage: svc [-v|--verbose] <command>
(exit 2)
$ svc --colour start
svc: unknown option: --colour
Usage: svc [-v|--verbose] <command>
(exit 2)
$ svc
svc: missing command
Usage: svc [-v|--verbose] <command>
(exit 2)
$ svc start extra
svc: start takes no arguments
Usage: svc [-v|--verbose] <command>
(exit 2)
--- a service that ignores SIGTERM
$ svc start
started
(exit 0)
$ svc -v stop
svc: SIGTERM ignored for 1s, sending SIGKILL
stopped
(exit 0)
$ svc status
not running
(exit 3)
The lifecycle behaves as described: status is 3 before start and after stop, a second start reports "already running" and exits 0, a second stop reports "not running" and exits 0. All four invalid-input cases (unknown command, unknown option, missing command, extra argument) exit 2 and print an error line followed by the usage text (the harness trims each case to the first two lines, so only the error line and the first usage line show). The last block shows the SIGKILL path, announced by -v.
The harness trims each case to two lines. Untrimmed, the usage text and the start extra error look like this (executed):
$ svc -h
Usage: svc [-v|--verbose] <command>
Commands:
start launch the service (exit 0 if it is already running)
stop stop the service (exit 0 if it is already stopped)
status exit 0 if running, 3 if not
Exit codes: 0 ok, 1 failure, 2 usage error, 3 not running (status only)
(exit 0)
$ svc start extra
svc: start takes no arguments
Usage: svc [-v|--verbose] <command>
Commands:
start launch the service (exit 0 if it is already running)
stop stop the service (exit 0 if it is already stopped)
status exit 0 if running, 3 if not
Exit codes: 0 ok, 1 failure, 2 usage error, 3 not running (status only)
(exit 2)
Exit-code contract
| Code | Meaning | Who relies on it |
|---|---|---|
| 0 | the requested state holds (start, stop) or the service is running (status) | scripts using if svc status |
| 1 | operational failure: cannot create the state directory, process exited immediately, would not die | alerting, deploy pipelines |
| 2 | usage error: wrong command, option or argument count | humans and CI authors |
| 3 | status only: not running | monitoring |
Pitfalls
- Never print usage and exit 0 on bad input; callers would treat a typo as success.
- Keep exit codes below 126, which shells use for "found but not executable" (126) and "command not found" (127); 128+N means killed by signal N.
- A
startthat returns before the process is actually serving only proves it did not crash in the first 0.2 seconds; a readiness probe (poll a port or health URL) is the next step.
How do you test shell scripts? Show a test for a script that parses arguments, handles a missing file and logs, including how you substitute external commands, and where integration tests in a container fit.
Sample Answer
Direct answer
Test shell scripts at three levels, and run them in this order in continuous integration (CI, the automated build that runs on every change): static analysis (ShellCheck, a linter that finds bugs in shell scripts, and shfmt, a shell formatter), unit tests with bats-core (Bash Automated Testing System) that run the script with its external commands replaced by fakes, and integration tests inside a throwaway container where the real tools and a real listener are present. To make a script testable, put the logic in functions, call main only when the file is executed (not when it is sourced), and give every failure a distinct exit code. Substitute an external command by putting a fake executable with the same name first in PATH for one test. (Terms differ slightly: a stub returns canned answers, a fake is a simplified working replacement, and a mock also checks how it was called. The curl replacement below records its arguments, so it is a stub with a call log; the text below uses "fake" for all of them.)
The script under test
It counts ERROR lines in a log file (arguments: -v for verbose, -u URL to post the count with curl, then the file), logs to stderr, and has distinct exit codes (0 success, 2 usage error, 66 input file missing, 69 upload failed; 66 and 69 are the sysexits.h values (a conventional list of exit codes from BSD Unix) for no input and service unavailable).
#!/usr/bin/env bash
# Usage: count-errors.sh [-v] [-u URL] LOGFILE
# Counts ERROR lines in LOGFILE and, with -u, posts the count to URL with curl.
# Exit codes: 0 ok, 2 usage error, 66 input file missing, 69 upload failed.
log() { printf '%s %s\n' "$1" "$2" >&2; }
usage() { echo "usage: ${0##*/} [-v] [-u URL] LOGFILE" >&2; }
main() {
local verbose=0 url="" opt
OPTIND=1
while getopts ':vu:' opt; do
case $opt in
v) verbose=1 ;;
u) url=$OPTARG ;;
*)
usage
return 2
;;
esac
done
shift $((OPTIND - 1))
(($# == 1)) || {
usage
return 2
}
local file=$1
[[ -f $file ]] || {
log ERROR "no such file: $file"
return 66
}
local count
count=$(grep -c ' ERROR ' -- "$file" || true)
((verbose)) && log INFO "found $count errors in $file"
echo "$count"
if [[ -n $url ]]; then
curl --silent --fail --data "errors=$count" "$url" >/dev/null || {
log ERROR "upload failed"
return 69
}
fi
}
if [[ ${BASH_SOURCE[0]} == "$0" ]]; then
set -euo pipefail # only when executed: sourcing for tests must not change the test shell
main "$@"
fi
Reading the argument parsing: getopts ':vu:' opt loops over the command-line options. In the option string, v is a flag with no value, u: (the colon after u) is an option that takes a value, and the leading colon turns on silent mode, where getopts reports problems through opt instead of printing its own message. Each pass puts the option letter in opt; for u the value goes in OPTARG, and OPTIND is the index of the next argument to process. In silent mode an unknown option sets opt to ? and a missing value sets it to :, so the one *) case catches both. After the loop, shift $((OPTIND - 1)) discards the options already handled so that $1 is the log file. Measured with the same option string (bash 5 in a container):
args: -v -u http://x/ app.log
opt=v OPTARG= OPTIND=2
opt=u OPTARG=http://x/ OPTIND=4
left over: app.log
args: -x app.log
opt=? OPTARG=x OPTIND=2
left over: app.log
args: -u
opt=: OPTARG=u OPTIND=2
left over:
Three idioms in the script: (($# == 1)) || { usage; return 2; } evaluates an arithmetic test ($# is the number of arguments) and runs the braces only when it is false; ((verbose)) && log INFO ... logs only when verbose is non-zero; and grep -c ... || true keeps the script going when there are no matches, because grep -c prints 0 but exits with status 1, which set -e would otherwise treat as a failure (measured: the output was 0 and the status 1, and || true made the status 0). BASH_SOURCE[0] is the file the running code lives in, while $0 is the script that was started: run directly they are equal, and when another file sources it they differ (measured: BASH_SOURCE[0]=./who.sh $0=caller.sh), which is what the guard tests.
Testability choices: functions (main, log, usage) instead of top-level code; the if [[ ${BASH_SOURCE[0]} == "$0" ]] guard, true only when the file is run directly, so a test can source it and call main without side effects; set -euo pipefail placed inside that guard so sourcing does not change the test shell's options; results on stdout and diagnostics on stderr so tests can assert on each.
Unit tests with fakes (bats)
bats_require_minimum_version 1.5.0
setup() {
SCRIPT="$BATS_TEST_DIRNAME/../count-errors.sh"
STUBS="$BATS_TEST_TMPDIR/stubs"
mkdir -p "$STUBS"
printf '%s\n' '2024-05-01T00:00:01Z INFO ok' '2024-05-01T00:00:02Z ERROR a' '2024-05-01T00:00:03Z ERROR b' \
> "$BATS_TEST_TMPDIR/app.log"
# Fake curl: record the arguments, exit with $FAKE_CURL_RC (default 0).
cat > "$STUBS/curl" <<'STUB'
#!/usr/bin/env bash
printf '%s\n' "$*" >> "$CURL_CALLS"
exit "${FAKE_CURL_RC:-0}"
STUB
chmod +x "$STUBS/curl"
export CURL_CALLS="$BATS_TEST_TMPDIR/curl.calls"
export PATH="$STUBS:$PATH" # the stub shadows the real curl for this test only
}
@test "counts ERROR lines" {
run "$SCRIPT" "$BATS_TEST_TMPDIR/app.log"
[ "$status" -eq 0 ]
[ "$output" = "2" ]
}
@test "-v logs to stderr, not stdout" {
run --separate-stderr "$SCRIPT" -v "$BATS_TEST_TMPDIR/app.log"
[ "$output" = "2" ]
[[ $stderr == "INFO found 2 errors in "* ]]
}
@test "missing file exits 66 and names the file" {
run "$SCRIPT" "$BATS_TEST_TMPDIR/nope.log"
[ "$status" -eq 66 ]
[[ $output == *"no such file: $BATS_TEST_TMPDIR/nope.log"* ]]
}
@test "no arguments and unknown options exit 2 with usage" {
run "$SCRIPT"
[ "$status" -eq 2 ]
[[ $output == usage:* ]]
run "$SCRIPT" -x "$BATS_TEST_TMPDIR/app.log"
[ "$status" -eq 2 ]
}
@test "-u posts the count through curl, exactly once" {
run "$SCRIPT" -u http://metrics.example/ingest "$BATS_TEST_TMPDIR/app.log"
[ "$status" -eq 0 ]
[ "$(wc -l < "$CURL_CALLS")" -eq 1 ]
[ "$(cat "$CURL_CALLS")" = "--silent --fail --data errors=2 http://metrics.example/ingest" ]
}
@test "a failing curl becomes exit 69" {
FAKE_CURL_RC=22 run "$SCRIPT" -u http://metrics.example/ingest "$BATS_TEST_TMPDIR/app.log"
[ "$status" -eq 69 ]
[[ $output == *"upload failed"* ]]
}
@test "without -u curl is never called" {
run "$SCRIPT" "$BATS_TEST_TMPDIR/app.log"
[ ! -e "$CURL_CALLS" ]
}
@test "sourcing the script does not run main, so functions can be tested directly" {
source "$SCRIPT"
run main -v "$BATS_TEST_TMPDIR/app.log"
[ "$status" -eq 0 ]
}
How the pieces work:
$BATS_TEST_TMPDIRis a temporary directory that bats creates fresh for each test, so tests cannot leak files into one another;$BATS_TEST_DIRNAMEis where the test file lives (both are documented bats variables).run cmdruns the command and captures$statusand$output;run --separate-stderralso fills$stderr(needsbats_require_minimum_version 1.5.0, declared at the top), which is how the-vtest proves that log lines do not pollute stdout.- Substituting
curl.setupwrites an executable namedcurlinto a per-test directory and puts that directory first inPATH. The fake records its arguments and exits with a status set byFAKE_CURL_RC. Tests can then assert exactly what the script would send (--silent --fail --data errors=2 <url>), force a failure (exit 22) to check the 69 path, and provecurlis not called without-u. No network is involved, so these tests are deterministic. - For shell functions called by your code (not external programs), define a function with the same name in the test; functions take precedence over executables.
Running the eight unit tests in the bats/bats container image printed ok for all eight:
1..8
ok 1 counts ERROR lines
ok 2 -v logs to stderr, not stdout
ok 3 missing file exits 66 and names the file
ok 4 no arguments and unknown options exit 2 with usage
ok 5 -u posts the count through curl, exactly once
ok 6 a failing curl becomes exit 69
ok 7 without -u curl is never called
ok 8 sourcing the script does not run main, so functions can be tested directly
Do the tests have teeth?
Three deliberate mutations of the script were run against the same suite: changing return 66 to return 65, removing --fail from the curl call, and posting errors=$((count+1)). Each made the suite fail (one failing test each), so the tests check the exit code, the flags and the payload.
Static analysis
docker run --rm -v "$PWD":/mnt koalaman/shellcheck:stable count-errors.sh
shfmt -d -i 4 -ci count-errors.sh
ShellCheck reported nothing for the script, and shfmt -d (which prints a diff when formatting differs) printed nothing after the script had been formatted with the same options. ShellCheck catches unquoted expansions, cd without a failure check, and useless uses of cat; it cannot catch logic errors, which is what the unit tests are for.
Integration tests in a container
Fakes prove how the script calls curl; they cannot prove the call works. The integration file starts a real HTTP listener and uses the real curl. The listener, sink.py, is a few lines of Python that accept an HTTP POST on 127.0.0.1:8099, append the request body to a file and reply 204 No Content; the test then reads that file to see what the script really sent:
# Runs against the real curl and a real HTTP listener: no stubs.
setup() {
printf '%s\n' '2024-05-01T00:00:02Z ERROR a' '2024-05-01T00:00:03Z ERROR b' '2024-05-01T00:00:04Z ERROR c' \
> "$BATS_TEST_TMPDIR/app.log"
cat > "$BATS_TEST_TMPDIR/sink.py" <<'PY'
import http.server, sys
out = sys.argv[1]
class H(http.server.BaseHTTPRequestHandler):
def do_POST(self):
body = self.rfile.read(int(self.headers["Content-Length"]))
open(out, "ab").write(body + b"\n")
self.send_response(204); self.end_headers()
def log_message(self, *a): pass
http.server.HTTPServer(("127.0.0.1", 8099), H).serve_forever()
PY
python3 "$BATS_TEST_TMPDIR/sink.py" "$BATS_TEST_TMPDIR/received" &
SINK=$!
for _ in 1 2 3 4 5 6 7 8 9 10; do curl -s -o /dev/null http://127.0.0.1:8099/ -X POST -d '' && break; sleep 0.2; done
: > "$BATS_TEST_TMPDIR/received"
}
teardown() { kill "$SINK" 2>/dev/null || true; }
@test "real curl delivers the count to a real server" {
run "$BATS_TEST_DIRNAME/../count-errors.sh" -u http://127.0.0.1:8099/ingest "$BATS_TEST_TMPDIR/app.log"
[ "$status" -eq 0 ]
[ "$(cat "$BATS_TEST_TMPDIR/received")" = "errors=3" ]
}
@test "real curl against a dead port gives exit 69" {
run "$BATS_TEST_DIRNAME/../count-errors.sh" -u http://127.0.0.1:1/ "$BATS_TEST_TMPDIR/app.log"
[ "$status" -eq 69 ]
}
Run it in a container with bats, curl and python3 installed. This three-line Dockerfile builds one from the official bats image (ENTRYPOINT [] clears the image's default command so the next command can start with bats):
FROM bats/bats:latest
RUN apk add --no-cache curl python3
ENTRYPOINT []
docker build -t bats-curl-py .
docker run --rm -v "$PWD":/code -w /code bats-curl-py bats test/integration.bats
1..2
ok 1 real curl delivers the count to a real server
ok 2 real curl against a dead port gives exit 69
--rm removes the container (and its anonymous volume) on exit, so every run starts from a clean image and nothing is left on the CI host. The container also gives you the target environment: if production runs GNU userland on Debian and your laptop is macOS, the integration job is the one that finds sed -i or date -d differences. The bats/bats image used here is built on Alpine Linux, whose userland is BusyBox (measured: /bin/grep is a link to busybox), so this particular job exercises BusyBox and musl behaviour; to mirror a Debian production, base the image on debian and install bats instead.
CI pipeline
| Step | Command | Catches |
|---|---|---|
| 1. Lint | shellcheck on every script | quoting, portability and common-bug patterns |
| 2. Format | shfmt -d | style drift |
| 3. Unit | bats test/count-errors.bats | logic, exit codes, argument handling, with fakes |
| 4. Integration | the container command above | real tool behaviour, environment differences |
Fail fast: the cheap steps run first, and the container job runs only if they pass.
Trade-offs
- A fake encodes your assumption about the tool, so a wrong assumption passes the unit test. The integration test is the check on that assumption; keep a few, covering the calls that matter, rather than duplicating every unit test.
- Do not assert on log wording more than needed: test the exit code and the one phrase an operator relies on.
- Scripts that are mostly one pipeline are better covered by a couple of input/expected-output tests than by mocks of each stage.
Implement argument parsing in Bash for a tool that supports both short and long forms of help, verbose and an output option, combined short flags, and any number of positional files. What are the limits of the built-in parsing facility, and what would you do instead?
Sample Answer
Direct answer
Mind the one-letter difference: getopts (with an s) is a builtin command inside Bash that parses short options only, while getopt (no s) is a separate program from the util-linux package that also understands long options. Bash's built-in getopts handles only single-character options: it cannot parse --help or --verbose, it has no optional option-arguments, and it stops at the first non-option, so tool file.log -v leaves -v as a positional (a positional argument is a plain argument such as a file name, as opposed to an option like -v). For a tool that needs long forms and positionals in any order, I would write a small while/case loop that walks "$@" itself (portable to Bash 3.2 and later, no external dependency), and use the util-linux getopt command only when I can guarantee it exists. The hand-rolled parser below supports -h/--help, -v/--verbose, -o/--output in three spellings, combined short flags such as -vo out.txt, -- and any number of file arguments anywhere.
The parser
#!/usr/bin/env bash
set -u
usage() {
cat <<'USAGE'
Usage: mytool [-hv] [-o FILE | --output=FILE | --output FILE] [--] [FILE...]
-h, --help show this help and exit
-v, --verbose print extra detail
-o, --output F write results to F
USAGE
}
die() { printf 'mytool: %s\n' "$1" >&2; usage >&2; exit 2; }
verbose=0
output=''
files=()
while (($#)); do
case $1 in
-h|--help) usage; exit 0 ;;
-v|--verbose) verbose=1 ;;
-o|--output)
(($# >= 2)) || die "option $1 needs a value"
output=$2; shift ;;
--output=*) output=${1#--output=} ;;
--) shift; files+=("$@"); break ;;
--*) die "unknown option: $1" ;;
-?*) # a cluster such as -vo, -vvh or -ofile
cluster=${1#-}
while [[ -n $cluster ]]; do
flag=${cluster:0:1}; cluster=${cluster:1}
case $flag in
h) usage; exit 0 ;;
v) verbose=1 ;;
o) if [[ -n $cluster ]]; then
output=$cluster # -ofile
else
(($# >= 2)) || die "option -o needs a value"
output=$2; shift # -o file (or the tail of -vo file)
fi
cluster='' ;;
*) die "unknown option: -$flag" ;;
esac
done ;;
*) files+=("$1") ;; # positional, including a lone "-"
esac
shift
done
printf 'verbose=%s output=%q files=%d' "$verbose" "$output" "${#files[@]}"
for f in ${files[@]+"${files[@]}"}; do printf ' [%s]' "$f"; done
printf '\n'
How it works:
while (($#))consumes one argument per iteration andshifts, so a value-taking option can take one more argument with an extrashift.casearms are ordered from specific to general: exact long options,--output=*, the--terminator, any other--*(rejected as unknown), then-?*for short clusters, then everything else as a positional. A lone-(conventionally stdin) is a positional because it does not match-?*.- A cluster such as
-vois walked one letter at a time with${cluster:0:1}.oneeds a value, so it takes either the rest of the cluster (-ofile.txt) or the next argument (-vo out.txt). - Four expansions do the string work, and each was executed to confirm its reading:
${1#-}removes one leading dash, so the argument-vobecomes the clustervo;${1#--output=}removes the prefix--output=from the front of$1(for--output=out.txtit leavesout.txt);${cluster:0:1}is the substring ofclusterstarting at position 0, one character long (the first letter), and${cluster:1}is everything from position 1 on, so forcluster=vofile.txtthey givevandofile.txt; loopingflag=${cluster:0:1}; cluster=${cluster:1}therefore peels one letter per pass. - The pattern
-?*matches a dash followed by at least one more character (?is exactly one character,*is any number), so-vomatches, while a lone-and a file name such asa.logdo not. - Errors go to stderr and exit with status 2, the usual "usage error" code;
--helpprints to stdout and exits 0. ${files[@]+"${files[@]}"}in the final loop reads as: iffilesis set, expand its elements, otherwise expand to nothing. It avoids an "unbound variable" error underset -ufor an empty array on Bash older than 4.4. Executed on Bash 3.2.57, a plain"${files[@]}"on an empty array stopped withfiles[@]: unbound variable, while the guarded form ran; Bash 5 accepts both.
Executed cases
The harness prints each command, the first two lines of output and the exit status:
t() { # show the command, the first two output lines, the exit status
printf '$ mytool'; printf ' %q' "$@"; printf '\n'
out=$(./mytool.sh "$@" 2>&1); rc=$?
printf '%s\n' "$out" | head -n 2 | sed 's/^/ /'
echo " (exit $rc)"
}
t -v -o out.txt a.log b.log
t -vo out.txt a.log
t -vofile.txt a.log
t --verbose --output=out.txt a.log
t a.log -v 'my file.log' --output out.txt
t -- -v.log
t -x
t --nope
t -o
t --help
$ mytool -v -o out.txt a.log b.log
verbose=1 output=out.txt files=2 [a.log] [b.log]
(exit 0)
$ mytool -vo out.txt a.log
verbose=1 output=out.txt files=1 [a.log]
(exit 0)
$ mytool -vofile.txt a.log
verbose=1 output=file.txt files=1 [a.log]
(exit 0)
$ mytool --verbose --output=out.txt a.log
verbose=1 output=out.txt files=1 [a.log]
(exit 0)
$ mytool a.log -v my\ file.log --output out.txt
verbose=1 output=out.txt files=2 [a.log] [my file.log]
(exit 0)
$ mytool -- -v.log
verbose=0 output='' files=1 [-v.log]
(exit 0)
$ mytool -x
mytool: unknown option: -x
Usage: mytool [-hv] [-o FILE | --output=FILE | --output FILE] [--] [FILE...]
(exit 2)
$ mytool --nope
mytool: unknown option: --nope
Usage: mytool [-hv] [-o FILE | --output=FILE | --output FILE] [--] [FILE...]
(exit 2)
$ mytool -o
mytool: option -o needs a value
Usage: mytool [-hv] [-o FILE | --output=FILE | --output FILE] [--] [FILE...]
(exit 2)
$ mytool --help
Usage: mytool [-hv] [-o FILE | --output=FILE | --output FILE] [--] [FILE...]
-h, --help show this help and exit
(exit 0)
Note that -v after a.log is still recognised (options and positionals can mix), my file.log stays one element, and -- -v.log makes a file whose name looks like an option.
The limits of getopts
parse() {
local opt OPTIND=1
while getopts ':hvo:' opt; do
case $opt in
h) echo help ;; v) echo verbose ;; o) echo "output=$OPTARG" ;;
\?) echo "rejected: -$OPTARG" ;; :) echo "missing value for -$OPTARG" ;;
esac
done
shift $((OPTIND - 1)); echo "left over: $*"
}
echo '$ parse -vo out.txt a.log'; parse -vo out.txt a.log
echo '$ parse --verbose a.log'; parse --verbose a.log
echo '$ parse a.log -v b.log'; parse a.log -v b.log
$ parse -vo out.txt a.log
verbose
output=out.txt
left over: a.log
$ parse --verbose a.log
rejected: --
verbose
rejected: -e
rejected: -r
rejected: -b
output=se
left over: a.log
$ parse a.log -v b.log
left over: a.log -v b.log
-vo out.txt works, which is the one thing getopts does well (including combined short flags). But --verbose is not understood as a word: getopts reads it as a cluster of single letters. The second dash is rejected as an illegal option, v is accepted as -v, then e, r and b are rejected, and o swallows the remaining se as its value (output=se). And once it meets a.log, it stops, so -v and b.log come back unparsed. In short: no long options, no mixed order, no optional arguments, and a silent misparse when a long option reaches it.
Alternatives
| Choice | Long options | Portability | Verdict |
|---|---|---|---|
getopts | no | every POSIX shell | fine for short-only scripts |
hand-rolled while/case | yes | Bash 3.2+ | my default for a tool people will run on laptops and servers |
util-linux getopt (enhanced) | yes, plus reordering | Linux; the old BSD/macOS getopt does not support long options | good on Linux-only fleets |
a real language (Python argparse, Go flag) | yes, with generated help | needs the runtime | once the script has subcommands and validation, switch |
The util-linux getopt normalises arguments into a canonical order that you then eval:
set -- -vo out.txt --verbose --output=x.txt a.log 'my file.log'
parsed=$(getopt -o hvo: -l help,verbose,output: -n mytool -- "$@") || exit 2
eval "set -- $parsed"
printf '[%s]' "$@"; echo
[-v][-o][out.txt][--verbose][--output][x.txt][--][a.log][my file.log]
The options come first (-vo is split into -v and -o, --output=x.txt into --output x.txt), then --, then the positionals with my file.log intact. It is compact, but it depends on the util-linux version of getopt and on eval of its output. The eval is safe because getopt wraps every word in single quotes (escaping any quote inside it) before printing, so eval sees data, never code. The raw text for the same arguments, then for a hostile value, executed:
echo "$parsed"
set -- -o "it's; rm -rf x" a
getopt -o o: -- "$@"
-v -o 'out.txt' --verbose --output 'x.txt' -- 'a.log' 'my file.log'
-o 'it'\''s; rm -rf x' -- 'a'
The ; and the quote in the second value stay inside the quoted word, so eval "set -- $parsed" just sets a positional parameter to that text.
Pitfalls
- A value that begins with
-(-o -weird) is accepted as the value in this parser; some tools instead reject it. Pick one rule and document it. - Do not parse
$*or an unquoted$@, because that splits file names with spaces. - Validate required options after the loop (for example "at least one file unless
--stdin") so error messages are specific.
Unlock Full Question Bank
Get access to all 40 Shell Scripting and Automation interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.