Version Control and Developer Tooling Questions
The everyday toolchain of software work: version control with Git (branching, merging, rebasing, conflict resolution, using git bisect to find a regression), command-line and shell proficiency for day-to-day navigation, log inspection, and troubleshooting, IDE and editor workflows, build systems and package/dependency management (npm, Maven, pip, Gradle, CocoaPods, and embedded/cross-compilation toolchains), and the growing practice of AI-assisted coding: using, reviewing, and verifying AI-generated code and tests. Deliberately generic across languages and stacks; language- and domain-specific frameworks live in their own categories. This topic covers a developer's individual command of these tools, not: writing durable shell automation and glue scripts (Shell Scripting and Automation owns that), producing, versioning, and publishing build artifacts or container images (Build Automation and Artifact Management owns that), release cadence and change governance (Release Management and Change Control owns that), or diagnosing a live production incident end to end (Performance Troubleshooting and Incident Response and the Observability topics own that).
You ask an AI assistant to modify a function in an existing codebase, but it quietly changes the function's signature and breaks several callers elsewhere in the code. What safeguards would you put in place to catch this before merging, and how would you preserve interfaces and backward compatibility when accepting AI-generated changes going forward?
Sample Answer
Direct answer
I would put safeguards in place before the change ever lands in the main branch, not rely on catching it in review by eye: an explicit "do not change the signature" constraint in the prompt, a search across the codebase for every caller of the function being touched, and a CI (continuous integration) check that actually exercises those callers. Catching a broken signature by carefully reading a diff does not scale, since the break is often invisible in the diff of the function itself, it only shows up as a failure somewhere else entirely.
Structured elaboration
Preventing it at the prompt level
- Constrain the request explicitly: "modify the internal logic but do not change the function's name, parameter order, parameter names, or return type," rather than a general "improve this function" that leaves the interface up for grabs.
- Provide the function's current callers as context, so the assistant can see what would break, rather than reasoning about the function in isolation.
Catching it before merge
- Search the codebase for every call site of the changed function before approving the change (a simple text or symbol search, or an IDE's find-references), and confirm each one still matches the new signature.
- Run the full test suite, including integration tests that exercise the function through its real callers, not only unit tests that call the function directly with hand-picked arguments a test author already expects to work.
- Add a lightweight contract check where the interface is easy to characterize: a type-checked function signature, or a small test that specifically asserts the function's signature has not changed, so a future AI-assisted change trips this test immediately rather than being caught downstream.
Preserving backward compatibility going forward
- Default to additive changes: add a new optional parameter with a default rather than changing an existing required one, so old call sites keep working unmodified.
- When a breaking change is genuinely necessary, use a wrapper or adapter function that preserves the old signature and forwards to the new one, and version the interface explicitly if it is public.
- Require an explicit compatibility note in any pull request that touches a function with more than a couple of callers: what changed, who is affected, and how it was validated, so a reviewer does not have to reconstruct that from the diff alone.
Preserving behavior after a pure refactor
A pure refactor (same signature, intended to be behavior-preserving, done purely for readability) has its own distinct checklist beyond signature preservation: confirm the output is identical on a representative set of real inputs, not just that the code compiles; check latency has not regressed, since a "cleaner" rewrite can accidentally introduce an expensive operation inside a loop that used to be outside it; and check error handling specifically, since a refactor is a common place for an exception type or an edge-case branch to silently disappear during a rewrite that otherwise looks like a pure simplification.
Worked example
An AI assistant is asked to "clean up" a function parse_config(path) used in six places across the codebase. It returns a cleaner implementation, but the new version returns None on a missing file instead of raising a FileNotFoundError the way the original did, a behavior change with no signature change at all, so a caller relying on the exception (a try/except FileNotFoundError block elsewhere) now gets a silent None instead, and fails later in a way that is much harder to trace back to this change. A signature-only check would miss this entirely, which is exactly why behavior and error-handling checks matter as a separate step from an interface check, and why "did the tests pass" is not sufficient if none of the six callers had a test covering the missing-file case.
Trade-offs and pitfalls
Contract tests and full-codebase caller searches add real time to every review, which can feel excessive for a function with one caller. The safeguards are worth scaling to the function's actual blast radius: a private helper used in one place needs a lighter check than a function called from a dozen places across the codebase, and treating every change with maximum ceremony trains people to skip the ceremony rather than actually protecting the changes that matter most.
Your team uses Conventional Commits. Describe how to implement client-side and server-side checks to enforce commit message structure before merges. Include tools you would use and where they run in the CI/CD pipeline.
Sample Answer
Direct answer
Enforce Conventional Commits (a convention where a commit message's first line has the shape type(scope): description, for example feat(auth): add SSO login) at two layers: a local commit-msg git hook that rejects a malformed message before the commit is even created, and a server-side check in CI that re-validates every commit in the pull request (PR). The local hook gives fast feedback; the CI check is the actual gate, because local hooks can be skipped or simply not installed.
Structured elaboration
What the convention requires: a first line like <type>[optional scope]: <description>, with recognized types such as feat, fix, chore, docs, refactor, test, ci, and a ! after the type or scope (or a BREAKING CHANGE: footer) marking a breaking change.
Client-side enforcement. Git runs a commit-msg hook automatically right after you write a commit message but before the commit object is created, passing the hook the path to a file containing that message. A linter (commitlint, configured with a conventional-commits ruleset) checks the message against the format and exits non-zero on a violation, which makes git abort the commit. Because git hooks live in .git/hooks and are not tracked or shared by git itself, distribute the hook via a tool like Husky or lefthook that installs it automatically when a contributor sets up the repo (e.g. on npm install), otherwise every new contributor silently has no hook at all.
Server-side enforcement. A CI job triggered on every push to the PR runs commitlint against the full commit range being merged (commitlint --from <base-sha> --to <head-sha>), independent of whether any individual contributor had the local hook installed or chose to bypass it with git commit --no-verify. Teams that squash-merge PRs should instead (or additionally) lint the PR title, since a squash merge collapses every commit into one, and only the title survives as the final commit message.
Where each check runs in the pipeline: the commit-msg hook runs on the developer's machine at commit time, before anything is pushed, fastest feedback, but bypassable. The CI job runs on every push to the PR, before merge is allowed, and should be wired into branch protection as a required check so the PR literally cannot merge until it passes, that's the real enforcement point.
Worked example
Local hook config (illustrative, matching commitlint's documented setup):
// .commitlintrc.json
{ "extends": ["@commitlint/config-conventional"] }
# .husky/commit-msg
npx --no -- commitlint --edit "$1"
CI step (illustrative GitHub Actions syntax):
- name: Lint commit messages
run: npx commitlint --from ${{ github.event.pull_request.base.sha }} --to ${{ github.event.pull_request.head.sha }}
These are standard commitlint usage patterns, not output from an executed run, verify against the current commitlint docs before shipping the exact config in a real repo.
Trade-offs and pitfalls
A local hook by itself is not a real gate, git commit --no-verify bypasses it in one flag, and a contributor who never ran the setup script simply doesn't have it. The CI check is what actually matters, and it must be a required status check in branch protection, not just a check that runs and reports. Squash-merge workflows that only ever lint individual commits are checking the wrong thing, since those messages get discarded, the PR title needs its own lint pass. Over-strict scope enforcement (rejecting valid but unlisted scopes) creates friction that trains people to write vague chore: stuff commits just to get past the linter, so keep the type list tight but the scope list either open or genuinely comprehensive. Conventional Commits' real payoff, an automatically generated changelog and semantic version bump via a tool like semantic-release, only works if enforcement is consistent; a repo that's 90% compliant still can't reliably automate off its history.
Describe how to redirect stdout and stderr to a file while also printing stdout to the terminal. Give the exact shell command(s) you would use in bash to: (1) capture both stdout and stderr to /tmp/run.log, and (2) simultaneously see stdout in the terminal. Explain why your approach works.
Sample Answer
Direct answer
Two building blocks solve this: 2>&1 merges stderr into stdout, and tee duplicates a stream to both the terminal and a file. Capture-only needs command > /tmp/run.log 2>&1. Capture-and-still-see-it-live needs command 2>&1 | tee /tmp/run.log.
Structured elaboration
# (1) capture both streams to the file, nothing shown on the terminal
your_command > /tmp/run.log 2>&1
# (2) capture both streams AND still see them live on the terminal
your_command 2>&1 | tee /tmp/run.log
Key points:
- File descriptor 1 (fd 1) is stdout, file descriptor 2 (fd 2) is stderr.
2>&1means "point fd 2 at wherever fd 1 currently points." - Order matters.
command > file.log 2>&1first redirects stdout to file.log, then makes stderr point at that same target, so both land in the file.command 2>&1 > file.logdoes the opposite: stderr is pointed at the terminal (wherever stdout currently was), and only then is stdout redirected to the file, so stderr stays on the terminal. This ordering mistake is the single most common bug with this idiom. teereads its own stdin and writes it to both stdout and the named file. That is why2>&1 | tee file.logworks: a pipe only ever forwards stdout by default, so stderr has to be merged into stdout with2>&1before it reaches the pipe if you wantteeto see it too.
Worked example
$ run_demo() { echo "stdout line"; echo "stderr line" >&2; }
$ run_demo 2>&1 | tee /tmp/run.log
stdout line
stderr line
$ cat /tmp/run.log
stdout line
stderr line
Both lines appear on the terminal and in the file, in the order the process wrote them. This was run exactly as shown.
Trade-offs & pitfalls
- Complexity is constant per line, this is OS-level file-descriptor plumbing, not a data scan, so cost tracks bytes written, not file size on disk.
- Ordering under a pipe: stdout is usually line-buffered when connected to a terminal, but becomes fully block-buffered once piped to
tee. If a process interleaves fast stdout and stderr writes, the captured order can differ slightly from an unpiped run. Force line buffering withstdbuf -oL your_commandwhen strict interleaving order matters. teeoverwrites the log file by default each run; usetee -a /tmp/run.logto append instead.- A plain redirect does not keep a process alive after you log out of SSH by itself, pair it with
nohupor a terminal multiplexer if the command needs to survive the session ending.
Compare GitFlow, GitHub Flow, and trunk-based development for a mid-sized company delivering multiple releases per quarter. For each strategy list advantages, disadvantages, and how it affects CI/CD pipeline design, release cadence, and hotfix handling.
Sample Answer
Direct answer
GitFlow, GitHub Flow, and trunk-based development are three different answers to "how do branches map to releases." GitFlow uses several long-lived branches (main, develop, release, hotfix) and suits teams that ship and support multiple distinct versions at once. GitHub Flow keeps one long-lived branch (main) plus short-lived feature branches merged through pull requests (PRs, the request to merge one branch into another after review), and suits teams that release frequently off a single deployable line. Trunk-based development pushes that further: everyone integrates into main at least daily, relying on feature flags (a runtime toggle that hides unfinished code) to keep incomplete work invisible to users. For a mid-sized company shipping multiple releases a quarter but supporting one live version, GitHub Flow or trunk-based development is usually the better fit; GitFlow's overhead pays for itself mainly when you must maintain several released versions in parallel.
Structured elaboration
| Strategy | Branch structure | CI/CD pipeline impact | Release cadence fit | Hotfix handling |
|---|---|---|---|---|
| GitFlow | main, develop, feature/*, release/*, hotfix/* | Needs separate pipelines for develop and each release/* branch; feedback is slower because features sit on develop before a release branch is even cut | Built for scheduled, versioned releases, comfortable when several versions need ongoing support | Branch a hotfix/* off main, fix, merge into both main and develop so the fix isn't lost at the next release |
| GitHub Flow | main + short-lived feature/* branches merged via PR | CI runs on every PR against main; CD can deploy straight from main after merge, one pipeline to maintain | Fits frequent, even continuous, releases off a single line; awkward if you must support multiple released versions simultaneously | Same path as any change: branch off main, PR, review, merge, deploy |
| Trunk-based | Single trunk (main); branches, if used at all, live under a day; feature flags hide incomplete work | CI must be fast and reliable since it gates every merge to trunk; release branches are cut from a known-green trunk commit when needed | Best for continuous delivery and frequent releases; release cadence is decoupled from code-complete because flags control exposure | Fix directly on trunk (fastest path) or cherry-pick the fix commit onto a short-lived release branch cut earlier |
Why the branch model changes CI/CD design: GitFlow's multiple long-lived branches mean the pipeline has multiple "definitions of green" to maintain in parallel, which is real ongoing infrastructure cost. GitHub Flow and trunk-based development collapse that to effectively one pipeline definition, at the cost of needing that one pipeline to be trustworthy enough to gate every merge.
Worked example
A mid-sized company shipping every 6 weeks, one supported production version, chooses GitHub Flow. A developer opens feature/add-2fa off main, opens a PR after a few days of work, CI runs unit tests plus a build, a teammate reviews, it merges into main, and main deploys to staging automatically then to production on a scheduled cadence. If the same company instead had to support v1, v2, and v3 of an on-prem product simultaneously (a genuinely different constraint), GitFlow's release/* branches would earn their overhead because each version needs its own long-lived line to receive fixes.
Trade-offs and pitfalls
The most common mistake is adopting GitFlow's ceremony (develop, release branches, the whole ritual) for a team that only ever ships one live version, that overhead becomes pure friction with no corresponding benefit. The opposite mistake is trunk-based development without real investment in feature flags and fast CI: without flags, "integrate to trunk daily" just means shipping half-finished features to users; without fast, reliable CI, trunk-based development turns every merge into a gamble. GitHub Flow's blind spot is multi-version support, since it has no native branch for "the previous release," teams that need it either bolt on release branches anyway (converging toward GitFlow) or accept that only the latest version gets fixes.
Before pasting AI-generated code into a production repository, what checklist do you use to validate correctness, security, style, licensing, and test coverage?
Sample Answer
Direct answer
I run through a fixed checklist covering correctness, security, style, licensing, documentation accuracy, and test coverage, in that rough priority order, before any AI-generated code enters a production repository. Correctness and security come first because they are the categories where a subtle mistake causes real damage; style and documentation matter, but a stylistically perfect, well-documented function that is subtly wrong is still a bad outcome.
Structured elaboration
Correctness
- Run the code on a small, known input and compare the actual output to the expected output by hand.
- Review edge cases explicitly: empty input, boundary values, unexpected types.
Security
- Look for unsafe deserialization (loading data in a way that can execute arbitrary code), shell or command injection (user-controlled input reaching a shell command unsanitized), secret leakage (a hardcoded key or token), and unvalidated file or network access.
Style
- Confirm the code follows the project's existing conventions: naming, logging patterns, and whatever the project's linter already enforces.
Licensing
- Verify that any snippet the assistant appears to have reproduced closely (a distinctive, non-trivial block that reads like it came from somewhere specific) is not lifted from a source with a restrictive license, since copied code can carry legal obligations the rest of the codebase does not have.
Tests
- Add unit tests covering the normal path, at least one failure path, and any regression case relevant to why the change was made in the first place.
Checking whether AI-written documentation actually matches the code
A specific checklist item worth calling out on its own: when an AI tool also generates docstrings or inline comments alongside the code, verify the documentation actually describes what the code does, not what the code was probably intended to do. A concrete way to check this: read the docstring first, form an expectation of the function's behavior purely from that description, then read the code and see if it actually matches. A mismatch (the docstring claims a default that the code does not actually apply, or describes a parameter the function no longer accepts after an edit) is a specific, easy-to-miss failure mode, since documentation reads as authoritative and reviewers tend to trust it rather than re-deriving behavior from the code every time.
Worked example
A pull request adds a function with an AI-generated docstring: "Returns the top N results, sorted by score descending. If N exceeds the number of available results, returns all of them." Reading only the code, the function actually raises an IndexError if N exceeds the available count, rather than gracefully returning everything, a real behavior gap between the stated contract and the implementation. Catching this before merge (by deliberately reading the docstring and the code as two separate sources of truth and comparing them) avoids the situation where downstream callers write code trusting a documented contract the function does not actually honor.
Trade-offs and pitfalls
For ML (machine learning) code specifically, add checks for data leakage, non-determinism, and version-specific API usage, since those are failure modes generic code review checklists do not usually cover. I do not trust generated code until it passes CI locally, has a readable diff, and I understand every dependency it introduces; for anything touching model training or inference, I also sanity-check that the change did not silently change resource usage (memory, latency) in a way that would only show up once it hits production traffic.
Unlock Full Question Bank
Get access to all Version Control and Developer Tooling interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.