Version Control and Developer Tooling Questions
The everyday toolchain of software work: version control with Git (branching, merging, rebasing, conflict resolution, using git bisect to find a regression), command-line and shell proficiency for day-to-day navigation, log inspection, and troubleshooting, IDE and editor workflows, build systems and package/dependency management (npm, Maven, pip, Gradle, CocoaPods, and embedded/cross-compilation toolchains), and the growing practice of AI-assisted coding: using, reviewing, and verifying AI-generated code and tests. Deliberately generic across languages and stacks; language- and domain-specific frameworks live in their own categories. This topic covers a developer's individual command of these tools, not: writing durable shell automation and glue scripts (Shell Scripting and Automation owns that), producing, versioning, and publishing build artifacts or container images (Build Automation and Artifact Management owns that), release cadence and change governance (Release Management and Change Control owns that), or diagnosing a live production incident end to end (Performance Troubleshooting and Incident Response and the Observability topics own that).
You need an AI assistant to generate a feature engineering function from product requirements and a data dictionary. How would you phrase the prompt to minimize ambiguity and prevent leakage or schema drift?
Sample Answer
Direct answer
I would make the prompt explicit about inputs, outputs, and forbidden behavior rather than describing the feature in prose and hoping the assistant infers the constraints. Ambiguous prompts tend to produce elegant-looking but unsafe feature logic, code that runs and looks reasonable but silently uses information that would not actually be available at prediction time (leakage), or assumes a column shape that will not hold once the upstream data changes (schema drift, meaning the input data's structure shifts over time in ways the function was never told to expect).
Structured elaboration
A well-specified prompt for this task includes:
- The exact business goal and feature definition taken directly from the product requirements, not paraphrased, so there is one source of truth for what the feature is supposed to represent.
- The full data dictionary: column names, types, allowed values, and timestamp semantics (when is each column actually populated relative to the event being predicted).
- An explicit role for every column: which are raw inputs, which are derived features, which are identifiers, and which are target-adjacent fields that must never be used as inputs.
- A stated leakage policy: no fields populated after the prediction point, no timestamps from the future relative to the prediction time, and no use of the label or anything derived from it.
- Schema-drift constraints: preserve a documented output column order and schema, and specify exactly how the function should behave on missing or unseen fields (fail loudly, default safely, or something in between), rather than leaving that undefined.
- The desired output shape: a pure function, a scikit-learn transformer (a reusable object implementing
fit/transformso the same logic runs identically in training and serving), or a plain function with accompanying tests, whichever matches how the feature will actually be consumed downstream. - An explicit instruction to state assumptions, so the assistant lists what it is inferring rather than silently inventing a new derived column that was never asked for.
Scikit-learn pipeline schema alignment
When the target output is specifically a scikit-learn Pipeline (a container that chains several preprocessing/model steps so they run in a fixed sequence) or ColumnTransformer (a step that applies different preprocessing to different columns, then combines the results) step, rather than a standalone function, the prompt needs one more layer of specificity: name exactly which columns the transformer will see at fit time (when it learns parameters from the training data, like a mean or a vocabulary) versus transform time (when it applies those already-learned parameters to new data, without learning anything new), whether the transformer needs to handle a column that is present in training but absent at serving time, and whether the output feature names and order need to stay stable across pipeline versions, since a churn-prediction pipeline that silently reorders its output columns between training and serving is a schema-drift bug that will not show up as an error, only as quietly wrong predictions.
Worked example
Bad prompt: "Write a function that computes days since last purchase as a feature." Better prompt: "Write a Python function days_since_last_purchase(df) that takes a DataFrame with columns customer_id (string), purchase_ts (UTC timestamp, may be null for customers with no purchase history), and prediction_ts (UTC timestamp, the point in time this feature is being computed for). Return one row per customer_id with a column days_since_last_purchase (integer, null if the customer has no purchase history before prediction_ts). Only use purchases where purchase_ts < prediction_ts; never use a purchase at or after prediction_ts, since that would leak future information. Preserve customer_id as the join key and do not drop customers with no purchase history. State any assumptions about timezone handling explicitly." The second version pins the leakage boundary (purchase_ts < prediction_ts), the null-handling behavior, and the join key, none of which the first version specified.
Trade-offs and pitfalls
Writing this level of detail takes real time upfront, which can feel like overkill for a throwaway exploratory feature. The trade-off is worth it for anything that will train a production model, because the cost of finding leakage after a model is already deployed and looking suspiciously good is far higher than the cost of a careful prompt. For quick exploration, a lighter-weight prompt is fine as long as everyone treats the output as scratch code, not production-ready.
Describe how to redirect stdout and stderr to a file while also printing stdout to the terminal. Give the exact shell command(s) you would use in bash to: (1) capture both stdout and stderr to /tmp/run.log, and (2) simultaneously see stdout in the terminal. Explain why your approach works.
Sample Answer
Direct answer
Two building blocks solve this: 2>&1 merges stderr into stdout, and tee duplicates a stream to both the terminal and a file. Capture-only needs command > /tmp/run.log 2>&1. Capture-and-still-see-it-live needs command 2>&1 | tee /tmp/run.log.
Structured elaboration
# (1) capture both streams to the file, nothing shown on the terminal
your_command > /tmp/run.log 2>&1
# (2) capture both streams AND still see them live on the terminal
your_command 2>&1 | tee /tmp/run.log
Key points:
- File descriptor 1 (fd 1) is stdout, file descriptor 2 (fd 2) is stderr.
2>&1means "point fd 2 at wherever fd 1 currently points." - Order matters.
command > file.log 2>&1first redirects stdout to file.log, then makes stderr point at that same target, so both land in the file.command 2>&1 > file.logdoes the opposite: stderr is pointed at the terminal (wherever stdout currently was), and only then is stdout redirected to the file, so stderr stays on the terminal. This ordering mistake is the single most common bug with this idiom. teereads its own stdin and writes it to both stdout and the named file. That is why2>&1 | tee file.logworks: a pipe only ever forwards stdout by default, so stderr has to be merged into stdout with2>&1before it reaches the pipe if you wantteeto see it too.
Worked example
$ run_demo() { echo "stdout line"; echo "stderr line" >&2; }
$ run_demo 2>&1 | tee /tmp/run.log
stdout line
stderr line
$ cat /tmp/run.log
stdout line
stderr line
Both lines appear on the terminal and in the file, in the order the process wrote them. This was run exactly as shown.
Trade-offs & pitfalls
- Complexity is constant per line, this is OS-level file-descriptor plumbing, not a data scan, so cost tracks bytes written, not file size on disk.
- Ordering under a pipe: stdout is usually line-buffered when connected to a terminal, but becomes fully block-buffered once piped to
tee. If a process interleaves fast stdout and stderr writes, the captured order can differ slightly from an unpiped run. Force line buffering withstdbuf -oL your_commandwhen strict interleaving order matters. teeoverwrites the log file by default each run; usetee -a /tmp/run.logto append instead.- A plain redirect does not keep a process alive after you log out of SSH by itself, pair it with
nohupor a terminal multiplexer if the command needs to survive the session ending.
Design a reproducible AI-assisted coding workflow where prompts, generated snippets, human edits, tests, and approvals are all traceable in version control. What artifacts would you store, and how would you review them during code review?
Sample Answer
Direct answer
I would treat AI-assisted changes like an auditable build pipeline rather than just a diff: version the prompt, the model's raw output, the human's edits on top of it, the tests, and the approval as linked artifacts, not only the final merged code. The goal is that months later someone can reconstruct not just what the code does, but why it ended up written this way.
Structured elaboration
Artifacts to store
- The prompt text, model name and version, and any relevant parameters (temperature, system instructions), with a timestamp.
- The raw generated snippet or diff, captured before any human edits, as its own artifact.
- The human-edited version on top of that, so the two diffs can be compared: what did the model produce, and what did a person change.
- The test plan and actual test results for the change.
- Approval metadata: reviewer identity, timestamp, and a short reason for sign-off.
Where this lives
- A pull request template with explicit sections for the prompt, the raw generated output, and the human changes on top, so this is a normal part of every AI-assisted PR rather than a special process.
- A lightweight audit folder (for example
ai-audit/) or PR-attached transcript for the raw prompt and output, since pasting a long prompt directly into a commit message is unwieldy. - A commit trailer convention, for example
AI-Assisted-By: <model-name>, mirroring how Git already supports aCo-authored-by:trailer, sogit log --grepcan find every AI-assisted change later without needing external tooling.
Code review approach
- Review the raw AI diff and the human's edits on top of it separately: what did the model get right, what did the human have to fix, and does that pattern reveal something worth flagging (a class of mistake this model or this prompt style keeps making)?
- Confirm the tests actually exercise the AI-introduced logic and its edge cases, not just that line coverage went up. Coverage without meaningful assertions is a false sense of safety.
- Treat a missing prompt or missing raw-output artifact as a reason to request changes, the same way a PR with no description would be, because it breaks the traceability the whole workflow exists for.
Notebook and script-specific handling. Notebooks do not diff cleanly, since the underlying file format is JSON that embeds cell outputs and execution counts alongside code, so a tiny logic change can produce a huge, noisy diff. I would either (a) require a pre-commit hook that strips outputs before commit (a tool like nbstripout) so the diff that lands in review is just the code, or (b) treat the notebook purely as scratch space and require the reproducible logic to land in a plain .py module before merge. Either way, the prompt log for a notebook exploration session goes into the same audit artifact as script-based work, not left as an ephemeral comment in a cell that gets overwritten the next time someone re-runs it.
Hyperparameter-tuning-specific handling. Reproducibility for a tuning run needs more than the workflow above: every source of randomness (data split, model initialization, the tuner's own sampling) needs a fixed, recorded seed; every trial's configuration and resulting metric needs to be logged to an experiment tracker (for example MLflow or Weights and Biases), not just the winning configuration, since "why didn't we pick option B" is a real question later; and the environment itself needs pinning (a dependency lockfile or a container image digest), because a library version bump can silently shift results even with every seed fixed.
Worked example
A team asks an AI assistant to draft a nightly batch inference job (a scheduled process that scores a whole day's accumulated data at once, rather than one prediction at a time) that scores yesterday's transactions for fraud risk. The prompt, model, and generated draft are logged in the PR's audit section. A human engineer adds input validation and a --dry-run flag the AI draft omitted; that diff is visible separately from the AI's original output. Tests cover the empty-input case and a schema-mismatch case (a column the model expects is missing). Before deploy, a reviewer approves after comparing yesterday's predictions from a shadow run of the new job (running the new job alongside the real one on the same data, without letting its output affect anything, purely to compare results) against the existing production job's predictions on the same data, checking they agree closely enough to trust. A rollback step is wired in before deploy: the last several model artifacts are retained, and reverting to the prior artifact and config is a single documented action, so if next-day monitoring shows a skew in prediction distribution, the team can revert immediately rather than debugging under pressure.
Trade-offs and pitfalls
Documenting every AI interaction in full has real overhead, so the rigor should scale with blast radius: a one-line lint fix does not need a full prompt audit trail, but a change to a training pipeline or a production-serving path does. A prompt log alone proves traceability, not correctness, so it never substitutes for tests. And prompts that include pasted context can themselves contain sensitive data (customer records, internal credentials), so the audit artifacts need the same access control as the data they might reference, not looser controls just because they are "documentation."
If an AI assistant proposes a novel model architecture or training trick that you do not fully understand, how do you decide whether to prototype it, ask for more evidence, or reject it outright?
Sample Answer
Direct answer
I use a simple decision framework built around four questions: how novel is the idea really, what is the realistic expected gain, what is the risk if it goes wrong, and how much evidence actually supports it right now. Not fully understanding an idea is not, by itself, a reason to reject it, engineers accept genuinely useful ideas they do not deeply understand all the time, but it is a reason to demand more evidence before committing real time or production risk to it.
Structured elaboration
Prototype it when
- The idea is compatible with existing infrastructure and the current data pipeline, so a test does not require a parallel system just to try it.
- The potential upside is large enough to justify a short, bounded experiment.
- A clean baseline and a clear success metric already exist, so the experiment can produce an unambiguous answer.
Ask for more evidence when
- The method is genuinely novel, but the explanation for why it should help is thin or hand-wavy.
- The assistant cannot articulate why this should outperform the current baseline, only that it might.
- The trick looks fragile, expensive to run, or hard to debug if it misbehaves in production.
Reject outright when
- It depends on unsupported infrastructure, unmaintainable custom operations, or math that nobody on the team, including the assistant, can actually explain clearly.
- It adds real training or serving complexity without a measurable, demonstrated gain to justify that cost.
- It conflicts with a compliance, latency, or reliability constraint the system already has to meet.
Worked example
Suppose the AI assistant proposes a new attention variant, claiming it should improve model quality, along with a rough implementation. Rather than accepting or rejecting on the description alone: I would ask for the specific mechanism and why it should help this particular task, not just "attention variants generally help." I would ask what its memory and compute cost looks like relative to the current approach, since a change that meaningfully increases training cost needs a correspondingly meaningful expected gain to be worth prototyping at all. Then I would run a small, controlled ablation: same data, same baseline configuration, only the attention mechanism changed, with a predetermined evaluation metric decided before running the experiment, not chosen afterward to make the result look better. If the assistant cannot explain the mechanism's failure modes (when would this plausibly hurt rather than help), that gap itself is useful information, since it tells me the risk of adopting it blind, even if the experiment result looks good, is higher than for a well-understood technique.
Trade-offs and pitfalls
A staff-level habit worth protecting: research curiosity is valuable, and "I don't fully understand this yet" should not automatically mean "no." But production systems do not reward cleverness for its own sake, they reward reliability and explainability when something eventually goes wrong. The failure mode to actively guard against is adopting something because a benchmark number looked good without understanding why, since that leaves you unable to debug it later, unable to predict when it will fail on a different distribution of data, and unable to explain it to anyone who asks. Bounding the cost of the experiment (a small, cheap ablation before any real commitment) is what makes it safe to stay curious without exposing production to something nobody on the team can reason about.
You accidentally deleted a branch on the remote that contained commits you need. Describe how you'd recover the branch using git reflog and remote commit hashes, including the exact commands to recreate the branch on the remote.
Sample Answer
Direct answer
git reflog is a local, per-repository log of every place HEAD and branch tips have pointed, kept for a limited time as a personal safety net; it is not shared or pushed to the remote. It lets you find the commit hash the deleted branch used to point at, as long as some local clone had that branch checked out or fetched recently. From there, recreating the branch and pushing it restores the remote.
Structured elaboration and worked example
- Find the commit hash locally:
git reflog show feature-x
This lists recent positions of feature-x's tip on this machine, most recent first, each with a commit hash. If the branch was never checked out or fetched on your machine, your local reflog won't have it; check whether a teammate's clone or your CI system's checkout does.
- If no local reflog has it, check whether the remote's own ref list or CI logs retained the commit hash (git's short-form identifier for a commit, sometimes called a SHA) even after the branch pointer itself was deleted:
git ls-remote origin | grep feature-x
This only helps if the ref still resolves; once it's truly deleted it won't show up here, which is why the reflog approach is usually the real answer.
- Confirm you have the right commit before recreating anything:
git show abc1234
- Recreate the branch locally at that commit and push it back to the remote:
git branch feature-x abc1234
git push origin feature-x
Trade-offs and pitfalls
The reflog only exists locally and only for a limited retention window; git periodically expires old reflog entries, and git gc can prune commits that are no longer reachable from anything, including an expired reflog entry. The longer it's been since the branch was deleted, the less likely this recovery path works. If the branch had already been merged before deletion, its commits are also still reachable from whatever branch it was merged into, which is a second, often easier recovery path worth checking first. Recreating and pushing is safe as described, but if the remote branch somehow still exists and has diverged, pushing over it needs --force-with-lease, not a bare --force, to avoid clobbering someone else's newer work.
Unlock Full Question Bank
Get access to all Version Control and Developer Tooling interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.