DevOps Culture and Delivery Practices Questions
The principles and engineering culture behind modern software delivery: what DevOps means as a shared-ownership way of working and how to ease development and operations friction, adopting DevOps, infrastructure-as-code and GitOps practices across teams, the software development lifecycle and its methodology trade-offs (waterfall, agile, trunk-based development and branching agreements), engineering velocity and how it is measured without perverse incentives, development standards and how they are rolled out, enforced, excepted and justified with a business case, code review policy, turnaround expectations and handling bypasses, governance of shared delivery assets and changes to production automation, and the role of a platform team in enabling faster, safer delivery. Focuses on why teams work this way and how to drive adoption, not on how to build the machinery. Pipeline and build mechanics, deployment and rollback techniques, infrastructure tooling, reliability practice, and incident handling are covered elsewhere.
Your org currently deploys by having CI push to production, and someone proposes pull-based GitOps. What changes in ownership, audit and who can change production, and is the move worth it for a 200-engineer company?
Sample Answer
Direct answer
I would say the move is worth it for a 200-engineer company if most workloads run on Kubernetes (a system that runs applications packaged as containers, which are self-contained bundles of a program and what it needs to run, across many machines) and audit or access control matters, and I would roll it out as a staged pilot, not a big bang. GitOps means the desired state of production lives in a Git repository, and an agent inside the environment (Argo CD and Flux are the common open-source ones) continuously pulls that state and makes reality match it. That matching step is called syncing or reconciling: the agent compares what Git says should run with what is running, and fixes any difference. Today CI pushes into production, so CI holds production credentials. In the pull model CI only builds and proposes a change; production changes itself.
What changes
| CI pushes to production (today) | Pull-based GitOps | |
|---|---|---|
| Who can change production | Anyone with CI permissions or the pipeline credentials | Anyone whose change is merged to the config repo branch, plus a small break-glass group (a few named people allowed to change production directly in an emergency, with every use logged and reviewed afterwards) |
| Credentials | CI holds production access | The agent holds access from inside the environment; CI needs no production credentials (it still needs write access to the config repository) |
| Audit | Pipeline logs, which can be scattered or expire | Git history: who proposed, who approved, when |
| Drift (production differing from what was intended) | Silent until someone notices | Detected; reverted automatically only if self-healing (automatic correction of drift) is switched on, otherwise flagged for a person |
| Rollback | Re-run an older pipeline | Revert a commit |
| Ownership | The pipeline owner owns deploys | Service teams own their config; a platform team owns the agent and conventions |
One caveat on the credentials row: it describes the common in-cluster model (Flux, or Argo CD installed in each cluster). Argo CD can also run as one central instance that holds credentials for many remote clusters; that setup still removes production access from CI, but the central instance becomes a high-value credential store, so it needs the same protection CI used to need.
Is it worth it at 200 engineers?
- In favour: removing production credentials from CI is a real security gain, review of config becomes the control, and audit evidence falls out of Git rather than being assembled by hand for each audit.
- Against: it adds a layer to learn and operate; you need a decision about where config lives (usually a separate repo from application code); secrets (passwords and API keys, for example a database password) need a separate solution because they must not sit in Git in plain text, and Git history is permanent; "merged but not yet synced" is a new state engineers must understand; and if you mostly run virtual machines or serverless functions, the tooling fit is weaker.
My recommendation
Yes if the Kubernetes share is high and there is audit pressure. Run a pilot with two or three willing teams for one quarter. Judge it by whether those teams' change lead time stays about the same, whether manual production changes drop, and whether the platform team can support it without becoming the new gate. If fewer than roughly a third of services are on Kubernetes, I would stay with push and tighten credentials instead; between a third and a half, the audit pressure decides, and above half I would adopt it. These thresholds are judgement calls, not research findings. The reasoning: the agent only manages Kubernetes workloads, so the platform team runs two deployment systems whatever the share, and the question is whether the credential and audit gain justifies running the second one. Below about a third that gain covers too few services to repay the operating cost; above half the pull model becomes the main path and push the exception. Two or three teams for a quarter is enough to see several full release cycles and a few incidents while limiting the damage if the conventions turn out wrong.
What would change my call: a compliance mandate for an audit trail (accelerates), or a team of five platform engineers already stretched (delays).
Your team's pull requests wait about two days for review and delivery is suffering. What would you change in policy, tooling and team habits to get that down, and what quality risk does each change introduce?
Sample Answer
Direct answer
I would first measure where the two days go, because most of the delay is usually waiting for a reviewer to start, not time spent reviewing. Then I would change a few things in each area (policy, tooling, habits), keeping the review's purpose intact: catching problems and spreading knowledge. Every speed-up has a cost to quality, so I name it and add a safeguard.
Measure first
Illustrative: of about 48 elapsed hours, a PR (pull request: a proposed change awaiting review) spends 44 hours waiting for a first response and 4 hours in actual review and revision. Queue time is 44 / 48 = about 92 percent. If that is the shape, the fix is making review start sooner, not making reviews shorter.
Changes and the quality risk each introduces
| Area | Change | Quality risk it introduces | Safeguard |
|---|---|---|---|
| Policy | First response within one business day (Google's engineering guidance uses this maximum) | Reviewers skim to meet the clock | Response means a genuine read; approval standard stays 'improves code health' |
| Policy | One required approver for routine changes | Fewer eyes on each change | Code owners still required for sensitive areas such as payments and security |
| Policy | Size guidance: changes small enough to review in one sitting | Splitting work into pieces that do not make sense alone | Each piece must leave the code working and be explainable on its own |
| Tooling | Automatic reviewer assignment (round robin: reviewers take turns in a fixed order; code-owner based: the person or team listed as owner of the changed files is chosen) | Reviewer may lack context | Author can request a specific person for tricky changes |
| Tooling | Formatters, linters and tests run before a human looks | Over-trust: 'green means good' | Reviewers still check design and tests' meaning, not only that they pass |
| Tooling | A visible list of PRs waiting longest | Pressure leads to rubber-stamping | Treat as a team queue, never an individual ranking |
| Habits | A set review slot each morning | Interrupts the reviewer's own focus time | Short slot, protected afternoons |
| Habits | Authors write context: what, why, how tested | Extra work for authors | A short template; saves reviewer questions later |
| Habits | A quick call or pairing for complex or contentious changes | Decisions made in a call are lost | Summarise the outcome back on the PR |
Order of rollout
- Publish the response expectation and turn on automatic assignment (cheap, immediate).
- Move mechanical checks to automation, so reviews get shorter.
- Encourage smaller changes, which compounds: smaller PRs are picked up sooner.
- After a month, re-measure queue time and also check escaped defects (bugs found after merging) so the quality side is visible.
Example author-context template (pasted into every PR description)
What: Add retry with backoff to the invoice export job.
Why: Exports failed about once a day when the storage service timed out.
How tested: Unit tests for 3 retry cases; ran an export of 5,000 invoices in staging.
Look closely at: the retry limit in export_job.py (is 3 attempts enough?).
Four short lines let a reviewer start without a conversation, which is where much of the waiting disappears.
What would change my call: if reviews are slow because only one person understands the code, no policy fixes that; the real work is spreading knowledge, for example by pairing and rotating reviewers.
Pitfall: declaring victory on review speed alone. If defects reaching production rise as waiting falls, the speed-up has been bought with quality.
You are setting up code review for a team that has none. What goes into the policy (who must approve, what gets checked, how long reviews may take), and how do you keep it from turning into a bottleneck?
Sample Answer
Direct answer
Code review means a teammate other than the author reads a proposed change (a pull request, or PR) before it joins the shared codebase. For a team starting from nothing I would write a short, one-page policy with three parts: who approves, what is checked, and how fast reviews must start. Then I would put the speed controls in place from day one, because a review process with no time expectation turns into a queue within weeks.
Draft policy (the excerpt I would publish to the team)
| Topic | Rule |
|---|---|
| Who approves | One approval from a teammate who is not the author. Changes in sensitive areas (payments, authentication, database schema) also need approval from the named code owner for that area (the person or group officially responsible for that part of the code, listed in a file the hosting tool reads). Authors never approve their own change. |
| What reviewers check | Does it work and is it tested? Is the design understandable and in the right place? Would a new teammate be able to maintain it? Are there security or data-privacy implications? |
| What reviewers do not check by hand | Formatting and style: a formatter (a tool that rewrites code into one standard layout) and a linter (a tool that flags likely mistakes and rule violations) run automatically, so humans spend time on design and correctness. |
| Approval standard | Approve once the change improves the overall health of the codebase, even if it is not how the reviewer would have written it. Style preferences are marked as optional comments ("nit", short for nitpick, meaning a minor point the author may ignore). In plain terms: if the change makes the code better than before and breaks nothing, do not hold it up over taste. This follows Google's published engineering practices. |
| Speed | First response within one business day (Google's guidance is that one business day is the maximum). First response means a real read, not a rubber stamp, and it is not the same as final approval. |
| Size | Aim for changes a reviewer can understand in one sitting; split larger work into steps that each leave the code working. |
Keeping review from becoming a bottleneck
- Automate the mechanical checks (tests, formatting, linting) so a human is never the one catching a missing semicolon.
- Spread the load: rotate or auto-assign reviewers instead of letting the most senior person become the gate.
- Small changes: a short change is reviewed faster and more carefully than a large one.
- Make the wait visible: a weekly look at how long PRs sit before a first response, treated as a team problem, not a ranking of individuals.
- Allow approval with minor follow-ups: a reviewer who trusts the author can approve while leaving small comments to be fixed.
Worked example (illustrative)
A team of 6 engineers each opens about one PR a day, so there are about 6 reviews to do per day. With 5 possible reviewers per PR (everyone except the author) that is 6 / 5 = 1.2 reviews each per day. That is manageable only if changes are small and reviews are shared evenly; if one lead reviews everything, that person faces 6 per day and becomes the bottleneck. In hours, at about 30 minutes per review: 1.2 reviews is 36 minutes a day each, but 6 reviews is 3 hours a day for the lead, which is why work starts to queue.
Pitfalls
- Two required approvers everywhere sounds safer but doubles the waiting for little added protection on routine changes.
- Using review turnaround to rank people encourages hasty approvals.
- A policy nobody can see enforced (no branch protection, the repository setting that blocks merging until the required approvals are in) lasts only until the first deadline.
A stakeholder asks your team to skip code review to hit a deadline. How do you respond, and what, if anything, would you relax to keep releases safe?
Sample Answer
Position. I would not skip review, but I would not answer a flat "no" either. I would turn "skip review" into "what is the cheapest safe way to hit the date". Code review means a teammate reads a change before it merges, to catch defects, share knowledge and keep to team standards.
Steps, in order
- Ask what the deadline is and what happens if it slips. A contractual launch date and a "would be nice by Friday" need different answers. Ask the stakeholder what they are actually worried about: usually it is waiting time, not review itself.
- Find where the time really goes. Often the delay is a pull request (PR, a proposed change awaiting review) sitting a day before anyone looks, not the 20 minutes of reading.
- Offer relaxations that keep the safety.
- Review turnaround: a same-day review rota (a schedule naming who is on duty to review each day) for the sprint (illustrative: first response within 4 hours).
- Smaller PRs: several 200-line changes are faster and safer to review than one 1,500-line change.
- Risk tiers: a deep review (two reviewers) only for code touching payments, authentication or data migrations; one reviewer for ordinary changes; a post-merge review (the change merges first and a teammate reads it afterwards, within one business day) for text or config tweaks. Risk tiers means sorting changes by how much damage a mistake could do and matching review effort to that.
- Pairing as live review: pairing means two people writing the risky part together at one screen, so the second person's reading happens live and replaces a later review.
- Keep the automated checks (tests, linting, and secret scanning, which automatically detects passwords or keys accidentally committed into code): they cost no human time.
- Decouple merging from releasing: put unfinished work behind a feature flag (a switch that hides code from users), so a merge does not mean users see it.
- Cut scope instead of cutting checks, which is the trade the stakeholder can own.
- Write down what was accepted. One line in the ticket: which tier was relaxed, who approved, when the post-merge review happens.
Worked example (illustrative). Launch in 10 working days, 6 PRs left, each building on the previous one so they cannot be reviewed in parallel. Assume an 8-hour day and about 2 hours of actual reading and fixing per PR (0.25 day). With a one-day wait for a first response, each PR takes 1 + 0.25 = 1.25 days, so 6 x 1.25 = 7.5 days of the 10. With a 4-hour rota (0.5 day wait), each takes 0.5 + 0.25 = 0.75 days, so 6 x 0.75 = 4.5 days. The rota saves 3 days without removing a single review. The two PRs touching billing keep two reviewers; the four copy and layout PRs get one reviewer, except that a change that only edits text may merge first and be read within one business day.
What would change my call. A live outage or a legal deadline justifies an emergency path: merge with one named approver and a mandatory post-merge review the next business day. I would never relax review on security-sensitive code without a compensating control.
Closing the loop. Tell the stakeholder what they get (date held) and what the risk is in plain words; after launch, report whether the relaxed tiers produced defects.
What I would actually say to the stakeholder (illustrative)
"I can hold the date, and I do not want to skip review, because review is what catches the bugs that would cost us a bad launch week. What I can change is the waiting. Today a change sits about a day before anyone looks at it. For the next 10 days I will set up a review rota so every change gets a first look within 4 hours, and I will split the work into smaller pieces. That saves about 3 days. The two billing changes keep two reviewers because a mistake there costs real money. The four copy and layout changes get one reviewer, and a change that only edits text is read within one business day after it merges. The remaining risk is that a small mistake in a text change reaches users a few hours before a second person reads it; I will have that read within one business day and tell you if anything turns up. Is there a part of the launch you are most worried about, so I can make sure it is covered?"
A non-technical executive asks what DevOps actually means and why it is not just buying tools or creating a DevOps team. How would you explain it, and what would you say it changes about how work and ownership flow between development and operations?
Sample Answer
Direct answer
I would tell the executive: DevOps is a way of working in which the people who build software also share responsibility for running it, so problems are found and fixed quickly instead of being passed between teams. Tools help, but they do not create that shared responsibility, and a separate "DevOps team" often just becomes a third group in the queue.
An analogy a non-technical executive can use
Imagine a restaurant where the chefs cook and then slide plates through a hatch to waiters who have never seen the kitchen. When a dish comes back wrong, the waiters blame the chefs and the chefs blame the waiters' handling. DevOps is the kitchen and the floor agreeing on one goal (the customer enjoys the meal), seeing each other's problems, and fixing them together.
In software terms: development writes the code, operations keeps it running. Historically developers were rewarded for shipping features, operations for stability, so they pulled in opposite directions and handed work over a wall.
What it changes about work and ownership
- Ownership: the team that builds a service also carries responsibility for how it behaves in production, including being reachable when it fails ("you build it, you run it"). Being reachable means being on call: carrying the pager, the alert device or phone app that wakes the responsible person when the service breaks. You might say: "The team that writes the payment code is the first to be woken when it breaks, so they build it to break less."
- Shared goals: development and operations are judged on the same outcomes, such as how often changes reach customers and how quickly service is restored, not on opposing ones.
- Feedback: production data (errors, usage) flows back to the builders quickly, so learning is in days, not quarters.
- Smaller, more frequent changes: small changes are easier to understand and undo than big ones, so more frequent releases can actually be safer.
- Blameless learning (a review after an incident that asks how the process allowed it, not who to blame): after an incident the team asks what in the system allowed it, not who to punish, so people report problems early.
- Operations expertise moves closer: reliability specialists coach teams and build shared platforms, rather than owning every release as a gatekeeper.
Why not just buy tools or create a team?
- Tools automate whatever process you already have. If the process still has a wall in it, you get a faster wall.
- A "DevOps team" that owns deployments for everyone recreates the hand-off: developers request, the team queues, and developers still do not feel production consequences.
- What helps is a platform team (a group that builds shared tools for the product teams) that provides self-service tooling, meaning teams can deploy, create environments or roll back themselves without filing a ticket, while product teams keep ownership of their services.
Worked example
Before: a payment bug appears at 2 a.m.; operations pages a developer who is asleep, the developer needs a ticket approved to deploy a fix, and the next release window is Thursday. After: the owning team gets the alert, rolls back (returns to the previous working version) with a button they control, and fixes it the next morning in a small change.
One sentence for the executive: "DevOps is about who owns the outcome, not which tools we own, so I would look for teams that can ship and fix their own work."
What I would ask the executive to look for
Not the tool list, but whether a team can release on its own, how quickly it recovers from a bad release, and whether people can describe who owns each service.
Unlock Full Question Bank
Get access to all 27 DevOps Culture and Delivery Practices interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.