DevOps Culture and Delivery Practices Questions
The principles and engineering culture behind modern software delivery: what DevOps means as a shared-ownership way of working and how to ease development and operations friction, adopting DevOps, infrastructure-as-code and GitOps practices across teams, the software development lifecycle and its methodology trade-offs (waterfall, agile, trunk-based development and branching agreements), engineering velocity and how it is measured without perverse incentives, development standards and how they are rolled out, enforced, excepted and justified with a business case, code review policy, turnaround expectations and handling bypasses, governance of shared delivery assets and changes to production automation, and the role of a platform team in enabling faster, safer delivery. Focuses on why teams work this way and how to drive adoption, not on how to build the machinery. Pipeline and build mechanics, deployment and rollback techniques, infrastructure tooling, reliability practice, and incident handling are covered elsewhere.
A non-technical executive asks what DevOps actually means and why it is not just buying tools or creating a DevOps team. How would you explain it, and what would you say it changes about how work and ownership flow between development and operations?
Sample Answer
Direct answer
I would tell the executive: DevOps is a way of working in which the people who build software also share responsibility for running it, so problems are found and fixed quickly instead of being passed between teams. Tools help, but they do not create that shared responsibility, and a separate "DevOps team" often just becomes a third group in the queue.
An analogy a non-technical executive can use
Imagine a restaurant where the chefs cook and then slide plates through a hatch to waiters who have never seen the kitchen. When a dish comes back wrong, the waiters blame the chefs and the chefs blame the waiters' handling. DevOps is the kitchen and the floor agreeing on one goal (the customer enjoys the meal), seeing each other's problems, and fixing them together.
In software terms: development writes the code, operations keeps it running. Historically developers were rewarded for shipping features, operations for stability, so they pulled in opposite directions and handed work over a wall.
What it changes about work and ownership
- Ownership: the team that builds a service also carries responsibility for how it behaves in production, including being reachable when it fails ("you build it, you run it"). Being reachable means being on call: carrying the pager, the alert device or phone app that wakes the responsible person when the service breaks. You might say: "The team that writes the payment code is the first to be woken when it breaks, so they build it to break less."
- Shared goals: development and operations are judged on the same outcomes, such as how often changes reach customers and how quickly service is restored, not on opposing ones.
- Feedback: production data (errors, usage) flows back to the builders quickly, so learning is in days, not quarters.
- Smaller, more frequent changes: small changes are easier to understand and undo than big ones, so more frequent releases can actually be safer.
- Blameless learning (a review after an incident that asks how the process allowed it, not who to blame): after an incident the team asks what in the system allowed it, not who to punish, so people report problems early.
- Operations expertise moves closer: reliability specialists coach teams and build shared platforms, rather than owning every release as a gatekeeper.
Why not just buy tools or create a team?
- Tools automate whatever process you already have. If the process still has a wall in it, you get a faster wall.
- A "DevOps team" that owns deployments for everyone recreates the hand-off: developers request, the team queues, and developers still do not feel production consequences.
- What helps is a platform team (a group that builds shared tools for the product teams) that provides self-service tooling, meaning teams can deploy, create environments or roll back themselves without filing a ticket, while product teams keep ownership of their services.
Worked example
Before: a payment bug appears at 2 a.m.; operations pages a developer who is asleep, the developer needs a ticket approved to deploy a fix, and the next release window is Thursday. After: the owning team gets the alert, rolls back (returns to the previous working version) with a button they control, and fixes it the next morning in a small change.
One sentence for the executive: "DevOps is about who owns the outcome, not which tools we own, so I would look for teams that can ship and fix their own work."
What I would ask the executive to look for
Not the tool list, but whether a team can release on its own, how quickly it recovers from a bad release, and whether people can describe who owns each service.
Your teams deploy several times a day but keep hitting merge conflicts and painful integration, and some engineers swear by long-lived release branches. What branching strategy would you recommend and what would you give up by choosing it?
Sample Answer
Direct answer
I would recommend trunk-based development: everyone integrates into one shared main branch (the trunk) in small changes, using branches that live a day or two at most, with unfinished work hidden behind feature flags. For a team deploying several times a day, long-lived release branches are the cause of the pain, not the protection against it. What you give up is the comfort of isolation: you need good automated tests, fast builds, and discipline about flags.
Definitions
- Trunk-based development: developers collaborate on a single branch and resist pressure to create other long-lived development branches. Short-lived branches (a couple of days) used for review are fine.
- Feature flag: a switch in the code that turns a feature on or off without redeploying, so half-finished work can be merged safely while staying invisible to users.
- Release branch: a copy of the code frozen for a release, used to stabilise it while development continues elsewhere.
Why long-lived branches hurt here
Illustrative arithmetic: 8 engineers each merge one change a day to main. The author of the branch is one of the 8 and commits to the branch, so the other 7 engineers merge to main. A branch that lives 15 working days must reconcile with 7 x 15 = 105 other changes when it finally merges. A branch that lives 2 days faces 7 x 2 = 14. That is roughly a 7.5-fold difference in drift (105 / 14). Conflicts grow with how far a branch has drifted, and integration problems appear late, when fixing them is expensive and the author has forgotten the details.
What the recommendation looks like in practice
- Branch from main, make a small change, open a pull request, merge within a day or two.
- Merge incomplete features behind a flag that defaults to off.
- The build and tests run on every merge; a broken main is fixed first, before other work.
- If a release really needs stabilising (for example a mobile app that goes through store review), cut a release branch from main just in time, fix only critical bugs there, and delete it afterwards.
What you give up
- Isolation: half-built work is in main, so you depend on flags and tests rather than separation.
- Flag discipline: flags are temporary; stale ones become tangled code and multiply the combinations to test. Give each a removal date.
- Slow, manual testing: if confidence comes from a two-week manual test cycle, daily merging is not safe until automation replaces it.
- Cultural comfort: some engineers like a long private branch; reviews must also be fast, because small changes waiting days defeat the model.
What would change my call
- If we had to maintain several shipped versions that customers run (for example on-premises software), maintained release branches are justified for those versions.
- If automated test coverage is weak, I would first invest in tests while shortening branches gradually, rather than switching overnight.
Pitfall: calling a branch 'short-lived' while it lasts three weeks. Measure the age of open branches.
Your team's pull requests wait about two days for review and delivery is suffering. What would you change in policy, tooling and team habits to get that down, and what quality risk does each change introduce?
Sample Answer
Direct answer
I would first measure where the two days go, because most of the delay is usually waiting for a reviewer to start, not time spent reviewing. Then I would change a few things in each area (policy, tooling, habits), keeping the review's purpose intact: catching problems and spreading knowledge. Every speed-up has a cost to quality, so I name it and add a safeguard.
Measure first
Illustrative: of about 48 elapsed hours, a PR (pull request: a proposed change awaiting review) spends 44 hours waiting for a first response and 4 hours in actual review and revision. Queue time is 44 / 48 = about 92 percent. If that is the shape, the fix is making review start sooner, not making reviews shorter.
Changes and the quality risk each introduces
| Area | Change | Quality risk it introduces | Safeguard |
|---|---|---|---|
| Policy | First response within one business day (Google's engineering guidance uses this maximum) | Reviewers skim to meet the clock | Response means a genuine read; approval standard stays 'improves code health' |
| Policy | One required approver for routine changes | Fewer eyes on each change | Code owners still required for sensitive areas such as payments and security |
| Policy | Size guidance: changes small enough to review in one sitting | Splitting work into pieces that do not make sense alone | Each piece must leave the code working and be explainable on its own |
| Tooling | Automatic reviewer assignment (round robin: reviewers take turns in a fixed order; code-owner based: the person or team listed as owner of the changed files is chosen) | Reviewer may lack context | Author can request a specific person for tricky changes |
| Tooling | Formatters, linters and tests run before a human looks | Over-trust: 'green means good' | Reviewers still check design and tests' meaning, not only that they pass |
| Tooling | A visible list of PRs waiting longest | Pressure leads to rubber-stamping | Treat as a team queue, never an individual ranking |
| Habits | A set review slot each morning | Interrupts the reviewer's own focus time | Short slot, protected afternoons |
| Habits | Authors write context: what, why, how tested | Extra work for authors | A short template; saves reviewer questions later |
| Habits | A quick call or pairing for complex or contentious changes | Decisions made in a call are lost | Summarise the outcome back on the PR |
Order of rollout
- Publish the response expectation and turn on automatic assignment (cheap, immediate).
- Move mechanical checks to automation, so reviews get shorter.
- Encourage smaller changes, which compounds: smaller PRs are picked up sooner.
- After a month, re-measure queue time and also check escaped defects (bugs found after merging) so the quality side is visible.
Example author-context template (pasted into every PR description)
What: Add retry with backoff to the invoice export job.
Why: Exports failed about once a day when the storage service timed out.
How tested: Unit tests for 3 retry cases; ran an export of 5,000 invoices in staging.
Look closely at: the retry limit in export_job.py (is 3 attempts enough?).
Four short lines let a reviewer start without a conversation, which is where much of the waiting disappears.
What would change my call: if reviews are slow because only one person understands the code, no policy fixes that; the real work is spreading knowledge, for example by pairing and rotating reviewers.
Pitfall: declaring victory on review speed alone. If defects reaching production rise as waiting falls, the speed-up has been bought with quality.
One engineer keeps bypassing pull-request review and pushing straight to the main branch, causing intermittent breakages. How do you handle it, both technically and with the person?
Sample Answer
Position. Fix the system first, then talk to the person. If one engineer can push to main, the real defect is that the control does not exist, not only that someone used the gap.
Technical fix (same day)
- Turn on branch protection on main: require a pull request (a proposed change others review), require at least one approval, require the automated checks to pass, and apply the rule to administrators too (so even people with admin rights cannot push around it).
- Add a CODEOWNERS file (a file naming who must review which paths) for sensitive directories.
- Make reverting cheap: a broken main is reverted first (a revert is a new change that undoes an earlier one, restoring the last good state) and debugged second, so breakages cost minutes.
- Provide a documented emergency path (a hotfix PR with one fast approver), because people bypass rules that have no legitimate fast lane.
With the person: a private 1:1 (a one-on-one meeting between a manager and one report), curious before corrective
- Open with facts, not a verdict. I would use SBI (Situation, Behavior, Impact, a feedback model from the Center for Creative Leadership): "On Tuesday afternoon (situation) you pushed commit a1b2c3 (a saved snapshot of code changes, identified by that id) directly to main without review (behavior); the nightly build failed and two teammates lost about an hour (impact)."
- Ask why. I might say: "Help me understand what was going on when you pushed that. What made going straight to main the easiest option at that moment?" Then I listen without interrupting. Common causes: reviews take a day, they work in a different time zone, they believed the change was trivial, or they were never told the rule. Each cause has a different fix; slow review is a team problem I would own.
- Agree expectations in writing: all changes go through PRs; if reviews are slow, say so in the team channel and I will unblock within a set time. I might say: "From now on every change to main goes through a pull request. If a review is stuck for more than half a working day, message me in the team channel and I will get it moving. Does that work for you? I will send this to you in writing after we talk."
- Follow up in a week and acknowledge the change.
Escalation. If it repeats after the technical block is gone and expectations are clear, it becomes a conduct issue (a behaviour problem rather than a skill gap) handled through the normal performance process (the company's formal, documented route for addressing a sustained problem with someone's work, with written expectations and follow-ups) with my own manager and HR involved. Repeated attempts to circumvent a control on purpose are treated seriously; one-off shortcuts under deadline pressure are coached.
What would change my call. If the root cause is a review queue that really is too slow, the priority shifts to fixing that queue. If the engineer is the on-call fixer of urgent incidents, the emergency path matters more than a reprimand.
How would you make the business case, with numbers, for stricter development standards and a delivery pipeline over the next year? What would you expect to improve, what could get worse, and how would you report honestly on it?
Sample Answer
Direct answer
I would present it as an investment with three kinds of return, a year-one dip that I state up front, and one line item I flag as the least certain. Standards (agreed rules for how code is written and reviewed) plus a delivery pipeline (the automated path from a merged change to production, often called CI/CD, continuous integration and continuous delivery) cost real money in year one. They pay back mainly through fewer failed releases, less manual release work, and fewer customer-facing incidents. The honest version of the case shows when it breaks even, and what happens if the shakiest assumption turns out wrong.
Worked example (all inputs illustrative, replace them with your own baseline)
Assume 120 engineers, a loaded cost (salary plus employer overheads such as benefits, equipment and office) of $150,000 per engineer per year (about $83 per hour: 150,000 / 1,800 working hours = 83.33), and 1,500 production deployments per year.
Costs
- Two platform engineers to build and run it: 2 x $150,000 = $300,000
- Tooling and licences: $60,000
- Adoption drag in the first half-year (3% of every engineer's time): 120 x $150,000 x 3% x 0.5 = $270,000
- Year-one cost: $630,000. Year two has no adoption drag: $360,000.
Benefits at full effect
- Fewer failed deployments. Change fail rate (the share of deployments that need immediate intervention, one of the DORA delivery metrics from DevOps Research and Assessment) falls from 20% to 12%. That is 300 failures down to 180, so 120 avoided. Each costs about 4 people x 5 hours = 20 hours (an illustrative assumption; use your own incident records), so 120 x 20 = 2,400 hours, about $200,000.
- Less manual release work. Each release drops from 3 hours of hand steps to 0.5 hours: 1,500 x 2.5 = 3,750 hours, about $312,500.
- Fewer customer-facing incidents. Suppose (an illustrative guess to be checked against your incident history) 25% of avoided failures would have become customer-facing incidents at $20,000 each: 120 x 25% x $20,000 = $600,000.
Full-effect benefit is $1,112,500 per year. Assume only half is realised in year one (teams adopt gradually): $556,250.
| Year 1 | Year 2 | |
|---|---|---|
| Benefit | $556,250 | $1,112,500 |
| Cost | $630,000 | $360,000 |
| Net | -$73,750 | +$752,500 |
A note on the hour-based lines: avoided hours are freed capacity, not cash. They turn into savings only if the time is redirected to other valuable work or headcount growth is avoided, so I would present them that way rather than as money returned to the budget.
The least certain line. The incident term is $600,000 of the $1,112,500, and it rests on a guess about incident cost. Drop it and the table becomes year one -$373,750, year two +$152,500, and about -$221,250 after two years. So I would say plainly: the hour-based savings alone (about $512,500 a year at full effect, and only capacity unless redirected) do not repay this in two years at this size, and the cash case without incident avoidance is weaker still; the case depends on incident avoidance, which I would validate from our own incident history before committing.
One basis caveat on the table above. It nets the adoption drag ($270,000 of engineers' time) against the benefit hours, so it is a capacity view: the drag and the freed hours are both time, not budget cash. On a cash-only view (spend of $360,000 a year against the one line that is real money, the avoided incident cost, at half effect in year one) the result is year one $300,000 - $360,000 = -$60,000, year two $600,000 - $360,000 = +$240,000, about +$180,000 over two years. The conclusion is the same either way: it stands or falls on incident avoidance, and the hour-based lines are upside to be redirected, not repayment. Also, every benefit line above comes from the pipeline; the standards half of the investment (review rules, conventions) is not priced here. I would say so and carry it as unquantified benefit (review consistency, onboarding time) rather than invent a figure.
What I expect to improve
- Change fail rate and deployment rework rate (unplanned deployments caused by a production problem) should drop.
- Change lead time (commit to production) and deployment frequency should improve once manual steps go.
- Review consistency and onboarding time for new engineers.
What could get worse
- Individual change lead time can rise at first, because gates and new conventions add steps.
- Adoption drag: some engineers lose focus time while learning.
- Workarounds: if the pipeline is slow or rigid, people route around it, and the dashboard then flatters reality.
- Standards ossify (harden): rules nobody revisits keep being followed as ritual, without anyone remembering their purpose.
Reporting honestly
- Record the baseline for each measure before starting, and write the targets down before seeing results.
- Report leading indicators (percentage of services on the pipeline, bypass count) beside outcomes (the five DORA metrics: change lead time, deployment frequency, failed deployment recovery time (how long it takes to recover from a deployment that needs immediate intervention), change fail rate, deployment rework rate).
- Say when a metric moved for other reasons (a big launch, a staffing change).
- Publish misses with the same prominence as wins, and set a quarterly checkpoint where the sponsor (the executive who funds and backs the work) can stop or resize the investment.
- Use DORA metrics for team learning, not to rank teams, since ranking invites gaming.
Unlock Full Question Bank
Get access to all 22 DevOps Culture and Delivery Practices interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.