DevOps Culture and Delivery Practices Questions
The principles and engineering culture behind modern software delivery: what DevOps means as a shared-ownership way of working and how to ease development and operations friction, adopting DevOps, infrastructure-as-code and GitOps practices across teams, the software development lifecycle and its methodology trade-offs (waterfall, agile, trunk-based development and branching agreements), engineering velocity and how it is measured without perverse incentives, development standards and how they are rolled out, enforced, excepted and justified with a business case, code review policy, turnaround expectations and handling bypasses, governance of shared delivery assets and changes to production automation, and the role of a platform team in enabling faster, safer delivery. Focuses on why teams work this way and how to drive adoption, not on how to build the machinery. Pipeline and build mechanics, deployment and rollback techniques, infrastructure tooling, reliability practice, and incident handling are covered elsewhere.
A non-technical executive asks what DevOps actually means and why it is not just buying tools or creating a DevOps team. How would you explain it, and what would you say it changes about how work and ownership flow between development and operations?
Sample Answer
Direct answer
I would tell the executive: DevOps is a way of working in which the people who build software also share responsibility for running it, so problems are found and fixed quickly instead of being passed between teams. Tools help, but they do not create that shared responsibility, and a separate "DevOps team" often just becomes a third group in the queue.
An analogy a non-technical executive can use
Imagine a restaurant where the chefs cook and then slide plates through a hatch to waiters who have never seen the kitchen. When a dish comes back wrong, the waiters blame the chefs and the chefs blame the waiters' handling. DevOps is the kitchen and the floor agreeing on one goal (the customer enjoys the meal), seeing each other's problems, and fixing them together.
In software terms: development writes the code, operations keeps it running. Historically developers were rewarded for shipping features, operations for stability, so they pulled in opposite directions and handed work over a wall.
What it changes about work and ownership
- Ownership: the team that builds a service also carries responsibility for how it behaves in production, including being reachable when it fails ("you build it, you run it"). Being reachable means being on call: carrying the pager, the alert device or phone app that wakes the responsible person when the service breaks. You might say: "The team that writes the payment code is the first to be woken when it breaks, so they build it to break less."
- Shared goals: development and operations are judged on the same outcomes, such as how often changes reach customers and how quickly service is restored, not on opposing ones.
- Feedback: production data (errors, usage) flows back to the builders quickly, so learning is in days, not quarters.
- Smaller, more frequent changes: small changes are easier to understand and undo than big ones, so more frequent releases can actually be safer.
- Blameless learning (a review after an incident that asks how the process allowed it, not who to blame): after an incident the team asks what in the system allowed it, not who to punish, so people report problems early.
- Operations expertise moves closer: reliability specialists coach teams and build shared platforms, rather than owning every release as a gatekeeper.
Why not just buy tools or create a team?
- Tools automate whatever process you already have. If the process still has a wall in it, you get a faster wall.
- A "DevOps team" that owns deployments for everyone recreates the hand-off: developers request, the team queues, and developers still do not feel production consequences.
- What helps is a platform team (a group that builds shared tools for the product teams) that provides self-service tooling, meaning teams can deploy, create environments or roll back themselves without filing a ticket, while product teams keep ownership of their services.
Worked example
Before: a payment bug appears at 2 a.m.; operations pages a developer who is asleep, the developer needs a ticket approved to deploy a fix, and the next release window is Thursday. After: the owning team gets the alert, rolls back (returns to the previous working version) with a button they control, and fixes it the next morning in a small change.
One sentence for the executive: "DevOps is about who owns the outcome, not which tools we own, so I would look for teams that can ship and fix their own work."
What I would ask the executive to look for
Not the tool list, but whether a team can release on its own, how quickly it recovers from a bad release, and whether people can describe who owns each service.
Leadership wants a single 'developer productivity' number for the quarter. How would you design a way to measure and report engineering velocity that teams cannot easily game, and what would you refuse to measure?
Sample Answer
Direct answer
I would not give leadership one number. I would give them a one-page report of a few balanced measures, tracked as trends per team, and I would say plainly why a single productivity number does not exist. The principle behind it is Goodhart's law: when a measure becomes a target, people optimise the measure instead of the thing it stood for. So the design goal is measures that are hard to improve without actually improving delivery.
The design
- Measure the system, not individuals. Report at team level. The DORA delivery measures (DORA is DevOps Research and Assessment, a long-running research programme on software delivery; the measures are change lead time, deployment frequency, failed deployment recovery time, change fail rate, deployment rework rate) come from deploy and incident records rather than self-reports, which makes them harder to inflate than story points. They are not ungameable, though: deployment frequency can be inflated by splitting changes into trivial deploys, and change fail rate can be lowered by reclassifying incidents, so the counting rules (what is a deploy, what is a failure) are fixed in advance and the measures are always read in pairs.
- Pair opposing measures. Never report speed without stability: deployment frequency next to change fail rate, lead time next to rework rate. Gaming one shows up in its partner.
- Add the human side. The SPACE framework (a 2021 paper by researchers including Nicole Forsgren) argues that developer productivity spans several dimensions, not one: Satisfaction and well-being, Performance, Activity, Communication and collaboration, and Efficiency and flow, and that satisfaction and well-being belong in the picture. A short quarterly developer survey (is it easy to get changes out, how much time is lost waiting) catches problems no pipeline log shows.
- Report trends against each team's own baseline. The first quarter is a baseline with no targets. Targets set before people trust the numbers get gamed.
- Keep the data away from performance reviews. The moment metrics decide pay or ranking, they stop being honest.
What I would refuse to measure, and why
- Lines of code or commit counts: they reward verbose changes and punish deletions and careful work.
- Story points or velocity compared across teams (a story point is a team's own rough unit for how big a piece of work is, and sprint velocity is the points finished per sprint, a short fixed work period of usually one to four weeks): points are a team's private estimating unit, and pressure to raise them just inflates estimates.
- Individual pull request counts or review counts: encourages splitting work into trivia and rubber-stamp reviews.
- Hours online, keystrokes or time in the office: they measure presence, not value.
- Bug counts per developer: it punishes people who take on hard, risky work and discourages honest bug reports.
Worked example of gaming (illustrative)
Leadership sets a goal: raise velocity. A team's sprint velocity goes from 30 points to 45 points, a 50 percent rise. Deployment frequency and lead time are unchanged. The conclusion is that the team re-estimated the same work as bigger, not that it got faster. Because the report also contained delivery measures, the discrepancy is visible rather than celebrated.
What the one-page report looks like (illustrative)
Team A Q1 baseline Q2 Trend
Deploys/wk 6 9 up
Fail rate 8% 8% steady
Lead time 5.5 days 4 days down (better)
Rework 5% 5% steady
Recovery 2 hours 1.5 h down (better)
Survey 3.4 / 5 3.8 up
Read it as pairs: deploys went up with fail rate steady, and lead time fell with rework steady, so speed did not cost stability or create extra repair work, and the survey agrees. If deploys rose to 12 while fail rate went to 16% and the survey fell to 3.0, the team is speeding past its safety. A survey score is read as a trend on a five-point scale, never as a target to push upward.
What I would tell leadership who still want one number
I would offer a summary headline that is explicitly composed of the pairs (for example, "deploys per week with fail rate held steady, and lead time falling") rather than a score, and I would agree beforehand what decisions the report will and will not be used for: guiding where to invest in tooling or process, never ranking people.
Pitfalls: starting from what is easy to count; changing the definitions mid-quarter; publishing a team league table.
Engineering wants to move to trunk-based development, and the product side worries about half-finished work reaching users. Walk through the team agreements, review changes and release-practice changes needed, and how you would get both sides comfortable.
Sample Answer
Position. Yes to trunk-based development, with one condition: unfinished work must be invisible to users. Trunk-based development means developers integrate into a single shared branch (the trunk) at least daily and avoid long-lived branches; short-lived branches exist only for review and checks. Feature flags (switches that hide unfinished code from users) are what let product trust it.
Team agreements
- Merge to trunk at least once a day, in small changes (aim for a PR, a pull request or proposed change awaiting review, that a reviewer can read in minutes).
- Trunk stays green (the automated build and tests all pass): if the automated build breaks, fixing or reverting it is the team's first priority.
- Incomplete work goes behind a flag that defaults to off.
- Flags have an owner and a removal date; a stale flag is a ticket.
Review changes. Smaller PRs and a fast review response (hours, not days). Review checks the flag defaults and tests; large refactors are split. Pairing is encouraged for the big pieces.
Release-practice changes. Separate deploy (code reaches production) from release (users see it). Product, not engineering, decides when a flag turns on, for a small percentage first, then everyone. Rollback becomes a flag switch.
Getting both sides comfortable
- For product: "half-finished work cannot reach users because it is off by default; you control the switch." Show a demo in staging (a production-like environment used for testing) with the flag on.
- For engineering: trial for six weeks on one team, with the current long-lived branch pain (late merge conflicts) as the baseline.
Worked example (illustrative). A new checkout page is built in 8 small PRs behind checkout_v2. Production runs it off. QA (quality assurance: testers who check the work) and product see it with the flag on for internal users; at launch it goes to 5% of users, then 100%; two weeks later the old code and the flag are deleted.
| Flag state | Who sees it |
|---|---|
| off | nobody (default) |
| internal | staff and testers |
| percentage | a sample of users |
| on | everyone |
What would change my call. If automated test coverage is weak, trunk-based development will break often; invest in tests first. A regulated product with formal release sign-off may use release branches (a copy of trunk frozen for stabilising one release) cut from trunk. Large structural changes use branch by abstraction (a seam in the code allowing old and new to coexist) rather than a long-lived branch.
Walk me through the phases of the software development lifecycle as your team actually runs them. For each phase, what is the concrete output that hands off to the next, and where do projects most often go wrong?
Sample Answer
Direct answer
The software development lifecycle (SDLC) is the sequence of stages a change travels through from "someone wants this" to "it runs in production and is looked after". I describe six phases, and for each one I can name the artifact that gets handed on. When the artifact is vague, the next phase starts by guessing, and that is where projects usually go wrong.
The six phases, their hand-off output, and the usual failure
| Phase | Concrete output handed to the next phase | Where it most often goes wrong |
|---|---|---|
| 1. Requirements and discovery | A written problem statement with acceptance criteria (the checkable conditions that mean the work is done) and a named owner who can answer questions | Solutions disguised as requirements; nobody checked whether users actually have the problem |
| 2. Design | A short design note: components touched, data changes, risks, and what is deliberately out of scope | Design done alone and never reviewed; hard choices (data model, API shape) deferred until they are expensive to change |
| 3. Implementation | Reviewed code merged to the shared branch, with tests | Large, long-running changes that are merged late and break integration |
| 4. Verification (testing) | Test results plus a go/no-go decision (an explicit yes or no on shipping) from someone other than the author | Testing squeezed to "whatever time is left"; only the happy path is checked |
| 5. Release | The change running in production behind a switch or a staged rollout, with a way to turn it back | Releases that are rare and large, so each one is risky and nobody remembers what is in it |
| 6. Operate and maintain | Monitoring, a named on-call owner (the person who is paged when it breaks outside working hours), and a feedback loop (usage data, bug reports) into phase 1 | Treating "shipped" as "done"; the team that built it has moved on and nobody owns the feedback |
How teams really run it
Most teams do not run these as a one-way waterfall. They loop through all six phases in small slices, for example a few days per slice, so each hand-off is a short conversation instead of a document dropped over a wall. The phases still exist; the unit of work shrinks.
Worked example: one slice, "let users export their invoices as CSV"
- Requirements: "Finance users can download the last 12 months of invoices as a CSV; done when the file opens in a spreadsheet with one row per invoice." Owner: the product manager.
- Design: a half-page note choosing a background job (large exports must not time out the web request) and listing what is out of scope (PDF export).
- Implementation: two small merged changes, each reviewed.
- Verification: tests include an account with zero invoices and one with thousands.
- Release: shipped to 10 percent of accounts first, then everyone.
- Operate: a dashboard shows export failures; the first week of support tickets goes back into the backlog (the team's ranked list of work not yet started).
Trade-offs and pitfalls
- Skipping a phase does not remove it; the work moves later where it is costlier (untested code is "verified" by customers).
- Over-documenting every hand-off turns the lifecycle into ceremony. The test is whether the next person can start without a meeting.
- Ownership gaps between phases (especially 5 and 6) are the commonest hidden risk.
Your company is forming a platform engineering team to serve dozens of product teams. How would you decide what it owns, how you would treat product teams as customers, and how you would measure whether developers are better off?
Sample Answer
Position. I would run the platform team as a product team with voluntary adoption and the thinnest viable platform: it owns what most teams need and nobody wants to rebuild, and it earns its users. A platform team builds internal services (deployment pipelines, service templates, observability defaults) so that product teams can ship without becoming infrastructure experts. Team Topologies (a book by Skelton and Pais on how to organise software teams) describes this as a compelling internal product, kept thin so it does not become a bloated bottleneck. A paved path is the well-supported, easiest route the platform offers, such as a template that sets up a new service with build, deploy and monitoring already working.
What it owns: three tests
- Do most teams need it?
- Is it undifferentiated, meaning it gives no competitive edge when each team does it differently?
- Does standardizing it reduce the cognitive load on developers (the amount a person must hold in their head to do their job)?
Example: a CI/CD template. Most teams need to build and deploy (yes); no customer picks us because our deploy script is unique (undifferentiated, yes); one template removes dozens of decisions a developer would otherwise make (lower load, yes). It belongs in the platform. A team's checkout pricing logic fails test 1 (only one team needs it) and test 2 (it is the business), so it stays with that team. The thinnest viable platform is therefore a short list like: a service template, a pipeline template, a way to create environments, and default logging and metrics, rather than a catalogue of every tool.
It would start with paved paths for new service scaffolding, CI/CD templates, environment provisioning, and logging and metrics defaults. It would not own product teams' business code or take over their on-call for their own services.
Treating teams as customers
- Interview 8 to 10 product teams in the first month and publish a roadmap built from their pain.
- Name a product manager or lead for the platform; keep a public backlog.
- Service level objectives (SLOs, targets for reliability) for the platform's own services, and a support channel with a response target.
- Adoption is voluntary at first; if teams must be forced, the product is wrong. Interaction mode (how two teams work together) is mostly "X-as-a-Service" (the platform is consumed like a service, with little back-and-forth), with temporary collaboration when a new capability is being built.
Measuring "developers better off"
- Delivery outcomes per team using current DORA (DevOps Research and Assessment) metrics: change lead time (committed to running in production), deployment frequency (how often you deploy), failed deployment recovery time (how long to recover from a deployment that needs immediate intervention), change fail rate (the share of deployments needing immediate intervention). A fifth, deployment rework rate, is the share of deployments that are unplanned because of a production incident. Compare before and after for adopting teams, not individuals.
- Time to first production deploy for a new service (illustrative: three weeks down to days).
- Developer survey on satisfaction and friction, run each quarter.
- Adoption of paved paths, and platform cost per team served.
Avoid counting commits or tickets, which can be gamed.
What would change my call. With few teams (as a rough heuristic of mine, not a published threshold, fewer than about five) a dedicated platform team is overhead; a shared-tooling working group suffices. Very different tech stacks may require two thinner platforms rather than one.
Unlock Full Question Bank
Get access to all 6 DevOps Culture and Delivery Practices interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.