Airbnb Senior Engineering Manager Interview Preparation Guide
Airbnb's Engineering Manager interview process for senior-level candidates typically spans 6-8 weeks across 6 rounds: recruiter screening, technical phone screen, and 4 onsite rounds. Each onsite session runs 45-60 minutes. The process evaluates three core dimensions: (1) technical depth and architectural judgment sufficient to lead engineering teams credibly, (2) people management and team development capability, and (3) cultural alignment with Airbnb's collaborative, values-driven ethos. Rounds are sequenced to assess technical foundation first, then scale to leadership and cultural dimensions. Unlike individual contributor interviews, Engineering Manager rounds emphasize your ability to multiply team impact, set technical direction, and create psychological safety.
Interview Rounds
Recruiter Screening
What to Expect
Initial conversation with Airbnb's recruiting team assessing basic qualifications, management experience, and preliminary cultural alignment. The recruiter explores your background progression from engineer to manager, your approach to leadership, familiarity with Airbnb's engineering culture and tech stack, and motivation for the specific role. This round also evaluates communication clarity and whether your values resonate with Airbnb's collaborative ethos and 'Be a Host' principle.
Tips & Advice
Articulate a clear narrative of your engineering-to-management journey. Be specific about your management philosophy—avoid generic answers about 'supporting the team.' Demonstrate genuine interest in Airbnb's mission and discuss how their 'Be a Host' value aligns with your leadership approach. Connect your background to the specific role: mention familiarity with marketplace challenges, scaling, or the technologies mentioned. Ask informed questions about team structure, technical challenges, and how the role contributes to broader engineering goals. Show awareness that you're managing technical teams, not just coordinating resources. Emphasize balance: how do you stay technically credible while scaling team impact?
Focus Topics
Technical Credibility & Stack Familiarity
Your hands-on technical background, comfort with Airbnb's tech stack (Ruby, Kotlin, TypeScript), and ability to stay credible while scaling management responsibilities
Practice Interview
Study Questions
Career Progression: IC to Engineering Manager
Your journey from individual contributor through senior engineering roles to management; key milestones, growth decisions, and why you're ready for senior management at Airbnb
Practice Interview
Study Questions
Airbnb 'Be a Host' Value & Personal Alignment
Your understanding of this core value and how it shapes your approach to leading, treating team members with care, and building belonging in engineering teams
Practice Interview
Study Questions
Management Philosophy & Leadership Style
Your approach to building teams, developing talent, making decisions, handling conflict, and creating psychological safety
Practice Interview
Study Questions
Technical Phone Screen
What to Expect
A focused technical interview conducted over video assessing your ability to solve algorithmic or system design problems. This may involve coding in a shared editor (algorithmic problem) or verbal system design discussion. The goal is to verify you maintain sufficient technical depth to lead engineering teams credibly, conduct meaningful code reviews, guide architectural decisions, and mentor engineers effectively. Problems are typically medium-to-hard difficulty, reflecting real Airbnb scenarios.
Tips & Advice
This round tests whether you're still technically sharp, not whether you're an expert problem solver. Think aloud, explain your reasoning, and discuss trade-offs rather than racing to a solution. For coding: write clean, readable code that handles edge cases; explain complexity analysis; show test-thinking. For system design: start with requirements clarification; outline multiple approaches before committing; discuss scalability, consistency, and operational concerns; connect decisions to business impact. As a manager, emphasize communication—can you explain technical concepts clearly to engineers with varying backgrounds? If you get stuck, break the problem down, ask clarifying questions, and show systematic thinking. This demonstrates the problem-solving approach you'd model for your team.
Focus Topics
Code Quality & Clean Code Principles
Knowledge of readability, maintainability, design patterns, naming conventions, function design, and avoiding common antipatterns
Practice Interview
Study Questions
Core Data Structures & Algorithm Complexity
Proficiency with arrays, trees, graphs, linked lists, hash tables; algorithmic approaches (DFS, BFS, sorting, dynamic programming); and complexity analysis (Big O notation)
Practice Interview
Study Questions
System Design Fundamentals
Scalability patterns, consistency models, caching strategies, database design, API design, message queues, load balancing, and architectural trade-offs (CAP theorem, latency vs. throughput)
Practice Interview
Study Questions
Onsite Round 1: System Design & Technical Architecture
What to Expect
Deep dive into designing a scalable system inspired by Airbnb's real challenges: property booking systems, search indexing, recommendation algorithms, real-time updates, or pricing engines. You'll architect a solution considering high availability, data consistency, latency, cost, and operational constraints. This round assesses your technical depth, ability to reason about trade-offs, and capacity to guide technical strategy across teams. Interviewers evaluate whether you can think architecturally—not just coding ability, but systems thinking appropriate for leading engineering.
Tips & Advice
Begin by clarifying requirements: how many concurrent users? What's the consistency model—strong or eventual? What's the latency requirement? Is cost a constraint? Outline your approach before diving deep. Present multiple design options with explicit trade-offs: Why Elasticsearch over a traditional database? When would Redis help? What's the cost of consistency vs. availability? For Airbnb-specific scenarios, discuss real technologies they use: Elasticsearch for search, Redis for caching, message queues for real-time updates, and their monolith-to-microservice migration philosophy. Connect every design decision to business impact: How does this design improve user experience? Reduce operational overhead? Enable faster booking? Interviewers want you to think like a technical leader, not just solve a puzzle. Iterate based on feedback; show flexibility and willingness to reconsider trade-offs. For senior roles, demonstrate awareness of organizational context: Can your team operate this system? Do we have the expertise? What's the migration path?
Focus Topics
API & Backend Design for Reliability & Scale
RESTful API design, idempotency, rate limiting, circuit breakers, graceful degradation, error handling, and monitoring/observability
Practice Interview
Study Questions
Distributed Systems Technologies: Search, Caching, Messaging
Elasticsearch for search relevance and scalability, Redis for caching and rate limiting, Kafka/message queues for event streaming, and when to apply each technology
Practice Interview
Study Questions
Airbnb Marketplace Architecture Patterns
Understanding two-sided marketplace challenges: reservation consistency, search index fan-out, inventory management, pricing systems, real-time notifications, and the interplay between host and guest workflows
Practice Interview
Study Questions
Scalability & Consistency Trade-offs
CAP theorem applications, eventual vs. strong consistency, database sharding strategies, denormalization patterns, handling millions of concurrent users, and identifying bottlenecks
Practice Interview
Study Questions
Onsite Round 2: Technical Leadership & Code Review
What to Expect
Evaluates your ability to improve code quality, mentor engineers, and provide constructive technical feedback. You'll review realistic code (typically in Airbnb's stack: Ruby, Kotlin, or TypeScript) and discuss improvements, architectural concerns, design patterns, and testing gaps. This round assesses balance: Can you elevate quality without being dogmatic? Do you understand the difference between MVP pragmatism and production excellence? As an Engineering Manager, this demonstrates whether you'll create a learning culture or a blame culture around code quality.
Tips & Advice
Approach holistically: correctness, performance, readability, maintainability, testability, and edge cases. Explain your feedback constructively, as if mentoring a junior engineer. For example: 'This function is doing two things—fetching data and transforming it. Splitting these could make testing easier and reuse more likely. What do you think?' Rather than: 'This code is messy.' Ask clarifying questions before judging: 'What constraints were you working under? Is this MVP code or production-critical?' Balance perfectionism with shipping velocity. Point out one or two high-impact improvements, not nitpicks. Show respect for the engineer's perspective. As a manager, emphasize psychological safety: code review should feel collaborative, not gatekeeping. Demonstrate curiosity about their reasoning. Suggest concrete improvements, not vague criticisms. This is how you'd want your team reviewing each other's work.
Focus Topics
Performance & Operational Readiness
Identifying performance bottlenecks, database query optimization, caching implications, monitoring and observability concerns, and operational burden of code changes
Practice Interview
Study Questions
Testing Strategy & Quality Assurance
Unit vs. integration vs. end-to-end testing trade-offs, appropriate test coverage, identifying gaps, and balancing testing investment with velocity
Practice Interview
Study Questions
Design Patterns & Architectural Consistency
Recognizing appropriate design patterns (MVC, Factory, Observer), spotting architectural antipatterns, refactoring opportunities, and maintaining consistency across codebases
Practice Interview
Study Questions
Mentoring & Psychological Safety in Review Process
Delivering feedback that accelerates growth, explaining 'why' not just 'what,' creating safe feedback loops, and seeing code review as a learning conversation
Practice Interview
Study Questions
Code Review Standards & Best Practices
Techniques for identifying issues early across languages (Ruby, Kotlin, TypeScript), consistency checks, common pitfalls, and balancing feedback with team morale
Practice Interview
Study Questions
Onsite Round 3: People Management & Team Development
What to Expect
Comprehensive assessment of your ability to manage, develop, and scale engineering teams. Interviewers explore your experience hiring engineers, conducting one-on-ones, mentoring junior and senior talent, handling underperformance, creating growth opportunities, managing conflicts, and building team cohesion. This round evaluates emotional intelligence, conflict resolution skills, alignment with Airbnb's collaborative values, and your track record multiplying team impact through people development.
Tips & Advice
Prepare 5-7 concrete STAR stories demonstrating: successful hiring of a high-impact engineer, mentoring someone to the next level, navigating a conflict between team members, addressing underperformance constructively, building team culture in a new role, and learning from a leadership mistake. For each story, emphasize outcomes: Did the hire succeed? Did the mentee grow? Was the team healthier after addressing conflict? For difficult conversations, show you listened to all perspectives, maintained empathy, and focused on improvement. Discuss your one-on-one philosophy—frequency, structure, how you identify growth opportunities. Show examples of creating stretch assignments that developed engineers. For conflict, show you resolved issues while preserving relationships. Avoid hero narratives; focus on how you enabled your team's success. Emphasize learning and growth mindset: 'Here's what I learned when that hiring decision didn't work out...' Show self-awareness about areas you're developing as a leader.
Focus Topics
Mentorship & Development of Future Leaders
Identifying high-potential engineers, creating growth opportunities, mentoring toward senior/staff roles, succession planning, and developing the next generation of leaders
Practice Interview
Study Questions
Conflict Resolution & Difficult Conversations
Addressing underperformance directly, managing interpersonal conflicts between team members, navigating disagreements about technical direction, and maintaining relationships while setting clear expectations
Practice Interview
Study Questions
Psychological Safety & Trust Building
Creating an environment where engineers feel safe taking risks, admitting mistakes, asking questions, proposing ideas, and challenging decisions respectfully
Practice Interview
Study Questions
Team Performance Management & Accountability
Setting clear expectations aligned with team and company goals, measuring performance regularly, providing constructive feedback, addressing gaps, and maintaining fairness
Practice Interview
Study Questions
One-on-One Meetings & Career Development
Regular one-on-ones to understand motivations and growth aspirations, creating development plans, identifying stretch assignments, providing feedback, and supporting career progression to next level
Practice Interview
Study Questions
Recruiting & Hiring Technical Talent
Your approach to identifying engineering talent, assessing technical fit and culture fit, building diverse teams, selling the opportunity, and ramping new hires effectively
Practice Interview
Study Questions
Onsite Round 4: Behavioral & Airbnb Values Alignment
What to Expect
Deep behavioral assessment exploring your past experiences, decision-making approach, resilience, and alignment with Airbnb's core values, particularly 'Be a Host.' You'll discuss major challenges you've overcome, cross-functional collaboration with product/design/data teams, ambiguous situations where you had to make calls with incomplete data, and what 'belong anywhere' and collaborative leadership mean to you personally. This round assesses cultural fit, leadership maturity, decision-making quality, and whether you'll thrive in Airbnb's values-driven environment.
Tips & Advice
Prepare 5-6 strong STAR stories covering: major technical or organizational challenge you led, navigating ambiguity and making decisions with incomplete information, collaborating across departments (product, design, data, infrastructure) on complex initiatives, learning from failure or mistake, conflict resolution with a peer or senior leader, and a time you embodied the values of collaboration and belonging. For each story, show the business impact, not just the problem solved. Connect stories to Airbnb values when relevant. For 'Be a Host,' be authentic—explain what belonging means to you and how you create it for your team. Share examples showing genuine empathy, not performative inclusion. When discussing decisions, show how you balanced data with judgment, considered perspectives from multiple teams, and took accountability for outcomes. Discuss learning orientation: 'Here's what I'd do differently now.' Ask thoughtful questions about how the team collaborates, how Airbnb supports manager development, and what success looks like in the first year. Avoid corporate platitudes—genuine, reflective answers resonate more than polished statements.
Focus Topics
Leadership Growth & Self-Awareness
How you've developed as a leader, feedback you've received and acted on, areas where you're still growing, coaching or mentors that shaped you, and commitment to continuous improvement
Practice Interview
Study Questions
Decision-Making & Judgment Under Ambiguity
Your approach to making significant decisions with incomplete information, balancing data and intuition, soliciting diverse perspectives, and owning outcomes—good or bad
Practice Interview
Study Questions
Overcoming Significant Challenges & Resilience
Stories navigating complex technical or organizational problems, leading through uncertainty, bouncing back from setbacks, learning from failures, and persisting through ambiguity
Practice Interview
Study Questions
Cross-Functional Collaboration & Influence
Effectively partnering with product, design, data, infrastructure, and other engineering teams; influencing without direct authority; aligning diverse perspectives; and building shared ownership
Practice Interview
Study Questions
Airbnb Core Value: 'Be a Host'
Understanding and embodying Airbnb's defining value: treating people (team members, colleagues, stakeholders) with genuine care, respect, and empathy; creating belonging; enabling others' success; and leading with hospitality mindset
Practice Interview
Study Questions
Frequently Asked Engineering Manager Interview Questions
Tell me about a time you handled a performance incident in production. Use the STAR format: Situation, Task, Action, Result. Focus on how you diagnosed the problem, what short-term mitigation you implemented, and what long-term changes you pushed afterwards.
Sample Answer
Direct answer: A strong answer to this question shows three things in sequence: how you triaged and made a decision before you had a confirmed root cause, what you shipped as a reversible stop-gap versus what you shipped as the durable fix, and what changed on the team afterward so the same failure mode doesn't come back. "I fixed it and it worked" tells the interviewer nothing.
Situation and Task: I was the on-call backend engineer for a checkout service. About twenty minutes after a routine deploy, our p95 latency (the 95th percentile response time: 95% of requests come in faster than this number, so it captures the slow tail an average would hide) jumped from around 300ms to over two seconds, and the 5xx error rate started climbing as some requests began timing out. My job in the first few minutes wasn't to find the root cause, it was to stop the customer-facing bleeding and figure out fast whether the deploy was even the cause.
Action and Result: I pulled the deploy timeline next to the latency graph and split the latency metric by build version. It was a rolling rollout, so twenty minutes in most instances were already serving the new build while a handful had not been updated yet: the new-build instances showed the spike and the not-yet-updated ones sat at their normal 300ms. That split is what let me act on twenty minutes of data instead of waiting for a root cause, and it is also why the version label on the latency metric matters more than it looks: without it the graph is one line that went up and every hypothesis stays alive. Rather than wait for a full root cause, I rolled the deploy back right away since it was low-risk and reversible, and asked a teammate to pull a thread dump and a slice of traces from one struggling instance before it was torn down, so we wouldn't lose the evidence. Latency recovered within a few minutes of the rollback finishing. Working from the preserved traces afterward, we found that the new deploy had added a synchronous call to a rarely used fraud-check service with no timeout set, so a handful of slow fraud-check responses were tying up request threads and building a backlog. Longer term, we added an explicit timeout and circuit breaker (a safeguard that stops sending requests to a dependency once its failure rate crosses a threshold, so it gets a cooldown to recover instead of being hit while already failing) around that call, added a load test that specifically exercises dependency slowness before merge, put a real canary stage in front of this service's rollouts (a deliberately small first slice of traffic on the new build, watched against the old build's percentiles before the rollout proceeds, so a regression like this one is caught at a few percent of exposure instead of after it has reached almost every instance), and turned the incident writeup into the template our team now uses for "mitigate first, diagnose second" decisions.
Trade-offs and pitfalls: A weaker version of this story stops at "I rolled it back and it got better," with no explanation of why the rollback was the right call at that moment or what evidence justified it. It's also a mistake to only tell the mitigation half of the story: an interviewer specifically wants to hear the long-term change, because a stop-gap that never gets followed up on is the same incident waiting to happen again.
Explain the difference between a stack and a queue and give a concrete example where each is the right choice. Then show how you would implement a queue using only two stacks (or a stack using only queues), and give the amortized cost per operation.
Sample Answer
Direct answer
A stack is last-in-first-out (LIFO): the most recently added item comes out first. A queue is first-in-first-out (FIFO): items come out in the order they arrived. Use a stack when you need to undo or backtrack in reverse arrival order, such as a browser's back button or a function call stack; use a queue when arrival order must be preserved, such as a task scheduler or a print spooler. You can build a queue out of two stacks: push is O(1) worst case, and pop is amortized (averaged over a sequence of operations) O(1) because each element only ever moves between the two stacks once over its lifetime.
Structured elaboration
| Stack (LIFO) | Queue (FIFO) | |
|---|---|---|
| Order returned | Most recent first | Oldest first |
| Concrete example | Undo history, expression parsing, recursive call stack | Print queue, request processing, breadth-first search frontier |
Queue from two stacks. Keep an in_stack that absorbs pushes and an out_stack that serves pops. Pushing always goes to in_stack in O(1). When a pop or peek is requested and out_stack is empty, drain all of in_stack into out_stack; this reverses the order, so the oldest element (which was at the bottom of in_stack) ends up on top of out_stack, ready to be returned first.
Why the amortized argument holds. Use the aggregate method: over any sequence of n operations, each element is pushed onto in_stack exactly once (cost 1), moved from in_stack to out_stack at most once in its lifetime (cost 1), and popped from out_stack exactly once (cost 1). No element is ever moved more than that, so the total work across the whole sequence is bounded by a constant multiple of n, which is what "amortized O(1) per operation" means, even though any single pop that triggers the drain costs O(n) by itself.
Worked example
class QueueFromStacks:
def __init__(self):
self.in_stack: list[int] = []
self.out_stack: list[int] = []
def push(self, x: int) -> None:
self.in_stack.append(x)
def _transfer(self) -> None:
if not self.out_stack:
while self.in_stack:
self.out_stack.append(self.in_stack.pop())
def pop(self) -> int:
self._transfer()
return self.out_stack.pop()
def peek(self) -> int:
self._transfer()
return self.out_stack[-1]
if __name__ == "__main__":
q = QueueFromStacks()
q.push(1)
q.push(2)
q.push(3)
seq = [q.pop(), q.peek()]
q.push(4)
seq += [q.pop(), q.pop(), q.pop()]
print(seq)
Running this prints [1, 2, 2, 3, 4]. The first pop() triggers a drain (in_stack [1,2,3] becomes out_stack [3,2,1], top popped is 1); peek() then reads 2 for free from the already-drained out_stack; pushing 4 goes straight to in_stack without disturbing out_stack; the remaining pops (2, 3) come from out_stack, and the last pop (4) triggers a second drain since out_stack had emptied.
Complexity
| Push | Pop / peek | |
|---|---|---|
| Worst case (single call) | O(1) | O(n) |
| Amortized (over n calls) | O(1) | O(1) |
Space: O(n) total across the two internal stacks, since every pushed element lives in exactly one of them at any time (no extra space is used beyond storing the n elements themselves).
Edge cases
- Calling
pop()orpeek()on an empty two-stack queue: in the reference implementation,_transfer()leavesout_stackempty when both stacks are empty, sopop()'sself.out_stack.pop()andpeek()'sself.out_stack[-1]both raise an unhandledIndexErrorinstead of failing cleanly. Guard this explicitly, for exampleif not self.in_stack and not self.out_stack: raise IndexError("pop from empty queue")before touchingout_stack, so the caller gets a clear, intentional signal rather than an incidental one. - A single push followed immediately by a pop: the drain moves that one element from
in_stacktoout_stackand it is returned, leaving both stacks empty again, which is the state the empty-queue guard above must handle correctly on the next call.
Trade-offs & pitfalls
The most common confusion is treating "amortized" as "always fast": a single pop can still cost O(n) when it triggers the drain. Note the asymmetry with building a stack out of a single queue by rotating on every push (dequeue-then-requeue the previous elements so the newest sits at the front): that rotation happens on every single push, not just occasionally, so it is genuinely O(n) per push with no amortization to appeal to, unlike the two-stack construction above where the expensive transfer is rare and each element only ever pays for it once.
Two people you mentor are in conflict with each other, and it's starting to affect the team's work. How do you handle it?
Sample Answer
Direct answer
Talk to each person privately before bringing them together, so you understand the facts and stakes from each side without an audience. Then classify the conflict as substantive (a genuine disagreement about the right call) versus interpersonal (friction dressed up as a substantive disagreement), because each needs a different resolution path. Bring them together around a shared goal and concrete evidence, not around who's right, and if it's genuinely undecidable in the room, use a time-boxed way to get more evidence rather than let the standoff continue to block the team.
Resolution framework
Never mediate cold in a group. Talk to each person separately first. You're listening for their read of the facts, what they think is at stake, and what "winning" would actually look like to them. This also surfaces things people won't say in front of the other person.
Classify before you intervene. A disagreement that looks technical or process-based on the surface is sometimes substantive and sometimes really about communication style or unresolved friction. Treating an interpersonal conflict as if it just needs more evidence wastes everyone's time; treating a real substantive disagreement as if it just needs better feelings management does too.
Reframe the joint conversation around the shared goal. Ask both people directly what evidence would change their mind. This shifts the conversation from defending a position to examining what's actually true, and it's a useful tell: someone who can't answer that question may be more attached to being right than to the outcome.
Use a time-boxed way to break a genuine deadlock. If the disagreement is real and evidence-based but neither side has enough information to concede, propose a small, bounded experiment or spike to settle it rather than let the argument continue indefinitely. If there's no time for that, make the call yourself and say plainly that you're doing so.
Follow through explicitly. Name who owns the resulting decision, document it somewhere durable, and check back in later to make sure the resolution actually held rather than just went quiet.
Worked example
Two people you mentor are at an impasse over a decision, and delivery is now stalled because of it. Separate conversations reveal the disagreement is mostly substantive, but there's real interpersonal friction layered on top, one of them has started talking over the other in shared meetings. You facilitate a joint session with explicit ground rules, focused on what evidence would resolve the substantive question, and propose a short timeboxed way to get that evidence rather than debate it further. Separately, and privately, you have a direct conversation with the person who'd been talking over the other about how that was landing on the team, independent of who turns out to be right on the substance.
Trade-offs and pitfalls
Mediating in a group before talking to each person privately risks blindsiding someone and getting performative, professional-sounding answers that hide what's actually going on.
Always pushing for consensus wastes time on disagreements that genuinely don't have a consensus answer. Sometimes the right move is a clean, owned decision, not more discussion.
Making the call yourself resolves the immediate block but has a cost: it can create resentment, or teach people to escalate disagreements to you instead of learning to resolve them with each other, so it's worth being deliberate about when you step in to decide versus when you keep facilitating.
A conflict that keeps recurring in slightly different forms is often a signal of a structural problem, unclear ownership boundaries between the two people, rather than a personality clash, and treating the symptom each time without noticing the pattern means you'll be back here again.
You must lead an urgent 60-minute architecture alignment with C‑level stakeholders, sales, legal and engineering to approve a risky cloud migration. Provide a minute-by-minute facilitation plan, prep materials to distribute beforehand, and criteria you will use to obtain a go/no-go decision by the end of the session.
Sample Answer
Direct answer
For a go/no-go meeting with executives, sales, legal and engineering, I would do most of the persuading before the meeting and use the 60 minutes to test the risks and take the decision. I would send a one-page decision brief and written go/no-go criteria in advance, speak to each executive for a few minutes beforehand, and run a tight agenda that ends with the named decider saying go, no-go or go with conditions. The risk of a "risky" migration is usually in the cutover (the moment traffic switches from the old system to the new one) and the rollback (switching back to the old system if the new one fails), so the criteria focus there. "C-level" means the chief officers, such as the CEO, CTO and CFO. ("Go/no-go" means a single decision to proceed or stop.)
Structured elaboration: prep materials
- One-page decision brief: what is being migrated, why now, what happens if we wait, the decision requested, and the decider named.
- Top risks: the five most serious, each with likelihood and impact in plain words (high, medium, low), mitigation and residual risk (what is still left after the mitigation).
- Before-and-after architecture diagram on one page, a rollback plan summary, and a customer impact list prepared with sales.
- Legal checklist: data location, contract commitments, regulatory notices.
- Pre-wiring (talking to people one-to-one before the meeting so nothing in it is a surprise): a short one-to-one with each executive to hear objections early.
Minute-by-minute plan
| Minutes | Block | Owner |
|---|---|---|
| 3 | Purpose, decision requested, who decides, how | Facilitator |
| 7 | The ask and the stakes (cost of acting, cost of waiting) | Sponsor |
| 12 | Risk walk: the 5 risks from the brief, about 2 minutes each, then 2 minutes on overall residual risk | Lead architect |
| 8 | Legal obligations and customer commitments (4 minutes each) | Legal, sales |
| 10 | Challenge and questions: round robin (each person speaks once in turn before anyone speaks twice), with a parking lot (a visible list where off-topic points are written down for later) | Facilitator |
| 10 | Criteria check: each owner says Met, Not met, or Conditional | Criteria owners |
| 7 | Decision and dissent recorded | Decider |
| 3 | Read-back of conditions, owners, dates | Scribe |
The blocks sum to 60 minutes.
Go/no-go criteria (written beforehand)
- The rollback has been rehearsed in a non-production environment, and it can restore service within the downtime window the business accepted (for example, 4 hours on a Sunday night; the rehearsal restored service in 2 hours 40 minutes).
- Legal has confirmed data location and contractual obligations are satisfied.
- The migration window does not collide with committed customer events, and sales has the list of customers to notify.
- The cutover team and on-call coverage are staffed and named.
- The cost stays within the approved range (for example, forecast $370,000 against an approved $400,000; illustrative figures).
Decision rule and worked example of the outcome
Rule agreed in advance: any criterion marked Not met means no-go for this window; Conditional is allowed only when the condition has an owner, a date and a default of no-go if it fails.
Four criteria are Met and the first, the rollback rehearsal, is Conditional because it ran once. The decider says "go, on condition that a second rehearsal passes by Thursday; the engineering lead verifies and reports to me". The conditions, owner and date are read back and recorded, and if the condition fails the default is no-go for that window.
Trade-offs and pitfalls
- In urgent meetings the temptation is to go straight to a decision. Written criteria stop the loudest voice from setting the bar in real time.
- A migration that can be reversed cheaply deserves a faster decision than one with an irreversible step. Name which steps are irreversible and add checks there.
- Agree before the meeting what happens if the decider is missing: no decision is itself a decision, so the default should be written down.
Your product currently serves 1 million monthly active users and sees 100,000 peak concurrent requests. You expect 10x growth over the next 12 months. Outline your capacity-planning approach: what telemetry you'd collect, how you'd model the growth, the kinds of architectural changes that would need to happen to support that scale, and your contingency plan if growth exceeds the forecast.
Sample Answer
Direct answer
Capacity planning for 10x growth isn't one calculation, it's figuring out what breaks first as load rises, buying enough lead time to fix each thing before it does, and having a plan for when reality outpaces the forecast. The most useful anchor number for a product manager or engineering manager to hold onto is the ratio of peak concurrent load to your total user base, because it turns an abstract user-growth target into a concrete load number engineering can actually plan against.
Structured elaboration
Telemetry to collect, and why each matters
- Business and usage signals: monthly active users (MAU, monthly active users), session length, requests per user, and which features drive the most usage. This tells you where growth is actually coming from, not just that it's happening.
- Traffic signals: peak concurrent requests, how bursty traffic is (a spike lasting seconds looks very different from sustained peak load), and geographic distribution. Bursty traffic needs headroom that average traffic doesn't.
- Performance signals: response time at the 95th percentile (P95, the value below which 95% of requests are faster; a better planning number than the average, because it reflects what your slower users actually experience) and error rates. These tell you where the system is already close to its limit today.
- Infrastructure signals: server utilization, database load, and how much of current capacity is autoscaled versus fixed. This tells engineering how much of a 10x jump can be absorbed by "turning a dial" versus requiring new design work.
Modeling the growth, in plain terms
Don't model user growth and system load as the same number; they're related but not identical. Build a small set of scenarios (conservative, expected, aggressive) for user growth, and separately translate each into a load number using your current ratio of peak load to user base as a starting anchor: if peak load has historically been roughly 10% of your monthly active user count, apply that same ratio to your growth target as a first estimate, while treating that ratio as something that can shift as the product evolves, not a law.
What kinds of architectural changes this usually forces
You don't need to design these yourself, but knowing the categories helps you ask the right questions of engineering and set a realistic timeline:
- Absorbing more load without touching the core system: a content delivery network (CDN, a network of servers positioned close to users that serves cached content) for anything that doesn't need to hit your servers on every request, and caching for data that's read far more often than it changes.
- Spreading read load: adding read replicas, additional copies of the database that can serve read traffic, so reads don't all compete with writes on one machine.
- Spreading write and storage load: sharding, splitting data across multiple databases, once a single database's capacity becomes the actual constraint rather than reads.
- Smoothing spikes: moving non-urgent work (sending a notification, generating a report) onto an asynchronous queue, so a traffic spike doesn't force every piece of work to happen synchronously in the request path.
Each of these is a real engineering project with its own timeline, not a switch you flip; the earlier you know which ones the forecast requires, the more lead time engineering has.
Contingency plan if growth exceeds the forecast
- Short term (hours): pre-agreed emergency levers, like temporarily disabling a non-critical, expensive feature, or throttling lower-priority traffic, to protect the core product experience.
- Medium term (days to weeks): accelerate whichever planned change (more caching, more read capacity) is already closest to ready, rather than starting something new under pressure.
- Long term (months): treat sustained over-forecast growth as a signal to revisit the architecture itself, not just add more of the same capacity.
- Have this agreed with engineering and leadership before you need it: what a service-level objective (SLO, an internal target for how the system should perform, for example a target response time) breach looks like, who decides to pull an emergency lever, and what budget is pre-approved for burst capacity so that isn't a debate happening during the incident itself.
Worked example
Assume the numbers given: 1,000,000 monthly active users (MAU) and 100,000 peak concurrent requests today. The ratio of peak load to user base:
1,000,000100,000=0.10
If that same ratio holds at 10x user growth (10,000,000 MAU), the projected peak concurrency is:
0.10×10,000,000=1,000,000 peak concurrent requests
This is a planning assumption, not a guarantee: it treats the 10% ratio as constant, which holds only if usage patterns per user don't change materially as the product grows. If the product becomes stickier (longer sessions, more features used per visit as it matures), the ratio itself could rise, which is why it should be re-measured from real telemetry each quarter rather than locked in once at the start of the 12-month window.
Trade-offs & pitfalls
- Treating the user-growth number and the load number as interchangeable is the most common planning mistake; a 10x user target does not automatically mean 10x load if usage intensity per user changes at the same time.
- Setting architectural targets around average load rather than P95 or peak load under-provisions for exactly the moments (launches, marketing pushes, viral moments) capacity planning is meant to protect against.
- A contingency plan that exists only as a document, without pre-approved budget and a named decision-maker for pulling emergency levers, tends to fail exactly when it's needed, because the debate about whether to act happens during the incident instead of before it.
- Re-forecasting only once, at the start of the 12-month window, means the plan quietly goes stale; the ratio and the scenarios both need periodic revisiting against real telemetry.
A team proposes an event-sourced architecture for a business-critical domain (for example, bookings or orders). Evaluate it: describe the main components (event store, projections, snapshots) and the operational challenges (replay cost, migration, storage growth). As the engineering manager, what organizational prerequisites (testing strategy, developer tooling, team readiness) would you require before approving this architecture?
Sample Answer
Direct answer
Event sourcing stores every change to a business entity as an immutable event ("BookingCreated", "BookingDateChanged", "BookingCancelled") and derives current state by replaying those events, instead of overwriting a row. For a business-critical domain like bookings it can be the right call when the business genuinely needs a complete, trustworthy history (audit, disputes, "what did we know at the time?") or needs many different read models of the same facts. It is expensive to operate: events are permanent, so schema mistakes live forever, rebuilding read models can take hours, and the team needs skills most CRUD (create, read, update, delete) teams do not yet have (event modelling, projection design, and replay/rebuild tooling among them; see the prerequisites below). As the engineering manager, I would approve it only if the team can name the business need that a normal database plus an audit table cannot meet, and only once specific testing, tooling and readiness prerequisites are in place. Otherwise I would push back towards a conventional design, possibly with event publishing for integration (publishing a small, deliberately stable set of events other teams can subscribe to, kept separate from whatever internal event shapes the service uses for its own logic).
The main components
flowchart LR
Cmd["Command: change booking date (a request to make a change)"] --> Agg[Booking aggregate: load events, check rules]
Agg -->|append new events| ES[(Event store: append-only log per booking)]
ES --> P1[Projection: booking detail view]
ES --> P2[Projection: daily occupancy report]
ES --> Snap[(Snapshots every N events)]
Snap --> Agg
P1 --> Q[Queries from UI and APIs]
- Event store: an append-only log, one ordered stream per entity (per booking). It must support "append these events only if the stream is still at version N" (this is called optimistic concurrency: instead of locking the stream first, you attempt the append and let the version check reject it if another write already happened), so two concurrent changes cannot both win. This can be a dedicated product or a relational table with a unique (stream id, version) constraint.
- Aggregate: the code that loads a booking's events, rebuilds its current state in memory, checks business rules ("cannot cancel after check-in"), and emits new events.
- Projections: read models built by consuming events, such as a booking-detail table for the UI or an occupancy report. Each is disposable: drop it and replay the events to rebuild it. This is closely related to CQRS (command query responsibility segregation: separate models for writes and for reads), which event sourcing almost always implies.
- Snapshots: a saved copy of an aggregate's state at event number N, so loading a long-lived entity replays only the events after the snapshot instead of all of them.
Operational challenges, with numbers
Assumptions for a mid-size booking platform: 200,000 bookings per day, an average of 8 events per booking over its life, 1 KB (1,000 bytes) per stored event including metadata.
Storage growth
- Per day: 200,000 × 8 × 1 KB = 1,600,000 KB = 1.6 GB.
- Per year: 1.6 GB × 365 = 584 GB.
- After 3 years: 1.752 TB, plus indexes and projections.
Events are never updated, so this only grows. The plan needs an archival policy (move closed bookings' streams older than the retention period to cheaper storage) and a clear answer to data-deletion requests: personal data inside immutable events is a real problem under privacy laws such as the GDPR (General Data Protection Regulation). The usual approach is to keep personal data out of events or encrypt it with a per-customer key that can be destroyed ("crypto-shredding").
Replay cost
After 3 years there are 200,000 × 8 × 365 × 3 = 1,752,000,000 events. If a new projection processes 20,000 events per second (an assumption; measure your own), a full rebuild takes 1,752,000,000 / 20,000 = 87,600 seconds, about 24.3 hours. That has consequences:
- A bug in a projection means up to a day of rebuilding while the old projection keeps serving.
- New read models need a blue-green approach: build the new projection alongside the old one, switch reads when it catches up.
- Partitioning the rebuild across workers (by stream) is needed to bring this down.
Aggregate load time and snapshots
Most bookings have about 8 events and need no snapshot. A long-lived entity such as a corporate account with 50,000 events would replay all of them on every command; snapshotting every 500 events caps a load at the snapshot plus up to 499 events.
Migration and schema evolution
You cannot run ALTER TABLE on history. When the meaning of an event changes you need:
- Versioned event types (
BookingCreated.v1,v2). - Upcasters: code that converts old event versions to the current shape when they are read.
- Occasionally a copy-and-transform migration of the whole store, which is a major, risky operation.
Every one of these must be tested against real historical events, not just new ones.
Eventual consistency
Projections lag behind the event store by milliseconds to seconds. A user who changes a booking and immediately reloads may see the old date unless the UI reads from the aggregate or waits for the projection to catch up. The product team must accept that behaviour or it must be designed around.
Organizational prerequisites I would require before approving
Testing strategy
- Given-when-then tests on aggregates: given these past events, when this command arrives, then these new events (or this rejection) result. These are the core of correctness and are fast.
- Projection tests that replay a fixed event sequence and assert the resulting read model.
- Upcaster tests against a sample of real production events from every historical version.
- A full-replay rehearsal in a staging environment on a production-sized copy, with a measured rebuild time, before go-live.
Developer tooling
- An event catalog: every event type, its versions, its schema, its owner.
- Tools to inspect a single stream ("show me every event for booking 123") for support and debugging.
- A projection rebuild tool with progress reporting and the ability to run a new projection alongside the old one.
- Monitoring of projection lag, with an alert when a projection falls behind.
Team readiness
- At least one or two engineers who have run an event-sourced system in production, or budgeted time for a proof of concept on a non-critical slice first.
- Agreement from support and product on eventual consistency and on how corrections are made (you append a compensating event (a new event that records the correction, rather than editing the old one); you never edit history).
- On-call runbooks for a stuck projection, a poison event (one that crashes a projection on every attempt), and a failed upcast.
The approval question itself
I would ask the team to answer, in writing: which requirement fails if we use a normal relational model plus an audit log table and publish integration events? If the answer is "none, but event sourcing is elegant", that is golden-hammer adoption (reaching for a favourite pattern regardless of fit) and I would decline. If the answer is "disputes require reconstructing exactly what the booking looked like at any moment, and we need five read models that disagree on shape", that is a real case.
Recommendation shape
- Approve for the booking aggregate only, not the whole platform; surrounding domains (customer profiles, content) stay conventional.
- Gate go-live on the staged replay rehearsal and the tooling above.
- Review after one quarter in production: projection lag, rebuild time, incident count, and how long a new engineer takes to ship their first change.
Pitfalls
- Event sourcing everything. Most domains do not need it; applying it platform-wide multiplies the operational cost.
- Using events as the integration contract. Internal events change as the model evolves; publish separate, stable integration events for other teams.
- Ignoring deletion and privacy until later. Immutable history and data-deletion obligations collide, and retrofitting encryption onto three years of events is painful.
- Treating the event store as a message queue. It is the system of record (the one authoritative place a fact is permanently and correctly stored, not a transient message that can be dropped or replayed); durability, backups and ordering guarantees matter more than throughput.
Partway through designing a system, you're told to plan for three possible curveballs: a region outage, an upstream schema change that breaks your data pipeline, and a sudden 10x traffic spike. How would you prioritize which to design for first, and how does each change your architecture?
Sample Answer
Direct answer
Prioritize by expected business impact combined with how quickly the failure mode compounds if unaddressed: a region outage first, because it's a full-availability event with no partial-degradation option; a sudden 10x traffic spike second, because it threatens availability but usually has partial mitigations (throttling, degraded modes) available immediately; and an upstream schema change third, because it's typically detectable and containable with fast rollback before it causes user-facing damage, even though it can silently corrupt data if left uncaught.
Structured elaboration
For each curveball, separate the immediate runbook response from the longer-term architectural change it justifies.
Region outage. Immediate: fail over reads and writes to a secondary region using health-checked traffic routing, and pause non-essential batch work to reduce write pressure during the transition. Architectural change: multi-region active-passive (or active-active) replication for the data layer, with regularly rehearsed failover drills; a design that was never built to fail over won't fail over correctly under real pressure, only under a rehearsed one.
Sudden 10x traffic spike. Immediate: autoscale the serving tier, shed or degrade non-critical functionality (serve cached or slightly stale results rather than fail outright), and throttle low-priority background jobs to protect the real-time path. Architectural change: pre-warmed capacity headroom, adaptive rate limiting, and a defined degraded mode that's tested before it's needed, not designed during the incident.
Upstream schema change breaking the data pipeline. Immediate: fail fast on schema-validation errors at ingestion rather than let malformed data propagate, quarantine the bad batch, and roll the downstream transform back to the last known-good schema. Architectural change: enforce a schema contract at the pipeline boundary (a strongly typed serialization format with a compatibility check, such as Avro or Protocol Buffers) so a breaking upstream change is caught at ingestion rather than discovered downstream after it has already corrupted derived data.
Worked example
An illustrative prioritization exercise, scoring each curveball on business impact (1 low to 5 high) and detectability/containability (1 hard to 5 easy) to make the ranking auditable rather than a gut call: region outage scores high impact (5/5: full outage, all users) and moderate containability (3/5: requires a rehearsed failover, not just a code fix); 10x traffic spike scores high impact if unmitigated (4/5) but higher containability (4/5: autoscaling and shedding are standard, fast-acting levers); schema break scores lower immediate user-facing impact (2/5: the pipeline can often keep serving stale-but-correct data while paused) but containability that depends entirely on whether validation exists at the ingestion boundary (2/5 without it), if it doesn't, undetected corruption can silently spread for a long time before anyone notices, which is exactly why validation is the priority architectural investment for that curveball specifically, even though it's ranked last for immediate response.
Trade-offs & pitfalls
- Ranking these purely by which is scariest in the abstract, rather than by business impact and how fast each compounds if left unaddressed, produces a plausible-sounding but ungrounded priority order; tie the ranking to a concrete criterion.
- A schema break that lacks ingestion-time validation is deceptively low-priority in the short term and highest-priority for silent, compounding damage; don't let "least immediately visible" become "least urgent to architect for."
- Building all three mitigations simultaneously from scratch during a single design pass is rarely realistic; sequence the architectural investments and say explicitly which curveball's mitigation ships first and why.
- Rehearsing failure (game days, chaos testing, restore drills) is what turns a runbook from theory into something that actually works under pressure; a runbook that has never been executed is a plan, not a capability.
Create a framework to measure and improve engineering team productivity that avoids gaming and incentivizes long-term quality. Define a balanced set of primary metrics and leading indicators, describe the review process with the team, and propose safeguards against metric manipulation.
Sample Answer
Framework overview
Measure productivity as sustained delivery of customer value + maintainability. Use a balanced scorecard: primary outcomes (lagging) and leading indicators.
Primary metrics (balanced)
- Cycle Time for completed customer-facing stories (median) — focuses on throughput of value
- Change Failure Rate (deploys causing incidents) — quality
- Escaped Defects per KLOC or per release — user-visible quality
- Technical Debt Index (trend from static analysis + backlog items) — maintainability
- Customer/Stakeholder Satisfaction (NPS or feature feedback)
Leading indicators
- PR Review Time and Review Coverage
- Test Coverage of critical modules and automated test pass rate
- Code churn on released files
- Percentage of work spent on maintenance vs new features
- Mean Time to Restore (MTTR) for incidents
Review process
- Monthly team review: present dashboard, contextual narratives, signal vs noise
- Quarterly retrospective with engineers + PMs to validate metric relevance and set improvement experiments
- Use anonymized team-level metrics for coaching; tie individual performance to qualitative peer feedback and technical impact, not raw counts
Safeguards against gaming
- Use medians and percentiles, not means; apply smoothing over time
- Combine automated signals with human audits (random code reviews, postmortems)
- Monitor for gaming patterns (sudden increases in minimal-size PRs, batch commits) via anomaly detection
- Reward behaviors: mentoring, refactors merged, reduction in tech debt, successful incident blameless postmortems
- Governance: cross-functional metrics committee to approve changes; require before/after validation of any metric-driven initiative
Why this works
Balances speed and quality, emphasizes trends and context, uses leading indicators to surface problems early, and embeds human review to deter manipulation while encouraging long-term health.
Outline a plan to scale a team from roughly 5 to 50 people (or from 3 to 12, for a smaller function) while preserving candor, autonomy, and psychological safety. Cover hiring criteria, organizational structure, onboarding, communication rituals, decision rights, and how you would propagate the culture and catch drift as the team grows.
Sample Answer
Direct answer
Scaling a team from roughly 5 to 50 people while preserving candor and psychological safety means deliberately converting practices that worked informally at small scale (everyone just knew the norms) into explicit, documented structures before the informal version breaks down, rather than waiting until it already has.
Structured elaboration
- Hiring criteria. Screen explicitly for candor and comfort with feedback, not just technical skill, since a small number of hires who are defensive about critique can quietly shift a team's norms faster than any process can counter. Include a structured interview stage that probes how a candidate has handled being wrong or challenged in the past.
- Organizational structure. Split into smaller sub-teams (pods or chapters of 5 to 8) before the whole-group size makes candor feel risky, since psychological safety is much easier to sustain in a group where everyone knows everyone than in a room of 50. Keep a clear owner for culture within each pod, not just at the top.
- Onboarding. Make the team's actual norms around candor and mistake-reporting an explicit part of onboarding, with real examples, rather than assuming new hires will absorb it by observation, since observation-only onboarding is exactly what breaks down as headcount grows and new hires increasingly onboard from peers who are also new.
- Communication rituals. Preserve at least one regular, small-group forum (not just all-hands) where junior members interact directly with senior leadership, since large-group settings systematically suppress the same voices that a 5-person team never had to worry about.
- Decision rights. Document who decides what as the team grows, since ambiguity about decision rights at scale creates exactly the kind of quiet frustration and unaddressed disagreement that erodes safety over time.
- Propagation and drift detection. Run a lightweight, anonymous pulse check periodically, segmented by pod or tenure, specifically to catch drift early (newer joiners or a particular pod reporting lower safety) before it becomes a pattern across the whole organization.
Worked example
At 8 people, the team relies on a single weekly meeting where anyone can raise anything, and it works because everyone already trusts everyone. At 25 people, that same meeting has quietly become a forum where only the four most senior people speak, so the team splits into pods of 6, each running its own version of that ritual, with a monthly all-pod sync led by rotating hosts rather than always the most senior voice. At 50 people, a pulse survey shows one newer pod reporting noticeably lower safety scores than the others; investigating finds that pod's lead came from a much more hierarchical background and had not been through the same onboarding on the team's norms, which gets addressed directly rather than assumed away.
Trade-offs and pitfalls
The main pitfall is assuming that what worked informally at small scale will simply continue to work if you just keep doing the same things, without noticing that the same practice (one big meeting, one set of unwritten norms) has different, worse effects at 10x the headcount. A second pitfall is over-formalizing too early, turning a small, trusted team into a bureaucracy before it needs one, which can suppress the very candor it is trying to protect.
As a security architect, you don't own another team's backlog, but you need your threat-modeling findings built into their design before they start coding. How do you get that prioritized without direct authority over their roadmap?
Sample Answer
Direct answer
As a security architect you rarely have line authority over another team's backlog, so you get findings prioritized by making them cheap to accept and costly to ignore: translate the finding into the other team's own vocabulary (a defect, a customer risk, a compliance control they must attest to) and attach it to a decision they are already about to make, rather than asking them to open a brand-new work item. You lead with a specific, demonstrated risk instead of a policy citation, offer a menu of remediation options at different costs, and use an existing recurring forum, like a design review or architecture council, so the tradeoff is made visible to the team's own stakeholders, not just to you.
Structured elaboration
- Translate, don't mandate: reframe the threat-modeling finding in terms the team already tracks (a customer-facing incident scenario, a compliance control, a defect class QA can reproduce) instead of a generic "security best practice."
- Time it to their planning cycle: bring a written finding before backlog grooming or sprint planning, not after code is merged, so accepting it is a normal prioritization decision instead of a rework request.
- Offer options, not a mandate: propose two or three remediation paths (a quick mitigating control now, a full fix next sprint, an explicit accepted-risk sign-off) so the team's own product owner makes an informed tradeoff instead of feeling overridden.
- Borrow a forum, don't invent one: attach the ask to a ritual the team already respects, like their design review, so it reads as peer-level influence rather than a unilateral security gate.
- Make patterns visible upward: when a team consistently deprioritizes findings, escalate the pattern, not the individual finding, to a shared forum with both engineering and security leadership present, so someone with authority over both sides makes the call.
Worked example (illustrative, adapt to your own experience)
A security architect threat-models a new payments feature two weeks before the product team's sprint planning. Instead of filing a ticket titled "add input validation" into the team's backlog and hoping it gets picked up, they write a one-page finding: the specific attack path, the customer-facing scenario it enables, and three remediation options ranked by effort. They bring it to the team's existing design review, present it alongside the team's own product owner, and let the team choose between a lightweight mitigating control shippable in the current sprint or a fuller fix in the next one. The team picks the lightweight option and schedules the fuller fix on their own board, because the tradeoff was made visible and owned by them, not imposed from outside.
Trade-offs and pitfalls
- Too formal (a mandatory sign-off gate) breeds resentment and workarounds; too informal (a message in passing) gets lost in someone else's priority queue.
- Offering remediation options is powerful but risks a team always choosing the cheapest option indefinitely, so track accepted-risk decisions somewhere durable so a pattern of chronic deferral becomes visible over time.
- Borrowing an existing ritual only works if that ritual has real teeth; if the design review itself gets skipped or ignored, attaching your ask to it just inherits its weakness.
What the interviewer probes next
They typically follow up on how you handle a team that keeps saying "next sprint" indefinitely, whether you would ever reach for a hard gate like a release-blocking scan instead of persuasion, and how this influence model holds up when you are supporting a dozen teams at once instead of just one.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Engineering Manager jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs