Airbnb Senior Systems Engineer Interview Preparation Guide
Airbnb's interview process for senior engineering roles follows a structured funnel approach. After initial recruiter screening, candidates proceed through technical phone screens to assess foundational knowledge, followed by an intensive onsite loop (Engineering Loop) consisting of multiple rounds covering system design, technical depth, infrastructure operations, and cultural alignment. For senior-level systems engineers, the process emphasizes architectural thinking, large-scale system design, infrastructure optimization, and the ability to handle complex multi-component system integration challenges.
Interview Rounds
Recruiter Screening
What to Expect
The initial screening typically involves a 15-20 minute conversation with an Airbnb recruiter. This round focuses on understanding your background, years of experience, motivation for joining Airbnb, and alignment with the organization's values. Recruiters assess your communication clarity, technical background depth, and cultural fit with Airbnb's collaborative ethos. For senior-level candidates, recruiters probe into your prior experiences leading infrastructure initiatives, working with cross-functional teams, and your philosophy on system design. They also assess your familiarity with Airbnb's business model (marketplace platform, real-time systems, global scale) and your interest in solving the company's infrastructure challenges.
Tips & Advice
Prepare clear, concise stories about your background and why you're interested in the role at Airbnb specifically. Research Airbnb's technical challenges and business model beforehand. Highlight 2-3 impactful infrastructure or systems projects you've led. Be ready to discuss your technical background and years of hands-on experience with systems engineering. Demonstrate knowledge of Airbnb's scale (global marketplace, millions of properties) and how that excites you. Ask thoughtful questions about their infrastructure stack and challenges. Show enthusiasm for mentoring and collaborating with teams.
Focus Topics
Cross-Functional Collaboration Experience
Examples of working effectively with product teams, backend engineers, data teams, and other infrastructure teams; your communication style and conflict resolution approach.
Practice Interview
Study Questions
Airbnb Values and 'Be a Host' Alignment
Understanding what 'belong anywhere' means and how you embody Airbnb's values of collaboration, openness, honesty, and creating belonging in your work and teams.
Practice Interview
Study Questions
Systems Engineering Philosophy
Your approach to designing reliable, scalable systems; your thoughts on infrastructure automation, observability, and the relationship between infrastructure and business impact.
Practice Interview
Study Questions
Motivation and Interest in Airbnb
Specific reasons for joining Airbnb, understanding of their business model (marketplace platform, hosts/guests), scale of operations, and what aspects of their systems engineering challenges interest you.
Practice Interview
Study Questions
Professional Background and Experience
Clear articulation of your career progression, years of systems engineering experience, and key accomplishments in infrastructure and systems roles.
Practice Interview
Study Questions
Technical Phone Screen 1: Systems Fundamentals and Troubleshooting
What to Expect
The first technical phone screen (45-60 minutes) evaluates your core knowledge of systems engineering concepts, infrastructure fundamentals, and your problem-solving approach. This round typically covers distributed systems basics, infrastructure components (servers, networking, storage), troubleshooting methodologies, and your ability to ask clarifying questions. For a senior systems engineer, expect questions that assess your depth of understanding in areas like system reliability, performance bottlenecks, failure modes, and your approach to designing fault-tolerant systems. This is an opportunity to demonstrate both technical depth and communication clarity.
Tips & Advice
Approach each question systematically: clarify requirements, identify constraints, and work through the problem step-by-step. Draw diagrams or pseudocode when helpful. For troubleshooting questions, explain your diagnostic approach—how you would identify the bottleneck or failure point. Discuss trade-offs in your solutions (performance vs. complexity, consistency vs. availability). Demonstrate knowledge of monitoring and observability. Don't just solve the problem; explain how you would verify your solution works. For senior roles, show strategic thinking about scalability and operational concerns from day one.
Focus Topics
System Design Trade-offs and Decision Making
Evaluating design choices based on requirements; understanding cost vs. performance, simplicity vs. power, and alignment with organizational priorities.
Practice Interview
Study Questions
Networking Concepts and Infrastructure Components
TCP/IP basics, DNS resolution, load balancing, routing, firewalls, VPC architecture, and understanding how network issues affect system performance.
Practice Interview
Study Questions
Distributed Systems Fundamentals
Core concepts including CAP theorem, consistency models, replication strategies, distributed consensus algorithms, and eventual consistency. Understanding trade-offs between consistency, availability, and partition tolerance.
Practice Interview
Study Questions
System Performance Analysis and Bottleneck Identification
Techniques for identifying performance bottlenecks using metrics, profiling tools, and monitoring systems. Understanding CPU, memory, disk I/O, and network bottlenecks.
Practice Interview
Study Questions
Failure Modes and Reliability Engineering
Understanding cascading failures, graceful degradation, fault isolation, retry strategies, circuit breakers, and building resilient systems. Disaster recovery and redundancy strategies.
Practice Interview
Study Questions
Technical Phone Screen 2: Advanced Infrastructure and System Integration
What to Expect
The second technical phone screen (45-60 minutes) assesses your ability to handle more complex infrastructure challenges and systems integration scenarios. This round often includes questions about orchestration, infrastructure as code, monitoring and observability at scale, and managing interdependencies between multiple systems. For senior engineers, expect scenarios involving real-world challenges like managing multi-region deployments, coordinating between different infrastructure components, security and compliance requirements, or optimizing costs at scale. This round emphasizes your strategic thinking and experience with large-scale infrastructure operations.
Tips & Advice
Demonstrate senior-level systems thinking by considering operational aspects from the start. When discussing infrastructure solutions, address: how would this be monitored? How would you handle failures? What are the operational runbooks? For senior roles, discuss team coordination and knowledge sharing—how would your systems team maintain this at scale? Talk about automation and reducing toil. Show experience with infrastructure tools and platforms (container orchestration, IaC, monitoring platforms). Discuss security and compliance considerations as integral design aspects, not afterthoughts. Reference patterns and solutions from your prior experience managing large-scale systems.
Focus Topics
Systems Security, Compliance, and Risk Management
Security architecture for infrastructure, compliance requirements, identity and access management, encryption strategies, and conducting security reviews of system designs.
Practice Interview
Study Questions
Multi-Region and Geographic Distribution
Challenges of operating systems across multiple regions or data centers, consistency across regions, failover mechanisms, and managing latency in globally distributed systems.
Practice Interview
Study Questions
Infrastructure as Code and Configuration Management
Tools and practices for managing infrastructure through code (Terraform, CloudFormation, Ansible, etc.), version control for infrastructure, and treating infrastructure configuration like software.
Practice Interview
Study Questions
Container Orchestration and Deployment Systems
Understanding Kubernetes, Docker, service meshes, deployment strategies (rolling deployments, canary releases, blue-green), and managing containerized infrastructure at scale.
Practice Interview
Study Questions
Observability, Monitoring, and Alerting at Scale
Designing effective monitoring strategies, metrics collection and aggregation, log management, distributed tracing, meaningful alerting thresholds, and observability architecture for complex systems.
Practice Interview
Study Questions
Onsite Round 1: Large-Scale System Design
What to Expect
The first onsite round (45-60 minutes) focuses on your ability to design large-scale systems from first principles. You'll be given a real-world infrastructure challenge (e.g., designing a property listing search infrastructure, building a booking reservation system backend, creating a real-time notification system, or designing a data warehouse for analytics). The interviewer expects you to ask clarifying questions, define requirements, discuss architectural choices, explain trade-offs, and consider operational concerns. For senior engineers, interviewers assess your ability to design for scale, reliability, security, and cost-efficiency while explaining your reasoning clearly. The focus is on your systems thinking, not just technical knowledge.
Tips & Advice
Start by asking clarifying questions about scale, geographic distribution, consistency requirements, failure tolerances, and business constraints. Define clear requirements and non-functional properties (throughput, latency, availability goals). Draw a high-level architecture diagram showing major components. Discuss technology choices for different layers and justify your decisions based on requirements. Address scalability through partitioning, caching, and replication strategies. Discuss monitoring, alerting, and operational procedures. For senior roles, demonstrate leadership in the conversation—guide the interviewer through your thinking, be open to feedback, and adjust your design based on their concerns. Discuss cost implications and ways to optimize. Show awareness of Airbnb's actual technology stack (e.g., they use Elasticsearch for search, caching layers, distributed databases).
Focus Topics
Real-Time Data Processing and Analytics Infrastructure
Designing systems for real-time event processing, message queues, streaming pipelines, data warehouse architecture, and handling high-volume data ingestion from operational systems.
Practice Interview
Study Questions
Caching Strategies and Performance Optimization
Understanding caching layers (Redis, Memcached), cache invalidation patterns, cache-aside vs. write-through, distributed caching, and optimizing performance through strategic caching.
Practice Interview
Study Questions
High-Availability and Disaster Recovery Design
Designing for fault tolerance, redundancy, graceful degradation, multi-region failover, data consistency during failures, and recovery time objectives (RTO) and recovery point objectives (RPO).
Practice Interview
Study Questions
Search and Indexing Infrastructure
Building search systems for large datasets, inverted indexes, search optimization, ranking algorithms, and handling complex search filters. Understanding full-text search, faceted search, and search performance.
Practice Interview
Study Questions
Marketplace Architecture Design
Designing systems for Airbnb's marketplace (properties, bookings, searches, payments). Understanding how property listings, reservations, and user searches interact at scale. Designing for peak load periods (vacations, weekends).
Practice Interview
Study Questions
Scalability Through Partitioning and Sharding
Database sharding strategies, partitioning schemes, handling hot partitions, consistent hashing, and managing data distribution across multiple nodes or clusters.
Practice Interview
Study Questions
Onsite Round 2: Infrastructure Architecture and Operations
What to Expect
The second onsite round (45-60 minutes) evaluates your hands-on knowledge of infrastructure technologies, deployment strategies, and operational excellence. You might be asked to design infrastructure for a specific service, optimize an existing system's resource usage, plan a migration strategy, or solve an infrastructure problem. This round emphasizes your practical experience with infrastructure tools, cloud platforms, containerization, networking, and your understanding of how infrastructure decisions impact operations and costs. For senior engineers, expect scenarios requiring thoughtful analysis of trade-offs between automation, maintainability, reliability, and cost. Interviewers assess your ability to solve real, messy infrastructure problems.
Tips & Advice
Demonstrate practical experience with infrastructure tools and platforms. When designing infrastructure, consider all layers: compute, storage, networking, and security. Discuss automation and how to reduce manual toil. Address operational runbooks and knowledge management for your team. For migrations or redesigns, discuss risk mitigation and rollout strategies. Show awareness of resource utilization and cost implications—explain how your design provides value. Discuss monitoring and alerting for the infrastructure you design. For senior roles, demonstrate ownership mentality: how would you hand this off to a team? How would you ensure long-term maintainability? What documentation and runbooks would you create?
Focus Topics
Network Architecture and Design
VPC design, subnetting, routing, load balancing strategies, DDoS protection, VPN and private connectivity, and network segmentation for security.
Practice Interview
Study Questions
Deployment Strategies and Release Management
CI/CD pipelines, canary releases, blue-green deployments, feature flags, rollback strategies, and minimizing deployment risk while enabling frequent releases.
Practice Interview
Study Questions
Storage and Database Infrastructure
Understanding different storage solutions (relational databases, NoSQL, object storage, cache layers), backup and recovery strategies, and choosing appropriate storage for different use cases.
Practice Interview
Study Questions
Infrastructure Security and Compliance Hardening
Security groups, firewalls, encryption at rest and in transit, secrets management, identity and access control, and meeting compliance requirements (GDPR, PCI-DSS, etc.).
Practice Interview
Study Questions
Cloud Infrastructure and Compute Optimization
Understanding cloud platforms (AWS, GCP, Azure), instance sizing, autoscaling policies, reserved instances vs. spot instances, and optimizing cloud spend while maintaining performance.
Practice Interview
Study Questions
Onsite Round 3: Technical Deep Dive and Problem Solving
What to Expect
The third onsite round (45-60 minutes) assesses your ability to solve complex, specific technical problems with depth and rigor. You might receive a real infrastructure problem Airbnb faces (troubleshooting a performance issue, optimizing resource utilization, improving reliability, designing a novel solution), and you'll work through it systematically with the interviewer. This round tests your problem-solving methodology, ability to form hypotheses, debug systematically, and validate solutions. For senior engineers, expect scenarios with ambiguity and multiple possible solutions. Interviewers look for your ability to think critically, consider multiple angles, and explain your reasoning clearly. This is where you demonstrate mastery in your domain.
Tips & Advice
Approach the problem methodically: clarify constraints, identify the root cause, and develop a solution. If given a troubleshooting scenario, walk through your diagnostic approach step-by-step (hypothesis formation, testing, elimination). Show your thought process, not just conclusions. For senior roles, demonstrate strategic thinking—how does this problem fit into the broader infrastructure picture? What's the long-term solution vs. a quick fix? Discuss trade-offs explicitly and justify your approach. Be comfortable saying 'I don't know' but then explain how you would find the answer. Engage with the interviewer, ask for hints or feedback, and adjust your approach based on their guidance. This round values your ability to think, not just your knowledge.
Focus Topics
Incident Response and Post-Mortem Analysis
Structured incident response processes, blameless post-mortems, root cause analysis, and extracting learning from incidents to prevent recurrence.
Practice Interview
Study Questions
Microservices and Distributed System Coordination
Managing dependencies between services, service discovery, distributed transactions, eventual consistency patterns, and debugging issues across service boundaries.
Practice Interview
Study Questions
Data Consistency and Eventual Consistency Patterns
Understanding consistency guarantees, handling stale data, eventual consistency, compensating transactions, and designing systems that tolerate inconsistency.
Practice Interview
Study Questions
Performance Optimization and Capacity Planning
Identifying and removing performance bottlenecks, optimizing resource utilization, forecasting capacity needs, and making trade-offs between performance and cost.
Practice Interview
Study Questions
Systems Troubleshooting and Debugging Methodology
Structured approach to diagnosing issues: forming hypotheses, collecting data, eliminating possibilities, and validating root causes. Using monitoring and observability tools effectively.
Practice Interview
Study Questions
Onsite Round 4: Infrastructure Operations and Runbooks
What to Expect
The fourth onsite round (45-60 minutes) evaluates your approach to operational excellence—how you would ensure infrastructure runs smoothly, how teams interact with systems, and how you scale operations. You might be asked to design operational processes, create runbooks for complex procedures, structure documentation, plan training for teams, or address operational challenges like coordinating across teams or improving incident response. This round is unique to senior-level roles and assesses your maturity as an engineer leader. Interviewers look for your ability to think beyond the technical solution to the human and organizational aspects of infrastructure. How would you ensure your team understands these systems? How would you prevent repeated mistakes?
Tips & Advice
Demonstrate that you think about infrastructure as a service to engineering teams. When discussing systems, include how teams would operate them: What monitoring is essential? What runbooks exist? How would someone on-call respond to alerts? Discuss knowledge sharing and documentation—how do you ensure operational knowledge spreads across the team rather than being siloed? Show awareness of human factors: fatigue, context switching, communication during incidents. For senior roles, discuss how you would mentor team members and build infrastructure expertise. Address cost management and business alignment—how does infrastructure decision impact engineering velocity and business value? Show examples from your prior experience of improving operational processes or reducing toil.
Focus Topics
Reducing Toil and Automation Strategy
Identifying repetitive work, deciding what to automate, improving tools and processes, and measuring impact of automation efforts.
Practice Interview
Study Questions
Knowledge Management and Organizational Learning
How to capture and share infrastructure knowledge across teams, preventing knowledge silos, mentoring junior engineers, and building shared understanding of complex systems.
Practice Interview
Study Questions
Cross-Team Coordination and Communication
Coordinating infrastructure changes across multiple teams, managing dependencies, communicating impact, and building consensus around infrastructure decisions.
Practice Interview
Study Questions
On-Call Practices and Incident Management
Designing sustainable on-call rotations, defining escalation paths, creating meaningful alerts that require human action, and supporting on-call engineers effectively.
Practice Interview
Study Questions
Operational Runbooks and Documentation
Creating clear, actionable runbooks for common operations and incident scenarios. Documentation that's accurate, accessible, and regularly updated.
Practice Interview
Study Questions
Onsite Round 5: Behavioral and Airbnb Values
What to Expect
The final onsite round (45-60 minutes) dives deep into your past experiences, how you work with teams, leadership style, and alignment with Airbnb's values. Expect behavioral questions about challenges you've overcome, how you've handled difficult situations, conflicts you've resolved, and projects you've led. Interviewers assess your communication style, emotional intelligence, and how you embody Airbnb's core values, particularly 'Be a Host' (hospitality), 'Belong Anywhere' (inclusive thinking), and collaborative ethos. For senior roles, interviewers focus on your impact as a leader: How have you grown others? How have you influenced decisions? How do you handle ambiguity and uncertainty? This round determines cultural alignment and leadership readiness.
Tips & Advice
Prepare 4-5 concrete stories demonstrating key competencies: overcoming a difficult technical challenge, conflict resolution, mentoring others, handling ambiguity, and working across teams. Use the STAR method (Situation, Task, Action, Result) but focus on what YOU did and learned. For senior roles, emphasize leadership moments—how you influenced decisions, guided teams, or shifted team direction. Connect your stories to Airbnb values. For 'Be a Host,' discuss how you support and enable your team. For 'Belong Anywhere,' show examples of inclusive decision-making or building psychological safety. Be authentic and show vulnerability—talking about failures and what you learned is more compelling than only successes. Prepare thoughtful questions about team dynamics, career growth, and technical challenges at Airbnb. Research recent news about Airbnb and be ready to discuss the company thoughtfully.
Focus Topics
Impact and Driving Results
Examples of initiatives you've led that had meaningful impact—either technical (improved reliability, reduced toil) or organizational (team growth, process improvement).
Practice Interview
Study Questions
Collaboration and Conflict Resolution
Examples of working effectively with difficult people, resolving disagreements, building consensus, and situations where you had to compromise. How you maintain relationships despite conflict.
Practice Interview
Study Questions
Airbnb 'Be a Host' Values and Belonging
How you create inclusive environments, support team members, embody hospitality in technical interactions, and ensure diverse perspectives are heard in decision-making.
Practice Interview
Study Questions
Leadership and Mentorship
Concrete examples of mentoring others, developing team members' skills, delegating effectively, and creating growth opportunities. How you've helped others advance their careers.
Practice Interview
Study Questions
Overcoming Technical Challenges and Ambiguity
Stories about facing novel problems with no clear solution, how you approached them, who you collaborated with, and what the outcome was. Demonstrating comfort with ambiguity.
Practice Interview
Study Questions
Frequently Asked Systems Engineer Interview Questions
You have to tell leadership that a high-visibility project is going to be late. Walk through how you'd deliver that message, the concrete next steps you'd share, and how you'd handle pointed questions afterward.
Sample Answer
Direct answer
Lead with a one-line factual headline, not a narrative buildup, then follow immediately with what you're doing about it. Leadership's first question is always "what happens next," and making them wait for it while you explain the backstory reads as stalling.
Structured elaboration
- Headline first: what's late, by roughly how much, stated plainly, no hedging language that makes people wonder if you're minimizing it.
- One or two sentences of cause, at a level leadership can act on, a dependency slipped, a scope surprise, not a blow-by-blow of every technical decision that led here.
- The plan: concrete next steps with owners and rough timing, even if the timing itself isn't final. "Here's how we'll know more by Friday" is a plan; "we're working on it" is not.
- Handling pointed questions: answer what you actually know, say plainly when you don't know something rather than guessing to sound in control, and commit to a specific follow-up instead of a vague "I'll keep you posted." Take the most senior or highest-stakes question first rather than letting the conversation drift to whoever's loudest.
- Close with a cadence: when you'll update again, and through what channel, so the room isn't left wondering whether this becomes a pattern of surprises.
Worked example
A high-visibility platform migration is going to miss its committed date by three weeks. In the leadership update, you open with "the migration will land three weeks past the committed date," not with a summary of everything that's gone right so far. You give the cause: "a dependency on the vendor's new authentication API turned out to need more integration work than their documentation implied." You lay out the plan: a revised, phased timeline with the riskiest piece de-risked first, and an owner and date for each phase. When someone asks "why didn't we know this two weeks ago," you say plainly: "we suspected it a week ago and confirmed it Tuesday, that's a valid gap, and here's what we're changing about how we track vendor dependencies going forward," rather than getting defensive or deflecting.
Trade-offs and pitfalls
- Leading with justification instead of the headline, explaining everything that went right first, reads as avoidance and makes people tune out before they hear the actual news.
- Promising a specific new date under pressure before you've actually validated it is the single most common way this conversation creates a second, worse version of itself two weeks later.
- Answering "I don't know" honestly feels risky in the room but is almost always received better than a confident guess that turns out wrong. Confidence you can't back up erodes trust more than admitting a gap does.
- Treating every pointed question as an attack invites a defensive tone that reads worse than the delay itself. Most pointed questions from leadership are about risk to other commitments, not about assigning blame.
Tell me about a migration you led from a legacy three-tier network to a spine-leaf fabric, covering stakeholder alignment, cutover and rollback, and the results.
Sample Answer
Direct answer (the story in one breath)
I led the move of 24 server racks from a legacy three-tier network (access, aggregation, core) to a spine-leaf fabric. I aligned stakeholders around outcomes they cared about (fewer change-related outages, room to add racks without redesign), migrated a few racks per maintenance window (a pre-agreed time when changes are allowed) with the old uplinks left in place so that every step could be reversed, and finished with a fabric whose capacity and failure behaviour I could state in numbers. The figures below come from the design itself; a real answer should replace them with your own measured numbers.
Situation
- Legacy: each rack had a 48 x 1G access switch with 2 x 10G uplinks to an aggregation pair, so each rack was oversubscribed 48 / 20 = 2.4:1 (oversubscription is the ratio of server-facing capacity to uplink capacity; 2.4:1 means the servers could ask for 2.4 times what the uplinks can carry, counting both uplinks as forwarding). A path between two racks crossed up to 5 switches (access, aggregation, core, aggregation, access).
- Pain: Spanning Tree (the protocol that deliberately blocks redundant Layer 2 links so frames cannot loop) blocked half the uplinks, so for a VLAN whose rack had one uplink blocked the usable ratio was 48 / 10 = 4.8:1 (per-VLAN root placement can spread VLANs across the two uplinks, so the real figure sat between 2.4:1 and 4.8:1), a Layer 2 loop in one VLAN had taken several racks down, and the aggregation pair was nearly out of ports.
- Target: a spine-leaf fabric (every leaf switch connects to every spine switch, so any two racks are leaf-spine-leaf apart) of 4 spine switches (spine-1 to spine-4, 32 x 100G each) and 24 leaf switches (48 x 10G server ports, 4 x 100G uplinks, one to each spine), plus 2 border leaves (leaf switches that connect the fabric to the legacy core and the outside, also with one link to each spine), routed point-to-point with ECMP (equal-cost multipath: traffic is spread over all equal-length paths).
Task and my role
I owned design, the migration plan and the go/no-go (the decision to proceed or stop and roll back) on each cutover; the server, security and application teams owned their own validation.
Actions
- Stakeholder alignment. I met the application owners first and asked for their constraints (batch windows, IP address dependencies, which servers cannot reboot). I presented the risk in their terms: the legacy design's single biggest outage risk was a Layer 2 loop, and the new design removes spanning tree from the core. I agreed a change calendar with the service owners and a freeze list of critical dates. In words, roughly: "I am not asking you to learn networking. I need three things from you: which servers cannot be rebooted, which ones have hard-coded IP addresses, and which weeks are off limits. In return, every move happens in a window you have approved, and if your health check fails I put your servers back within the same window."
- Design numbers I could defend. New leaf oversubscription: 48 x 10G = 480G down against 4 x 100G = 400G up, which is 1.2:1, down from 2.4:1. Losing one spine removes 1 of 4 uplinks (25% of capacity); losing one aggregation switch in the old design removed 1 of 2 (50%). Spine ports used: each spine connects to 24 leaves plus 2 border leaves, so 26 of 32 (81%), leaving 6 ports and room for 6 more racks without buying spines.
- Parallel build. Built and tested the fabric alongside the old network in a lab and then in production with no servers attached, and connected it to the legacy core through the pair of border leaves so both networks could talk during migration.
- Cutover plan. 24 racks at 3 racks per window is 8 windows. For each rack: pre-stage the leaf, confirm routes and the gateway on the new fabric (the gateway cutover: the server's default gateway address moves from the old aggregation switches to the leaf, so that is the moment traffic starts using the new fabric), move the server patching rack by rack, test with a checklist (gateway reachable, DNS, a service-level synthetic check), then keep the legacy access switch powered and cabled for the next window.
- Rollback. Defined before the first window, with a decision point at a fixed time into each window: if a checklist item fails and is not fixed by then, move the patch cables back and re-enable the old gateway. Because the old uplinks were not removed until the whole group was validated, rollback needed only cabling and a gateway switch, not a rebuild.
- Decommission last. Only after a full cycle (including a month-end batch) did I remove the legacy aggregation links.
Result
- Capacity: oversubscription went from 2.4:1 physical (between 2.4:1 and 4.8:1 in effect while spanning tree blocked links) to 1.2:1 with every uplink forwarding; worst-case path 5 switches down to 3 (leaf, spine, leaf).
- Resilience: a spine failure costs 25% of leaf uplink capacity rather than 50%, with no spanning tree convergence involved.
- Growth: 26 of 32 spine ports used (81%), with room for 6 more racks.
- Operations: one standard leaf template, which is what made later automation possible.
What I would do differently
I would automate the pre-checks and the post-move validation from the first window, not by the third; I would have run a real rollback drill in the lab before window one, and I would have involved the application teams in the synthetic tests earlier, because two of the early delays came from unknown IP-based dependencies found only during cutover.
Pitfalls to avoid in your own version
Generic virtues ("communication, teamwork") without a decision you made; claims of precise improvements you cannot derive or measure; and a story with no rollback path.
Someone you're mentoring has been stuck on a hard problem for a while and asks for help. Walk through how you decide whether to pair with them, give a hint, or step in directly.
Sample Answer
Direct answer
Default to a diagnostic question or a hint first, since that's the cheapest intervention and preserves ownership of the solution. Escalate to pairing when hints aren't moving them or they're clearly missing a building block they can't discover alone in reasonable time. Reserve stepping in directly for cases bounded by a hard constraint: a real deadline, cost, safety issue, or someone else being blocked by their block.
Decision framework
Start with a diagnostic question, not a hint. "What have you tried, and what's your current hypothesis?" tells you whether they're missing information, missing a concept, or just haven't structured their attempts yet. This costs almost nothing and often unblocks people on its own.
Escalate to pairing when the pattern repeats. If they're cycling through the same failed approach without adjusting, or they're missing a conceptual piece they genuinely can't discover unaided in the time available, sit with them. Let them keep driving; you're there to redirect attention, not take over.
Escalate to stepping in directly only under a real constraint. A hard deadline, a cost or safety issue, someone else waiting on this to move, or clear signs of demoralization (not just frustration) are the legitimate triggers. "I could solve this faster myself" is not one of them; that's true of almost every delegation ever made.
Time-box the struggle explicitly. Instead of leaving it open-ended, agree on a checkpoint: "take another thirty minutes with this angle, then let's regroup regardless of where you land." This protects both their learning and the actual delivery timeline.
Debrief after any intervention, at any level. Even a small hint deserves a quick "here's the reasoning trap you were in" afterward, so the moment converts into a transferable lesson instead of just an unblock.
Worked example
Someone you're mentoring has been stuck for a while and comes to you for help. You ask what they've tried and what they currently believe is going wrong. Their answer reveals a specific reasoning gap, not a knowledge gap, so you give a pointed hint rather than the answer itself. They make progress but hit a second wall later, closer to a real deadline, and this time you sit down and pair with them directly, letting them stay at the keyboard while you ask redirecting questions. Once it's resolved, you debrief separately from the fix itself: what was the actual reasoning trap, and what's the general takeaway for the next similar problem, distinct from the specific bug.
Trade-offs and pitfalls
Defaulting to stepping in because it's faster erodes the person's own problem-solving muscle over time and can create a pattern where they escalate immediately instead of trying, because they've learned help arrives fast if they ask.
Refusing to intervene out of a rigid "let them struggle" stance burns real time and morale, and can backfire if they land on a fragile or outright wrong solution through persistence rather than understanding, and you didn't catch it.
The honest trade-off with hints: they preserve the person's ownership of the solution, but they slow things down and risk letting someone loop past the point where struggle is still productive into the point where it's just frustration with no learning attached.
A subtler failure mode worth naming: a "hint" that's actually the answer in disguise. It looks like coaching and feels generous, but the person doesn't actually earn the insight, and you won't be able to tell the difference from watching them succeed.
An enterprise needs eventual consistency between service A and service B using events. Design an idempotent event processing and reconciliation strategy that guarantees convergence and supports replays, while preserving ordering where necessary.
Sample Answer
Direct answer: To make eventual consistency between service A and B idempotent and reconciliation-friendly, service A publishes events with a stable event ID (or a monotonic sequence number per entity), service B's consumer deduplicates on that ID before applying any change, and a periodic reconciliation job independently compares A's and B's views to catch and repair anything that slipped through despite the idempotency guarantees.
Structured elaboration
Idempotent event processing on the consumer side. Every event from A carries a stable identifier; B's consumer checks (atomically, alongside applying the event) whether that ID has already been processed, using the same "dedup record plus the actual state change in one transaction" discipline as any idempotent write. This is what makes at-least-once delivery (which any reasonable messaging setup between A and B will actually provide) safe: redelivery is a no-op rather than a duplicate application.
Preserving ordering where necessary. If events for the same entity must be applied in order (e.g. "created" before "updated" before "deleted"), B's consumer needs either a strictly-ordered delivery channel per entity (partition by entity ID) or an explicit sequence number in each event that B checks against the last-applied sequence for that entity, rejecting or buffering an out-of-order arrival rather than applying it prematurely.
Supporting replays. Because B's state can, despite everything, still drift from A's (a bug, an extended outage, a schema-migration mistake), the design should support REPLAYING A's full event history into B from scratch (or from a checkpoint) to rebuild B's view, which requires A to retain (or be able to regenerate) its event history for at least as long as any realistic replay window, and requires B's apply logic to be safe to run repeatedly over the same events (which it already is, by the idempotency design above).
Reconciliation as the safety net, not the primary mechanism. A periodic job independently compares A's and B's data (via checksums, row counts, or a full diff on a schedule appropriate to the data's size and criticality) and either auto-repairs small, well-understood divergences or flags larger ones for human review. This is deliberately a SEPARATE mechanism from the event-driven sync path, its job is to catch failures of that path (a dropped event no retry ever recovered, a bug in the consumer's apply logic), not to be the primary way B stays in sync (that would defeat the point of event-driven propagation in the first place).
Worked example. Service A (an Orders service) publishes OrderUpdated{order_id, sequence, payload} events. Service B (a search index) consumes them, checking (order_id, sequence) against the last sequence it applied for that order, skipping (as an idempotent no-op) any event with a sequence it's already seen or older, and buffering (briefly) any event that arrives out of order, applying it once the gap is filled or timing it out into a "request full replay for this order_id" fallback if the gap doesn't close. Nightly, a reconciliation job compares a sample (or full set, for smaller datasets) of orders between A's source of truth and B's index, flagging any order where B's data doesn't match A's for investigation, this is how the team discovered a bug where B's consumer was silently dropping events during a brief scaling event, well before any customer noticed stale search results.
Trade-offs and pitfalls. Skipping the reconciliation job because "the event pipeline is reliable" is a common and risky shortcut, event-driven consistency mechanisms fail in ways that are often invisible until reconciliation (or a customer complaint) surfaces them, since a missed event usually produces no error, just quietly stale data.
Design an on-call escalation system for an organization with multiple teams that need to coordinate coverage across time zones. How do you route pages, prevent alert-noise from cascading into unnecessary escalations, and decide who gets pulled in for a revenue-impacting versus a data-sensitive incident?
Sample Answer
An escalation system for a multi-team, multi-timezone org needs three separate mechanisms working together: routing (getting an alert to the right on-duty person without a human deciding that in the moment), noise suppression (so one root cause doesn't fan out into ten pages), and a severity model that decides who gets pulled in and how fast, because a revenue-impacting outage and a data-sensitive incident need different people in the room, not just different urgency.
Core building blocks
- Alert gateway: every alert is deduplicated by fingerprint, enriched with service/team/severity tags, and checked against maintenance windows before it's allowed to page anyone.
- Routing table: maps service ownership and team schedule (including timezone-local business-hours windows) to whoever is currently on duty, kept in the paging tool as the single source of truth rather than a wiki page someone forgets to update.
- Severity model: decides who gets paged first and how many people, based on impact type, not just raw error rate.
- Escalation ladder: a fixed sequence of who gets paged next if no one acknowledges, with a hard time budget at each step.
Severity and routing matrix
| Impact type | First page | Ack SLA | If unacked | Extra routing |
|---|---|---|---|---|
| Revenue-impacting (checkout, payments down) | Primary on-call for the affected service | 5 min | Escalate to secondary, then service lead | Auto-opens a major-incident bridge if still unacked at 15 min |
| Data-sensitive (PII exposure, access-control gap) | Primary on-call and security/compliance on-call, paged together | 5 min | Escalate both chains in parallel | Legal/compliance notified regardless of ack status, on a fixed clock, not gated on resolution |
| BI/dashboard degradation (stale or broken dashboards, no customer-facing impact) | Data platform on-call only | 30 min | Escalate to data platform lead | No bridge; tracked as a ticket unless it crosses a staleness threshold (e.g. data older than its documented freshness SLA) |
Escalation flow
flowchart TD
A[Alert fires] --> B[Gateway: dedupe, enrich, tag severity]
B --> C{Severity type}
C -->|Revenue-impacting| D[Page Primary, ack SLA 5m]
C -->|Data-sensitive| E["Page Data on-call AND Compliance lead in parallel; notify Legal/Compliance on fixed clock, independent of ack"]
D --> F{Acked by 5m?}
F -->|No| G[Escalate to Secondary, ack SLA +10m]
G --> H{Acked by 15m?}
H -->|No| I[Escalate to Service Lead, open incident bridge]
F -->|Yes| J[Primary mitigates]
H -->|Yes| J
E --> K{Acked by 5m?}
K -->|No| L[Escalate both chains in parallel, ack SLA +10m]
K -->|Yes| M[Data on-call + Compliance mitigate]
L --> M
Preventing alert-noise from cascading into unnecessary escalations
Most alert storms come from one root cause tripping many downstream checks at once (a database going down pages every service that depends on it). The gateway groups alerts by a correlation key (same root dependency, same time window) before routing, so the escalation ladder above runs once for the incident, not once per symptom. Escalation timers also only start on the first page for a correlated group; late-arriving duplicates reset nothing.
Keeping the matrix trustworthy
A routing matrix that's wrong is worse than no matrix, because it creates false confidence. Two things keep it honest: primary/backup contacts are pulled live from the scheduling tool rather than hand-maintained, and if the on-duty person marks themselves absent (leave, travel) in that same tool, pages route to the next person automatically instead of timing out silently first. A silent timeout during a real on-call absence is exactly the failure mode that erodes trust in the whole system.
Trade-offs and pitfalls
Stricter deduplication reduces noise but risks folding two genuinely unrelated incidents into one correlation group if the correlation key is too broad; the fix is scoping correlation to a real dependency graph, not just a time window. A common wrong turn is building one severity ladder for everything, which either pages security teams for routine downtime or under-escalates a compliance-relevant incident because it didn't look revenue-critical on the dashboard. The severity model has to be impact-type aware, not just impact-size aware.
Support teams will soon need to interpret alerts from a new monitoring dashboard. How would you get them ready, and how would you collect feedback in the first month?
Sample Answer
Direct answer
I would prepare support around the decisions they will make, not the dashboard's features: for each alert, what it means, what to check first, and when to escalate. Then rehearse with realistic examples before go-live, and run a structured feedback loop in the first month, measuring whether their escalations are accurate.
Before launch (two to three weeks)
- Alert cards. One short card per alert: meaning in plain words, severity, the first thing to check, who to escalate to and how, and an example. Link the card from the alert itself.
- Training on real cases. A 60-minute session using past incidents replayed on the dashboard: "what does this alert mean, what do you do?" Include the alerts that look scary but are harmless.
- Dry run. A week where support watches the dashboard alongside engineering with no obligations, then a few simulated alerts to test the escalation path.
- One-page cheat sheet and a named engineering contact for the first weeks.
First month: collecting feedback
- A quick "was this alert clear?" thumbs up or down and comment on each alert card.
- A dedicated channel for questions, tagged by alert type.
- A 15-minute weekly review: top misread alerts, unclear cards, missing alerts, noisy alerts.
- Fix cards and alert wording the same week and tell support what changed.
What I would measure
- Escalation accuracy: escalations that engineering judged necessary divided by total escalations. Illustrative: 40 escalations, 28 necessary gives 28/40 = 70 percent. Track its trend by week.
- Time from alert to first support action.
- Number of "what does this mean" questions per week (should fall).
- Missed alerts: alerts that mattered but were not escalated (found in review).
Worked example: an alert card (illustrative)
Alert: Checkout error rate high
Means: More payment requests are failing than usual.
Check first: Is the payment provider's status page showing an outage?
Escalate if: still high after 10 minutes or provider status is normal -> page on-call payments engineer.
Ignore if: single spike under 2 minutes.
Pitfalls
- Training on every feature; support only needs actions per alert.
- Only measuring satisfaction. Use accuracy to see actual understanding.
- Noisy alerts teach people to ignore the dashboard. Feed noise reports back and tune alert thresholds.
Design an archival policy that moves older data from a warm or hot storage tier into a cheaper cold or archival tier, without breaking jobs that occasionally still need to read snapshots of that older data. Cover the lifecycle transition rules you would set, how you would catalog or index archived data so it can still be found, the restore workflow and its expected latency, the trade-off between retrieval cost and access speed, and how you would avoid unexpected restore failures or surprise costs when older data actually gets requested back.
Sample Answer
Direct answer
Build the policy around three separate pieces that people usually collapse into one: (1) transition rules that move data between tiers automatically based on age or access pattern, (2) a catalog that always knows which tier and which retrieval class a given piece of data currently lives in, independent of where it physically sits, and (3) a restore path whose retrieval-speed class is chosen up front, at tiering time, to match the worst-case restore-time SLA (service-level agreement: a promised or contractually required target, here the maximum acceptable time to get requested data back) that data might ever need, not improvised after someone requests it back. Verify the whole thing works with real, scheduled restore drills rather than trusting the design on paper.
Lifecycle transition rules
- Trigger type: age-based (simplest: "move to cold after 365 days," configured as a native bucket/table lifecycle policy, e.g. S3 Lifecycle rules or Azure Blob lifecycle management, so the storage service itself enforces it rather than relying on an application job someone eventually forgets to run) or access-frequency-based (more precise: track actual read frequency and demote data that has genuinely gone cold, e.g. S3 Intelligent-Tiering's automatic monitoring). Age-based is the right default; add frequency-based only where access patterns don't correlate cleanly with age.
- Event-based overrides: some data should move immediately on a business event (a case closes, a record is finalized) rather than waiting for an age threshold; model this as an explicit catalog write that fires the transition, not a special case bolted onto the age rule.
- Prefer declarative, storage-service-native lifecycle rules over custom migration code wherever the storage system supports it: a misconfigured lifecycle rule fails loudly (data doesn't move, and you can see that in the catalog's per-tier byte counts); a custom migration job can fail silently and nobody notices until a restore request comes in for data that was never actually archived.
Cataloging and indexing archived data
The moment data leaves the hot, natively-queryable store, you need a durable catalog that maps a logical identifier (a file path, a partition key, a record range) to where it currently lives: tier, storage/retrieval class, the physical object key, its size, a checksum recorded at archive time, and when it becomes eligible for expiry. Without this, "find it" degenerates into enumerating every tier by hand.
- Implement it as a small, always-hot, highly available table (a Postgres table, a DynamoDB table, or, if the archived data is itself tabular, an open-table-format manifest such as Apache Iceberg or Delta Lake: a structured metadata file format used in data-lake/analytics stacks that tracks a table's underlying data files, schema, and partitions), never as something that itself gets archived: the catalog is on the critical path of every restore, so it must always be instantly readable.
- Update the catalog as part of the same operation that performs the tier transition (or reconcile it against the transition job reliably), not as an unrelated best-effort side effect; a catalog that is out of sync with reality is worse than no catalog, because it answers confidently and wrong at exactly the moment (an audit, an incident) when correctness matters most.
Restore workflow and retrieval-speed class
Cold-tier restores aren't uniform: choose among near-instant, expedited, or bulk retrieval classes based on the restore-time SLA the specific data actually needs, decided when the data is tiered, not guessed at restore time.
- Near-instant (e.g. S3 Glacier Instant Retrieval): millisecond-latency GETs, priced closer to a warm tier. Use it for cold data that is rarely read but must never make a caller wait.
- Expedited (e.g. S3 Glacier Flexible Retrieval's Expedited option): typically 1 to 5 minutes, at a premium per-request and per-GB retrieval price. Use it for data bound by a tight, hard restore-time SLA that near-instant pricing can't justify for the whole dataset.
- Bulk: the cheapest per-GB retrieval option, typically hours (and up to about two days on the deepest archival classes). This is the right default for the vast majority of archived data, which is rarely if ever requested back and has no hard restore-time bound.
The mistake this guards against: assuming any cold tier can be rehydrated "fast enough" and only discovering the real restore-time bound during an actual audit or incident, when it's too late to move the data into a faster-retrieval class.
After a restore completes, verify it, don't just trust that the read call returned success: recompute a checksum on the restored object and compare it against the checksum recorded in the catalog at archive time.
Avoiding restore failures and surprise costs
- Restore-drill verification: schedule periodic real test restores, not synthetic health checks. Pick a sample of archived partitions, run the actual restore workflow end to end, verify row counts and checksums against the catalog's recorded values, and log the observed restore duration as evidence the assigned retrieval class is genuinely meeting its SLA target. This is what catches a wrong retrieval-class assignment before a real, time-boxed request does.
- Automated compliance reporting: a recurring report or dashboard that reads the catalog and surfaces per-tier byte counts (so data that silently failed to migrate shows up as an unexpected size anomaly in the wrong tier), upcoming retention expirations (so purges happen exactly on schedule, neither late, which is a compliance finding, nor early, which is evidence destruction), and legal-hold status per record.
- Surprise-cost guardrail: expedited and near-instant retrieval are priced at a premium per GB and per request; a misconfigured or accidental bulk job that restores a whole cold partition at expedited speed instead of bulk can spend a meaningful fraction of a month's storage budget in one run. Enforce a retrieval-class allowlist per caller (only the narrow incident-response path is allowed to request expedited; scheduled batch jobs are pinned to bulk) and alert on retrieval-request volume.
- Legal hold and lifecycle expiry are two independent gates on deletion, and both must block it: a lifecycle rule alone will happily delete data that is under active litigation hold unless the hold is checked first.
Worked example: 7-year audit-log retention with a 4-hour rehydrate requirement
A system ingests 200 GB/day of audit logs, subject to a 7-year regulatory retention requirement, and an audit process that can demand a named partition rehydrated within 4 hours.
gb_per_day, years, hot_days = 200, 7, 30
warm_days = 365 - hot_days
total_gb = gb_per_day * 365 * years
hot_gb = gb_per_day * hot_days
warm_gb = gb_per_day * warm_days
cold_gb = gb_per_day * 365 * (years - 1)
# Illustrative US East list prices, fetched live from aws.amazon.com/s3/pricing/
p_std, p_ia, p_deep, p_flex = 0.023, 0.0125, 0.00099, 0.0036
sla_frac = 0.10 # share of the cold tier subject to the 4h rehydrate SLA
cold_sla_gb = cold_gb * sla_frac
cold_bulk_gb = cold_gb * (1 - sla_frac)
cost_split_cold = cold_sla_gb * p_flex + cold_bulk_gb * p_deep
total_monthly = hot_gb * p_std + warm_gb * p_ia + cost_split_cold
baseline_all_std = total_gb * p_std
print(f"total 7yr volume: {total_gb:,} GB")
print(f"tiers: hot={hot_gb:,} GB, warm={warm_gb:,} GB, cold={cold_gb:,} GB "
f"(of which {cold_sla_gb:,.0f} GB is SLA-bound)")
print(f"monthly cost, tiered with the SLA split: ${total_monthly:,.0f}")
print(f"monthly cost, all-Standard baseline: ${baseline_all_std:,.0f}")
print(f"savings: {(1 - total_monthly/baseline_all_std):.0%}")
Output:
total 7yr volume: 511,000 GB
tiers: hot=6,000 GB, warm=67,000 GB, cold=438,000 GB (of which 43,800 GB is SLA-bound)
monthly cost, tiered with the SLA split: $1,523
monthly cost, all-Standard baseline: $11,753
savings: 87%
The concrete decision the SLA forces: S3 Glacier Deep Archive's retrieval options are Standard (9-12 hours) and Bulk (up to 48 hours); neither meets a 4-hour bound. The 10% of cold data actually subject to that bound has to live in a faster-retrieval class (here, Glacier Flexible Retrieval, whose Expedited option is typically 1-5 minutes) instead, at roughly 3.6x Deep Archive's per-GB price, an explicit, budgeted trade for meeting the SLA rather than an assumption that gets discovered wrong during a real audit. The restore drill for this slice runs monthly: pick one random SLA-bound partition, run an Expedited restore, record the wall-clock time to completion, verify its row count and checksum against the catalog, and post the result to the compliance dashboard, so drift toward missing the 4-hour bound (for instance during a regional capacity crunch) is caught by the drill instead of by an auditor.
Trade-offs and pitfalls
- Defaulting everything to near-instant or expedited retrieval defeats the purpose of tiering: most archived data is never requested back, and the savings tiering exists to capture come from most of it staying in the cheapest, slowest-to-restore class. Reserve the premium retrieval class narrowly, for the specific data actually bound by a hard SLA, the way the worked example does with its 10% split.
- A catalog is only as trustworthy as its consistency with the actual data movement; treat catalog updates as part of the transition operation, not an afterthought.
- "The restore call returned success" and "the restored data is correct" are different claims; skip the checksum/row-count verification and a silently truncated or corrupted restore is indistinguishable from a good one until someone reads the data and notices.
- Don't wait for a real audit to discover a retrieval-class mismatch. A restore drill that only checks "did the API call succeed" without timing it against the SLA target isn't actually validating the thing that matters.
Describe how Horizontal Pod Autoscaler (HPA) can scale based on a custom metric such as queue length. Which components are required (metrics adapter, exporter), how do you expose the metric to the cluster, and what operational pitfalls should you watch for when autoscaling on custom metrics?
Sample Answer
The Horizontal Pod Autoscaler (HPA) never talks to Prometheus, your app, or a cloud queue directly: it only ever queries one of three Kubernetes metrics APIs (metrics.k8s.io, custom.metrics.k8s.io, external.metrics.k8s.io), so scaling on queue length requires something that exposes the queue's depth through one of those APIs. The standard pipeline is: the application or a sidecar exposes the metric, a metrics system collects it, and a metrics adapter (most commonly prometheus-adapter, or a project like KEDA, Kubernetes Event-Driven Autoscaling, which ships its own adapter and is now the more common current choice specifically for external event sources like queues) is registered with the API server as an aggregated API and translates queries into that metric API's shape.
The three metrics APIs
| API group | What it serves | Tied to a Kubernetes object? |
|---|---|---|
metrics.k8s.io | CPU and memory only, from metrics-server | Yes (per pod/node) |
custom.metrics.k8s.io | Any metric associated with a specific Kubernetes object (a Deployment, a Service) | Yes |
external.metrics.k8s.io | Any metric not tied to a Kubernetes object at all | No |
A message queue's depth (say, a managed queue service or a Kafka consumer-group lag) is not a property of any Kubernetes object, so it belongs under external.metrics.k8s.io and the HPA metric type: External, not Pods or Object.
Required components and how the metric gets exposed
- Instrumentation: the application (preferred) or a sidecar exporter publishes a metric, e.g. a Prometheus gauge
myapp_queue_length{queue="orders"}on/metrics. - Prometheus scrapes it on a scrape job.
prometheus-adapteris deployed with rules mapping a PromQL query to an external metric name Kubernetes will expose, and it registers anAPIService(apiregistration.k8s.io) sokubectl get --raw /apis/external.metrics.k8s.io/v1beta1returns real data. This registration needs its own RBAC: aClusterRolegranting the HPA controller's service account (system:kube-controller-managerreaching through the aggregation layer) permission to read the external metrics API, plus the adapter's own service account needing permission to read Prometheus.- The HPA references the metric by name:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
spec:
minReplicas: 2
maxReplicas: 20
metrics:
- type: External
external:
metric: {name: queue_length_orders}
target: {type: AverageValue, averageValue: "100"}
(autoscaling/v2 has been the stable API since Kubernetes 1.23; the older v2beta2 was removed in 1.26, so any manifest or tooling still referencing it is out of date.)
Operational pitfalls
- A silently wrong zero is worse than a visible failure. If the adapter genuinely can't reach its data source, the HPA does not quietly treat that as "no load": it sets the
ScalingActivecondition toFalsewith reasonFailedGetExternalMetricand shows the metric's current value as<unknown>inkubectl describe hpa, holding replica count steady. The real danger is the opposite case: an adapter that, on a query error, returns a literal0instead of erroring. That looks healthy to the HPA and drives a real scale-down during an actual outage in the metrics path, which is why adapter error-handling is worth testing explicitly rather than assumed. - Cardinality. High-cardinality labels on the underlying metric (one series per customer ID, for instance) can make Prometheus memory blow up long before the HPA ever sees a problem; keep the label set the adapter maps from small and stable.
- Flapping. A noisy queue-length signal causes replica oscillation; use
behavior.scaleDown.stabilizationWindowSeconds(part of theautoscaling/v2HPA behavior fields) or pre-aggregate with a Prometheus recording rule rather than reacting to raw noise. - Latency in the chain. Scrape interval, adapter caching, and the HPA's own sync period all stack up between a real queue-depth change and a scaling action; a 15s scrape interval plus a slow adapter cache can easily add tens of seconds of lag, which matters for a bursty queue.
- Cost. Recomputing an expensive PromQL query on every HPA sync (default every 15 seconds) across many HPAs can meaningfully load a Prometheus instance; pre-aggregate with recording rules for anything non-trivial.
Validate end to end in staging with synthetic queue load before trusting this in production, watching the full chain (producer to queue to exporter to Prometheus to adapter to HPA) rather than any single hop in isolation.
Trade-off note
Hand-rolling prometheus-adapter rules gives full control over the PromQL mapping but is fiddly YAML to maintain; KEDA trades some of that flexibility for purpose-built scalers for dozens of common event sources (queues, streams, schedules) and is usually less operational overhead for exactly this "scale on queue depth" scenario, at the cost of being one more component to run alongside (or instead of) prometheus-adapter.
Write an OPA/Rego policy that rejects any resource missing a tags map with owner and environment keys, and explain how you'd wire that policy into CI so a bad terraform plan can't get merged.
Sample Answer
Direct answer
A tags policy denies any Terraform resource change whose after state has a missing tags attribute, or a tags map without both owner and environment keys, evaluated against Terraform's JSON plan output. Wired into CI as a conftest (a CLI that runs Rego policies against JSON/YAML input) or opa eval step that runs after terraform plan, it fails the pipeline (non-zero exit) before merge, so terraform apply never runs against a non-compliant resource.
Approach
Read Terraform's plan as JSON (terraform show -json), walk resource_changes, skip anything being deleted (a destroy has no future tags to check), and for everything else assert after.tags exists and contains owner and environment. Emit one deny message per violation so CI can print exactly which resource and which key failed, not just "policy failed."
Policy (Rego)
package terraform.tags
import rego.v1
deny contains msg if {
some rc in input.resource_changes
some action in rc.change.actions
action != "delete"
attrs := rc.change.after
tags := object.get(attrs, "tags", null)
not is_map(tags)
msg := sprintf("resource %v missing tags map", [rc.address])
}
deny contains msg if {
some rc in input.resource_changes
some action in rc.change.actions
action != "delete"
attrs := rc.change.after
tags := object.get(attrs, "tags", null)
is_map(tags)
not tags.owner
msg := sprintf("resource %v missing tags.owner", [rc.address])
}
deny contains msg if {
some rc in input.resource_changes
some action in rc.change.actions
action != "delete"
attrs := rc.change.after
tags := object.get(attrs, "tags", null)
is_map(tags)
not tags.environment
msg := sprintf("resource %v missing tags.environment", [rc.address])
}
is_map(x) if {
x != null
type_name(x) == "object"
}
Key points
- Iterate
resource_changes, not the raw HCL, so the policy sees what will actually be created or changed, including values interpolated from variables and modules. - Exclude pure
deleteactions withsome action in rc.change.actions; action != "delete": this is an existential check, so a replace (["delete", "create"]) still gets validated on itscreatehalf, while a plain destroy (["delete"]) is correctly skipped. - Default the tags lookup with
tags := object.get(attrs, "tags", null)before ever callingis_mapon it. Do not callis_map(attrs.tags)directly: whenattrs.tagsdoes not exist, referencing it produces no value at all, and in Rego, negating a call whose argument is itself undefined never resolves to true or false, it just never fires.object.getalways returns a defined value (nullwhen the key is absent), sonot is_map(tags)becomes a normal negation over a normal function result and correctly denies the "no tags key at all" case. - One
denyrule per failure mode (missing map, missing owner, missing environment) so the CI output tells the developer exactly what to fix.
Complexity and edge cases
Evaluation is O(number of resource_changes) with constant work per resource: no recursion, no external calls. Edge cases worth naming explicitly: resource types that do not support tags at all (need an allowlist of resource types, otherwise this policy false-positives on them); for_each/count resources, which appear once per instance in resource_changes and are each checked independently, so a module with ten instances produces ten independent checks; and a tags value that depends on another resource not yet created, which the plan reports as unknown rather than present. A naive check treats "unknown at plan time" the same as "missing," which is a false positive; a stricter version would inspect rc.change.after_unknown.tags and only deny when the value is genuinely absent, not merely unresolved yet.
Wiring into CI
- Generate the plan in CI:
terraform plan -out=plan.binary && terraform show -json plan.binary > plan.json. - Evaluate it:
conftest test --policy ./policy plan.json, oropa eval --input plan.json --data policy.rego "data.terraform.tags.deny". - Treat any non-empty
denyset as a pipeline failure and make that check required on the PR, not advisory, so a merge is mechanically blocked, not just discouraged. - Surface the
denymessages as a PR comment or check annotation so the resource address and missing key are visible without digging into CI logs.
Trade-offs & pitfalls
- A policy this strict on day one breaks every existing untagged module. Roll it out in warn-only mode first, generate a report of current violations, and flip to blocking once the backlog is cleared, not the other way around.
- Checking the JSON plan catches drift-introducing changes before apply, but it cannot catch tags removed later by someone with direct cloud console access. That needs a separate periodic scan against live resources (a detective control), not just this preventive one.
- The
after_unknowncase above is easy to miss in a first pass and is exactly the kind of thing that turns into a noisy false positive once real modules with cross-resource references start hitting the policy. - Rego's "negating a call over an undefined argument never resolves" behavior is a well-known footgun: a
not f(x)guard that looks correct in review can silently never deny anything for the exact input it was written to catch. Always test the "attribute completely absent" case with the realopabinary, not just the "attribute present but wrong" case, since the two paths through a guard function can diverge. - This policy is written in OPA v1 syntax (
deny contains msg if { ... }), which has been the default parser since OPA 1.0 (2024). An older pinned binary that still defaults to v0 needs either an upgrade or the--v0-compatibleflag; don't ship v0-only syntax and call it current.
Design a self-healing telemetry ingestion pipeline: it should detect a failed collector or processor, reroute telemetry to a healthy instance, replay buffered data after a failure, and auto-scale under load, all while exposing its own health so platform engineers can tell when the observability system itself is degraded. What state would you need to track to do this safely, and what stops the remediation logic itself from causing a cascading failure?
Sample Answer
Direct answer
Put a durable, replicated buffer (a log like Kafka, not an in-memory queue) between collectors and processors so a failed component can be detected and rerouted around without losing the data that was in flight. Track three pieces of state to do this safely: per-partition consumer offsets (so replay resumes from the right place, not from zero), per-component health signals (lag, error rate, heartbeat) that drive rerouting decisions, and an idempotency key on every event so a replay after a failure doesn't get double-counted downstream. What stops the remediation logic itself from cascading is a hard budget on how many auto-remediation actions any one component can trigger in a window, with a circuit breaker that pages a human once that budget is exhausted instead of retrying forever.
Structured elaboration
State to track:
- Consumer offsets per partition per consumer group, checkpointed durably, so a restarted processor resumes exactly where it left off.
- Health signals per component: consumer lag, processing error rate, heartbeat/liveness. These feed the detection logic, not just an external dashboard.
- Idempotency keys on events (a stable ID derived from source + sequence number, not regenerated on replay), so downstream stores can deduplicate when the same event is delivered twice after a replay.
Detection and reroute:
- Liveness and readiness probes remove an unhealthy processor from its consumer group; the group rebalances so the partitions it owned are picked up by healthy consumers.
- Rising consumer lag or error rate on a partition, even without an outright liveness failure, is itself a detection signal (this catches degraded-but-technically-alive processors, not just crashed ones).
- Collectors buffer locally and retry with backoff when they can't reach a processor, rather than dropping data immediately.
Replay:
- On restart, a processor resumes from its last committed offset; the durable log's retention window bounds how far back replay can reach.
- For a deliberate, controlled replay (not just crash recovery), a replay controller can reset a consumer group's offset to an earlier point and replay into an isolated processing path first, so a bad replay doesn't double-write into the live output the way replaying directly into the primary consumer group would.
- Stateful processors restore their state from a changelog or external state store before resuming, so replay doesn't start from a stale in-memory state.
Autoscaling: scale on consumer lag and error rate together, not CPU alone. A processor that's CPU-idle but falling behind (for example, downstream I/O-bound) needs more replicas even though CPU utilization alone wouldn't trigger a scale-up.
Degraded-mode operation: when the pipeline itself is under load it can't fully process (for example a partial outage reducing available processor capacity), fail toward reduced fidelity rather than data loss: temporarily increase sampling (drop a configured fraction of lower-priority telemetry) to keep the pipeline within capacity while preserving full-fidelity handling for high-priority signals, then return to full sampling once capacity recovers.
A specific, common failure mode worth naming: a collector with a slow memory leak. This doesn't trip a liveness check the way a crash does; it degrades gradually. The remediation here isn't "restart on OOM" (that's just a symptom-fix that fires only after the leak has already caused damage), it's tracking a per-collector memory trend and proactively rotating (drain, restart) a collector whose memory is trending toward its limit before it OOMs and drops whatever was buffered in-process at the time.
flowchart LR
C[Collector] -- push --> BUF[(Durable log: partitioned, replicated)]
BUF --> P1[Processor pod]
P1 -- heartbeat/lag --> HM[Health manager]
HM -- unhealthy: rebalance --> BUF
HM -- budget exceeded --> CB[Circuit breaker: page human]
HM -- lag rising --> AS[Autoscaler]
AS -- scale replicas --> P1
RC[Replay controller] -- controlled replay --> BUF
RC --> ISO[Isolated replay path]
Worked example
Sizing the durable buffer to survive an outage. Assume the pipeline sustains r=200,000 events/sec at 300 bytes/event, and the target is to absorb up to a 15-minute processor outage without dropping data (detect + reroute + recover within that window):
bufferBytesNeeded=r×(15×60)×300×1.5=200,000×900×300×1.5=81 GB(The 1.5x factor is headroom for more than one component being degraded at once, not just a single clean outage.) With replication factor 3 for durability:
bufferBytesWithRF=81×3=243 GBThat's the concrete number that goes into sizing the durable log's disk footprint for a 15-minute recovery SLO at this ingestion rate: roughly a quarter-terabyte of replicated buffer, not an arbitrary "keep some retention" guess.
Bounding the remediation loop so it can't cascade. Use a token-bucket budget per component: capacity 3 auto-remediation actions (for example, restarts), refilling 1 token every 10 minutes. In the worst case, a flapping component can trigger:
worstCaseActionsPerHour=3+⌊1060⌋=3+6=9 actions/hourbefore the bucket is empty and the circuit breaker opens, escalating to a human instead of continuing to retry. Pair this with capped exponential backoff on reroute retries (base 2s, doubling, capped at 60s):
2,4,8,16,32,60,60,60(sum=242s before giving up on a single reroute attempt)The cap matters here specifically: without it, a component that's actually gone for good would have retry intervals growing unbounded, delaying the eventual "stop retrying, alert a human" decision far longer than necessary.
Trade-offs & pitfalls
| Design choice | Prevents | Costs |
|---|---|---|
| Durable buffer with 15-min absorption (243GB, RF=3) | Data loss during outages up to the target window | Disk and replication cost scale with the target outage window; longer targets get expensive fast |
| Remediation token bucket (9 actions/hr worst case) | Cascading restart storms | A genuinely flapping component still gets 9 restart attempts before paging, which is 9 chances to make things worse if the restart itself is the problem |
| Degraded-mode sampling | Total pipeline overload during partial capacity loss | Silently dropping lower-priority telemetry is only safe if "lower-priority" was actually decided in advance, not improvised during the incident |
Common wrong turns: replaying directly into the live consumer group after a failure instead of an isolated path first, which can double-process events downstream if the failure that triggered the replay wasn't a clean crash (partial writes already landed); sizing the durable buffer for the average outage duration instead of the target recovery SLO, which quietly fails during the outages that matter most; and building auto-remediation without a hard action budget, on the assumption that "it can only help," when a bad remediation action (for example restarting a processor whose real problem is a poison-pill message) can itself be the thing that turns a single-partition issue into a rebalancing storm across the whole consumer group.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Systems Engineer jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs