Backend Developer (Mid-Level) Interview Preparation Guide for Google
Google's backend developer interview process for mid-level candidates typically consists of a recruiter screening phase followed by technical phone screens and onsite rounds. The process assesses algorithmic problem-solving, system design capabilities, coding quality, architectural thinking, and cultural fit. Expect 5-7 total rounds spanning 4-6 weeks, with emphasis on scalability, distributed systems, and production-grade code quality.
Interview Rounds
Recruiter Screening
What to Expect
Initial conversation with a recruiter to verify background, discuss role expectations, and assess cultural fit and motivation. This may be combined with or followed by a brief technical competency screen. The recruiter will discuss your experience with backend systems, technologies you've worked with, and reasons for interest in the role.
Tips & Advice
Be clear about your backend experience and the scale of systems you've worked on. Mention specific technologies and architectural challenges you've tackled. Show genuine interest in Google's infrastructure and problems. Be honest about your level—they're looking for mid-level candidates who can own projects but still have room to grow. Have 2-3 questions ready about the team, tech stack, and infrastructure. Highlight any experience with distributed systems, cloud platforms, or large-scale services.
Focus Topics
Motivation for Google Backend Role
Articulate why you're interested in this specific role at Google, referencing infrastructure challenges, technology stack, or team impact.
Practice Interview
Study Questions
Scale and Complexity Examples
Prepare 2-3 examples of backend systems you've built or worked on, emphasizing scale (traffic, users, data), complexity, and your role.
Practice Interview
Study Questions
Background and Experience Summary
Concise overview of your backend development experience, projects owned, and technologies mastered at mid-level (2-5 years experience).
Practice Interview
Study Questions
Technical Phone Screen
What to Expect
First technical assessment conducted over video/phone with a backend engineer. Usually a single algorithmic or data structure problem with moderate difficulty. The focus is on your problem-solving approach, code quality, communication, and ability to handle feedback. Expect questions involving arrays, strings, trees, graphs, or simple system design concepts.
Tips & Advice
Think aloud to show your thought process. Start with a brute-force solution and optimize incrementally. Discuss time and space complexity. Write clean, readable code as if it will be reviewed in production. Ask clarifying questions before coding. Handle edge cases explicitly. If stuck, ask for hints—it's better than silent struggling. Practice on platforms like LeetCode (Medium difficulty). Use your preferred backend language but ensure syntax is accurate.
Focus Topics
Problem-Solving Communication
Articulating your approach, explaining trade-offs, asking clarifying questions, and discussing complexity analysis clearly.
Practice Interview
Study Questions
Code Quality and Best Practices
Writing clean, readable, production-grade code with proper naming, error handling, comments, and testability in backend context.
Practice Interview
Study Questions
Data Structures and Algorithms Fundamentals
Solid understanding of arrays, linked lists, trees, graphs, hash tables, heaps, and sorting/searching algorithms. Ability to choose appropriate data structures for problems.
Practice Interview
Study Questions
System Design Round - Core Architecture
What to Expect
In-depth system design interview focusing on designing a scalable backend system. You'll be given a real-world scenario (e.g., design a rate limiter, notification system, payment processor, or collaborative editing backend) and asked to architect a solution from scratch. Expect 45-60 minutes. You should cover requirements gathering, high-level architecture, API design, database schema, scaling strategies, and trade-offs. This is critical for mid-level roles—you're expected to own system design end-to-end, not just implement components.
Tips & Advice
Start by clarifying requirements and constraints—ask about scale (QPS, data volume), consistency requirements, and latency expectations. Draw architecture diagrams showing components, databases, caches, and message queues. Discuss specific technologies (e.g., PostgreSQL vs. MongoDB, Redis for caching). Identify bottlenecks and propose solutions (sharding, replication, caching, load balancing). Be prepared to deep-dive on any component. Explain trade-offs (consistency vs. availability, latency vs. cost). Practice designing: rate limiters (token bucket, sliding window), payment systems (idempotency, saga pattern), notification systems (message queues, fan-out), and collaborative editors (OT vs. CRDTs). Reference real systems like Google's infrastructure patterns.
Focus Topics
Trade-off Analysis and Justification
Articulating decisions between consistency vs. availability, latency vs. throughput, cost vs. performance, and justifying choices based on requirements.
Practice Interview
Study Questions
System Reliability and Failure Handling
Designing for fault tolerance, implementing retry logic with exponential backoff, circuit breakers, graceful degradation, dead letter queues, and disaster recovery strategies.
Practice Interview
Study Questions
Database Design and Optimization
Choosing appropriate databases (SQL vs. NoSQL), designing schemas, indexing strategies, query optimization, sharding, replication, and backup/recovery approaches.
Practice Interview
Study Questions
Distributed Systems and Scalability
Horizontal scaling, load balancing, caching strategies (Redis, Memcached), message queues (Kafka, RabbitMQ), asynchronous processing, and handling eventual consistency.
Practice Interview
Study Questions
API Design (REST and gRPC)
Designing RESTful APIs with proper resource modeling, HTTP methods, status codes, pagination, versioning, rate limiting headers, and idempotency. Understanding when to use gRPC for internal services.
Practice Interview
Study Questions
System Design Round - Advanced Concepts
What to Expect
Second system design round (or deep-dive follow-up) focusing on advanced backend concepts. May involve designing for specific constraints (high throughput, strict consistency, real-time sync), implementing complex features, or discussing infrastructure at scale. Topics might include exactly-once semantics, event sourcing, CQRS pattern, distributed transactions, microservices architecture, or real-time collaboration. This round assesses depth of understanding and ability to handle complex trade-offs.
Tips & Advice
Be comfortable discussing advanced patterns and their trade-offs. For example, understand idempotency keys for exactly-once processing, saga patterns vs. two-phase commit for transactions, and event sourcing for audit trails. Know when to use microservices vs. monolith. Discuss monitoring, observability, and incident response. Reference patterns used in real systems (Google's Spanner, Dremel, Colossus). Explain how you'd debug and optimize a system under load. Be ready to discuss security (authentication, authorization, data encryption) and compliance considerations.
Focus Topics
Observability, Monitoring, and Incident Response
The four golden signals (latency, traffic, errors, saturation), structured logging, distributed tracing (OpenTelemetry), metrics, alerting, and on-call best practices.
Practice Interview
Study Questions
Distributed Transactions and Consistency Patterns
Saga pattern, two-phase commit, eventual consistency, compensation logic, and handling partial failures in distributed transactions.
Practice Interview
Study Questions
Security and Compliance in Backend Systems
Authentication (OAuth, JWT), authorization (RBAC, ABAC), data encryption at rest and in transit, PCI compliance for payments, audit logging, and secure API design.
Practice Interview
Study Questions
Event-Driven Architecture and Message Queues
Understanding publish-subscribe patterns, event sourcing, message queues (Kafka, Pub/Sub), dead letter queues, and asynchronous communication between services.
Practice Interview
Study Questions
Exactly-Once Semantics and Idempotency
Designing systems that guarantee exactly-once processing using idempotency keys, understanding replay semantics, and handling duplicate requests in distributed systems.
Practice Interview
Study Questions
Coding Round - Advanced Problems
What to Expect
On-site coding interview testing advanced algorithmic and backend-specific problem-solving. Problems may involve graph algorithms, dynamic programming, concurrent data structures, stream processing, or realistic backend scenarios (e.g., rate limiting implementation, dependency resolution, data processing pipelines). Expect 45-60 minutes to solve 1-2 problems. Emphasis is on clean, production-quality code, proper error handling, edge cases, and efficiency. You may need to discuss testing approach and code organization.
Tips & Advice
Focus on writing code that's correct, efficient, and maintainable. Discuss approach before coding. Walk through examples and edge cases. Handle errors explicitly (null checks, invalid inputs). Write comments for non-obvious logic. Optimize after getting a working solution. Be prepared to discuss test cases. For backend-specific problems, think about concurrency, error handling, and scalability. If you get stuck, explain your thought process and ask for hints. Practice on LeetCode Hard problems and backend-specific scenarios.
Focus Topics
Error Handling and Edge Cases
Anticipating failure scenarios, handling null/empty inputs, boundary conditions, timeouts, and writing robust error messages.
Practice Interview
Study Questions
Dynamic Programming
Recognizing DP problems, memoization, bottom-up approaches, and solving optimization problems efficiently.
Practice Interview
Study Questions
Concurrency and Thread Safety
Implementing thread-safe data structures, handling race conditions, deadlocks, using locks/mutexes, and understanding concurrent programming paradigms.
Practice Interview
Study Questions
Graph Algorithms and Problem Solving
BFS, DFS, topological sorting, shortest path algorithms, cycle detection, and applying graphs to real problems like dependency resolution.
Practice Interview
Study Questions
Behavioral and Culture Fit Round
What to Expect
Interview focusing on past experiences, teamwork, learning ability, problem-solving approach, and alignment with Google's culture and values (innovation, collaboration, user focus). You'll be asked about specific projects, challenges overcome, conflicts resolved, and situations demonstrating initiative and growth. Expect questions about your biggest achievement, a failure and lessons learned, how you handle disagreement, and why you want to work at Google. This round assesses soft skills, maturity, and cultural fit critical for mid-level roles where you own projects and collaborate across teams.
Tips & Advice
Prepare STAR-format stories (Situation, Task, Action, Result) from your experience demonstrating: ownership of complex projects, mentoring or helping junior colleagues, handling ambiguity, making tough technical decisions, collaborating with cross-functional teams, and recovering from failures. For mid-level, emphasize project ownership and technical leadership, not individual contribution. Be specific with numbers and impact. Practice saying these stories concisely (2-3 minutes each). Research Google's values and culture—be genuine about alignment. Ask thoughtful questions showing you've researched the team and role. Be authentic about both strengths and growth areas. Avoid canned answers; be conversational.
Focus Topics
Alignment with Google Culture
Understanding Google's values (innovation, collaboration, user focus) and providing genuine examples of how your approach aligns.
Practice Interview
Study Questions
Learning and Growth from Challenges
Stories about failures, mistakes made, or challenging situations, focusing on what you learned and how you applied those lessons.
Practice Interview
Study Questions
Collaboration and Mentoring
Examples of working effectively with teammates, mentoring junior engineers, resolving technical disagreements, and contributing to team decisions.
Practice Interview
Study Questions
Project Ownership and End-to-End Impact
Stories demonstrating ownership of backend projects from design through deployment, making key decisions, and delivering measurable impact.
Practice Interview
Study Questions
Technical Decision-Making and Trade-offs
Examples of technical decisions made (technology choices, architecture decisions), reasoning, trade-offs evaluated, and outcomes.
Practice Interview
Study Questions
Frequently Asked Backend Developer Interview Questions
Request threads verify tokens with keys that a background job rotates in memory. How do you make sure readers always see a consistent set, with no pause in traffic and no half-updated state?
Sample Answer
Direct answer
Never edit the key set that readers are using. Build a complete new key set off to the side, then publish it with one atomic swap of a single reference. Each request reads that reference once and uses the resulting immutable snapshot (an object that is never modified after it is built) for the whole verification. Readers take no lock, so traffic never pauses, and because a snapshot cannot change, no reader can ever see a half-updated set.
Why in-place updates fail
Suppose the rotation job does keys.clear() and then keys.put(newKid, newKey) on a map the request threads are reading. Between those two calls a reader can find the map empty and reject a perfectly valid token. With a plain HashMap it is worse: a lookup racing a resize is a data race (two threads touch the same memory with no ordering between them and at least one writes), which can return wrong answers or loop forever. Wrapping every read in a lock fixes consistency but makes every request queue behind the rotation, which is the pause you were told to avoid.
The pattern, step by step
- Immutable value. One object holds everything a verification needs: the id of the current signing key and a map from key id to key (the key id,
kidin the code, is the short label a token carries to say which key signed it;newKidis the label of the key being introduced). Build it withMap.copyOfso nothing can change it later. - One shared mutable cell. A single
volatilereference to the current snapshot.volatileis the Java keyword that makes every read and write of that one field go straight to shared memory in a well-defined order, instead of being cached by a thread. In Java, a write to avolatilefield happens-before every later read of that field (Oracle'sjava.util.concurrentpackage documentation), and "happens-before" means anything the writer did before the write is guaranteed visible to a reader that sees the new value. So the reader sees every field of the fully built snapshot, not a partly constructed one: the writer finished building the snapshot, then did the volatile write, so a reader that sees the new reference sees the finished object behind it. - Reader reads the reference once.
Snapshot s = ring;then usesfor the lookup and the check. Readingringtwice in one request can return two different snapshots, which re-creates the inconsistency you were avoiding. - Writer builds, then swaps. Create the new snapshot completely, validate it (key parses, id not already in use), and only then assign
ring = next. If building fails, the old snapshot simply stays live. With one rotation thread a plain assignment is enough; with several writers useAtomicReference.updateAndGet, which retries a compare-and-set loop so no update is lost: it computes a new value from the current one and installs it only if the reference still holds that same current value (compare-and-set); if another writer got in first, it recomputes from the newer value and tries again. - Overlap the keys. Tokens signed with the previous key are still in flight, so the new snapshot keeps the previous key until the longest token lifetime plus allowed clock skew (the difference between two machines' clocks, which makes a token look slightly older or newer than it is) has passed. Order matters across machines too: publish the new verification key to verifiers before the signer starts using it.
Worked example, run in a Linux container
The program performs 19,998 rotations (the loop runs n = 2 up to 19,999) while four reader threads check, on every loop, that the key id they just read can be found. It runs the snapshot design and the naive clear-then-put design side by side.
import java.util.*;
import java.util.concurrent.*;
import java.util.concurrent.atomic.*;
public class KeyRing {
// Immutable snapshot: never modified after construction.
record Snapshot(String currentKid, Map<String, String> keysByKid) {
Snapshot { keysByKid = Map.copyOf(keysByKid); }
}
// One volatile reference is the only shared mutable cell.
static volatile Snapshot ring = new Snapshot("k1", Map.of("k1", "secret-1"));
static void rotate(int n) {
Snapshot old = ring;
String newKid = "k" + n;
Map<String, String> next = new HashMap<>();
next.put(newKid, "secret-" + n);
next.put(old.currentKid(), old.keysByKid().get(old.currentKid())); // keep previous key for the grace window
ring = new Snapshot(newKid, next); // single atomic publish
}
// Naive version: mutate one shared map in place.
static final Map<String, String> naive = new ConcurrentHashMap<>(Map.of("k1", "secret-1"));
static volatile String naiveCurrent = "k1";
static void naiveRotate(int n) {
String prev = naiveCurrent;
String prevSecret = naive.get(prev);
naive.clear(); // readers can see an empty map here
naive.put(prev, prevSecret);
naive.put("k" + n, "secret-" + n);
naiveCurrent = "k" + n;
}
public static void main(String[] a) throws Exception {
final int READERS = 4, ROTATIONS = 20000;
AtomicBoolean stop = new AtomicBoolean();
AtomicLong snapChecks = new AtomicLong(), snapBad = new AtomicLong();
AtomicLong naiveChecks = new AtomicLong(), naiveBad = new AtomicLong();
List<Thread> ts = new ArrayList<>();
for (int i = 0; i < READERS; i++) {
Thread t = new Thread(() -> {
while (!stop.get()) {
Snapshot s = ring; // read the reference ONCE
snapChecks.incrementAndGet();
if (s.keysByKid().get(s.currentKid()) == null) snapBad.incrementAndGet();
// naive reader: look up the key id it just read
String kid = naiveCurrent;
naiveChecks.incrementAndGet();
if (naive.get(kid) == null) naiveBad.incrementAndGet();
}
});
t.start(); ts.add(t);
}
for (int n = 2; n < ROTATIONS; n++) { rotate(n); naiveRotate(n); }
stop.set(true);
for (Thread t : ts) t.join();
System.out.println("snapshot reads with missing current key: " + snapBad.get());
System.out.println("naive reads with missing current key > 0: " + (naiveBad.get() > 0));
}
}
Run with docker run --rm -v "$PWD":/w -w /w eclipse-temurin:21-jdk java KeyRing.java (JDK 21):
snapshot reads with missing current key: 0
naive reads with missing current key > 0: true
The snapshot count is 0 because a reader only ever holds a complete snapshot. The naive design printed true, meaning at least one reader found the map empty or missing its own key id (the exact count varies run to run, so the program prints only whether it happened).
Trade-offs and pitfalls
- Cost. Every rotation allocates a small new map. Rotations happen every few hours or days and the map holds a handful of keys, so this is negligible. Copy-and-swap is the wrong tool only for large sets that change on every request.
- Old snapshots live while someone holds them. In a garbage-collected runtime that is free: the old snapshot is reclaimed when the last in-flight request drops it. In C or C++ you need an equivalent, such as
std::atomic<std::shared_ptr>(C++'s reference-counted pointer that can be swapped atomically; the old snapshot is freed when the last holder lets go) or RCU-style deferred freeing (read-copy-update: wait until every reader that could hold the old pointer has finished), otherwise you free a snapshot a reader is still using. - Unknown key id. A token naming a key id not in the snapshot should fail fast. If you refresh keys from a remote endpoint on a miss, rate-limit that refresh, or an attacker can force a fetch per request.
- Emergency revocation. If a key is compromised, publish a snapshot without it immediately, accepting that its tokens fail. That is the one case where you deliberately skip the overlap.
- Zeroing secrets. You cannot overwrite an old key's bytes while readers may still be using that snapshot. Drop the reference and let the runtime reclaim it, and keep keys out of logs and heap dumps (a heap dump is a file containing a copy of the process's memory, which would include any key still in it).
- What would change this choice. If the key set must change on every request, or readers need to see a write the moment it happens, use a concurrent map or a lock instead.
What is the difference between 'culture fit' and 'culture add', and which do you think better describes you as a candidate? Give one concrete example of a perspective, skill, or way of working you would bring to a team that is not already well represented there.
Sample Answer
Direct answer
Culture fit asks whether you already share a team's existing norms and behaviors; culture add asks what you would bring that the team does not already have. I would describe myself mostly as a culture add: I share the fundamentals a team needs to trust me (reliability, candor, respect for other people's time), but the useful thing I offer beyond that is a genuinely different working background rather than a mirror of the team that is already there.
Structured elaboration
- Define both terms precisely before answering for yourself. Culture fit is about alignment on shared behaviors and values: does this person operate the way we already operate. Culture add is about complementary difference: does this person's background, working style, or perspective fill a gap the team doesn't currently have.
- Explain why the distinction matters, not just define it. A team optimized purely for fit tends toward groupthink: everyone reasons the same way, so blind spots go unchallenged and the same kinds of mistakes recur. A team that only adds without any shared fit becomes uncoordinated: people can't predict each other's reasoning enough to move fast together. The healthy target is fit on a small number of load-bearing behaviors (honesty, follow-through, respect) plus deliberate add on everything else.
- Give a genuine, specific example of your own add, not a generic trait. Vague claims ("I bring diverse perspectives") are the single most common failure mode here; a strong answer names the concrete gap and the concrete evidence.
- Anticipate the natural follow-up: how do you know your difference is actually useful, versus just different for its own sake. The answer is to point at a specific decision, disagreement, or piece of feedback that changed because of the difference you brought, not just a credential or background fact.
Worked example
Suppose your last two teams were both product engineering teams building consumer-facing features, and the team you're interviewing for is mostly staffed by engineers with that same background. Your own prior role was on a data-platform team, closer to the systems that feed those consumer features than to the features themselves. A concrete add-story: in a past project, a product team wanted to ship a new recommendation feature quickly; because of your platform background, you asked a question the rest of the team hadn't raised (whether the upstream data pipeline's freshness guarantees actually matched what the feature's UI implied to users), which surfaced a real gap between a 24-hour batch refresh and a UI copy that said "updated just for you." The team fixed the copy and adjusted the refresh cadence before launch rather than after a user complaint. That is a genuine add: a different background produced a question the existing team composition was less likely to ask on its own, and it changed a real outcome.
Trade-offs & pitfalls
The common failure is answering only the definitional half (correctly explaining fit versus add) and then, when asked for a personal example, retreating to generic self-description ("I'm a good communicator", "I care about quality") that any candidate could say and that does not actually demonstrate difference. A second pitfall is overcorrecting into implying you don't fit at all; the strongest answers are explicit that you also share the small set of behaviors every functioning team needs, and that add is about everything on top of that baseline, not a replacement for it.
You're picking an approximate nearest-neighbor index for a semantic search feature with 50 million embeddings. How do you trade off recall, query latency, and memory when tuning it?
Sample Answer
Direct answer
Approximate nearest-neighbor (ANN) indexes trade exact recall for speed and memory by avoiding a full comparison against every vector. The tuning knobs let you buy back recall by spending more query latency and memory, so the right setting is whichever point on that curve meets your product's minimum acceptable recall at the lowest cost.
Structured elaboration
Exact (brute-force) search gives 100% recall but scales linearly with corpus size, too slow at 50 million vectors for interactive search. Graph-based ANN indexes like HNSW (Hierarchical Navigable Small World: a graph structure where each stored vector links to a handful of its nearest neighbors, so a query can hop from node to node toward the target region instead of comparing against every vector) approximate the search, with a search-width parameter controlling how much of the graph is explored per query, raising both recall and latency together as it increases. Index memory scales with vector count and dimensionality, since it holds both the vectors and the graph edges.
Worked example
768-dimensional embeddings stored as 32-bit floats:
bytes per vector=768×4=3,072 bytes≈3KB
raw vector memory=50,000,000×3KB≈143GB
Assuming a graph-based index adds 30% overhead for its structure:
index memory≈143GB×1.3≈186GB
a real decision between one high-memory machine and a sharded cluster. If a lower-recall setting lets you quantize vectors to 8-bit integers instead of 32-bit floats:
raw vector memoryint8=4143GB≈36GB
roughly a 4x memory reduction, at the cost of some recall loss on top of whatever the ANN approximation already gave up.
The recall-versus-latency side of the trade-off is just as traceable with illustrative figures (not measured against a real dataset here, since the actual curve depends on corpus and hardware, but self-consistent so the shape is real): if each additional graph node the search visits costs a fixed per-candidate distance computation, say roughly 0.02ms, then a low search-width setting of 64 nodes costs about 64×0.02ms≈1.3ms and typically lands recall in the low 90s (~90%); a higher setting of 256 nodes costs about 256×0.02ms≈5.1ms and typically reaches the high 90s (~98%); pushing further to 512 nodes roughly doubles the cost again to 512×0.02ms≈10.2ms for only a small further recall gain (~99.3%), the same flattening-curve pattern discussed below.
Trade-offs and pitfalls
Chasing a small recall gain near the top of the curve, say 97% to 99%, often costs a disproportionate amount of latency and memory, because ANN recall curves flatten sharply. Measure recall against a labeled ground-truth set rather than eyeballing results, and re-measure after any re-embedding, since a new embedding model changes what "correct" nearest neighbors even are.
What the interviewer probes next
Expect questions on how you'd build a ground-truth recall evaluation set, sharding once the index outgrows one machine's memory, and how you'd handle index updates when embeddings change incrementally rather than in one batch.
Explain Command Query Responsibility Segregation (CQRS). As a data engineer, when is CQRS valuable for analytics or operational workloads? Discuss trade-offs including complexity, eventual consistency of read-models, and strategies to make reads 'fresh' when required.
Sample Answer
Direct answer
Command Query Responsibility Segregation (CQRS) is the pattern of using a different model for writes (commands that change state) than for reads (queries that return state), instead of forcing one schema to serve both well. As a data engineer, it earns its added machinery when the write side's transactional shape and the read side's analytical or lookup shape diverge enough that one schema serves neither well, for example narrow row-level operational writes versus wide, pre-aggregated reporting reads. Treat it as a deliberate trade: schema simplicity for the ability to scale, model, and store reads and writes independently.
Structured elaboration
What CQRS actually separates
- Command side: an authoritative model (a normalized transactional store, or an event log) that enforces write-time invariants and produces state changes.
- Query side: one or more purpose-built read models (denormalized tables, search indexes, in-memory caches), each shaped for a specific access pattern rather than for correctness enforcement.
- A projector connects the two asynchronously: it consumes the write side's changes (domain events, or change-data-capture (CDC) records) and updates the read model(s).
When CQRS is valuable for analytics workloads
- Reporting or business intelligence (BI) queries need aggregation shapes (rollups by category, hour, region) that would otherwise require expensive joins or full scans against the transactional schema.
- Several independent consumers need different projections of the same data (a finance rollup, a fraud-detection view, a customer dashboard); three denormalized read models are cheaper to operate than three sets of ad-hoc joins against the online transaction processing (OLTP) store.
- Analytical queries would otherwise contend for locks and I/O with operational writes on the same tables.
When CQRS is valuable for operational workloads
- Write throughput and read throughput need to scale independently and at different rates (write-heavy ingestion feeding a low-cardinality operational dashboard).
- Write-side invariants are complex enough (state machines, multi-step validation) that mixing them with read-optimization concerns would make the write model harder to reason about.
- Not valuable: a small application with one read pattern that already matches the write schema. There CQRS adds a projector, extra storage, and extra failure modes with no offsetting benefit.
Trade-off: complexity
You now operate an additional pipeline (the projector), additional storage (one or more read stores), and additional failure modes: projector lag, projector crashes mid-batch, and schema drift between the write shape and the read shape.
Trade-off: eventual consistency of read models
Because the read model updates asynchronously, a query issued immediately after a write can observe stale data. The size of that staleness window is a direct function of projector throughput and batching, not something a team can design around by ignoring it.
Strategies to make reads "fresh" when required
- Read-your-writes for the writer: return enough state in the command response (or a version/sequence number) that the client who just wrote never needs to trust the read model for its own write.
- Tighten the pipeline: smaller batches and event-driven push instead of periodic batch pull shrinks the staleness window, at the cost of more frequent projector invocations.
- Expose staleness explicitly: attach a last-updated version or timestamp to read-model responses so callers can judge whether the data is fresh enough, instead of the system silently presenting stale data as current.
- Selective synchronous update: for a small, well-identified set of critical fields, update the read model synchronously in the write path (accepting some coupling) while everything else stays asynchronous.
Worked example
An order system accepts writes at 500 orders per minute (about 8 to 9 orders per second) into a transactional order table. A "revenue by category, per hour" read model is built by a projector that drains the order-events stream every 60 seconds and applies that batch of updates.
- Worst-case staleness for that read model equals the batch interval: 60 seconds. An order committed just after a batch run will not appear until the next run.
- Average staleness is roughly half the batch interval, about 30 seconds, if orders arrive close to uniformly across the minute.
If the product requirement is "the dashboard must reflect a new order within 10 seconds," this projector cadence fails outright: 60 seconds worst case exceeds the 10-second bound. The fix is either to drop the batch interval below 10 seconds, or to read the specific "orders placed today" counter synchronously from the write side while the rest of the dashboard stays on the 60-second cadence.
Trade-offs and pitfalls
- Common wrong turn: adopting CQRS because it sounds architecturally sophisticated for a workload that has a single read pattern already matching the write schema. That is pure overhead with no payoff.
- Common wrong turn: treating "eventually consistent" as a detail to sort out later. Staleness needs an explicit, stated bound (or an explicit "no bound" with a user-facing affordance for it) decided at design time, not discovered in production when a user cannot see the order they just placed.
- Senior signal: naming a concrete staleness budget and matching the pipeline's cadence to it, rather than discussing CQRS only in the abstract.
A mentee becomes defensive, or pushes back hard, whenever you give them feedback, and stops acting on your suggestions. How do you handle it?
Sample Answer
Direct answer
When a mentee gets defensive and stops acting on feedback, the fastest way to make it worse is to double down with more direct feedback. Slow down, diagnose why the message isn't landing (the content, the delivery, or something the mentee brings into the room), then rebuild the conversation as a two-way one instead of a one-way correction. If the pattern doesn't shift after a genuine attempt at that, it needs to be named and escalated, not quietly tolerated.
Diagnose before you re-deliver
- Separate "defensive because of how I said it" from "defensive because of what's underneath it." Workload, unclear expectations, a confidence hit, or feedback that reads as a character judgment rather than a specific behavior all produce the same surface symptom (pushback, non-action) for different reasons.
- Ask, don't assume: open with a genuinely curious question rather than a repeat of the critique. "Walk me through how that landed for you" gets you information; "you need to stop being defensive" gets you more defensiveness.
Use motivational interviewing instead of more direct pressure
- Motivational interviewing is built for exactly this: someone who may intellectually agree but is resisting behaviorally. Instead of arguing for the change, reflect their own stated goals back to them and let them articulate the gap ("You mentioned you want to lead the next project. How does this pattern affect that?"). People act on reasons they generate themselves far more than reasons handed to them.
- Keep the ratio of affirmation to correction visible. If every interaction is corrective, the mentee starts hearing footsteps before you speak, which is what produces reflexive defensiveness.
Rebuild the mechanism, not just the next conversation
- Shrink the ask: instead of a broad critique, propose one small, concrete, reversible change and a short check-in window.
- Make feedback bidirectional: ask what kind of feedback has landed well for them before, and adjust format (written vs. verbal, immediate vs. batched) accordingly.
Know when coaching has run its course
- If, after two or three honest attempts using the above, the pattern is unchanged (commitments still not acted on, same defensiveness), that's a signal the issue may be outside what coaching alone fixes: a skill gap being misread as attitude, a values or fit mismatch, or a factor you're not positioned to see.
- At that point, loop in the mentee's manager, or HR if the dynamic has become adversarial, rather than continuing to privately absorb it. Frame it factually: what you tried, what changed, what didn't. This isn't giving up on the mentee; it's recognizing some situations need authority or context you don't have.
Worked example
A mentee kept missing agreed follow-ups on code review comments and would get visibly short in Slack whenever it came up. The instinct was to restate the same feedback more firmly. Instead, the better move: open the next 1:1 with "I want to understand how the review feedback has been landing for you, not go through it again," and listen first. It turned out the mentee had inherited a legacy module nobody had explained well, and every review comment felt like it was pointing out someone else's mess. The fix wasn't more feedback, it was pairing on the module once and shrinking the ask to one file at a time. If that hadn't worked, the next honest step would have been raising the pattern with the mentee's manager, not repeating the same conversation a fourth time.
Trade-offs and pitfalls
- The junior mistake is treating defensiveness as a discipline problem and pushing harder; that reliably produces more resistance, not less.
- Over-correcting the other way (going silent on real issues to avoid triggering defensiveness) just delays the same conversation and lets performance drift.
- Escalating too early, before you've tried adjusting your own approach, reads as offloading a coaching problem; escalating too late lets a stalled dynamic damage trust or delivery. The senior move is trying a genuine adaptation first, timeboxing it, and being honest about whether it moved anything.
Walk through the trade-off between synchronous and asynchronous replication. What does each cost you in write latency, and what does each risk during a failover?
Sample Answer
Synchronous replication waits for the replica (or a quorum of replicas) to acknowledge a write before telling the client the write succeeded, so it costs extra write latency in exchange for near-zero data loss (a near-zero RPO, recovery point objective: how much data, measured in time, you could lose in a failure). Asynchronous replication acknowledges the write as soon as it's durable on the primary and ships it to replicas afterward, so writes stay fast but a failover can lose whatever hadn't shipped yet.
Comparing the two
| Dimension | Synchronous | Asynchronous |
|---|---|---|
| Write latency | Local write + round-trip to replica(s) before ack | Local write only; replication happens after the client is told "done" |
| RPO on failover | Near-zero for acknowledged writes (they're already on the replica) | Bounded by replication lag at the moment of failure |
| Throughput | Bounded by the slowest replica in the acknowledgment path | Not bounded by replica speed; primary can run at its own pace |
| Behavior under partition | Can block writes entirely if the required replica/quorum is unreachable (trades availability for durability) | Keeps accepting writes on the primary; risks divergence if the primary later turns out to be on the wrong side of the partition |
| Typical use | Financial ledgers, inventory decrements, anything where losing an acknowledged write is unacceptable | Read replicas, cross-region DR copies, analytics/logging pipelines, caches |
Worked example: latency and RPO, with pinned assumptions
Pin a local write (fsync to disk) at 2 ms, a round-trip time to a same-region, cross-AZ replica at 4 ms, and a round-trip time to a cross-region replica at 70 ms (all stated as inputs for this comparison, not measurements of any specific vendor).
Synchronous, cross-AZ:
write latency=2ms (local)+4ms (RTT to replica)=6msThat's 3x the async latency of 2 ms. Acceptable for most OLTP systems.
Synchronous, cross-region:
write latency=2ms (local)+70ms (RTT to replica)=72msThat's 36x the async latency, which is why synchronous replication across regions is rare in practice for user-facing writes; the pattern that actually ships is synchronous within a region (to survive an AZ failure with RPO≈0) and asynchronous across regions (to survive a regional disaster, accepting a small RPO).
Quorum framing (this is where "synchronous" gets more precise than "one replica acks"): with N=3 replicas requiring a write quorum of W=2 (majority), a write only needs to wait for the fastest W−1=1 of the 2 non-primary replicas to ack, not all of them, which caps the latency cost at the RTT to whichever replica answers first rather than the slowest one. That's the practical reason quorum-based sync replication (Raft, Paxos-style commit) is preferred over "wait for every replica": it keeps the durability guarantee while bounding the latency tail.
Asynchronous RPO: if replication lag under normal load is 2 seconds but backs up to 30 seconds under a write burst, a failover during that burst loses up to 30 seconds of acknowledged-to-the-client-but-not-yet-replicated writes, i.e. RPO≈replication lag at failure time, not a fixed number, which is exactly why teams monitor lag continuously rather than relying on the steady-state figure.
Trade-offs and pitfalls
The pitfall in the synchronous column isn't just latency, it's availability: a strict "wait for every replica" policy means a single slow or unreachable replica can stall every write on the primary, which is why real systems use quorum semantics (wait for a majority, not all) instead. The pitfall on the async side is treating "eventually consistent" as "eventually correct": if the primary accepts writes during a partition and then loses a leader election, those writes can simply vanish, so any system using async replication for anything beyond caches or analytics needs a defined reconciliation or conflict-resolution story, not just "replication will catch up." A common wrong turn is picking one mode globally instead of matching it to the data: a payments write path and an analytics event stream in the same system usually deserve different replication modes, not the same one applied uniformly for simplicity.
Write a query to find duplicate rows on a natural key (for example, the same email or the same combination of columns appearing more than once). Show both the GROUP BY / HAVING COUNT(*) > 1 form and the ROW_NUMBER() OVER (PARTITION BY ... ORDER BY ...) form that lets you keep exactly one canonical row per group, and explain when you would reach for each.
Sample Answer
Both forms answer "which rows are duplicates" but only the window-function form also tells you which single row to keep, so use `GROUP BY`/`HAVING` for a quick existence check and `ROW_NUMBER()` when you need to actually resolve to one canonical row.
GROUP BY / HAVING form
```sql
SELECT email, COUNT() AS n
FROM accounts
GROUP BY email
HAVING COUNT() > 1;
```
ROW_NUMBER window-function form
```sql
WITH ranked AS (
SELECT *, ROW_NUMBER() OVER (PARTITION BY email ORDER BY created_at DESC) AS rn
FROM accounts
)
SELECT account_id, email, created_at FROM ranked WHERE rn = 1;
```
The `PARTITION BY email` groups rows the same way `GROUP BY` would, but `ROW_NUMBER()` additionally assigns a rank within each group, so filtering to `rn = 1` (with `ORDER BY created_at DESC` to break ties by recency) keeps exactly one row per email and discards the rest, which the `GROUP BY` form alone cannot express in a single pass. The same window-function query can be dropped straight into a `DELETE ... WHERE rn > 1` (via a CTE or subquery, depending on the engine) to physically remove the duplicates.
Worked example
Given `accounts` with two rows for `a@x.com` (created 2026-01-01 and 2026-01-02) and one row each for `b@x.com` and `c@x.com`: the `GROUP BY` form returns one row, `(a@x.com, 2)`. The `ROW_NUMBER` form returns three rows: `a@x.com`'s later (2026-01-02) record plus `b@x.com` and `c@x.com`, correctly keeping the most recent `a@x.com` row and dropping the older one.
Trade-offs and pitfalls
An exact-string duplicate definition misses near-duplicates caused by casing or whitespace (`'A@x.com'` vs `'a@x.com '`); normalize with `LOWER(TRIM(email))` in the partition key if that's a real risk in your data. Ties on the ORDER BY column (two rows with the identical `created_at`) make the `rn = 1` choice arbitrary unless you add a deterministic tie-breaker, like the primary key, as a second ORDER BY term. Finally, this check generalizes well: parametrize the table name and key columns and you can run the same query across many tables rather than writing one bespoke check per table.
Given 5 replicas per partition, model read and write availability under three quorum configurations: majority (R=3,W=3), read-optimized (R=2,W=4), and write-optimized (R=4,W=2). Assuming an independent per-node failure probability, express read and write availability for each configuration and explain the practical latency and consistency implications of the choice.
Sample Answer
With 5 replicas and R + W > 5 for every configuration worth choosing, all three quorum configurations give the same consistency guarantee: an acknowledged read is guaranteed to overlap an acknowledged write in at least one replica. The choice is really about where you want the availability and latency cost to land, not about correctness. Read-optimized (R=2, W=4) makes reads cheap and highly available at the cost of write availability and latency; write-optimized (R=4, W=2) is the mirror image; majority (R=3, W=3) splits the cost evenly. The size of that trade-off is a binomial-tail calculation, worth doing rather than eyeballing.
The model
Let a be the probability a single replica is both alive and reachable from the coordinator for a given operation, independent and identically distributed across the 5 replicas. The probability that at least k of n replicas are available is the binomial upper tail:
A(k,n,a)=i=k∑n(in)ai(1−a)n−iRead availability for a configuration is A(R,5,a); write availability is A(W,5,a).
Consistency: for all three configurations, R+W=6>n=5, so a successful read set and a successful write set are guaranteed to share at least one replica; that shared replica is what lets an acknowledged read observe the most recent acknowledged write, assuming the coordinator correctly resolves conflicting versions on read (by timestamp or vector clock).
Worked computation
Pin a=0.95: a single replica has an independent 5% chance of being down or unreachable for a given request. Then A(2,5,0.95) derives as:
A(2,5,0.95)=(25)0.9520.053+(35)0.9530.052+(45)0.9540.051+(55)0.955 =0.001128+0.021434+0.203627+0.773781=0.999970The other five values in the table below follow the same expansion with different bounds:
| Configuration | A_read = A(R,5,0.95) | A_write = A(W,5,0.95) | Read unavailability | Write unavailability |
|---|---|---|---|---|
| Majority (R=3, W=3) | 0.998842 | 0.998842 | 0.1158% | 0.1158% |
| Read-optimized (R=2, W=4) | 0.999970 | 0.977407 | 0.0030% | 2.2593% |
| Write-optimized (R=4, W=2) | 0.977407 | 0.999970 | 2.2593% | 0.0030% |
At this per-replica reliability, majority sits in the middle at about 0.12% unavailability on both sides. Read-optimized cuts read unavailability by roughly 40x versus majority (0.0030% vs 0.1158%) but write unavailability rises nearly 20x (2.2593% vs 0.1158%); write-optimized is the exact mirror. For a read-heavy workload like a product catalog or a social feed, read-optimized buys a meaningfully more available read path at a write-availability cost that matters far less because writes are rare; for a write-heavy workload like event ingestion, write-optimized does the same in reverse.
R and W also set how many replicas an operation must wait for, not just how many must be reachable: R=2 only needs the two fastest responses, so read tail latency tracks the second-fastest replica, while R=4 tracks the fourth-fastest, which is usually a much bigger latency cost than the availability arithmetic alone suggests.
Trade-offs and pitfalls
This model assumes independent, identically distributed replica failures, which is the least realistic part of it: a rack or region outage takes down multiple replicas at once, and if replicas are spread out to make that unlikely, cross-region network latency then dominates real-world quorum latency far more than the independent-availability math predicts. Correlated failure modes shrink the effective advantage of any of these configurations relative to what the independent-failure formula suggests, sometimes sharply. Mean availability also hides tail latency: a configuration with high average availability can still have poor P99 (99th percentile) read latency if the quorum happens to include a consistently slow replica, something the availability formula does not capture at all. Finally, R + W > n only buys the overlap guarantee for a single key's operations as seen by a correct coordinator; it says nothing about cross-key transactions, and it assumes the coordinator resolves conflicting versions correctly in the first place.
An enterprise needs eventual consistency between service A and service B using events. Design an idempotent event processing and reconciliation strategy that guarantees convergence and supports replays, while preserving ordering where necessary.
Sample Answer
Direct answer: To make eventual consistency between service A and B idempotent and reconciliation-friendly, service A publishes events with a stable event ID (or a monotonic sequence number per entity), service B's consumer deduplicates on that ID before applying any change, and a periodic reconciliation job independently compares A's and B's views to catch and repair anything that slipped through despite the idempotency guarantees.
Structured elaboration
Idempotent event processing on the consumer side. Every event from A carries a stable identifier; B's consumer checks (atomically, alongside applying the event) whether that ID has already been processed, using the same "dedup record plus the actual state change in one transaction" discipline as any idempotent write. This is what makes at-least-once delivery (which any reasonable messaging setup between A and B will actually provide) safe: redelivery is a no-op rather than a duplicate application.
Preserving ordering where necessary. If events for the same entity must be applied in order (e.g. "created" before "updated" before "deleted"), B's consumer needs either a strictly-ordered delivery channel per entity (partition by entity ID) or an explicit sequence number in each event that B checks against the last-applied sequence for that entity, rejecting or buffering an out-of-order arrival rather than applying it prematurely.
Supporting replays. Because B's state can, despite everything, still drift from A's (a bug, an extended outage, a schema-migration mistake), the design should support REPLAYING A's full event history into B from scratch (or from a checkpoint) to rebuild B's view, which requires A to retain (or be able to regenerate) its event history for at least as long as any realistic replay window, and requires B's apply logic to be safe to run repeatedly over the same events (which it already is, by the idempotency design above).
Reconciliation as the safety net, not the primary mechanism. A periodic job independently compares A's and B's data (via checksums, row counts, or a full diff on a schedule appropriate to the data's size and criticality) and either auto-repairs small, well-understood divergences or flags larger ones for human review. This is deliberately a SEPARATE mechanism from the event-driven sync path, its job is to catch failures of that path (a dropped event no retry ever recovered, a bug in the consumer's apply logic), not to be the primary way B stays in sync (that would defeat the point of event-driven propagation in the first place).
Worked example. Service A (an Orders service) publishes OrderUpdated{order_id, sequence, payload} events. Service B (a search index) consumes them, checking (order_id, sequence) against the last sequence it applied for that order, skipping (as an idempotent no-op) any event with a sequence it's already seen or older, and buffering (briefly) any event that arrives out of order, applying it once the gap is filled or timing it out into a "request full replay for this order_id" fallback if the gap doesn't close. Nightly, a reconciliation job compares a sample (or full set, for smaller datasets) of orders between A's source of truth and B's index, flagging any order where B's data doesn't match A's for investigation, this is how the team discovered a bug where B's consumer was silently dropping events during a brief scaling event, well before any customer noticed stale search results.
Trade-offs and pitfalls. Skipping the reconciliation job because "the event pipeline is reliable" is a common and risky shortcut, event-driven consistency mechanisms fail in ways that are often invisible until reconciliation (or a customer complaint) surfaces them, since a missed event usually produces no error, just quietly stale data.
Explain the architectural difference between a RESTful API and an RPC-style API: how URLs are used, whether the HTTP verb carries meaning, statelessness, resource orientation, and what each style implies for caching and for how tightly the client and server are coupled. Give one situation where you would prefer RPC over REST in a backend service.
Sample Answer
Direct answer. REST models a system as a set of addressable resources you act on with a small, uniform set of HTTP verbs; RPC models a system as a set of named procedures you call directly (createOrder, cancelOrder), with the URL identifying the ACTION rather than a resource. Choose RPC when the operations you are exposing genuinely are not resource-shaped (a computation, a workflow trigger, a batch job kickoff) rather than forcing them into an artificial noun.
URL design. REST: POST /orders/{id}/cancel (a resource, acted on via a sub-action) or, more purely, a state change expressed as PATCH /orders/{id} with a status field. RPC: POST /cancelOrder with the order id in the body, where the URL itself names the verb.
HTTP verb usage. REST assigns meaning to the verb itself: GET always reads, DELETE always removes. RPC typically uses POST for nearly everything, since the "meaning" lives entirely in the URL's procedure name, not in which HTTP method was used; the HTTP verb becomes just a transport detail, not semantic information a client or an intermediary can reason about.
Statelessness. Both styles can be, and should be, stateless; this is not actually a distinguishing factor between them, statelessness is a property of how you design the API, independent of resource-vs-procedure framing.
Resource orientation. REST insists that everything be modeled as a resource, even things that are awkward as nouns (a "calculation" or a "search" is forced into being a resource you POST to create); RPC has no such constraint, an operation is simply named for what it does, which can be a more natural fit for a genuinely action-shaped API (a recommendation-generation service, an image-processing job).
Caching implications. Because REST's GET requests map directly onto HTTP caching semantics, a well-designed REST read endpoint gets CDN and browser caching essentially for free; an RPC-style API that funnels everything through POST (even reads) forfeits that entirely, since HTTP caching semantics are built around GET being safe and cacheable.
Client-server coupling. RPC tends to couple the client more tightly to the server's specific set of named operations (the client has to know the exact procedure name and its exact parameter shape); REST's uniform interface means a client that understands "resources + a few verbs" can reason about a NEW resource it has never seen before, without learning a new procedure name for it.
When to prefer each. Prefer REST for a public, resource-centric API (an e-commerce catalog, a user-management API) where caching and broad client tooling compatibility matter. Prefer RPC for an internal service exposing genuinely action-shaped operations (trigger a batch export, run a recommendation computation) where forcing the action into a resource-and-verb shape would be more awkward than useful, and where every caller is a service you control that can simply be told the procedure name and shape directly.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Backend Developer jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs