Caching Strategies and Distributed Caching Questions
Using caches to reduce latency and load: cache-aside, read-through, write-through, and write-behind patterns, TTLs, eviction policies, and distributed caches such as Redis or Memcached. Covers cache invalidation, stampede and thundering-herd protection, and the consistency tradeoffs of caching. Focuses on where and how to cache across tiers.
Design a caching approach for GraphQL APIs where clients can request arbitrary fields across entities. Discuss normalized client-side caches (e.g., Apollo cache normalization), server-side response caching, persisted queries, caching at field vs object level, and invalidation strategies for mutations that touch nested data.
Sample Answer
Direct answer
GraphQL's flexible field selection makes whole-response caching largely ineffective (every distinct query shape produces a distinct response), so effective GraphQL caching happens at a finer grain: normalized entity caching on the client, persisted queries to make server-side caching by query identity practical, and field- or object-level caching strategies that survive arbitrary client query shapes.
Structured elaboration
- Normalized client-side caches: a library like Apollo Client's cache normalizes responses by entity ID (a
Product:123object is stored once, regardless of which query fetched it, or which fields were requested across different queries), so overlapping queries share cached entity data even when their exact field selections differ, and a mutation affectingProduct:123can update every query result referencing it without a full refetch. - Server-side response caching: because two clients requesting even slightly different field sets produce different response bodies, naive whole-response caching (keyed on the raw query string) has poor hit rates; this works better when combined with persisted queries (see below), which constrain the space of distinct queries actually seen.
- Persisted queries: instead of sending the full query text on every request, the client sends a hash referencing a pre-registered query (registered at build time); this both reduces request size AND makes server-side response caching practical again, since the server now sees a small, known set of distinct queries rather than an unbounded space of ad-hoc client-constructed ones.
- Field vs. object-level caching: caching at the object level (an entity like a specific product, regardless of which query fetched it) is more reusable across different client queries than caching at the whole-response level; some GraphQL server implementations support field-level caching directives to fine-tune this further for expensive individual fields.
- Invalidation for mutations touching nested data: a mutation updating one entity can affect many different cached query results that reference it (directly, or through a relationship); a normalized cache's entity-level invalidation (update
Product:123once, every referencing query result updates) is what makes this tractable, versus trying to enumerate and invalidate every affected whole-response cache entry.
Worked example
An Apollo Client cache normalized by entity ID means a query fetching Product(id: 123) { name, price } and a separate query fetching Product(id: 123) { name, description } both read and update the SAME normalized Product:123 entry for the name field they share, and a mutation updating Product:123's price automatically reflects in both queries' results on their next render, without either query needing to know about the other or trigger an explicit refetch.
Trade-offs and pitfalls
Whole-response server-side caching for GraphQL without persisted queries tends to have a poor hit ratio in practice, because the space of distinct query shapes a flexible client can construct is effectively unbounded; persisted queries are usually a prerequisite for server-side caching to be worthwhile at all, not an optional add-on. Normalized client caching requires entities to have stable, consistent IDs across every query that returns them; a schema that returns the same logical entity with inconsistent ID fields across different query paths breaks normalization silently.
Explain the trade-offs between client-side caching (in-process or browser) and a server-side shared cache (Redis). From an operations standpoint, what concerns differ between the two approaches?
Sample Answer
Direct answer
Client-side (in-process or browser) caching is faster and needs no network hop but is invisible and hard to invalidate from the server; a server-side shared cache (Redis) is slower per read (a network hop) but is centrally observable, controllable, and consistent across every consumer.
Structured elaboration
- In-process caching: fastest possible access (no network at all), but exists as N independent copies across N instances, invisible to any centralized monitoring or invalidation mechanism unless you specifically build one (e.g., pub/sub to every instance).
- Browser caching: similarly fast for the end user, but the server has essentially no direct control once data leaves it; "invalidation" really means waiting out a time-to-live (TTL) or changing a versioned resource identifier.
- Server-side shared cache (Redis): a single, centrally-managed copy; invalidation is a single operation that immediately affects every consumer, and its state is directly observable via standard monitoring, at the cost of a network round-trip on every access.
- Operational concerns for a shared server-side cache: capacity planning, replication/failover for availability, and security (access control, encryption) are all first-class operational concerns that a purely in-process or browser cache does not have in the same way, since those are either per-instance (in-process, trivially "available" as long as the instance is up) or entirely outside your infrastructure (browser).
- Service-mesh sidecar caching: a variant worth naming, a sidecar proxy running alongside each service instance can cache responses transparently at the network layer; this gets some of in-process caching's speed benefit without embedding caching logic in application code, but invalidation and consistency questions are similar to any other local-cache design (each sidecar has its own copy).
- A hybrid, fast-tier-plus-persistent-tier design: combining a fast in-memory tier for hot items with a persistent, larger tier for warm items (promotion/demotion between them based on access) gets some of both worlds' benefits, at the cost of managing two tiers' worth of operational complexity.
Worked example
An application caching a rarely-changing configuration value both in-process (near-zero latency access on every request) and in Redis (as the source the in-process cache warms from, and as what gets explicitly invalidated on a config change) gets the speed benefit of in-process access for the common case while still having a centrally controllable invalidation point; the in-process copies use a short TTL as a backstop in case a pub/sub-based invalidation signal is missed by a given instance.
Trade-offs and pitfalls
Relying purely on in-process caching for data that changes and needs prompt, reliable invalidation across many instances is a common design mistake; without a centralized signal (or at least a short backstop TTL), some instances can serve stale data indefinitely after a change. From an operations standpoint, a shared server-side cache is a dependency you must monitor, secure, and plan capacity/failover for, in a way an in-process cache simply is not; that operational surface is a real cost, not just a technical detail.
Explain the purpose of caching in distributed systems. Define cache hit, miss, and hit ratio, and describe the typical benefits and tradeoffs of adding a cache to a service. Give concrete examples of workloads that benefit from caching (and why) and workloads where caching could be harmful.
Sample Answer
Direct answer
A cache is a fast, smaller copy of data kept close to where it is used, so repeated reads can be served without redoing expensive work (a database query, a network call, a heavy computation) every time. The core trade-off is speed and reduced load in exchange for the risk that the cached copy becomes stale relative to the real source of truth.
Structured elaboration
- What problems caching solves: latency (serving from memory or a nearby node is far faster than recomputing or fetching from a distant source), throughput/load (fewer requests reach the expensive backend), and cost (fewer database queries or external application programming interface (API) calls, which often cost money directly).
- Where caches typically live: client/browser (closest to the user, zero network cost on a hit), content delivery network (CDN) / edge (near the user but shared across many users), application in-memory (fast, but local to one process/instance), and a shared distributed cache like Redis or Memcached (shared across all instances in a region, one network hop away).
- Primary trade-offs: staleness (the cached copy may not reflect the latest write), added complexity (invalidation logic, cache-miss handling, monitoring another moving part), and memory cost (caches are not free storage).
- When caching helps: read-heavy workloads where the same data is requested repeatedly and can tolerate at least a little staleness, or where the underlying computation/fetch is expensive relative to a cache read.
- When caching is harmful: data that changes on every read (no repeated value to cache), workloads that are already write-heavy with low read repetition (cache churn without benefit), or correctness-critical data where any staleness is unacceptable and the added complexity of cache invalidation introduces more risk than the latency win is worth.
Worked example
An API endpoint that computes a dashboard aggregate over a large dataset in 800ms, requested 200 times per minute by the same handful of users, is a strong caching candidate: a 30-second time-to-live (TTL) cache serves nearly every request from cache after the first, cutting both latency (single-digit milliseconds instead of 800ms) and backend load by over 95 percent, for a staleness window most dashboard users will not notice. The same 30-second TTL applied to an account balance shown right after a deposit would be actively harmful, showing the user a stale, "wrong" number at the exact moment they are checking it.
Trade-offs and pitfalls
The most common mistake is caching by default rather than by evaluating whether the read pattern and staleness tolerance actually justify it; caching adds a second place data can be wrong (the cache disagreeing with the source of truth), and that failure mode does not exist at all if you never cache the data in the first place.
You want to measure the real-world impact of adding a caching layer on user-facing latency and backend load. Design an experiment (canary or A/B) and list the metrics you would collect, your sampling method, how you'd ensure statistical significance, and how to attribute improvements specifically to caching.
Sample Answer
Direct answer
Design a canary or A/B experiment that isolates the caching change as the only variable, measures user-facing latency percentiles and backend load together, and uses a large enough sample and duration to be statistically confident the observed improvement is real and attributable to caching, not to normal traffic variance.
Structured elaboration
- Traffic selection: split traffic randomly (not by time-of-day, which confounds with natural traffic pattern variance) into a control group (no new caching) and a treatment group (the new caching layer), ideally at the request or session level so the same user consistently lands in one group.
- Metrics to collect: client-observed p50/p95/p99 (50th/95th/99th percentile) latency (what users actually experience, not just server-side timing), backend requests per second (RPS) and CPU utilization (the load-reduction claim), and cache hit ratio (to explain WHY the latency/load changed, not just that it did).
- Sampling method: a consistent hashing of user or session ID into control/treatment buckets keeps assignment stable across a user's session, avoiding a user flip-flopping between groups mid-session, which would muddy the comparison.
- Statistical significance: with latency distributions typically non-normal, use a test appropriate for percentiles/distributions (bootstrap confidence intervals are common) rather than assuming a simple mean-based t-test applies cleanly; run long enough to cover a full daily and, ideally, weekly traffic cycle, since a short window can catch an unrepresentative slice of traffic.
- Attributing improvement specifically to caching: because the treatment group differs from control ONLY in the caching change (everything else held constant by the randomized split), a statistically significant difference in the treatment group's metrics can be attributed to the caching layer, rather than to some other simultaneous change, as long as no OTHER change ships selectively to one group during the experiment window.
Worked example
Running the experiment for two full weeks (covering weekday/weekend traffic patterns) with a 50/50 randomized split: if the treatment group's p95 latency is a statistically significant 40ms lower than control's (using a bootstrap confidence interval that excludes zero difference) and backend CPU utilization for the treatment group's traffic is meaningfully lower, both changes correlate with a treatment-group cache hit ratio of 85 percent, giving a coherent causal story (caching reduced backend calls, which reduced both load and the latency contributed by those calls).
Trade-offs and pitfalls
Running the experiment for too short a window (a few hours) risks capturing an unrepresentative traffic slice and either overstating or understating the real effect; always cover at least a full daily cycle, ideally longer. Shipping any OTHER change to the treatment group during the experiment window (even an unrelated one) breaks the attribution, since you can no longer isolate caching as the sole cause of any observed difference.
Describe how to implement negative caching safely for non-existent resources and how to set TTLs to balance reduced DB load with the risk of false negatives. Explain mechanisms to detect and recover from incorrectly cached negatives and how to prevent poisoning of the cache.
Sample Answer
Direct answer
Negative caching (caching the fact that something does NOT exist) reduces repeated backend load for lookups that legitimately miss often, but needs a shorter time-to-live (TTL) than positive caching and explicit protection against caching a TRANSIENT error as if it were a permanent absence.
Structured elaboration
- When negative caching helps: an endpoint frequently queried for resources that legitimately do not exist (a sparse product search, a lookup by an ID that is often invalid or not-yet-created) benefits from caching "not found" so repeated identical misses do not all hit the backend.
- TTL sizing: negative entries should generally use a SHORTER TTL than positive entries, since a resource that does not exist now might be created moments later (a product about to be listed, a user about to be created), and a long negative-cache TTL would delay legitimate new data from becoming visible.
- Avoiding poisoning from transient errors: a backend timeout or a temporary error must NOT be cached as "not found"; only cache an explicit, confirmed "this does not exist" response (e.g., an actual 404 from a healthy backend), never an ambiguous failure, or a temporary outage will look like every queried resource permanently vanished.
- Detecting and recovering from an incorrectly cached negative: monitor the negative-cache hit rate for anomalies (a sudden spike suggesting something is being wrongly cached as absent), and provide an explicit purge mechanism for a specific key if a false negative is identified.
- Preventing cache poisoning attacks: without care, an attacker (or a misbehaving client) could probe many nonexistent IDs to fill the negative cache with entries, which is generally low-risk for negative caching specifically (it just wastes some memory) but worth being aware of if negative-cache capacity is shared with more sensitive cached data.
Worked example
A product-search endpoint receiving many queries for discontinued or misspelled product IDs: caching "not found" for 30 seconds (much shorter than the 5-minute TTL used for actual product data) absorbs repeated identical misses for the same nonexistent ID within that window, while still allowing a newly-listed product with that same ID to become visible within 30 seconds rather than being stuck behind a long negative-cache TTL.
Trade-offs and pitfalls
Caching a timeout or a 500-level error as a negative ("not found") result is the single most damaging mistake in negative caching; it turns a transient backend problem into an apparent, cache-durable "this doesn't exist" for every subsequent request until the negative TTL expires, actively making an outage worse and longer-lasting than it needed to be. Setting a negative TTL as long as the positive TTL "for simplicity" delays visibility of genuinely new data unnecessarily; size them independently based on how quickly a false negative needs to self-correct.
Unlock Full Question Bank
Get access to all Caching Strategies and Distributed Caching interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.