Caching Strategies and Distributed Caching Questions
Using caches to reduce latency and load: cache-aside, read-through, write-through, and write-behind patterns, TTLs, eviction policies, and distributed caches such as Redis or Memcached. Covers cache invalidation, stampede and thundering-herd protection, and the consistency tradeoffs of caching. Focuses on where and how to cache across tiers.
How do you decide whether to introduce a cache for a given service endpoint? Describe the signals and measurements you would collect, the tests you would run (load, latency, profiling), and the criteria that justify adding an in-process cache, a shared cache (Redis), or a CDN. Include considerations for cost, operational complexity, and correctness.
Sample Answer
Direct answer
Decide whether to add a cache by measuring the actual read pattern (how often the same value is requested, how expensive it is to produce, and how much staleness is tolerable), not by defaulting to caching every endpoint; a cache with a low repeat-read rate or zero staleness tolerance is a cost with no real benefit.
Structured elaboration
- Signals to collect: request rate for the same key/query (does the same data actually get read repeatedly, or is nearly every read unique), the cost of producing the value (a fast, cheap lookup gains little from caching even if repeated), and the data's staleness tolerance (how quickly must a change be visible).
- Tests to run: a load test comparing latency and backend load with and without a proposed cache, and a profile of the actual query/computation to confirm it is genuinely a meaningful cost worth caching against.
- Criteria for choosing a cache tier: an in-process cache fits data that is cheap to duplicate per instance and does not need cross-instance consistency; a shared cache (Redis) fits data that benefits from being consistent across instances or too large to duplicate per instance; a content delivery network (CDN) fits public, non-personalized content that benefits from being close to users geographically.
- Cost: weigh the infrastructure and operational cost of adding a caching layer (a new dependency to monitor, secure, and keep available) against the actual load/latency benefit measured above; a marginal benefit may not justify the added complexity.
- Operational complexity: caching adds invalidation logic, a new failure mode (cache unavailable), and another thing to monitor; these costs are real even when the caching decision is otherwise sound, and should be weighed explicitly.
- Correctness: if the data's staleness tolerance is effectively zero (a value that must always reflect the absolute latest state, with no acceptable delay), caching adds risk without benefit, since any caching mechanism introduces at least a small window of potential staleness.
Worked example
An endpoint returning a real-time stock quote, requested uniquely per symbol per user with essentially no repeat reads within any meaningful window, and requiring zero staleness tolerance: this fails on both the "does the same value get read repeatedly" test and the "can staleness be tolerated" test, making it a poor caching candidate regardless of how expensive the underlying computation is. Contrast with a product description, read thousands of times per hour by different users for the same handful of popular items, changing rarely: this passes both tests clearly.
Trade-offs and pitfalls
Caching by default, without measuring the actual read-repetition rate, either wastes cache capacity on data that gets no benefit or, worse, introduces a staleness risk on data that could not tolerate it; always start from measurement, not habit. The decision is not binary per endpoint; the same service can have some data that benefits enormously from caching and other data (even on the same page) that should never be cached, and treating the whole endpoint uniformly misses that nuance.
In Node.js (TypeScript), implement a simple cache-aside wrapper function fetchWithCache(key, fetchFn, ttlSeconds) that: (1) checks Redis for key, (2) on miss calls fetchFn() to get fresh value, (3) stores result in Redis with TTL, (4) handles Redis errors by falling back to fetchFn, and (5) ensures JSON serialization. Provide the function and then describe concurrency considerations for simultaneous cache misses.
Sample Answer
Approach
Implement fetchWithCache as a thin wrapper: check the cache first, and on a miss (or an error reading the cache) call the underlying fetch function directly so a cache outage degrades to "no caching" rather than an outright failure. Store successful results with the given time-to-live (TTL) after JSON-serializing them, and treat a cache write failure as non-fatal (log it, still return the fresh value).
import { createClient, RedisClientType } from "redis";
type FetchFn<T> = () => Promise<T>;
export async function fetchWithCache<T>(
redis: RedisClientType,
key: string,
fetchFn: FetchFn<T>,
ttlSeconds: number
): Promise<T> {
try {
const cached = await redis.get(key);
if (cached !== null) {
return JSON.parse(cached) as T;
}
} catch (err) {
// Redis unavailable or a bad payload: fall through to origin, do not fail the request.
console.error("cache read failed, falling back to origin", { key, err });
}
const fresh = await fetchFn();
try {
await redis.set(key, JSON.stringify(fresh), { EX: ttlSeconds });
} catch (err) {
console.error("cache write failed, returning fresh value anyway", { key, err });
}
return fresh;
}
Key points
Cache read and cache write failures are both caught separately and treated as non-fatal, so a Redis outage degrades to hitting the origin on every request rather than breaking the caller. fetchFn is only ever called once per request (on a cache miss), which keeps the origin call count predictable per request even though it is not yet protected against concurrent misses on the same key.
Complexity
Time is O(1) plus the cost of fetchFn on a miss (dominated by whatever fetchFn does, typically a database or network call); space is O(1) per cached key, bounded by the serialized payload size.
Concurrency considerations for simultaneous cache misses
As written, N concurrent requests that all miss the same key will all call fetchFn independently, which is correct but not efficient: it does not prevent a stampede. To fix that, add request coalescing: acquire a short-lived per-key lock (e.g., SET key:lock value NX PX 5000) before calling fetchFn; if the lock is already held, either wait briefly and retry the cache read, or call fetchFn directly as a fallback if waiting exceeds a bounded timeout, so a crashed lock-holder cannot permanently stall other requests.
Edge cases
A fetchFn that throws should propagate the error rather than being silently swallowed, since callers need to know the underlying fetch failed; only cache read/write errors should be treated as non-fatal, never a fetch failure. Values that fail to JSON.parse (a corrupted or incompatible cached payload) should be treated the same as a cache miss, not as a fatal error.
Describe cache warming and prepopulation strategies. Provide at least three distinct approaches and explain their pros/cons, safety concerns, and how to prioritize keys to warm first.
Sample Answer
Direct answer
Cache warming trades a bounded amount of upfront (or background) load for avoiding a much larger, uncontrolled cold-start spike; the right approach depends on whether you can predict the hot keys ahead of time and how much control you have over the deployment timing.
Structured elaboration
- Proactive/bulk warming: precompute and load the known hot keys before traffic hits the new instance or cluster (a scheduled job at deploy time). This gives the strongest guarantee against cold-start misses but requires knowing which keys will actually be hot ahead of time, and adds deploy-time work.
- Lazy/on-demand warming: let the cache populate naturally as real traffic arrives; simplest to implement (it is just the normal cache-aside miss path), but produces a real, uncontrolled burst of origin load right after a deploy or flush, exactly a cold-start stampede if not otherwise protected.
- Hybrid/synthetic traffic warming: replay a sample of recent real traffic (or synthetic requests mimicking it) against the new instance before it receives production traffic, warming the cache using realistic access patterns without waiting for actual users to trigger it.
- Prioritizing which keys to warm: use recent access logs to identify the actual hot-key distribution rather than guessing; warming the wrong keys wastes the warming budget on cold data while the truly hot keys still cold-start.
- Safety concerns: warming itself generates load against the origin; rate-limit or throttle the warming process so it does not become its own mini-stampede, and have a way to abort a warming job cleanly (restoring previous cache state) if it is causing unexpected load rather than letting it run to a bad outcome.
Worked example
A microservice serving 100,000 requests per second (RPS) with heavy skew to 1 percent of keys (the classic Zipfian hot-key pattern common in real traffic): warming just that top 1 percent of keys from recent access logs, prioritized by request volume, captures the overwhelming majority of the benefit for a small fraction of the warming effort a full-catalog preload would require; a full preload of every key, most of which are rarely read, would waste warming time and origin load on data unlikely to matter for cold-start performance.
Trade-offs and pitfalls
Warming based on stale or unrepresentative access logs (e.g., from before a major traffic-pattern shift) warms the wrong keys and provides a false sense of readiness; refresh the "what's hot" list close to the actual warming event. A warming job that itself is unthrottled can overload the origin just as badly as the cold-start stampede it is meant to prevent; treat warming as a controlled, rate-limited operation with the ability to abort and roll back cleanly, not a fire-and-forget script.
A distributed cache cluster shows inconsistent data across nodes after a network partition. Explain the plausible root causes and design a recovery plan that minimizes downtime and data loss. Discuss the trade-off between restoring availability quickly and reconciling consistency.
Sample Answer
Direct answer
Inconsistent data across cache nodes after a network partition means two sides of the partition may have kept accepting writes independently (split-brain) or replication simply fell behind (partial replication/stale writes); recovery has to first stop the bleeding (restore a single source of truth) before reconciling the divergence, and that choice trades availability against correctness.
Structured elaboration
- Plausible root causes: split-brain (both sides of a partition believed they were the primary and accepted writes independently, producing two diverging histories), partial replication (replication was in progress but not complete when the partition occurred, so some nodes have older data than others), and stale writes (a client wrote to a node that was, unknown to it, no longer the authoritative primary).
- Restoring availability quickly versus reconciling consistency: choosing to bring the cluster back online immediately (picking one side's data as authoritative and discarding the other's divergent writes) restores service fast but can silently lose data from the discarded side; choosing to carefully reconcile both sides first (comparing and merging where possible) is safer for correctness but extends the outage.
- A recovery plan: first, stop accepting further writes on whichever side is NOT chosen as authoritative (preventing further divergence); second, decide the authoritative side (usually the side with quorum, or the side that was serving the majority of traffic, per your topology's own quorum rules); third, reconcile or discard the non-authoritative side's diverged writes, documenting what was lost if anything; fourth, resync the non-authoritative side from the authoritative one before bringing it back into service.
- Minimizing downtime and data loss together: this is rarely a clean binary choice; a common middle ground is bringing read traffic back online quickly from the authoritative side (restoring availability for reads) while writes stay paused or degraded until reconciliation completes, rather than an all-or-nothing restoration.
Worked example
A 6-node Redis cluster splits into a 4-node majority partition and a 2-node minority partition due to a network issue; if the cluster's quorum configuration correctly prevented the minority side from accepting writes during the partition (a properly configured min-replicas-to-write or Sentinel quorum), the majority side is unambiguously authoritative and the minority side simply needs to resync from it once the network heals, with no actual data divergence to reconcile. If quorum protection was misconfigured and the minority side DID accept writes, those writes are now in conflict with the majority's history and must be explicitly reconciled or, if reconciliation is not possible, discarded and documented as lost.
Trade-offs and pitfalls
The best outcome here comes from PREVENTING split-brain in the first place (properly configured quorum requirements), not from having a great recovery plan for when it happens; a recovery plan is a backstop, not a substitute for prevention. Discarding a non-authoritative side's writes without first checking whether any of them can be safely reconciled (a value that was never touched on the authoritative side, for example) is faster but can lose more than strictly necessary; weigh the reconciliation effort against how much data is actually at stake.
You want to measure the real-world impact of adding a caching layer on user-facing latency and backend load. Design an experiment (canary or A/B) and list the metrics you would collect, your sampling method, how you'd ensure statistical significance, and how to attribute improvements specifically to caching.
Sample Answer
Direct answer
Design a canary or A/B experiment that isolates the caching change as the only variable, measures user-facing latency percentiles and backend load together, and uses a large enough sample and duration to be statistically confident the observed improvement is real and attributable to caching, not to normal traffic variance.
Structured elaboration
- Traffic selection: split traffic randomly (not by time-of-day, which confounds with natural traffic pattern variance) into a control group (no new caching) and a treatment group (the new caching layer), ideally at the request or session level so the same user consistently lands in one group.
- Metrics to collect: client-observed p50/p95/p99 (50th/95th/99th percentile) latency (what users actually experience, not just server-side timing), backend requests per second (RPS) and CPU utilization (the load-reduction claim), and cache hit ratio (to explain WHY the latency/load changed, not just that it did).
- Sampling method: a consistent hashing of user or session ID into control/treatment buckets keeps assignment stable across a user's session, avoiding a user flip-flopping between groups mid-session, which would muddy the comparison.
- Statistical significance: with latency distributions typically non-normal, use a test appropriate for percentiles/distributions (bootstrap confidence intervals are common) rather than assuming a simple mean-based t-test applies cleanly; run long enough to cover a full daily and, ideally, weekly traffic cycle, since a short window can catch an unrepresentative slice of traffic.
- Attributing improvement specifically to caching: because the treatment group differs from control ONLY in the caching change (everything else held constant by the randomized split), a statistically significant difference in the treatment group's metrics can be attributed to the caching layer, rather than to some other simultaneous change, as long as no OTHER change ships selectively to one group during the experiment window.
Worked example
Running the experiment for two full weeks (covering weekday/weekend traffic patterns) with a 50/50 randomized split: if the treatment group's p95 latency is a statistically significant 40ms lower than control's (using a bootstrap confidence interval that excludes zero difference) and backend CPU utilization for the treatment group's traffic is meaningfully lower, both changes correlate with a treatment-group cache hit ratio of 85 percent, giving a coherent causal story (caching reduced backend calls, which reduced both load and the latency contributed by those calls).
Trade-offs and pitfalls
Running the experiment for too short a window (a few hours) risks capturing an unrepresentative traffic slice and either overstating or understating the real effect; always cover at least a full daily cycle, ideally longer. Shipping any OTHER change to the treatment group during the experiment window (even an unrelated one) breaks the attribution, since you can no longer isolate caching as the sole cause of any observed difference.
Unlock Full Question Bank
Get access to all Caching Strategies and Distributed Caching interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.