Multi-Region and Geo-Distributed Systems Questions
Running a system across regions and continents: multi-region replication, data residency and sovereignty, geo-routing and CDN edge distribution, cross-region consistency and quorum placement, and conflict resolution when two regions accept writes. Covers regional failover and split-brain prevention, recovery objectives (RTO/RPO), region-by-region rollout and blast-radius containment, and the latency, cost, and consistency tradeoffs of going global. Global distribution strategy across the service and data tiers.
Design DNS and traffic routing for a multi-region service that requires low latency and safe automated failover. Discuss TTL strategies, health checks, DNS provider features (GeoDNS, Anycast), and how DNS failover should coordinate with application-level health and data failover to avoid serving stale writes.
Sample Answer
Direct answer
Treat DNS failover as the outermost, slowest layer of a three-layer failover stack, and
never let it move traffic faster than the data layer can safely serve it. The sequencing
that matters: application-level health must fail first and fast (seconds), DNS must follow
only once the target region is verified both healthy and holding current data, and the data
layer's own failover (promoting a new primary, replaying the write-ahead log) must complete
before DNS sends write traffic there, or you serve reads and accept writes against stale
data.
Design
Health checks, layered. A shallow health check (does the load balancer answer) is not
enough. Run a deep health check that a region can only pass if it (a) can reach its local
database replica, (b) that replica's replication lag is under a defined threshold, and (c)
dependent services report healthy. DNS failover should key off this deep check, not just
"is the box up."
TTL strategy. Use a short time-to-live (TTL, how long a DNS answer may be cached, e.g.
30-60 seconds) on the records that need to fail over, and a longer TTL on stable records
(the domain's NS records, CDN endpoints that rarely change) to keep query volume down where
speed doesn't matter. Accept that TTL is a floor, not a guarantee: some resolvers cache
longer, so failover completion time should be measured, not assumed from the configured
value.
Provider features. Most managed GeoDNS/Anycast providers support health-check-linked
failover records natively: point a record at region A, configure a health check against A's
deep health endpoint, and have the provider automatically fail the answer over to region B
when A stops passing. This removes the need to hand-roll a failover controller for the DNS
layer itself.
Coordinating with data failover to avoid stale writes. This is the part that is easy to
get wrong: if DNS fails over the moment app health looks bad, but the database in the new
region is still catching up on replication, you will route writes to a replica that either
rejects them (safe, but ugly) or, worse, accepts them and later conflicts with writes still
in flight to the old primary (unsafe, silent data loss or divergence). The safe sequence is:
- Detect the primary region is unhealthy.
- Fence the old primary (stop it from accepting further writes: revoke its write
credentials or force it read-only) so it cannot keep diverging. - Promote the new region's replica to primary only after confirming it has applied
everything it can from the last known good replication position (or accepting a defined,
bounded window of data loss if the old primary is truly unreachable, which is an explicit
recovery point objective (RPO) decision made ahead of time, not improvised during the
incident). - Only then flip DNS (or the load balancer's target) to send write traffic to the newly
promoted region.
Reads can often fail over faster and more loosely than writes, since serving a slightly
stale read is usually a smaller problem than accepting a write that later has to be
reconciled or discarded.
Worked example
Primary is in us-east, replica in us-west, asynchronous replication with typical lag
under 2 seconds. us-east's deep health check starts failing (database unreachable). The
failover controller immediately marks us-east read-only at the application layer (fencing)
so no new writes can land there even if the process is technically still up. It checks
us-west's replication position: it is 1.4 seconds behind the last write us-east
successfully committed before going dark, which is within the pre-agreed RPO of 5 seconds.
us-west is promoted to primary. Only now does the DNS record (30-second TTL) flip to point write traffic at us-west. Two
different clocks are worth separating when you report this number. The time from detection to the
cutover decision is entirely data layer: fence, read the replication position, promote, a few
seconds end to end, with DNS contributing nothing. The time from that decision to the last client
actually writing to us-west is entirely DNS: 30 seconds of TTL at best, and longer for the
resolvers that ignore it, which by the numbers in this example makes DNS the larger share of the
wall-clock outage even though it moves last. Both facts hold at once, and quoting only the first is
how a team reports a 10-second failover that users experienced as a minute. The ordering is still
deliberate, DNS should never be the fastest-moving part of this sequence, but it has a design
consequence: if write traffic has to cut over faster than DNS can drain, move the flip to a layer
you control end to end, a load balancer target or a stable anycast frontend, and leave DNS pointing
at that frontend rather than at a region.
Trade-offs and pitfalls
- The most common defect is flipping DNS before the data layer is ready, because DNS
failover is the easiest piece to automate and teams reach for it first. Gate DNS failover
on a signal from the data layer, not a timer. - Fencing the old region is not optional. Without it, a "split-brain" (two regions both
believing they are primary and accepting writes) can occur if the old region recovers
network connectivity mid-failover, producing writes that conflict and are expensive or
impossible to reconcile automatically. - DNS failover coordinating with a distributed denial-of-service (DDoS) attack is a
related failure mode worth planning for separately: an attacker flooding one region can
make its health checks fail for capacity reasons rather than a real outage, triggering a
failover that then overloads the second region too. Mitigate by scrubbing malicious
traffic upstream (at the anycast or CDN edge, before it reaches your health-checked
origin) rather than treating "traffic volume exceeded capacity" the same as "region is
down," and by rate-limiting failover flips so a brief health blip doesn't cause repeated
flapping between regions.
You must connect three regions with expected inter-region traffic 10 Gbps and occasional spikes to 20 Gbps. Compare cloud dedicated interconnect (Direct Connect/ExpressRoute), site-to-site VPN, and public internet peering. Propose an architecture balancing cost, reliability, and security for consistent replication and workload movement.
Sample Answer
Direct answer
Provision dedicated interconnects (a cloud provider's private circuit product, such as AWS
Direct Connect or Azure ExpressRoute) sized for the 20 Gbps burst, not just the 10 Gbps
steady state, using two bonded 10 Gbps ports per critical path (a link aggregation group,
or LAG, which combines multiple physical connections into one logical one) rather than a
single larger port. That gives you 20 Gbps of capacity with a graceful-degradation property
that a single circuit doesn't: if one port fails, the remaining 10 Gbps still fully covers
steady-state traffic, only the burst headroom is lost until the port is repaired.
Sizing the topology
With 3 regions needing to exchange replication and workload-movement traffic, first
establish whether the traffic pattern is full mesh (every region talks meaningfully to every
other region) or hub-and-spoke (one region is the primary, the other two mostly talk to it).
The question's framing (consistent replication and workload movement across three regions)
points toward full mesh being the safer default unless you know the actual traffic pattern
is asymmetric. For a full mesh, that means 3 pairwise links, each sized for the same 10 Gbps
steady / 20 Gbps burst profile unless per-pair traffic is known to differ.
Say out loud which reading of "10 Gbps expected, 20 Gbps spike" you are sizing to, because
the two readings differ by a factor of three in circuit cost and the question does not settle
it. If those figures are the aggregate across all inter-region traffic, then three 20 Gbps
LAGs provision 60 Gbps of capacity for a 20 Gbps worst case, which is a lot of money for
headroom you will never use at once. If they are the per-pair figures, the same build is
correctly sized. Ask. If nobody knows, start from the aggregate reading (a 2x10 Gbps LAG on
the busiest pair, VPN or a single port on the others) and let measured per-pair utilization
pull you toward the full build, rather than buying three full-size circuits on day one.
For each pairwise link: 2x 10 Gbps dedicated ports in a LAG gives 20 Gbps aggregate
capacity. At steady 10 Gbps, the link runs at roughly 50% utilization, leaving headroom for
the stated burst to 20 Gbps without needing to provision a third port. If one port in the
LAG fails, the remaining single 10 Gbps port still fully absorbs the steady 10 Gbps load,
the system degrades from "handles bursts" to "handles steady state only," rather than
failing outright.
Comparing the three connectivity options for this load
Dedicated interconnect (Direct Connect / ExpressRoute). Fits this workload well: the
traffic is continuous, predictable in its steady/burst envelope, and sensitive to jitter for
replication consistency. A reserved circuit gives a bandwidth guarantee and a stable latency
profile that a shared path can't, which matters when the workload is "consistent
replication," not occasional best-effort transfers.
Site-to-site VPN. Could technically carry this traffic, but VPN throughput is bounded by
the encryption/decryption capacity of the endpoint devices doing the tunneling, and at
10-20 Gbps sustained, that endpoint capacity (not the link) is very often the actual
bottleneck; VPN is a better fit as a backup path than as the primary carrier here.
Public internet peering. Cheapest to start, but offers no bandwidth guarantee and
variable latency, both of which are risky for a workload described as needing "consistent"
replication; reserve this as a fallback path for non-critical or lower-priority traffic, not
the primary link for the replication workload itself.
Recommended architecture
Primary: dedicated interconnect (2x10 Gbps LAG per region pair, full mesh) carrying
replication and workload-movement traffic. Secondary: a site-to-site VPN path per pair as an
automatic failover if the dedicated link degrades or is fully down, accepting reduced
throughput during that window rather than losing connectivity entirely. Balance cost by not
over-provisioning: a 20 Gbps LAG per pair matches the stated burst exactly rather than
padding further, and utilization monitoring should trigger a capacity review (adding a third
port, or re-evaluating whether full mesh is actually needed) if sustained traffic trends
toward the burst ceiling rather than staying an occasional spike.
Trade-offs and pitfalls
- Don't size only for the steady state. A single 10 Gbps port per pair meets the average
but leaves zero headroom the moment the described 20 Gbps spike occurs, causing queuing,
increased replication lag, or dropped connections exactly when the system is under the
most load. - A LAG of two same-speed ports is not the same as one port at double the speed
operationally: it adds the resilience property (partial capacity survives a single port
failure) that a bigger single port does not, at a similar aggregate cost, which is usually
worth the design choice. - Full mesh costs more circuits than hub-and-spoke, and the gap widens fast with region
count. At three regions the difference is modest: full mesh needs 3 pairwise circuits
against hub-and-spoke's 2, so 50% more, not three times as many. The reason to settle the
traffic pattern anyway is what happens next: full mesh grows as N(N-1)/2 while
hub-and-spoke grows as N-1, so the same decision at six regions is 15 circuits against 5.
If traffic really is concentrated through one primary region, confirm that before
committing to 3 separate pairwise circuits; building unnecessary full mesh is a real,
ongoing cost for capacity you don't use, and it is the decision that compounds as regions
are added. The counterweight is that hub-and-spoke makes the hub a single point of
failure for region-to-region traffic and adds a hop of latency to every spoke-to-spoke
path, so the honest framing is cost against blast radius, not cost alone.
Design a global ingress and routing strategy that uses BGP anycast to minimize latency and support fast failover. Explain DNS TTL choices, BGP propagation considerations, how to avoid traffic blackholing during failover, and mechanisms for DDoS protection in an anycast deployment.
Sample Answer
Direct answer
Anycast works by announcing the same IP address from every region using BGP (Border
Gateway Protocol, the protocol routers use to exchange reachability information across the
internet); each user's traffic is pulled toward whichever announcing region is
topologically closest according to internet routing, not physically closest. Fast failover
comes from withdrawing the route for a failed region, which propagates in seconds to tens
of seconds; the two things that actually go wrong in practice are blackholing during that
propagation window and denial-of-service (DoS) traffic landing disproportionately on one
node before the network naturally spreads it.
How the pieces work
DNS TTL choices. Anycast itself doesn't depend on DNS time-to-live (TTL) the way
region-based DNS routing does, since the IP address doesn't change, only which physical
location answers to it. But you should still keep the DNS record for that anycast IP at a
short-to-moderate TTL (a few minutes) so that if you ever do need to change the anycast IP
itself (a rare, larger operational event: re-numbering, provider migration), you aren't
stuck with a stale record for a long time.
BGP propagation. When a region's router stops announcing the anycast prefix (because a
health check fails or it's manually withdrawn), that withdrawal has to propagate through
every intermediate autonomous system (AS, an independently operated network) between here
and every other network on the internet. Convergence is typically in the tens of seconds,
but is not uniform: some networks apply route-flap damping (deliberately delaying
re-acceptance of a route that has changed too many times recently, to protect global
routing table stability), which can make your failover look slower from some vantage points
than others.
Avoiding blackholing during failover. Blackholing here means traffic keeps being routed
toward a region that can no longer serve it, silently dropped. This happens when the
region's network path is still being announced (BGP hasn't converged yet) but the
application behind it is already down. The fix is to keep withdrawal fast and decisive:
tie the BGP announcement to the same deep health check that gates application traffic
(don't keep announcing a route to a dead application), and prefer graceful, deliberate route
withdrawal over waiting for the network to notice a hard failure, since a router that simply
stops responding can take longer to be detected than one that explicitly says "don't route
here anymore."
DDoS protection. Anycast's structural advantage against a distributed denial-of-service
attack is that it naturally spreads inbound traffic across every announcing location, since
BGP routes different attacking sources to different regions based on their own network
proximity, rather than concentrating the entire attack on one data center the way a single
unicast IP would. Combine that natural spread with a scrubbing layer at each point of
presence (traffic-pattern analysis that drops malicious packets before they reach your
actual application servers) so that even the traffic that does land at one region doesn't
overwhelm it.
Worked example
A service anycasts 203.0.113.10 from 6 regions. A network fault takes down the application
in the Frankfurt region. The health-check system detects it within a few seconds and
triggers withdrawal of the BGP announcement from Frankfurt's router. Networks in Western
Europe that were routing to Frankfurt begin reconverging toward the next-closest
announcement (say, London) over roughly the following 10-30 seconds; during that window, a
subset of users may experience dropped or blackholed connections until their upstream
network's routing table updates, which is the unavoidable physical cost of BGP convergence.
Note what the withdrawal also did to capacity: Frankfurt's share of legitimate traffic did not
disappear with it, it landed on London and the other neighbours, so every remaining point of
presence now carries its own load plus a slice of the failed one. An anycast fleet sized so each
node runs near its ceiling has nowhere to put that traffic, and the failover turns one regional
outage into a wider one. Size for the loss of your busiest announcing location, not for the steady
state. Meanwhile, a volumetric DDoS attack originating from many geographies is naturally split
across the 5 regions still announcing the prefix, roughly in proportion to where the attacking
hosts are network-adjacent, rather than concentrating 100% of the attack traffic on a single
unicast target.
Trade-offs and pitfalls
- Convergence time is not fully in your control. You can make your own withdrawal fast,
but you cannot make every intermediate network on the internet reconverge instantly; plan
for a real, measured propagation window (test it, don't assume a number), not zero. It is also
not one number: the same withdrawal converges at different speeds for different observers, so
the honest form of the metric is a distribution across vantage points rather than a single
figure you can put in a runbook. - Fronting the deployment with a content delivery network (CDN) is the common shortcut. CDN
providers already operate a large anycast edge, so renting theirs means their existing presence
absorbs both the routing complexity and a large share of attack traffic before it reaches your
origin. The trade is that your failover behaviour becomes theirs: you inherit their convergence
characteristics and their health-check semantics instead of running your own. - Anycast doesn't preserve session state across a route change. A long-lived TCP
connection or a UDP session can be silently rerouted to a different backend mid-flight if
the network path changes for reasons unrelated to your failover (routine internet routing
churn happens constantly, not just during your outages). Terminate connections close to
the anycast edge and hand off session state through your application layer, don't assume
the network will keep a "session" pinned to one location. - Don't rely on anycast alone as your only DDoS defense. It reduces concentration, but a
sufficiently large attack can still saturate an individual point of presence; pair it with
a dedicated scrubbing service.
Compare private direct links, VPN tunnels, and internet peering/IX for inter-region connectivity. For each option, discuss bandwidth, latency, cost, security, SLAs, and operational considerations. Recommend a strategy for a SaaS provider handling sensitive customer data across regions.
Sample Answer
Direct answer
Private direct links give you the most predictable performance and the strongest security
posture, at the highest fixed cost and longest lead time to provision. Site-to-site VPN
(virtual private network, an encrypted tunnel over the public internet) gets you most of the
security benefit quickly and cheaply but inherits the public internet's variable latency and
congestion. Public internet peering or an internet exchange (IX, a physical facility where
many networks interconnect directly) is the cheapest and fastest to set up but offers the
weakest guarantees on all the other axes. For a SaaS provider handling sensitive customer
data, the right default is a private direct link (or VPN layered on top of one) for the
primary inter-region and cloud-to-customer paths, with VPN as the pragmatic fallback for
lower-volume or newly-onboarded connections.
Comparison
| Private direct link | Site-to-site VPN | Internet peering / IX | |
|---|---|---|---|
| Bandwidth | Dedicated, guaranteed capacity you provision (e.g. 1/10/100 Gbps ports) | Shared with whatever else is on the public internet path; effective throughput varies | Shared, best-effort; can be high but is not reserved for you |
| Latency | Most predictable: dedicated path, no public internet contention | Variable: rides the public internet, plus encryption overhead | Variable: depends entirely on internet routing and congestion at the time |
| Cost | Highest fixed cost (a physical circuit, often billed as a port plus a recurring fee) | Low: mostly the cost of the endpoints doing the encryption, using existing internet connectivity | Lowest: often just standard internet transit or peering fees |
| Security | Strongest: traffic never traverses the public internet at all | Strong: encrypted, but the ciphertext still transits shared public infrastructure | Weakest by default: needs application-layer encryption (TLS) since the link itself offers no confidentiality |
| Service-level agreement (SLA, a contract that guarantees a minimum level of performance or uptime, often with penalties if it is not met) | Typically the strongest, formal uptime and latency guarantees from the provider | Depends on the underlying internet connection's SLA, if any | Usually none, or a very loose one |
| Operational considerations | Longest lead time to provision (physical cross-connects, sometimes weeks); needs redundant paths to avoid a single point of failure | Fast to stand up; needs careful key/certificate rotation and tunnel monitoring | Fastest to set up; needs your own encryption and monitoring layered on top since the network gives you nothing |
Recommendation for sensitive customer data
For a SaaS handling sensitive customer data (personally identifiable information, health,
or financial records), the baseline should be: private direct links for the high-volume,
persistent paths that matter most (inter-region replication, and connections to large
enterprise customers who have their own private-link requirements as part of their
compliance posture), with encryption still applied at the application layer even over the
private link (defense in depth: a private link keeps traffic off the public internet, it
doesn't by itself guarantee confidentiality end to end). Use VPN for lower-volume paths and
for customers or regions where a private circuit isn't justified by volume, and treat plain
internet peering as acceptable only for genuinely public, non-sensitive traffic (serving
public marketing pages, for instance), never for the data path itself.
A common variant of this question is cloud-to-on-premises rather than cloud-to-cloud:
a SaaS customer with an on-premises data center needing a private path into your cloud
regions. The same hierarchy applies, but the private-link option there is typically a
provider's dedicated interconnect service into the customer's own network (for example a
cloud provider's dedicated interconnect product terminating at the customer's router), and
VPN is very often the actual default for smaller customers, since not every enterprise
customer's volume or compliance requirement justifies a dedicated circuit. Segment the
decision per customer by their data sensitivity and volume, not as one blanket policy.
Trade-offs and pitfalls
- Don't treat "private" as "encrypted." A private direct link is private in the sense of
not traversing shared public infrastructure, but it typically does not encrypt traffic by
default; keep TLS on top for genuinely sensitive payloads regardless of the transport. - A single private link is a new single point of failure unless you provision a redundant
second path (a different physical route or provider); many teams learn this only after
their one circuit goes down. - VPN's biggest hidden cost is throughput ceiling from encryption overhead on the
endpoints doing the encrypting, not the link itself, at high sustained volumes this can be
the binding constraint before internet congestion is.
Given an online multiplayer game with tight latency requirements and servers in 8 regions, compare DNS-based routing, anycast, and cloud global-load-balancer approaches. Recommend a routing design to minimize player-perceived latency and reduce cross-region jitter, explaining how you'd handle affinity for UDP traffic.
Sample Answer
Direct answer
For a real-time multiplayer game, don't route the actual game traffic through DNS
geo-routing, anycast, or a generic global load balancer at all; use one of those only for
the initial matchmaking/discovery step, and have the client pick its game server by
measuring latency directly, then connect to that specific region's server IP for the
duration of the match. Session affinity for UDP (User Datagram Protocol, a connectionless
transport protocol with no built-in session concept) is the crux of the problem: none of
the three generic routing approaches guarantee that every packet in an ongoing match keeps
landing on the same backend, and for a live game, a mid-match reroute is far worse than the
extra step of client-side region selection.
Why the generic approaches fall short here
DNS geo-routing picks a region once, at connection setup, based on the client's
resolver location, which is frequently wrong for gamers on VPNs, mobile carriers with
distant DNS infrastructure, or shared corporate resolvers, and it can't react mid-match if
that region degrades.
Anycast routes packet-by-packet based on internet topology at the time each packet is
sent. That's fine for stateless request/response traffic, but a live match is inherently
stateful: if the network path shifts mid-match (which happens for reasons unrelated to any
failure, ordinary internet route churn), the player's UDP packets can suddenly land on a
different physical server that has no idea a match is in progress, silently breaking the
session with no retry semantics to fall back on the way TCP has.
A managed global load balancer solves the affinity problem better (it can pin a
connection), but it still routes based on network path from the load balancer's vantage
point, not the actual round-trip time (RTT) the player experiences, and adds a
provider-operated hop into the latency-critical path.
Recommended design
- Matchmaking and discovery over HTTP/TCP, routed by a global load balancer or
geo-routed DNS. This traffic is infrequent, tolerant of a few hundred milliseconds, and
benefits from being simple and reliable. - Client-side active latency probing. Before or during matchmaking, the client sends
lightweight pings to a small set of candidate regions (not all 8, to avoid excess
overhead, typically the 3-4 geographically plausible ones) and measures real round-trip
time and jitter (the variation in that round-trip time between packets, which for a game
matters as much as the average, since inconsistent delay causes visible stutter even when
average latency looks fine). - Server selection by measured latency, with a jitter-aware tiebreak. Pick the region
with the lowest combination of RTT and jitter, not just the lowest average RTT: a region
with 40ms average but wildly variable jitter can feel worse to a player than a 55ms
region with tight, consistent timing. - Direct UDP connection to that region's specific server IP for the duration of the
match, bypassing anycast or load-balancer routing entirely once the match starts. Be
precise about what this fixes: the destination is pinned, not the path. Ordinary internet
route churn can still change which links the packets traverse mid-match, but with a
unicast destination that reroute delivers to the same physical server instead of handing
the flow to a different one, which is the property the session actually depends on. - Cross-region jitter reduction for multi-region matches (players from different
regions in one match, which is common once player pools are thin): host the authoritative
game server in whichever region minimizes the maximum RTT across all participants (not
necessarily any single player's home region), and consider a relay/overlay network for
the worst-connected players rather than forcing everyone onto the public internet path.
Worked example
A player in Brazil starts matchmaking. The client pings the 3 closest of the 8 regions
(São Paulo, Miami, and a European fallback for redundancy) and measures 18ms/2ms jitter,
115ms/4ms jitter, and 190ms/6ms jitter respectively. São Paulo wins clearly.
Sanity-check latency figures like these against the physical floor before you trust them,
because a probe that reports an impossible number is a broken measurement, not a fast
network. Light in single-mode fiber travels at roughly two thirds of its speed in vacuum,
about 200,000 km per second, which works out to about 1ms of round trip for every 100km of
one-way great-circle distance. São Paulo and Miami are about 6,570km apart, so the absolute
floor is roughly 66ms round trip, and real submarine cable routes run about 1.4 to 2 times
the great-circle distance, which puts the honest expectation anywhere from 92ms at the
straightest plausible routing to 131ms at the most indirect. Then narrow that band with what
you know about the specific path, rather than quoting something three-quarters as wide as the
floor itself: this pair is not served by a direct cable, the routes hug the South American
coast and land through the Caribbean, so the realistic slice is the upper half of the band,
roughly 110 to 130ms, and the 115ms probe sits inside it. Keep the two checks separate when
you report them, because they catch different things: "above the physical floor" would wave
through a 95ms reading, and "consistent with the route that actually exists" is the one that
would not. The same check
puts São Paulo to central Europe (about 9,800km) at a 98ms floor and roughly 190ms in
practice, which is what the third probe shows. Notice that the two probes have to be
consistent with each other: a 70ms reading to Miami would imply a fiber path only 1.07 times
the great-circle distance while the European reading implies about 1.9 times, and no pair of
real routes behaves that differently. The
matchmaking service, having gathered similar probes from other players queued for the same
match, picks São Paulo as host since it also minimizes the worst-case RTT across the small
group of players it has matched together. The client then opens a direct UDP session to the
São Paulo server's specific IP, not an anycast address, for the rest of the match: if an
internet route between the player and São Paulo shifts mid-match, the destination IP is
unchanged and the packet still reaches the same physical server, which anycast could not
have guaranteed.
Trade-offs and pitfalls
- Client-side probing adds a small delay to matchmaking (typically under a second for a
handful of pings) in exchange for a materially better in-match experience; that's almost
always the right trade for a latency-sensitive game. - Region-of-origin isn't the same as latency-optimal region. A player near a regional
boundary, or on a network with an unusual peering path, can genuinely have a lower RTT to
a "farther" region; trust measurement over geography. - Cross-region matches force a compromise host location, which means no participant gets
their personal best latency; be explicit in matchmaking rules about how much RTT spread
you'll tolerate before splitting into separate lobbies instead.
Unlock Full Question Bank
Get access to all 6 Multi-Region and Geo-Distributed Systems interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.