A read replica is an asynchronous, read-only copy of a primary database. The primary streams its stream of committed changes (the binary log in MySQL, the write-ahead log, or WAL, in PostgreSQL) to one or more replicas, which replay those changes to stay current. The replica accepts read queries but never writes, and it is always at least slightly behind the primary because replication is asynchronous: the primary does not wait for the replica before acknowledging a write.
How it works in managed offerings
- Amazon RDS (MySQL/PostgreSQL/MariaDB): uses the engine's own native replication mechanism, binlog-based for MySQL/MariaDB, WAL streaming for PostgreSQL, to ship changes to up to several read replicas per primary. Each replica is a fully separate instance with its own storage and its own endpoint.
- Amazon Aurora: replicas share the same underlying cluster storage volume as the writer instead of copying data over the network, so replica lag is typically much lower than RDS's engine-native replication under normal load, low enough that AWS markets Aurora's replicas as "low-latency." Aurora also lets a replica be promoted automatically as part of Multi-AZ failover (Multi-AZ means the provider keeps a synchronized standby copy of the database in a second availability zone; failover is automatically switching traffic to that standby, or here, to a promoted replica, when the primary fails), which RDS's plain read replicas cannot do.
- Google Cloud SQL: replicates asynchronously per-transaction from the primary to one or more read replicas, including optional cross-region replicas, using the same engine-native mechanisms under the hood.
In all three, the replica is provisioned and monitored by the provider (you don't manage the replication process by hand), but the replication itself is still asynchronous and lag is not bounded to zero.
What problems read replicas solve
- Read scaling: route SELECT-heavy traffic to one or more replicas so it does not compete with writes for CPU and I/O on the primary.
- Workload isolation: point a specific class of query, for example an analytics dashboard or a nightly reporting job, at a dedicated replica so a slow or expensive query can't degrade the primary's write latency.
- Reduced primary load and latency: fewer connections and less query volume on the primary generally means more headroom and lower p95/p99 latency (the response time for the slowest 5% and slowest 1% of requests, a way of measuring worst-case performance rather than the average) for the writes and reads that must hit the primary.
- A promotable target: an RDS read replica (not a Multi-AZ standby) can be manually promoted to a standalone writable instance, which is useful for cross-region disaster recovery or for splitting off a subset of traffic to its own database.
Limitations and risks to watch
- Replication lag is not bounded. Under normal load it can be small enough to ignore; under a write burst, a long-running DDL (data definition language, for example an ALTER TABLE) statement, or a large batch job on the primary, it can grow to seconds or minutes, and the replica has no way to signal "I'm behind" to a client unless you build that check yourself.
- Eventual consistency, not read-your-writes. A client that writes to the primary and immediately reads from a replica can see stale data, or even see the write "disappear" if it lands on a replica that hasn't caught up yet. This is the single most common production bug with read replicas: a user submits a form, the confirmation page reads from a replica, and the user's own just-written data appears missing.
- No cross-replica consistency guarantee. Two replicas of the same primary can be at different points in the replication stream at the same instant, so two requests hitting different replicas can see different, mutually inconsistent snapshots.
- Promotion is a manual, disruptive operation, not an automatic failover mechanism. Promoting a read replica breaks its replication link permanently, and every other reader and writer in the application needs to be repointed. This is a materially different (and slower) recovery path than a Multi-AZ standby failover.
- Real added cost. A read replica is a full running instance, not a lightweight cache, so it roughly multiplies your compute bill by the number of replicas you run.
Worked example: how a write burst turns into visible lag
Assume a primary normally emits 12 MB/s of write-ahead log under steady state, a replica can apply changes at 20 MB/s, and a 10-minute batch job pushes the primary's write rate to 30 MB/s.
backlog growth rate during burst=30−20=10 MB/s
Over the 600-second burst that backlog reaches 10×600=6,000 MB of unreplayed write-ahead log sitting on the replica the instant the burst ends.
It is tempting to convert that 6,000 MB straight into a "seconds behind" figure by dividing by the replica's apply rate (6,000 / 20 = 300 seconds), but that is not the number a real replication-lag metric would show at that instant, and it is a common mistake when reasoning about replication under load. What tools like MySQL's seconds_behind_master or Aurora/RDS's ReplicaLag CloudWatch metric actually report is how old the transaction the replica is currently replaying is, which you get by comparing cumulative bytes applied against cumulative bytes produced, not backlog divided by capacity. At the moment the burst ends (600 seconds in), the replica has applied 20×600=12,000 MB while the primary has produced 30×600=18,000 MB since the burst started. The 12,000 MB the replica has applied so far is data the primary had already produced after just 12,000/30=400 seconds of the burst, so the replica is 600−400=200 seconds, about 3 minutes 20 seconds, behind at the instant the burst ends, not 300 seconds.
Lag keeps growing for a while even after the burst is over, because the replica is still capacity-constrained at 20 MB/s working through the backlog while the primary keeps producing new data (now at the lower 12 MB/s rate) that pushes "how old is what I'm applying right now" further out. Lag peaks about 5 minutes after the burst ends, at 300 seconds (5 minutes) behind, then shrinks back toward zero as the replica keeps applying at 20 MB/s against a primary now producing only 12 MB/s, a net drain rate of 20−12=8 MB/s. The 6,000 MB backlog itself (still sitting there right when the burst ends) fully drains 6,000/8=750 seconds, 12.5 minutes, after the burst ends, which is also when lag finally returns to zero.
Total elapsed time with a nonzero backlog and elevated lag: the 600-second burst plus the 750-second drain, 1,350 seconds, 22.5 minutes, not the 17.5 minutes you get from naively adding a 300-second instant-lag figure to the drain time. Lag grows for the first 900 seconds (15 minutes: the full 10-minute burst plus 5 more minutes while the primary is still outpacing the still-catching-up replica), peaks at 5 minutes behind, then shrinks over the remaining 450 seconds (7.5 minutes). For that whole 22.5-minute window, anything reading from that replica was looking at data that was seconds to minutes old. This is exactly the kind of window where a "read after write" bug shows up, and it is why replica lag needs to be an explicit, monitored metric (RDS and Aurora both expose it as a CloudWatch metric) rather than an assumption, or worse, a hand calculation done wrong under incident pressure.