Real-Time and Streaming System Design Questions
Designing client-facing, always-on live delivery systems: real-time communication transports (WebSockets, server-sent events, long-polling, WebRTC, MQTT), connection lifecycle and scaling for millions of persistent connections, presence, pub/sub fan-out to online users, chat and notifications, live feeds and tickers, real-time collaboration (CRDT vs OT, offline sync and reconciliation), and live and on-demand video delivery (ingest, transcoding, adaptive bitrate, CDN delivery, low-latency protocols, playback entitlement and content protection). Covers latency budgets, per-client ordering, delivery guarantees across disconnect and reconnect, per-client backpressure, capacity estimation for connection fleets, authentication, authorization, and revocation for real-time channels and premium content, and operational readiness (SLOs and error budgets, observability and incident response, rate limiting, safe rollout, and load and chaos testing) for live-delivery platforms. Data-pipeline stream processing (Kafka or Flink jobs, windowed aggregation, exactly-once pipelines, real-time analytics ingestion) is out of scope.
You're on a sales call: a customer says they need 'real-time' user presence and in-app notifications for 500k monthly users but has a limited budget and prefers cloud-managed services. What discovery questions do you ask to uncover performance, durability, retention, compliance, and UX expectations? Draft an MVP architecture and explain how you'd balance cost, functionality, and ability to iterate.
Sample Answer
Direct answer
"Real-time" is a feeling, not a requirement, so the first job on the call is to turn it into numbers: how fast an update must appear, what happens if one is lost, how long anything is kept, and which regulations apply. For 500k monthly users on a tight budget, the answer is almost always a fully managed, pay-per-use MVP (minimum viable product: the smallest version that proves value): a managed WebSocket gateway, serverless handlers (small units of code that run only when triggered and bill per invocation, with no server for you to provision or patch), a key-value table with automatic expiry for presence, and the phone platforms' own push services for users who are offline. At the volumes below it costs on the order of a few hundred dollars a month in gateway charges, and it can be replaced piece by piece once real usage data exists.
Terms used below
- Presence: showing whether a user is online. In-app notification: a message shown inside the app while it is open; a push notification reaches a phone even when the app is closed, via Apple Push Notification service (APNs) or Firebase Cloud Messaging (FCM).
- WebSocket: a persistent two-way connection so the server can send without the client polling.
- Durability: whether a message survives a crash or a disconnected user. Retention: how long it is kept.
- MAU / DAU: monthly / daily active users. Peak concurrency: how many are connected at the busiest moment.
Discovery questions, grouped by what they decide
| Area | Question I ask | What the answer changes |
|---|---|---|
| Performance | "When user A does something, how soon must user B see it: under 1 second, under 5, or 'within a minute is fine'?" | Under ~1 s needs a persistent connection; tens of seconds can be polling, which is cheaper and simpler. |
| Performance | "What is your busiest hour, and what fraction of monthly users are on at once then? Any spikes like a launch or a live event?" | Sizes peak connections and the new-connections-per-second quota. |
| Performance | "Who needs presence: everyone, a friends list, a team of 10?" | Presence fan-out (delivering one update to every person watching that user's status) grows with list size; a 5,000-member org presence is a different product. |
| Durability | "If a notification is sent while the user is offline, must they see it later, or is it fine to miss it?" | Missable means fire-and-forget; must-see means a stored inbox and replay on reconnect. |
| Durability | "Is any notification tied to money, safety or a legal deadline?" | Those need delivery confirmation and a fallback channel (email or SMS), not just a socket. |
| Retention | "How long should notification history be visible? Must you be able to prove later what was sent?" | Sets expiry on stored items and whether an audit log is needed. |
| Compliance | "Which regions are your users in? Any data residency rules, GDPR (EU General Data Protection Regulation) obligations, HIPAA (US health data law), or regulated industries?" | Region choice, what may appear in a push payload (lock screens are visible to anyone), deletion workflows. |
| Compliance | "What identity provider do you use today?" | We reuse its tokens instead of building login. |
| UX | "Does 'online' mean app open, or active in the last 5 minutes? Do you need typing indicators or read receipts?" | Each extra signal multiplies message volume; read receipts also need storage. |
| UX | "What should the user see when their connection drops: a banner, silent retry, stale data?" | Reconnect and catch-up behaviour, which is where most real-time bugs live. |
| Budget and team | "What monthly spend is comfortable, and who operates this after launch?" | Confirms managed services and rules out a self-run fleet. |
I close discovery by reading back a one-line spec, for example: "Presence for friends lists of up to 200, visible within 5 seconds; in-app notifications within 2 seconds while open, delivered as push when closed, kept 30 days; EU data stays in the EU." The customer agreeing to that sentence is the real output of the call.
MVP architecture (assuming an AWS shop; the same shape exists on other clouds)
flowchart LR
App[Mobile or web app] -->|WebSocket| GW[Managed WebSocket gateway]
GW -->|connect, disconnect, ping| FN[Serverless handlers]
FN --> CT[(Connections and presence table, TTL)]
FN --> IN[(Notification inbox table, TTL)]
SRC[Customer backend events] --> FN
FN -->|online: send to connection| GW
FN -->|offline: push| PUSH[APNs and FCM]
- Gateway: Amazon API Gateway WebSocket APIs. It holds the connections, so the customer runs no servers. Relevant limits from its documentation: a connection lasts at most 2 hours, idle connections are closed after 10 minutes, and the default is 500 new connections per second per account per region (raisable).
- Handlers: AWS Lambda functions (code run per event, billed per invocation) for connect, disconnect and ping.
- Presence: a DynamoDB (managed key-value database) row per connection with a TTL (time-to-live: the database automatically deletes the row once this deadline passes, with no manual cleanup needed) a few minutes ahead, refreshed by each ping. The gateway's disconnect event is best-effort, so the TTL, not the disconnect event, is the source of truth for "offline".
- Notifications: written to an inbox table first (durable, with a TTL matching the retention answer), then sent to any live connection. If none, sent as a push via APNs/FCM. On reconnect, the app asks for everything after its last seen notification ID.
- Identity: tokens from the customer's existing provider, verified on connect.
Sizing and cost (assumptions stated, arithmetic shown)
Assumptions to confirm with the customer: 20% of MAU are daily users; each daily user is connected 30 minutes a day; peak concurrency is 10% of DAU; client pings every 5 minutes (inside the 10-minute idle limit); 20 in-app notifications and 20 presence updates received per daily user per day; all messages under 32 KB. 30-day month.
- DAU = 500,000 x 0.2 = 100,000. Peak concurrent = 100,000 x 0.1 = 10,000 connections, far below any gateway limit.
- Connection minutes = 100,000 x 30 min x 30 days = 90,000,000. At the listed US East price of $0.25 per million connection minutes: 90 x $0.25 = $22.50/month.
- Messages: pings 100,000 x 6 per day x 30 = 18,000,000; notifications 100,000 x 20 x 30 = 60,000,000; presence 100,000 x 20 x 30 = 60,000,000. Total 138,000,000. At $1.00 per million messages (metered in 32 KB units): $138/month.
- Gateway total is about $160/month. Lambda invocations and DynamoDB reads and writes are billed separately; I would price them in the AWS pricing calculator with the customer's confirmed numbers rather than guess on the call.
The useful insight for the customer: presence and pings, not notifications, are more than half of the message bill. That is the lever for cost, not the choice of database.
Balancing cost, functionality and iteration
| Ship in the MVP | Defer until usage proves it | Why |
|---|---|---|
| Presence for small friend lists, 5-minute ping | Typing indicators, "last seen 2 min ago" | Each is a new message stream with its own fan-out cost. |
| Durable inbox with replay on reconnect | Read receipts synced across devices | Cross-device sync adds writes and conflict rules. |
| Push fallback for offline users | Multi-region active-active (running fully independent copies of the system in more than one region, each simultaneously accepting writes) | One region meets the latency target for most single-market customers. |
| Metrics: connect failures, delivery latency, messages per user | Custom self-hosted gateway | Only worth it when the per-message bill exceeds the cost of engineers running servers. |
How I keep the ability to iterate:
- One message envelope (
type, id, ts, payload) from day one, so new event types are additive. - Feature flags (switches that turn a feature on for a subset of users without a deploy) for presence granularity, so we can measure the cost of a richer signal on 5% of users before rolling it out.
- A clean seam between "decide who gets what" (customer's logic) and "deliver it" (the gateway), so the delivery layer can later be swapped for a managed real-time SaaS (software-as-a-service) product or a self-run fleet without rewriting business logic.
- An agreed revisit trigger: for example, revisit the architecture if peak concurrency exceeds 100,000 or the monthly gateway bill passes a figure the customer names.
Pitfalls I steer the customer away from
- Promising "instant" without a number: every later dispute comes from that word.
- Treating the disconnect event as reliable presence; users look online forever after a network drop.
- Putting message content in push payloads when compliance answers say it is sensitive; send "You have a new message" and fetch the content after unlock.
- Over-building for scale they do not have: 10,000 peak connections is a small system, and a self-hosted cluster would cost more in engineering time than the managed bill.
What would change the design: sub-200 ms latency or large collaborative rooms (move to a dedicated real-time service or self-hosted WebSocket fleet), strict must-deliver notifications (add a delivery-receipt state machine, a tracked set of stages per message such as sent/delivered/read so you can prove and query exactly where each one got stuck, and email/SMS fallback), or data residency across several regions (one stack per region with per-region user homing, meaning each user is permanently assigned to one region's stack).
A client asks whether to use a managed real-time platform (e.g., AWS API Gateway WebSocket, AppSync, Pusher) versus building and operating their own WebSocket fleet. Create a decision matrix comparing cost, time-to-market, operational risk, customization and feature flexibility, vendor lock-in, and compliance. Provide recommendations for a bootstrapped startup and for a regulated large enterprise.
Sample Answer
Direct answer
For a bootstrapped startup, buy: a managed real-time service gets them to market in days at tens of dollars a month, and an in-house WebSocket fleet costs an engineer's attention they cannot spare. Pick AppSync Events if they are already on AWS, Pusher if they want the most batteries-included SDKs regardless of cloud. For a regulated large enterprise, choose a cloud-provider managed service running inside their own cloud account (AppSync or API Gateway WebSocket), which keeps data, keys and logs inside their compliance boundary. Build your own fleet only when a hard requirement exceeds the managed limits (connection lifetime, protocol, fan-out economics or on-premises residency).
The options
- Third-party SaaS, software as a service (Pusher Channels): a hosted publish/subscribe service with client SDKs, channels, presence (who is online) and authentication hooks. Priced by plan with caps on concurrent connections and daily messages.
- Cloud-managed (AWS API Gateway WebSocket APIs, AWS AppSync): API Gateway keeps WebSocket connections open and forwards each message to your backend (typically AWS Lambda functions: code that runs on demand without you managing a server, billed per invocation); to send to a client, your backend calls a per-connection endpoint. AppSync is AWS's managed GraphQL (a query language and API style where the client specifies exactly which fields it wants in one request) service with real-time subscriptions; AppSync Events is its simpler channel-based publish/subscribe mode.
- Self-built fleet: your own WebSocket servers behind a load balancer, plus a broker (a separate service that passes messages between your own server processes), for example Redis (an in-memory data store often used this way) or Kafka (a distributed log built for high-throughput message streaming), for fan-out (delivering one message out to every subscribed connection) between nodes.
Verified limits and prices that shape the decision
From AWS and Pusher documentation (list prices, US East for AWS; check your region):
- API Gateway WebSocket: $1.00 per million messages (first volume tier; higher tiers are cheaper), $0.25 per million connection-minutes, messages metered in 32 KB increments. Quotas: 2-hour maximum connection duration, 10-minute idle timeout, 128 KB message payload with a 32 KB frame size, and a default of 500 new connections per second per account per Region (adjustable).
- AppSync: GraphQL real-time updates $2.00 per million; AppSync Events $1.00 per million operations; connection-minutes $0.08 per million for both.
- Pusher Channels: plan-based, for example Startup $49/month (500 connections, 1 million messages/day) and Pro $99/month (2,000 connections, 4 million messages/day), with enterprise plans above the listed tiers.
Worked cost examples
Startup: 2,000 concurrent connections all month, 1 million messages/day.
- Connection-minutes: 2,000 × 60 × 24 × 30 = 86.4 million.
- API Gateway: 86.4 × $0.25 = $21.60 plus 30 million messages × $1.00 per million = $30.00, $51.60/month.
- AppSync Events: 86.4 × $0.08 = $6.91 plus $30.00, about $36.91/month.
- Pusher Pro: $99/month flat, within both caps.
- Self-built: two small servers for redundancy plus a broker costs little in infrastructure, but the real cost is engineering and on-call time. As a labelled ESTIMATE: at a fully-loaded engineer cost of roughly $180,000/year ($15,000/month; adjust to your market), even 1% of one engineer's month, about 1.7 hours spread across that month (under half an hour a week, since a month holds roughly 4.3 weeks) spent on maintenance and on-call, is $150/month, which already exceeds all three managed figures above ($51.60, $36.91 and $99).
Enterprise: 200,000 concurrent connections all month, 20 messages per connection per hour.
- Connection-minutes: 200,000 × 43,200 = 8.64 billion. API Gateway: 8,640 × $0.25 = $2,160. AppSync Events: 8,640 × $0.08 = $691.20.
- Messages: 200,000 × 20 × 720 h = 2.88 billion. At $1.00 per million that is at most $2,880 (volume tiers lower the API Gateway figure).
- Monthly: API Gateway at most about $5,040, AppSync Events at most about $3,571. At this scale cost does not decide the question; compliance and control do.
- A hidden constraint: API Gateway's 2-hour limit forces every connection to reconnect at least every 2 hours. 200,000 ÷ 7,200 s ≈ 28 reconnects per second in steady state, well under the default 500/s quota, but clients must handle it silently and the quota must be raised before a mass-reconnect event.
Decision matrix
Scores are 1 (poor) to 5 (strong) and are judgement calls; weights differ by persona. Rerun them with the client's own weights.
| Criterion | SaaS (Pusher) | Cloud-managed (AppSync / API Gateway) | Self-built fleet | Weight: startup | Weight: enterprise |
|---|---|---|---|---|---|
| Cost at their scale | 4 | 4 | 2 | 0.25 | 0.10 |
| Time-to-market | 5 | 4 | 1 | 0.30 | 0.10 |
| Operational risk (who is paged) | 4 | 4 | 2 | 0.20 | 0.25 |
| Customization and feature flexibility | 2 | 3 | 5 | 0.10 | 0.15 |
| Vendor lock-in (5 = least) | 2 | 2 | 5 | 0.05 | 0.10 |
| Compliance fit | 2 | 4 | 4 | 0.10 | 0.30 |
| Weighted score | startup 3.80, enterprise 3.00 | startup 3.80, enterprise 3.65 | startup 2.35, enterprise 3.25 |
Example: the startup's cloud-managed score is 0.25×4 + 0.30×4 + 0.20×4 + 0.10×3 + 0.05×2 + 0.10×4 = 3.80.
Criterion by criterion
- Cost: managed wins at startup scale because the minimum spend is tiny; at enterprise scale self-built infrastructure can be cheaper, but only after paying for a team.
- Time-to-market: SaaS is fastest (SDKs, presence and auth out of the box); cloud-managed needs backend glue; self-built needs months.
- Operational risk: managed services absorb scaling, patching and DDoS (distributed denial-of-service) handling; self-built means you own reconnect storms, capacity and kernel tuning (adjusting the operating system's own connection and resource limits, which it does not raise by default).
- Customization: self-built allows custom protocols, binary framing (a compact byte-level message format you design yourself, instead of WebSocket's default text/JSON), connections longer than 2 hours and custom fan-out. API Gateway fan-out is one backend call per recipient connection, so broadcasting to 100,000 clients means 100,000 calls; AppSync and Pusher fan out channels for you.
- Vendor lock-in: mitigate in either managed option by keeping your own message schema and a thin internal publish interface, so the transport can be swapped.
- Compliance: cloud-managed runs in the enterprise's own account, under its encryption keys, network controls and logging, and inside the provider audit reports they already rely on. A third-party SaaS adds a new vendor to assess (their SOC 2 report, an independent audit of security controls, data-processing agreement: a contract spelling out how the vendor may handle the enterprise's data, required under regimes like GDPR, data residency; Pusher lets you choose a cluster region). Self-built gives full control but the enterprise then carries all evidence gathering itself.
Recommendations
- Bootstrapped startup: the matrix ties SaaS and cloud-managed at 3.80, so break the tie on context. Already on AWS: AppSync Events (pay-per-use, channel fan-out built in). Not on AWS, or needs presence and SDKs today: Pusher. Revisit when the bill or a missing feature starts hurting.
- Regulated enterprise: cloud-managed inside their own account (3.65). Choose AppSync Events for channel-style broadcast, API Gateway WebSocket when each message must hit custom backend logic. Build a fleet (3.25) only for a documented requirement the managed limits cannot meet, and budget the platform team explicitly.
Pitfalls
- Comparing infrastructure cost with no line for engineering and on-call time.
- Missing the 2-hour and 10-minute limits until clients start dropping in production.
- Building on API Gateway for broadcast-heavy workloads and discovering the per-connection fan-out cost later.
That is every published Real-Time and Streaming System Design question for Technical Product Manager so far. Browse the other topics in this category, or practice this one interactively.