IoT and Edge Device System Architecture Questions
Designing systems for large fleets of connected devices: device connectivity and provisioning, telemetry ingestion at scale, edge-versus-cloud processing splits, intermittent connectivity, and firmware/config rollout. Covers the constraints of constrained devices and the ingestion pipeline behind them. Distributed architecture where the edge is physical hardware.
Design a device twin replication and reconciliation protocol across geographic regions that guarantees eventual convergence. Compare using CRDTs versus last-writer-wins approaches for complex state such as configuration maps, and describe how you would surface and resolve conflicting operator intents in a user-friendly and auditable way.
Sample Answer
Clarify requirements & constraints
- Devices are constrained (flash/ram/CPU), intermittent connectivity, multi-region control planes must converge eventually, operator intents (configuration maps) must be auditable and resolvable with minimal device overhead.
High-level protocol
- Each device stores a local device-twin version vector and a compact change-log (opids + lightweight checksum).
- Regions replicate deltas via an anti-entropy mesh using push/pull; each delta includes: op id (uuid), origin region, timestamp, causal metadata (vector clock compressed), and operation payload (e.g., "set key X -> value V").
- Reconciliation runs at region-level and device-level: regions merge using convergence rules, devices pull reconciled state or incremental ops when online.
CRDT vs LWW for configuration maps
- CRDT (observed-remove map with LWW registers or multi-value register per key):
- Pros: Strong convergence without coordinator, preserves concurrent updates, merge is deterministic.
- Cons: Higher metadata per-key (tombstones, causal metadata) — pain on constrained devices. More complex implementation in C for safety.
- Use when preserving all operator intents and automatic merge semantics matter (e.g., sensor calibration offsets).
- LWW (last-writer-wins using region-preferred tie-breaker):
- Pros: Minimal metadata (timestamp + origin id), simple to implement, small footprint.
- Cons: Loses concurrent intents, relies on clock-sync/trust, surprising to operators for complex maps.
- Use when simplicity and resource constraints dominate.
Recommendation: Use CRDTs in the cloud/region replication layer and expose a compact, LWW-applied projection to devices. Devices receive small reconciled snapshots or idempotent ops to apply.
Surfacing & resolving conflicting operator intents
- Represent conflicts as first-class audit records: store op history and merge decisions in an append-only log (region-level).
- UI/CLI shows:
- Conflicting keys, list of competing ops (who, when, region), and automated-merged result with rationale.
- “Propose resolution” actions: pick op, pick region-priority, apply manual override (creates new op).
- Provide explainability: for each merged key include merge-policy (CRDT rule or LWW tie-break), contributing ops, and hash of resulting state for device verification.
- For safety-critical fields, mark as “manual-approval-required”: block auto-deploy to devices until operator resolves.
Auditing & device verification
- Each applied state change includes op id and signature; devices store last-applied op id. Region can prove which ops were applied — audit trail is compact (hash-chained).
- Provide a reconciliation API for devices to report divergent local state; server returns missing ops or full snapshot.
Trade-offs & practical notes for embedded dev
- Keep device agent minimal: apply ops idempotently, maintain last-applied op id and occasionally a lightweight bloom filter of seen ops.
- Offload heavy metadata & CRDT merging to cloud; devices only receive compact merged outcomes.
- Ensure OTA and rollback paths for bad resolutions; test merge semantics in low-memory C code, and simulate concurrent operator workflows.
Explain the LoRaWAN architecture and its device classes A, B, and C. For each class describe typical use cases, downlink timing behavior, expected power consumption, and the impact on gateway duty cycles and network capacity in unlicensed bands.
Sample Answer
Brief overview
LoRaWAN is a star-of-stars LPWAN: end devices talk to gateways using LoRa PHY; gateways forward packets to a Network Server (handles MAC, de-dup, ADR) and an Application Server. Devices are organized by class (A/B/C) that trade off downlink latency vs. power.
Class A — lowest power, uplink-initiated downlinks
- Downlink timing: two short RX windows after every uplink (Rx1 ~1 s after uplink, Rx2 ~2 s by default); downlink only possible following an uplink.
- Use cases: battery sensors, telemetry, meters, rare commands.
- Power: minimal — device sleeps most of the time; simple low-power modes and short wakeups.
- Gateway/network impact: lowest impact on duty cycle and capacity because downlinks are tightly coupled to uplinks and infrequent.
Class B — scheduled downlinks with beacons
- Downlink timing: device synchronizes to network beacons and opens periodic “ping slots” allowing scheduled downlinks between beacons.
- Use cases: semi-regular control, firmware update scheduling, time-synced actuations.
- Power: higher than A because device must wake for beacon sync and ping slots (requires RTC/accurate clock).
- Gateway impact: moderate — more predictable downlinks but increases overall downlink load and duty-cycle consumption.
Class C — low-latency, always-on receive
- Downlink timing: device keeps RX open almost continuously (closed only during transmission).
- Use cases: actuators requiring near-immediate command, gateways emulation in devices, testing.
- Power: highest — continuous RX consumes significant current; usually mains powered.
- Gateway impact: highest — frequent immediate downlinks increase duty-cycle usage in unlicensed bands (e.g., EU868 1%/0.1% depending on channel) and reduce network capacity; careful scheduling and gateway/channel planning required.
Embedded developer considerations
- Implement precise timers/RTC for Rx windows (Class A) and beacon sync (Class B).
- Optimize radio state transitions to minimize wake time and current draw.
- Respect regional duty-cycle / LBT rules; when designing firmware for large fleets, prefer Class A for battery devices and offload time-critical control to mains-powered Class C or alternative links.
- Monitor ADR and data-rate strategies to reduce airtime and gateway duty-cycle pressure.
Write a debounced button handler in C suitable for an embedded system where a tick function is called every 1 ms. The handler should avoid blocking calls, maintain state for each button, detect rising and falling edges, and be safe for use with an RTOS (i.e., ISR-safe signaling to a task). Provide the function prototype and a brief explanation of the state machine used.
Sample Answer
Approach (brief)
Use a per-button state machine sampled every 1 ms by tick(). Debounce requires a stable period (e.g., 20 ms). No blocking; ISR-safe signaling uses a volatile flag or RTOS safe fromISR notification (example uses a volatile event flag set in tick, to be read/cleared by task).
Function prototype
void button_tick(uint8_t button_id, bool raw_level);
Code (per-button structure + tick handler)
#include <stdint.h>
#include <stdbool.h>
#define MAX_BUTTONS 8
#define DEBOUNCE_MS 20
typedef enum { BTN_STABLE_LOW, BTN_DEBOUNCE_RISING, BTN_STABLE_HIGH, BTN_DEBOUNCE_FALLING } btn_state_t;
typedef struct {
btn_state_t state;
uint16_t counter; // ms counter
bool level; // last stable level
volatile uint8_t events; // bit0 = rising, bit1 = falling (ISR/task-safe)
} button_t;
static button_t buttons[MAX_BUTTONS];
void button_init(void){
for(int i=0;i<MAX_BUTTONS;i++){ buttons[i].state=BTN_STABLE_LOW; buttons[i].counter=0; buttons[i].level=false; buttons[i].events=0; }
}
// Called every 1 ms for a specific button with raw GPIO level (true = high)
void button_tick(uint8_t button_id, bool raw_level){
if(button_id >= MAX_BUTTONS) return;
button_t *b = &buttons[button_id];
switch(b->state){
case BTN_STABLE_LOW:
if(raw_level){ b->state = BTN_DEBOUNCE_RISING; b->counter = 0; }
break;
case BTN_DEBOUNCE_RISING:
if(raw_level){
if(++b->counter >= DEBOUNCE_MS){
b->state = BTN_STABLE_HIGH; b->level = true; b->events |= 0x01; // rising event
}
} else { b->state = BTN_STABLE_LOW; }
break;
case BTN_STABLE_HIGH:
if(!raw_level){ b->state = BTN_DEBOUNCE_FALLING; b->counter = 0; }
break;
case BTN_DEBOUNCE_FALLING:
if(!raw_level){
if(++b->counter >= DEBOUNCE_MS){
b->state = BTN_STABLE_LOW; b->level = false; b->events |= 0x02; // falling event
}
} else { b->state = BTN_STABLE_HIGH; }
break;
}
}
// Task-side: read and clear events (safe because single-writer tick, single-reader task)
uint8_t button_get_events(uint8_t button_id){
if(button_id >= MAX_BUTTONS) return 0;
uint8_t ev = buttons[button_id].events;
buttons[button_id].events = 0;
return ev;
}
State machine explanation
- Two stable states (LOW/HIGH) and two transient debounce states for edges.
- On candidate edge, enter debounce state and count ms; if level remains steady for DEBOUNCE_MS, commit state and set event.
- Events are stored in a volatile byte so tick (ISR or periodic) can set them without blocking; task polls button_get_events() to consume events. For RTOS, replace event flag updates with xTaskNotifyFromISR()/semaphore giveFromISR as appropriate.
Design a firmware architecture for an MCU with 256KB flash and 64KB RAM that must support multiple sensor drivers, a lightweight network stack, OTA updates, and a minimal bootloader. Describe module boundaries, a recommended memory layout (bootloader, OTA slots, application, scratch), inter-module interfaces (HAL vs drivers vs app), and strategies to minimize RAM usage (zero-copy, DMA, stack sizing).
Sample Answer
High-level approach
Design a layered, modular firmware: minimal bootloader (responsible for integrity, rollback, OTA trigger), a stable HAL + drivers layer, a lightweight network/OTA manager, and the application. Keep clear boundaries and simple RPC-style interfaces between layers.
Module boundaries
- Bootloader: verify signatures, swap/rollback, jump to app. Minimal hardware init.
- HAL: low-level registers, clocks, GPIO, UART, SPI, I2C, DMA APIs.
- Drivers: sensor drivers built on HAL (init, read, configure) — return buffer pointers, status codes.
- Network/OTA manager: small TCP/UDP or CoAP stack, OTA protocol handler, flash write abstraction.
- Application: business logic, sensor scheduling, power management.
Recommended flash/RAM layout (256KB Flash, 64KB RAM)
- Bootloader: 16 KB (0x0000 — 0x3FFF)
- OTA Slot A (primary app): 112 KB (0x4000 — 0x1BFFF)
- OTA Slot B (secondary): 112 KB (0x1C000 — 0x33FFF)
- Scratch/metadata (swap area, signatures, manifest): 8 KB (0x34000 — 0x35FFF)
- Note: reserve last 8 KB for factory data / fuses if needed.
RAM: reserve 4 KB for bootloader/interrupts, heap 12 KB, stack per thread (RTOS) 2 KB, driver buffers DMA-mapped 4 KB, rest for app (~42 KB total budgeting).
Inter-module interfaces
- HAL -> expose sync APIs and ISR registration; keep re-entrant and small.
- Drivers -> init(config*), read(buffer*, len), set_callback(cb*); avoid heap allocations.
- Network/OTA -> use flash abstraction (flash_write(addr, src, len), flash_erase), and driver read callbacks; OTA manager exposes start_update(manifest*), get_status().
- Application -> uses driver APIs and OTA manager; never touches flash directly.
RAM minimization strategies
- Zero-copy: drivers return pointers into DMA buffer or memory-mapped sensor FIFO instead of copying.
- DMA: use DMA for sensor and network RX/TX to avoid CPU copies; ring buffers in SRAM.
- Static allocation: prefer static buffers per peripheral; avoid malloc; use fixed-size pools.
- Stack sizing: analyze worst-case with static stack usage tools; set thread stacks small (e.g., 1–2 KB) and use stack-usage testing.
- Compression/chunked OTA: write OTA in chunks with small RAM window (e.g., 512B).
- Power/CPU: sleep between bursts to reduce transient memory needs.
Trade-offs
- Two full app slots double flash cost but enable safe rollback. If flash tight, use single-slot with incremental upgrade.
- More static RAM reduces flexibility but increases determinism.
This design yields safe OTA, clear module boundaries, and low RAM footprint suitable for 256KB/64KB constrained MCUs.
Design a geo-distributed telemetry ingestion system that guarantees at-least-once delivery and supports reconstructing time-ordered event streams for devices that may reconnect from different regions. Address partitioning strategy, watermarking for event windows, storage tiering for hot and cold data, and approaches to limit duplication and repair out-of-order events during processing.
Sample Answer
Clarify requirements & constraints (embedded perspective)
- Guarantee at-least-once ingestion from globally mobile/roaming devices that may reconnect from different regions.
- Must reconstruct time-ordered streams per device (event-time ordering), tolerate network partitions and device resource limits (flash, power).
- Reasonable latency for “hot” recent telemetry; long-term cold storage for analytics.
High-level architecture
- Devices → regional edge gateways / ingestion proxies → geo-replicated persistent queue (per-region Kafka or Kinesis) → stream processors (per-partition) → hot time-series store + cold object store (S3) + archival.
Partitioning strategy
- Partition (shard) by device-id (e.g., device UUID) using consistent hashing so all events for a device route to the same logical partition.
- Sticky routing at edge: use device-id → partition map cached at gateway; on partition rebalance, use rendezvous hashing to minimize movement.
- For load balancing, support per-device fan-out: if a single device produces extremely high rate, split by device-id + sub-shard using a device-local monotonic counter’s high bits.
Device-side identifiers & ordering
- Devices include:
- device-id (stable UUID)
- monotonic sequence number (persisted in flash; wrap/overflow handled)
- event timestamp (wall-clock or monotonic + epoch sync)
- Sequence number is primary ordering key; timestamp used for event-time semantics and watermarking.
- Because embedded devices can lose state, sequence numbers support duplicate detection and ordering; devices should persist last-seq on flash and increment atomically.
At-least-once delivery
- Devices retry until ACK from regional gateway. Gateways persist to durable queue (replicated to at least 2 replicas) before ACKing device — ensures at-least-once.
- Use idempotence keys: (device-id, sequence-number) as dedupe key downstream.
Watermarking & event windows
- Use event-time watermarks per partition to progress windowed processing:
- Watermark = max_event_time_seen_for_partition − allowed_lateness (configurable, e.g. 30s–5min depending on expected clock skew & reconnection characteristics).
- Maintain per-device max-event-time and per-partition max; emit window results when partition watermark passes window end.
- To handle clock skew, rely primarily on device sequence numbers for ordering; timestamps inform watermarks and late event detection.
Limiting duplication
- Maintain a recent dedupe index in the stream processor keyed by (device-id, sequence-number):
- Recent window stored in in-memory LRU + durable local state (RocksDB) to survive restarts.
- For long-term, compacted store (Cassandra/leveldb) keeps highest-seen sequence per device for coarse dedupe.
- Use Bloom filters to cheaply filter obvious duplicates at gateway for very recent events (tunable false-positive rate).
Repairing out-of-order events
- Buffering window per device in the partitioned stream processor:
- Buffer events up to allowed_lateness or until sequence-gaps are filled.
- Use sequence-number ordering first; when sequence gap occurs, wait for gap-fill until gap timeout; if timeout expires, emit with a gap marker and accept reorder repair later.
- Emit correction events if a late arrival requires changing prior outputs (compaction or tombstone/correction messages). Downstream readers subscribe to correction stream to apply adjustments.
- Store raw events (immutable) in cold store so reconstruction can always replay and rebuild correct order offline.
Storage tiering
- Hot tier (recent hours/days): fast time-series DB (e.g., TimescaleDB, Cassandra TS) partitioned by device-id + time window for low-latency queries and recent-telemetry dashboards. Keep indexed last-seq per device.
- Warm tier: compacted columnar store for query workloads (Parquet on SSD-backed object store).
- Cold tier: cold-object store (S3 Glacier or similar) with raw event packs (daily/monthly) for full replay/reconstruction.
- Pipeline: processors write to hot store; periodically batch/compact hot → warm → cold with dedupe/compaction.
Cross-region concerns
- Ingestion at region-local gateway; replicate persisted queues asynchronously to global cluster for analytics. For ordering per device, rely on partitioning guarantee — all events for a device map to the same logical partition (even if physically handled in different regions, logical ownership is consistent via hashing + routing). Use a global partition map service replicated to gateways.
- For low-latency regional consumption, route reads to nearest hot replica; reconcile via background compaction.
Operational considerations & metrics
- SLOs: ingestion ack latency, end-to-end availability, reorder latency (max time to repair), duplicate rate.
- Monitor sequence gaps, late event counts, watermark lag, disk-backed buffer sizes.
- Backpressure: gateways enforce rate-limits; devices implement exponential backoff and persist telemetry to circular flash buffer if offline.
Trade-offs
- At-least-once simplifies durability but requires dedupe. Deduplication state costs memory/storage; Bloom filters and compaction reduce cost.
- Strict exactly-once ordering globally is expensive across regions; use per-device ordering via sequence numbers and logical partitioning for practical correctness.
This design keeps device-side logic minimal (stable id, persisted monotonic seq, event timestamp), routes all of a device’s events to the same logical shard for ordering, uses watermarks with allowed-lateness to produce timely windows, tiers storage for cost/performance, and combines in-memory + durable dedupe with repair/correction for late/out-of-order events.
Unlock Full Question Bank
Get access to all IoT and Edge Device System Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.