Google Embedded Developer (Mid-Level) Interview Preparation Guide
Google's Embedded Software Engineer interview process for mid-level candidates combines technical depth with practical problem-solving. The process includes an initial recruiter screening, a technical phone screen focused on embedded systems and coding, and multiple onsite rounds covering low-level programming, system design, hardware-software integration, real-time systems optimization, and behavioral assessment. Interviews emphasize C programming proficiency, embedded systems concepts, bit manipulation, driver development, and practical experience with hardware constraints.
Interview Rounds
Recruiter Screening
What to Expect
Initial conversation with a Google recruiter to assess your background, motivation, and fit for the Embedded Developer role. The recruiter will discuss your experience with embedded systems, hardware-software integration projects, and clarify role expectations. This is a soft evaluation to ensure basic qualifications and cultural fit before technical interviews.
Tips & Advice
Prepare a clear 2-3 minute pitch about your embedded systems experience. Highlight specific projects involving microcontrollers, firmware, or hardware interaction. Ask questions about the specific embedded domain at Google (IoT, hardware platforms, performance constraints). Be enthusiastic about low-level systems work and demonstrate genuine interest in embedded development rather than general software engineering.
Focus Topics
Background & Career Journey
Clear articulation of your embedded systems experience, progression, and motivation for Google
Practice Interview
Study Questions
Google Role Understanding
Knowledge of what Google Embedded Developer role involves and alignment with your career goals
Practice Interview
Study Questions
Relevant Project Experience
Concrete examples of embedded projects (microcontrollers, firmware, drivers, IoT) with measurable outcomes
Practice Interview
Study Questions
Technical Phone Screen
What to Expect
A 60-minute technical interview conducted over video/phone with a Google engineer. This round assesses your embedded systems knowledge and coding ability under pressure. Expect a combination of practical embedded problems, bit manipulation questions, basic data structures, and C programming fundamentals. Problems are more practical and hardware-oriented than standard algorithmic challenges. You'll be coding in a shared document or IDE.
Tips & Advice
Focus on C programming correctness and understanding data types. Practice bit manipulation (bit shifting, masking, flags) as this is heavily tested. Think aloud about memory implications and hardware constraints even in simple problems. Be prepared to handle embedded-specific scenarios like register manipulation, interrupt handling concepts, or power optimization. Don't over-engineer solutions; simplicity and correctness matter. If stuck, explain your thought process rather than guessing.
Focus Topics
Arrays, Strings & Basic Data Structures
Practical usage of arrays, strings, linked lists, and simple data structures in embedded context
Practice Interview
Study Questions
Memory Management & Optimization
Stack vs heap, memory constraints in embedded systems, buffer management, and avoiding memory waste
Practice Interview
Study Questions
Embedded Problem Solving
Practical problems involving hardware constraints, register manipulation, interrupt handling concepts, or device communication
Practice Interview
Study Questions
Bit Manipulation & Bitwise Operations
Bit shifting, masking, setting/clearing bits, flag operations, and practical register manipulation
Practice Interview
Study Questions
C Programming & Fundamentals
Core C syntax, pointer manipulation, memory management, and understanding of data types for embedded contexts
Practice Interview
Study Questions
Onsite Round 1: Low-Level Programming & Embedded Fundamentals
What to Expect
First onsite technical round focusing on low-level C programming, microcontroller programming concepts, and embedded systems architecture. You'll work through problems that involve understanding hardware registers, bit-level operations, interrupt handling, and peripheral communication. This round evaluates your comfort with assembly-level thinking and hardware-software interaction.
Tips & Advice
Draw diagrams if discussing register layouts or memory mapping. Explain your understanding of how code translates to hardware operations. Be specific about data types chosen and why (int vs uint32_t, etc.). Discuss trade-offs between performance and code clarity. If you mention RTOS or specific microcontroller experience, be prepared to go deep. Think about edge cases like integer overflow in embedded contexts.
Focus Topics
Peripheral Communication Protocols (I2C, SPI, UART)
Understanding serial communication protocols used in embedded systems, protocol basics, and troubleshooting communication issues
Practice Interview
Study Questions
Interrupts & Exception Handling
Interrupt service routines (ISRs), interrupt priorities, edge cases in interrupt handling, and atomic operations
Practice Interview
Study Questions
Code Optimization for Embedded Constraints
Optimization techniques for limited memory, CPU cycles, and power consumption specific to embedded platforms
Practice Interview
Study Questions
Microcontroller & Processor Fundamentals
Understanding microcontroller architecture, CPU registers, memory layout, and basic processor operation relevant to embedded systems
Practice Interview
Study Questions
Register Manipulation & Hardware Abstraction
Direct register access, volatile qualifiers, memory-mapped I/O, and understanding Hardware Abstraction Layers (HAL)
Practice Interview
Study Questions
Onsite Round 2: Device Drivers & Hardware-Software Integration
What to Expect
Technical interview focused on driver development, hardware-software integration, and practical use-case problems. You may be asked about driver architecture, device initialization sequences, register configuration, handling hardware quirks, and integration patterns. This round tests your ability to work at the interface between software and hardware.
Tips & Advice
If you mentioned specific IP or hardware in your resume, be thoroughly prepared to discuss driver implementations. Discuss real challenges you faced (timing issues, hardware bugs, version differences). Explain how you debug hardware-software integration problems. Be comfortable discussing both simple and complex peripherals. Mention any experience with device trees, kernel drivers, or bootloader code. Show understanding of the full lifecycle from hardware spec to functional driver.
Focus Topics
Device Initialization & Configuration
Sequence of steps to initialize hardware, configuration register settings, clock setup, and reset handling
Practice Interview
Study Questions
Hardware-Software Integration & Debugging
Debugging hardware-software interactions, using oscilloscopes/analyzers, understanding common integration issues, and troubleshooting strategies
Practice Interview
Study Questions
Handling Hardware Quirks & Edge Cases
Working around hardware limitations, version-specific behavior, race conditions, and manufacturing variations
Practice Interview
Study Questions
Device Driver Architecture & Design
Basic driver structure, layered driver design, device discovery, initialization sequences, and driver state management
Practice Interview
Study Questions
Hardware Specification Analysis
Reading and interpreting hardware datasheets, understanding register maps, timing diagrams, and hardware capabilities
Practice Interview
Study Questions
Onsite Round 3: Real-Time Systems & Operating Systems Concepts
What to Expect
Technical interview evaluating understanding of real-time operating systems (RTOS), task scheduling, synchronization primitives, and embedded OS concepts. Questions may involve multi-tasking scenarios, interrupt handling in OS context, mutex/semaphore usage, priority inversion, and real-time constraints. This assesses your ability to design systems that meet timing requirements.
Tips & Advice
Understand the difference between RTOS and general-purpose OS. Be clear about preemption, context switching, and deterministic behavior. Discuss trade-offs between simplicity and robustness. If you have RTOS experience (FreeRTOS, Zephyr, ThreadX, etc.), discuss specific scenarios. Explain how you debug timing issues and race conditions. Show comfort with concurrent programming in embedded context. Discuss priority-based scheduling and how it applies to real systems.
Focus Topics
Interrupt Handling in OS Context
ISR design in OS environments, interrupt priorities, interrupt nesting, and interaction with task scheduling
Practice Interview
Study Questions
Memory Management in Embedded OS
Static vs dynamic allocation in RTOS, memory pools, fragmentation concerns, and stack/heap management
Practice Interview
Study Questions
Real-Time Constraints & Timing Analysis
Understanding deadline requirements, response time analysis, and designing systems to meet timing constraints
Practice Interview
Study Questions
Real-Time Operating Systems (RTOS) Fundamentals
Task scheduling, context switching, preemption, determinism, and real-time constraints in embedded OS
Practice Interview
Study Questions
Synchronization & Concurrency in Embedded Systems
Mutex, semaphore, event flags, message queues, and avoiding race conditions in embedded multi-tasking
Practice Interview
Study Questions
Onsite Round 4: System Design & Architecture
What to Expect
System design round appropriate for mid-level embedded developers. Rather than distributed system design, this focuses on embedded system architecture: designing microcontroller-based solutions, choosing appropriate components, handling data flow, power management strategies, and scaling embedded systems. You'll discuss trade-offs between performance, power, cost, and complexity. Questions may involve designing IoT devices, sensor systems, or embedded subsystems.
Tips & Advice
Structure your answer: clarify requirements, discuss hardware choices, explain software architecture, address power and memory constraints. Draw block diagrams and data flow. Discuss sensor selection, communication protocols, data processing pipeline. Consider edge cases like sensor failures or network outages. Show understanding of the full system from sensors through processing to actuators. Discuss why specific design choices were made (cost, performance, reliability). Be realistic about embedded constraints.
Focus Topics
Scalability & Modularity in Embedded Design
Designing embedded systems that scale from prototype to production, modularity, abstraction layers, and managing complexity
Practice Interview
Study Questions
Sensor Integration & Data Acquisition
Choosing sensors, ADC/DAC configuration, sampling rates, noise filtering, and reliable data collection
Practice Interview
Study Questions
IoT & Connectivity Design
Choosing communication protocols, edge vs cloud processing, data transmission optimization, and connectivity reliability
Practice Interview
Study Questions
Power Management & Optimization
Low-power design, sleep modes, power domains, energy budgeting, and optimizing power consumption
Practice Interview
Study Questions
Embedded System Architecture & Design Patterns
Architectural patterns for embedded systems, component selection, and designing systems with strict resource constraints
Practice Interview
Study Questions
Onsite Round 5: Behavioral, Collaboration & Google Culture
What to Expect
Behavioral interview assessing your communication, teamwork, problem-solving approach, and alignment with Google values (boldness, responsibility, collaboration, user-focus). Expect questions about past projects, handling conflicts, learning from failures, and how you work with hardware engineers and cross-functional teams. This round evaluates cultural fit and soft skills essential for mid-level roles.
Tips & Advice
Prepare 4-5 concrete examples using STAR format (Situation, Task, Action, Result). Focus on collaborative projects, technical challenges you solved, and learning experiences. Emphasize how you communicated with hardware engineers and non-technical stakeholders. Discuss a time you failed and what you learned. Show curiosity and ownership. Talk about your growth as an embedded developer. Mention how you stay current with embedded systems knowledge. Demonstrate understanding of Google's mission and how embedded systems support it.
Focus Topics
Learning & Growth Mindset
Continuous learning, adapting to new technologies, seeking feedback, and mentoring junior engineers
Practice Interview
Study Questions
Ownership & Accountability
Taking responsibility for projects, following through on commitments, and driving solutions to completion
Practice Interview
Study Questions
Problem-Solving & Debugging Methodology
Systematic approach to debugging, hypothesis testing, persistence with tricky bugs, and learning from failures
Practice Interview
Study Questions
Technical Communication & Documentation
Explaining complex embedded concepts clearly, documenting drivers and code, and creating design specifications
Practice Interview
Study Questions
Collaboration with Hardware Engineers
Communication patterns, understanding hardware perspectives, joint problem-solving, and navigating software-hardware disagreements
Practice Interview
Study Questions
Frequently Asked Embedded Developer Interview Questions
What is a counting semaphore? Describe wait and signal, how it differs from a binary semaphore, and a realistic use such as limiting concurrent access to a scarce resource. What goes wrong with fairness or misuse?
Sample Answer
Direct answer
A counting semaphore is a counter plus a wait queue, managed by two operations. Wait (also called P, down or acquire) blocks until the counter is above zero, then decrements it. Signal (V, up, release or post) increments the counter and wakes a waiter if there is one. The counter's initial value is the number of "permits" you have. A binary semaphore is the same thing limited to 0 or 1. Unlike a mutex, a semaphore has no owner: any thread may signal it. Typical use: limit concurrent access to a scarce resource, such as at most 3 simultaneous database connections.
How it behaves
With initial count 3:
| Step | Count | Effect |
|---|---|---|
| Thread 1 waits | 3 to 2 | proceeds |
| Thread 2 waits | 2 to 1 | proceeds |
| Thread 3 waits | 1 to 0 | proceeds |
| Thread 4 waits | 0 | blocks |
| Thread 1 signals | 0 to 1, then thread 4 takes it | thread 4 proceeds |
The C++ reference says exactly this about std::counting_semaphore: acquire decrements the counter or blocks until it can, release increments it and unblocks acquirers, and unlike std::mutex it is not tied to threads of execution, so acquiring and releasing can happen on different threads.
Counting vs binary
| Counting | Binary | |
|---|---|---|
| Range | 0 to N | 0 or 1 |
| Typical use | Pool of N identical resources; counting events so none is lost | One-shot signal between two threads; mutual exclusion in the simplest case |
| Signalling twice before anyone waits | Count becomes 2: both events remembered | The second signal is not remembered as a second event, so one event is remembered. What the API does with the extra signal depends on the platform: std::binary_semaphore makes a release beyond its maximum a precondition violation (undefined behaviour, so never rely on it), FreeRTOS's give call returns a failure code, and POSIX has no binary semaphore type, so sem_post would simply count up to 2 |
std::binary_semaphore is just std::counting_semaphore<1>.
A realistic use: limiting concurrent access
#include <atomic>
#include <chrono>
#include <cstdio>
#include <semaphore>
#include <thread>
#include <vector>
using namespace std::chrono_literals;
std::counting_semaphore<8> slots(3); // at most 3 threads may use the scarce resource
std::atomic<int> inside{0}, max_inside{0};
void use_resource(bool fail_early_bug) {
slots.acquire(); // wait (P): count 3 -> 2 -> 1 -> 0, then block
int now = ++inside;
int seen = max_inside.load();
while (now > seen && !max_inside.compare_exchange_weak(seen, now)) {}
std::this_thread::sleep_for(20ms);
--inside;
if (fail_early_bug) return; // BUG: returns without giving the permit back
slots.release(); // signal (V): count + 1, wakes one waiter
}
int main() {
std::vector<std::thread> ts;
for (int i = 0; i < 10; i++) ts.emplace_back(use_resource, false);
for (auto& t : ts) t.join();
std::printf("10 threads, 3 permits: max concurrent = %d\n", max_inside.load());
// Misuse: three callers hit an error path that skips release(). The pool is now empty.
std::vector<std::thread> bad;
for (int i = 0; i < 3; i++) bad.emplace_back(use_resource, true);
for (auto& t : bad) t.join();
bool got = slots.try_acquire_for(200ms);
std::printf("after 3 leaked permits, a new acquire succeeds = %s\n", got ? "yes" : "no (timed out)");
// Binary semaphore: a one-shot signal between threads, released by a different thread.
std::binary_semaphore ready(0);
std::thread producer([&] { std::this_thread::sleep_for(10ms); ready.release(); });
ready.acquire();
producer.join();
std::printf("binary semaphore signalled by another thread: ok\n");
}
Compiled with g++ -std=c++20 -O2 -Wall -Wextra -pthread sem.cpp (and again with -fsanitize=thread, same output, no reports) in a GCC 14 container:
10 threads, 3 permits: max concurrent = 3
after 3 leaked permits, a new acquire succeeds = no (timed out)
binary semaphore signalled by another thread: ok
Ten threads compete for three permits and never more than three are inside at once. The second line shows the classic misuse: three callers hit an early-return path that skips release, and the permits are gone for good, so every later caller waits until it times out.
ISR-to-task signalling vs resource-pool counting
An ISR (interrupt service routine: code the hardware runs when an interrupt fires) must never block, so it cannot take a lock. It can signal a semaphore, and a task blocked on that semaphore wakes to do the longer processing. The POSIX specification makes sem_post async-signal-safe, meaning it may be called from a signal handler (the user-space cousin of an interrupt handler), which is what makes signalling from such a context legitimate; on an RTOS (real-time operating system) use the ISR-specific give call the RTOS documents.
- Binary semaphore (initial 0): "an event happened". If two interrupts fire before the task runs, the second signal is not counted, so the task handles one event. That is fine when the task re-reads the device state anyway.
- Counting semaphore (initial 0): each interrupt adds one, so the task runs once per event and none is lost, provided the count cannot overflow its maximum.
- Resource pool (initial N): tasks wait to claim one of N buffers and signal to return it. Here the count means "free items".
Fairness and misuse
- Who wakes up. POSIX says
sem_post, under a real-time scheduling policy (the POSIX real-time scheduling policies SCHED_FIFO, where a thread runs until it blocks or yields, and SCHED_RR, which adds a fixed time slice among threads of equal priority; SCHED_SPORADIC, where supported, follows the SCHED_FIFO rule), unblocks the highest-priority waiter, and the one that has waited longest among equals; under other policies POSIX leaves the choice to the implementation. The C++ reference documents no wake-up order at all, so do not rely on first-come-first-served. A waiter can in principle be passed over repeatedly (starvation); if you need strict arrival order, queue tickets explicitly (each waiter takes a number and is served in number order). - Leaking a permit (shown above): release in a scope guard (a small object whose destructor gives the permit back on every exit path), not by hand on each path.
- Signalling without waiting inflates the count above the real number of resources, so more users get in than exist.
- A semaphore as a mutex. With no owner, the system does not know whom to boost to avoid priority inversion (a high-priority thread stuck behind a low-priority lock holder that a medium-priority thread keeps preempting), and any thread can release it by mistake.
- Waiting while holding another resource can deadlock: a thread holding permit A waits for permit B while another does the reverse.
You are choosing a scheduling policy for a flight-control computer that must be certified. Would you pick fixed-priority or earliest-deadline-first, and how would you defend the choice to a certification authority?
Sample Answer
Recommendation: fixed-priority preemptive scheduling, rate-monotonic order, proved by response-time analysis
For a certified flight computer I would choose fixed-priority scheduling (certified here means approved by an aviation authority, typically against DO-178C, the standard for airborne software, so the timing argument is part of the evidence package). Each task gets a constant priority, shorter period meaning higher priority (rate-monotonic assignment; deadline-monotonic is the same idea ordered by shortest deadline, used when deadlines are shorter than periods), and the scheduler always runs the highest-priority ready task. Earliest-deadline-first (EDF) instead gives each job a priority from its absolute deadline (fixed when the job is released), so the highest priority goes to whichever ready job has the earliest deadline and priorities change from job to job. EDF is better on paper: it can schedule any single-processor set of independent, preemptible tasks with deadlines equal to periods and total utilisation up to 100 percent, whereas rate-monotonic scheduling is only guaranteed up to n(2^(1/n) - 1), which is about 69.3 percent as the number of tasks n grows. I would still not choose it here, because a certification authority does not ask which policy packs the CPU best. It asks whether you can show, with evidence, that every deadline is met, and what happens when something goes wrong.
The argument to the certifier
- The analysis is exact and reproducible. Response-time analysis computes each task's worst-case response R by iterating R = C + sum over higher-priority tasks of ceil(R / Tj) x Cj until it stops changing (C is the worst-case execution time, Tj the period of a higher-priority task). A reviewer can redo it by hand from the table of C and T values, and it is a necessary and sufficient test for fixed priorities when deadlines equal periods (necessary and sufficient: a set passes if and only if it is schedulable, with no pessimism). For the
healthtask (C = 17, T = 80) with the four faster tasks above it: R = 17; then 17 + ceil(17/5) x 1 + ceil(17/10) x 2 + ceil(17/20) x 4 + ceil(17/40) x 6 = 17 + 4 + 4 + 4 + 6 = 35; then 17 + 7 + 8 + 8 + 6 = 46; and the iteration continues 61, 72, 76, 77, 77. It stops at 77, below 80, so the task passes. This is the shape of the response-time table a timing report hands to the certifier. This kind of analysis supports the timing-margin evidence DO-178C asks for (DO-178C objective 6.3.4.f, on source code being accurate and consistent, includes worst-case execution timing among the things reviews and analyses address; whichever method you use, measured times need an argument that the worst case was actually exercised, and the compiler, linker and hardware effects on those times need to be accounted for). - Utilisation is not the question. The 69.3 percent bound is sufficient, not necessary (a set under it is safe, but a set over it may still be schedulable). The code below uses harmonic periods (each period divides the next), which flight loops often have. The set has total utilisation 0.9625, far above the 0.7435 bound for five tasks, and still passes the exact test.
- Overload stays local. Priorities are a contract: apart from the blocking term in item 5, a high-priority task is never delayed by a lower one, so a late or overrunning task damages only itself and the tasks below it. Under EDF the set is at 96 percent utilisation with no slack, so the extra time an overrunning job takes comes straight out of every other job's margin; and once the overrunning job's own deadline has passed it is still ahead of every job released later (its deadline is earlier than theirs), so well-behaved tasks start missing too. The simulation below shows both.
- It matches existing practice and structure. ARINC 653, the avionics partitioning standard, uses a fixed cyclic schedule of time windows for partitions (isolated applications, each given guaranteed slices of processor time and its own protected memory) and, inside each window, preemptive fixed-priority scheduling of processes. Using fixed priorities inside a partition follows that structure.
- Locking has a proven bound. With the priority ceiling protocol, a task is blocked at most once, by at most one critical section of a lower-priority task, and deadlock is impossible (Sha, Rajkumar and Lehoczky, 1990). That blocking time is added to the analysis as a term B.
The exact test and the overload comparison, as code
The task set: five tasks, periods 5, 10, 20, 40 and 80 ms with worst-case execution times 1, 2, 4, 6 and 17 ms. The script computes the utilisation, both bounds, the exact response times, and a discrete-time simulation (1 tick is 0.1 ms) in which the third job of one task overruns by the stated amount. The simulator builds one job per release (release time, absolute deadline, remaining work, priority). On each tick it runs the ready job with the best key (smallest priority number for fixed priority, earliest deadline for EDF) and removes one tick of its work. A deadline miss is recorded when a job is still unfinished at its deadline tick and again if it later completes late, so the set() at the end removes the duplicate and each miss counts once. The overrun is injected by adding extra_ms to the job with k == 2, the third release of the named task. Run in a python:3.12-slim container with python fp_vs_edf.py:
from math import ceil
# (name, C, T) in ms; deadline = period; harmonic periods, as flight loops often are
tasks = [("attitude", 1.0, 5), ("actuators", 2.0, 10), ("nav", 4.0, 20),
("guidance", 6.0, 40), ("health", 17.0, 80)]
n = len(tasks)
U = sum(C / T for _, C, T in tasks)
print(f"U = {U:.4f} RMS bound n(2^(1/n)-1) = {n * (2 ** (1 / n) - 1):.4f} EDF bound = 1")
# exact fixed-priority test: response-time analysis, rate-monotonic order
for i, (name, C, T) in enumerate(tasks):
R = C
while True:
R2 = C + sum(ceil(R / Tj) * Cj for _, Cj, Tj in tasks[:i])
if R2 == R or R2 > T:
break
R = R2
print(f" {name:9s} R = {R2:5.1f} ms deadline {T:3d} ms {'ok' if R2 <= T else 'MISS'}")
# discrete-time preemptive simulator, 1 tick = 0.1 ms, one overrun injected
TICK = 10
def simulate(policy, overrun_task, extra_ms, horizon_ms=160):
jobs, misses = [], []
for name, C, T in tasks:
for k in range(int(horizon_ms // T)):
c = round(C * TICK) + (round(extra_ms * TICK) if (name == overrun_task and k == 2) else 0)
jobs.append({"task": name, "rel": k * T * TICK, "dl": (k + 1) * T * TICK, "left": c,
"prio": [t[0] for t in tasks].index(name)})
for now in range(horizon_ms * TICK):
ready = [j for j in jobs if j["rel"] <= now and j["left"] > 0]
if ready:
key = (lambda j: j["dl"]) if policy == "EDF" else (lambda j: j["prio"])
j = min(ready, key=lambda j: (key(j), j["rel"]))
j["left"] -= 1
if j["left"] == 0 and now + 1 > j["dl"]:
misses.append((j["task"], j["dl"] // TICK))
for j in ready:
if now + 1 == j["dl"] and j["left"] > 0:
misses.append((j["task"], j["dl"] // TICK))
return sorted(set(misses), key=lambda m: (m[1], m[0]))
from collections import Counter
for who, extra in (("guidance", 14), ("nav", 12)):
for policy in ("FP", "EDF"):
m = Counter(task for task, _ in simulate(policy, who, extra))
print(f"{who} job overruns by {extra} ms, {policy}: missed deadlines per task = {dict(m)}")
U = 0.9625 RMS bound n(2^(1/n)-1) = 0.7435 EDF bound = 1
attitude R = 1.0 ms deadline 5 ms ok
actuators R = 3.0 ms deadline 10 ms ok
nav R = 8.0 ms deadline 20 ms ok
guidance R = 18.0 ms deadline 40 ms ok
health R = 77.0 ms deadline 80 ms ok
guidance job overruns by 14 ms, FP: missed deadlines per task = {'guidance': 1, 'health': 1}
guidance job overruns by 14 ms, EDF: missed deadlines per task = {'actuators': 2, 'attitude': 2, 'nav': 2, 'guidance': 1}
nav job overruns by 12 ms, FP: missed deadlines per task = {'nav': 1, 'guidance': 1, 'health': 2}
nav job overruns by 12 ms, EDF: missed deadlines per task = {'actuators': 4, 'attitude': 6, 'nav': 3, 'guidance': 1}
Reading it: every task passes the exact test, and the slowest (health) has only 3 ms of margin (77 against 80), so the set is schedulable but tight. When guidance overruns by 14 ms, fixed priority loses one guidance deadline and one health deadline (the task below it) while attitude, actuators and nav are untouched. EDF loses deadlines in four tasks (attitude, actuators, nav and guidance), including attitude and actuators, which did nothing wrong. In this simulation a late job keeps running to completion, and each missed deadline is counted once.
What the answer leaves out unless you add it
The test above assumes no blocking, no release jitter and zero context-switch cost. A real submission adds B for shared resources (ceiling protocol), release jitter from interrupts, the cost of each switch and timer interrupt, and worst-case execution times with margin from a justified method. It also adds budget enforcement (a run-time guard): a timer that stops a task at its allotted worst case, so an overrun becomes a detected fault rather than a silent delay.
What would change the choice
Choose EDF if the task set cannot be made schedulable under rate-monotonic or deadline-monotonic priorities even with the exact test (for instance non-harmonic periods near full utilisation and no way to shed load), and pair it with budget enforcement so that the overload behaviour is bounded. Also consider a time-triggered (fully static cyclic) schedule if the certifier wants every execution order fixed in advance; it gives the simplest evidence at the price of flexibility.
Implement functions to pack an array of signed 12-bit ADC samples into a byte stream and to unpack them. Sign-extension must be correct on unpack. Signatures:
void pack12(const int16_t *samples, size_t n, uint8_t *out);
void unpack12(const uint8_t *in, size_t n, int16_t *samples);
Optimize for speed on a 32-bit MCU, handling odd sample counts, and minimize temporary memory.
Sample Answer
Approach (brief)
Pack signed 12-bit samples (two's complement) tightly: 12 bits per sample => 3 bytes for every 2 samples. Work in 32-bit words on a 32-bit MCU for speed: process pairs with bit ops, handle odd final sample. Ensure unpack sign-extends 12-bit values to int16_t.
Code (optimized, portable C)
#include <stdint.h>
#include <stddef.h>
void pack12(const int16_t *samples, size_t n, uint8_t *out) {
size_t i = 0, o = 0;
while (i + 1 < n) {
uint32_t a = (uint16_t)samples[i] & 0x0FFF; // lower 12 bits
uint32_t b = (uint16_t)samples[i+1] & 0x0FFF;
uint32_t w = (a) | (b << 12); // 24 bits: sample0 | sample1<<12
out[o++] = (uint8_t)(w & 0xFF);
out[o++] = (uint8_t)((w >> 8) & 0xFF);
out[o++] = (uint8_t)((w >> 16) & 0xFF);
i += 2;
}
if (i < n) { // odd sample left
uint32_t a = (uint16_t)samples[i] & 0x0FFF;
out[o++] = (uint8_t)(a & 0xFF);
out[o++] = (uint8_t)((a >> 8) & 0xFF);
}
}
void unpack12(const uint8_t *in, size_t n, int16_t *samples) {
size_t i = 0, o = 0;
while (o + 1 < n) {
uint32_t w = (uint32_t)in[i] | ((uint32_t)in[i+1] << 8) | ((uint32_t)in[i+2] << 16);
uint16_t a = w & 0x0FFF;
uint16_t b = (w >> 12) & 0x0FFF;
// sign-extend 12-bit to 16-bit
samples[o] = (int16_t)((a ^ 0x0800) - 0x0800);
samples[o+1] = (int16_t)((b ^ 0x0800) - 0x0800);
i += 3; o += 2;
}
if (o < n) { // last single sample
uint16_t a = (uint16_t)in[i] | ((uint16_t)in[i+1] << 8);
a &= 0x0FFF;
samples[o] = (int16_t)((a ^ 0x0800) - 0x0800);
}
}
Key points
- Use masking 0x0FFF to keep 12 bits.
- Sign-extension trick: (x ^ 0x0800) - 0x0800 converts 12-bit two's-complement to signed 16-bit efficiently without branching.
- Processes pairs to reduce loop overhead; minimal temporaries; works for odd count.
Complexity
- Time O(n), space O(1) extra. Fast on 32-bit MCU due to 32-bit ops and sequential memory access.
Edge cases
- Ensure out buffer size = ceil(n*12/8). Handle unaligned buffers, endianness assumed little-endian byte order for stream; document if big-endian required.
Alternative
- Use compiler intrinsics or unaligned 32-bit writes for slightly faster throughput if MCU supports.
A deployed Cortex-M device shows sporadic HardFault exceptions. Describe how you would use a JTAG/SWD debugger to capture the fault context: locate the stacked registers on the exception stack frame, inspect fault status registers (CFSR/HFSR/MMFAR/BFAR), map the faulting PC back to source code when optimizations are enabled, and outline steps to reproduce the fault without masking timing-dependent causes.
Sample Answer
Situation & immediate goal
I want to capture the CPU state at the moment of HardFault so I can identify the instruction and reason (memory fault, bus fault, usage fault) and map it back to source even when optimizations are enabled.
Steps to capture the fault context with JTAG/SWD
- Halt on fault: configure the debugger to stop on exception entry (e.g., "Stop on HardFault" or enable vector catch / fault halt). If it runs past, set a breakpoint at HardFault_Handler.
- Locate stacked registers: read EXC_RETURN in LR and the current SP (MSP on exception entry for handlers). The 8 pushed registers are at SP on entry:
- r0, r1, r2, r3, r12, lr, pc, xPSR (lowest address = SP)
- Use the debugger’s memory/read registers command to dump 32-bit words at SP and label them.
- Inspect fault status registers (SCB):
- Read SCB->CFSR (Usage/Bus/MemoryFault subfields), SCB->HFSR, SCB->MMFAR, SCB->BFAR
- Interpret bits: e.g., CFSR MemManage Fault Status bits indicate instruction vs data access, MMFAR holds faulting address if valid.
- Map PC to source with optimizations:
- Extract faulting PC from stacked PC. Mask Thumb bit (PC & ~1).
- Use objdump/addr2line with the linked ELF and the exact address: addr2line -e firmware.elf 0xADDRESS
- If inlined/optimized, also use objdump -d to see surrounding function and compiler map file; enable debug-symbol-rich builds (-g) and keep link map to correlate section offsets.
- Reproduce without masking timing:
- Avoid heavy logging or debug I/O that changes timing.
- Use non-invasive trace: enable hardware break-on-fault or SWV/ITM for lightweight logs, or ETM/trace if available.
- Add conditional breakpoints/watchpoints on the suspected fault address or memory region (use hardware watchpoints).
- Create stress test that exercises same code paths with same interrupts and timing (use timers/NMI to mimic).
- If race suspected, use reduced optimization or instrumented builds with minimal impact (cycle-accurate toggles, GPIO toggles) to capture timing without altering it too much.
- Repeat with increased monitoring (watchpoints) until captured.
Notes / best practices
- Remember to clear SCB fault status bits after reading to avoid stale info.
- When reading stacked PC, remember Thumb LSB and pipeline offset; disassemble from (pc - 4) to see the likely faulting instruction.
- Keep a copy of the exact firmware ELF and map file used on the device for reliable addr2line mapping.
What are the main techniques to prevent and detect buffer overflows in C beyond swapping in safer-looking library functions? Cover both coding practice and build-time or runtime defenses.
Sample Answer
Direct answer
Defense for buffer overflows has four layers: write code where the length always travels with the pointer and is checked before every copy; get the compiler to warn and to insert checks (-Wall -Wextra, -D_FORTIFY_SOURCE, -fstack-protector-strong); catch the bugs you did not see in review with AddressSanitizer plus fuzzing before release; and keep OS-level mitigations on as a last line that turns many exploits into crashes. (A fuzzer is a program that feeds a target huge numbers of generated or mutated inputs looking for crashes; fuzzing is running one.) No single layer is enough, and the safe-looking library function swap is the weakest of them.
Why "safer functions" alone fall short
strncpy does not terminate when the source fills the count (https://en.cppreference.com/w/c/string/byte/strncpy). Any bounded function still needs the correct bound, and the common bug is passing the wrong number (the pointer's size, the source length, or the buffer size without the terminator). The fix is structural: make the size impossible to forget.
Layer 1: coding practice
- Pass pointer and length together (a small
struct buf { uint8_t *p; size_t len; }), and checklenbefore copying, not after. - Use
size_tfor sizes and check arithmetic before it happens:if (n > cap - used) return ERR;rather thanused + n > cap, which can wrap. - Prefer functions that report truncation (
snprintfreturns the length it wanted), and check the result. - Allocate with the size computed from the same constant used to bound the copy, so the two cannot drift apart.
- Parse by validating the length field against what you actually received, then copy.
Layer 2: build-time defenses
-Wall -Wextrafor static diagnostics. They are not complete: in the example below, a plainstrcpyinto an 8-byte array fromargv[1]compiled with no warning.-fstack-protector-strongputs a canary (a guard value) between local arrays and the saved return address and checks it on function exit. Picture the stack frame asbuf[8], then the canary, then the saved return address (the place the function jumps back to when it ends): an overflow that runs pastbuftoward the return address must overwrite the canary first, and the check at exit sees the changed value and aborts instead of returning to an attacker-chosen address. GCC defines it as-fstack-protectorplus functions that have local arrays or reference local frame addresses (GCC manual: https://gcc.gnu.org/onlinedocs/gcc/Instrumentation-Options.html).-D_FORTIFY_SOURCE: glibc and the compiler add lightweight checks to string and memory functions when the destination size is known. The recommendation below uses level 2. Level 1 needs-O1or higher, level 2 adds more checks (some conforming programs can fail), and level 3 adds checks for buffers whose size is only known at run time (for example frommalloc(n)) and needs GCC 12 or later with glibc 2.33 or later (https://man7.org/linux/man-pages/man7/feature_test_macros.7.html).
Layer 3: test-time detection
-fsanitize=address instruments memory accesses to detect out-of-bounds and use-after-free bugs. Combine with -fsanitize=undefined, then run unit tests and a fuzzer so the sanitizer sees hostile inputs. This is the layer that finds the bug rather than merely containing it.
Worked example: one overflow, four builds
#include <stdio.h>
#include <string.h>
int main(int argc, char **argv)
{
char buf[8];
if (argc < 2) return 1;
strcpy(buf, argv[1]); /* no length check: overflows for inputs of 8 or more chars */
printf("copied: %s\n", buf);
return 0;
}
Run with a 32-character argument of A in a Linux arm64 container (GCC 14.4.0):
gcc -O0 -fno-stack-protector -> copied: AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA (exit 0, silent corruption)
gcc -O0 -fstack-protector-strong -> *** stack smashing detected ***: terminated (Aborted)
gcc -O2 -D_FORTIFY_SOURCE=2 -fno-stack-protector -> *** buffer overflow detected ***: terminated (Aborted)
gcc -O1 -g -fsanitize=address -> ERROR: AddressSanitizer: stack-buffer-overflow ... WRITE of size 33
The unprotected build overflowed an 8-byte array with 33 bytes (the 32 A characters plus the terminating zero byte that strcpy also writes, which is where WRITE of size 33 in the sanitizer line comes from), printed success and exited 0: this is why overflows are dangerous, since nothing signalled the corruption. The other three builds stopped it, at different costs and with different information (the sanitizer names the access, the production flags just abort).
Layer 4: OS and platform mitigations
Two mitigations are normally on by default. A non-executable stack marks stack memory as data, so an attacker who injects machine code into an overflowed buffer cannot simply run it. Address-space randomization (ASLR) loads the program, libraries, stack and heap at different addresses on each run, so an attacker cannot hard-code the address of the code they want to jump to. They make exploiting a surviving overflow harder; they are not a fix and can be bypassed (for example, an information leak reveals the randomized addresses), so treat them as damage limiters.
Recommendation
For a new C codebase handling untrusted data: length-carrying buffer types and checked arithmetic in the code, -Wall -Wextra -fstack-protector-strong -D_FORTIFY_SOURCE=2 in every release build, and a CI job that runs tests and a fuzzer under -fsanitize=address,undefined. If it is a hot, high-exposure parser, consider a memory-safe language for that component. What would change this: on a tiny embedded target without room for canaries, lean harder on coding rules, static analysis and host-side sanitizer testing of the same source.
Pitfalls
- Hardening flags are not tests: a program that aborts on overflow has still got the bug.
- Sanitizers slow programs and use more memory, so they belong in test builds, not production.
- Heap overflows, off-by-one writes inside a struct, and overruns that stay within one allocation may not trip a stack canary at all.
Explain SPI protocol fundamentals: master/slave roles, signals (SCK, MOSI, MISO, CS/SS), CPOL/CPHA meaning (modes 0–3), full-duplex vs half-duplex transfers, and considerations for bus topology when multiple slaves are present. Describe how to choose SPI mode for a device and how chip-select setup/hold timing affects reliable transfers.
Sample Answer
SPI fundamentals (brief)
SPI is a synchronous serial bus with a single master and one or more slaves. Master generates SCK (clock) and selects a slave via CS/SS (chip/slave select). Data is exchanged on MOSI (master out, slave in) and MISO (master in, slave out).
Signals & roles
- SCK: clock from master.
- MOSI: master → slave data.
- MISO: slave → master data.
- CS/SS: active-low select per slave; when asserted the selected slave drives MISO.
CPOL/CPHA and modes
- CPOL = clock idle polarity (0 = low, 1 = high).
- CPHA = clock phase: whether sampling occurs on first or second edge after CS.
Modes: - Mode 0: CPOL=0, CPHA=0 (sample rising edge).
- Mode 1: CPOL=0, CPHA=1.
- Mode 2: CPOL=1, CPHA=0.
- Mode 3: CPOL=1, CPHA=1.
Choose mode to match device datasheet; mismatch causes bit shifts or jitter.
Full- vs half-duplex
- Full-duplex: MOSI and MISO simultaneous (standard SPI).
- Half-duplex: single data line toggled direction or tri-stated (used to reduce pins).
Bus topology & multiple slaves
- Use one CS per slave; keep CS asserted only for active transfer. Avoid shared MISO driving—use pull-ups and tri-state or buffers. Consider series resistors for signal integrity and matched trace lengths for high speed.
CS timing
- Observe device setup/hold in datasheet: ensure CS asserted before first clock by tCSS, maintain after last clock for tCSH, and observe required inter-transfer delay. Violating these causes corrupted frames or device misinterpretation. In firmware, assert CS, wait tCSS, clock N bits, wait tCSH, deassert CS.
On an 8-bit AVR microcontroller, show how to safely read and write a 16-bit shared variable that can be modified in an ISR. Provide concise C examples using cli()/sei() (or equivalent atomic block macros) and discuss pros and cons of disabling interrupts briefly versus other synchronization techniques.
Sample Answer
Approach (brief)
On 8-bit AVR a 16-bit variable is not atomic; an ISR can interrupt a multi-byte access. Safest simple method is to disable interrupts around the read/write or use AVR-provided atomic macros.
Example — disable/restore interrupts (cli()/sei())
#include <avr/io.h>
#include <avr/interrupt.h>
volatile uint16_t shared = 0;
uint16_t read_shared(void) {
uint16_t val;
uint8_t sreg = SREG; // save global interrupt flag
cli(); // disable interrupts
val = shared; // atomic read of 16-bit
SREG = sreg; // restore interrupts (restores I-bit)
return val;
}
void write_shared(uint16_t v) {
uint8_t sreg = SREG;
cli();
shared = v;
SREG = sreg;
}
Example — using <util/atomic.h>
#include <util/atomic.h>
uint16_t read_shared2(void) {
uint16_t val;
ATOMIC_BLOCK(ATOMIC_RESTORESTATE) {
val = shared;
}
return val;
}
Pros/Cons
- Disabling interrupts: simple, minimal code, deterministic. Con: increases interrupt latency; avoid long critical sections.
- Atomic macros: safer (handles SREG), clearer intent.
- Alternatives: use double-buffering + version counters or message queues/flags to avoid long disable periods; use minimal critical sections or task-level synchronization in RTOS. Those reduce latency but add complexity and possibly more memory.
Write a concise skeleton of a character device driver (pseudo-C) for an embedded RTOS that supports open, close, read, write and handles a hardware interrupt for data-ready. Show how you would defer processing from the ISR to a worker thread and protect shared buffers from concurrent access.
Sample Answer
Approach
- Use a small ring buffer protected by a mutex for reader/writer.
- ISR signals a worker via a semaphore/queue to do non-ISR work.
- Expose open/close/read/write with proper locking and sleep/wakeup when empty/full.
Pseudo-C skeleton
// pseudo-C for RTOS (posix-like primitives)
#include <rtos.h>
#define BUF_SIZE 256
static uint8_t buf[BUF_SIZE];
static size_t head=0, tail=0;
static mutex_t buf_lock;
static sem_t data_sem; // signalled by ISR; worker waits
static cond_t read_wait; // readers wait when buffer empty
static bool open_count=false;
static thread_t worker_thread;
static inline size_t buf_used() { return (head - tail) % BUF_SIZE; }
static inline size_t buf_free() { return BUF_SIZE - buf_used() - 1; }
void isr_data_ready(void)
{
// minimal ISR: acknowledge hardware
hw_ack_interrupt();
// notify worker (use ISR-safe API)
sem_give_from_isr(&data_sem);
}
static void worker(void *arg)
{
while (1) {
sem_take(&data_sem, WAIT_FOREVER);
// do deferred processing
mutex_lock(&buf_lock);
while (hw_has_data() && buf_free()) {
buf[head] = hw_read_byte();
head = (head + 1) % BUF_SIZE;
}
mutex_unlock(&buf_lock);
cond_broadcast(&read_wait); // wake readers
}
}
int dev_open(void)
{
if (!open_count) {
open_count = true;
mutex_init(&buf_lock);
sem_init(&data_sem, 0);
cond_init(&read_wait);
worker_thread = thread_create(worker, NULL);
}
return 0;
}
int dev_close(void) { open_count = false; return 0; }
ssize_t dev_read(uint8_t *dst, size_t len)
{
size_t copied=0;
mutex_lock(&buf_lock);
while (buf_used()==0) {
mutex_unlock(&buf_lock);
cond_wait(&read_wait, &buf_lock); // atomically unlock and wait
mutex_lock(&buf_lock);
}
while (copied < len && buf_used()) {
dst[copied++] = buf[tail];
tail = (tail + 1) % BUF_SIZE;
}
mutex_unlock(&buf_lock);
return copied;
}
ssize_t dev_write(const uint8_t *src, size_t len)
{
size_t written=0;
mutex_lock(&buf_lock);
while (written < len && buf_free()) {
buf[head] = src[written++];
head = (head + 1) % BUF_SIZE;
}
mutex_unlock(&buf_lock);
// optionally notify hardware to send data
return written;
}
Key points
- ISR does minimal work and uses ISR-safe sem give.
- Worker performs slow operations and moves data to buffer.
- Mutex + condition variable protect and coordinate access.
- Ring buffer avoids copies and supports concurrent readers/writers safely.
You have a microcontroller with no OS and very limited debug equipment, and you need to find out where the CPU time is actually going. How would you go about measuring execution time and finding hotspots with what you've got?
Sample Answer
Framing
"Very limited debug equipment" still usually means something is available: a spare GPIO (general-purpose input/output) pin, maybe a scope or logic analyzer, and the MCU's (microcontroller unit's) own hardware timer. The approach depends on which of those you actually have, and on whether you already suspect where the hotspot is or need to find it.
If you suspect a specific region: toggle a GPIO pin around it
void suspect_function(void) {
GPIO_SET(DEBUG_PIN);
do_the_work();
GPIO_CLEAR(DEBUG_PIN);
}
Probe DEBUG_PIN with a scope or logic analyzer, the high pulse width read directly off the trace is the execution time of do_the_work(). No debugger stop needed, and it works while the rest of the system runs in real time, useful for catching timing that only shows up under real workload rather than a single-stepped debugger.
If the MCU has a free-running hardware counter
Read the counter (e.g. an ARM Cortex-M DWT cycle counter (data watchpoint and trace, a debug hardware unit that can count CPU cycles), or any general-purpose timer configured to just count up) before and after the region; the difference, converted using the known clock frequency, gives elapsed time. This needs no external equipment at all, just two counter reads and a way to report the result, UART or even an LED.
If you don't yet know where the hotspot is
Set up a periodic timer interrupt, say every 1 ms, and inside it, record the current program counter (or a return address off the stack). Over a long run, build a histogram of which addresses show up most, that's a poor-man's sampling profiler using only a timer peripheral, with the linker's map file used afterward to translate addresses back into function names.
Trade-off between the two
GPIO toggling and hardware-counter timing need almost zero overhead but require picking the region ahead of time. Timer-interrupt sampling finds hotspots you didn't already suspect, at the cost of being statistical (needs enough samples to be reliable) and using one timer plus interrupt overhead.
You are assigned a product with 32KB of SRAM. Subsystems: network stack (10KB nominal), sensor buffer (6KB worst case), logging ring buffer (4KB), OS and stacks (6KB), and misc variables. Decide which subsystems should use static allocation, dynamic allocation, or a hybrid. Explain your choices and describe fallback behaviors if memory pressure occurs at runtime.
Sample Answer
Recommendation: everything static except the network buffers, which come from a fixed-block pool (the hybrid), and no general-purpose heap at all. With 32 KB there is no room to absorb a bad surprise, so the rule is that every byte has an owner at link time and the only runtime decision is "which free block", never "is there memory".
Start with the arithmetic. The demands total 10 + 6 + 4 + 6 = 26 KB (26,624 bytes) before misc variables, leaving 6,144 bytes for misc and margin. The table below turns that into a budget that adds up exactly.
| Subsystem | Strategy | Bytes |
|---|---|---|
| OS and stacks | static (RTOS objects, meaning the kernel's task, queue and semaphore records, and task stacks as arrays) | 6,144 |
| Sensor buffer | static, sized for the 6 KB worst case | 6,144 |
| Logging ring | static ring buffer (a fixed array used as a circular queue: when the write position reaches the end it wraps to the start) | 4,096 |
| Network stack | hybrid: static pool of fixed-size blocks | 10,240 (80 blocks of 128) |
| Misc variables | static (.data and .bss) | 4,096 |
| Reserve | unassigned margin | 2,048 |
| Total | 32,768 |
The same numbers as compiled code, with _Static_assert (a compile-time check: the build stops with the message if the condition is false) making the build fail if the sum drifts from the part's RAM:
#include <stdio.h>
#define KB(x) ((x) * 1024u)
#define SRAM_TOTAL KB(32)
#define OS_AND_STACKS KB(6) /* static: RTOS objects and task stacks */
#define SENSOR_BUF KB(6) /* static: sized for the worst case */
#define LOG_RING KB(4) /* static: overwrite-oldest ring */
#define NET_BLOCK_SIZE 128u
#define NET_BLOCK_COUNT 80u
#define NET_POOL (NET_BLOCK_SIZE * NET_BLOCK_COUNT) /* hybrid: fixed-block pool */
#define MISC_DATA_BSS KB(4) /* budget for everything else in .data and .bss */
#define RESERVE KB(2) /* unassigned margin */
#define ASSIGNED (OS_AND_STACKS + SENSOR_BUF + LOG_RING + NET_POOL + MISC_DATA_BSS + RESERVE)
_Static_assert(ASSIGNED == SRAM_TOTAL, "budget must add up to the 32 KB part");
int main(void)
{
printf("OS and stacks %5u\n", OS_AND_STACKS);
printf("sensor buffer %5u\n", SENSOR_BUF);
printf("log ring %5u\n", LOG_RING);
printf("network pool %5u (%u blocks x %u)\n", NET_POOL, NET_BLOCK_COUNT, NET_BLOCK_SIZE);
printf("misc .data/.bss%5u\n", MISC_DATA_BSS);
printf("reserve %5u\n", RESERVE);
printf("total %5u of %u\n", ASSIGNED, SRAM_TOTAL);
printf("stated demands (10+6+4+6 KB) = %u bytes, leaving %u for misc and margin\n",
KB(10) + KB(6) + KB(4) + KB(6), SRAM_TOTAL - (KB(10) + KB(6) + KB(4) + KB(6)));
return 0;
}
Compiled with gcc -O2 -Wall -Wextra -fsanitize=address,undefined budget.c in a gcc:14 container (GCC 14.4.0, aarch64), it prints:
OS and stacks 6144
sensor buffer 6144
log ring 4096
network pool 10240 (80 blocks x 128)
misc .data/.bss 4096
reserve 2048
total 32768 of 32768
stated demands (10+6+4+6 KB) = 26624 bytes, leaving 6144 for misc and margin
If someone raises RESERVE to 3 KB without taking the space from anywhere, the same file no longer compiles and the build reports error: static assertion failed: "budget must add up to the 32 KB part", so a budget overrun is found at build time, not on the device.
Why each choice.
- Sensor buffer: static at the worst case. Its worst case is known, so there is nothing to negotiate at runtime. Allocating it dynamically would only add a way to fail at the moment the data arrives.
- Logging ring: static. A ring buffer has a fixed size by definition, and logging must never be able to take memory from something more important.
- OS and stacks: static. Task stacks and kernel objects live for the whole run, so there is no benefit from allocating them dynamically. Most RTOSes let you supply these as arrays (FreeRTOS calls this static allocation).
- Network stack: hybrid. Packet buffers are the one place where counts and lifetimes vary: a packet arrives, waits, is processed and is freed in an unpredictable order. A fixed-block pool lets the stack take and return buffers at run time while keeping the total at exactly 10 KB and making fragmentation impossible (any free block fits any buffer). If a third-party stack insists on
malloc, give it a private bounded heap of the same 10 KB, not the whole RAM. - Misc and reserve. The question does not size "misc variables", so the budget caps them at 4 KB and keeps 2 KB unassigned. The real figure comes from the linker map (
.dataplus.bss), and the build should fail if it exceeds the cap. The reserve exists for the numbers that are estimates, chiefly stack depths and the interrupt stack.
Fallback behaviour under memory pressure. Decide each one in advance, per subsystem.
- Network pool exhausted: drop the newest inbound packet (tail drop: refuse the packet that just arrived instead of discarding ones already queued) and count it. Keep a few blocks (for example 8 of 80) reserved for control traffic so acknowledgements and keep-alives (small periodic packets that show a connection is still alive) can still be sent when data buffers are gone. Stop reading from the radio or shrink the receive window (the amount of unacknowledged data the receiver tells the sender it will accept) so the peer slows down.
- Logging ring full: overwrite the oldest entries and increment a "dropped" counter that is written into the next entry, so the reader can see that entries are missing. Never block the caller on logging.
- Sensor buffer overrun: the buffer is sized for the worst case, so an overrun means the consumer is stalled. Keep the newest samples, set an overrun flag in the data stream and raise a fault counter rather than corrupting a block in flight.
- Stack or invariant failure (a stack guard value, a known pattern kept at the end of a stack, found overwritten, or a pool free list found corrupt): do not continue. Record a reason code in a retained register (a register or RAM area that keeps its value across a reset) or flash and reset through the watchdog (a hardware timer that resets the chip unless software keeps restarting it, and that software can also trigger on purpose) into a known state.
- Always: keep per-pool high-water marks, send them in telemetry (the device's own status reports to the outside), and resize the budget from those measurements.
The reserved-blocks rule and the drop counter fit in a few lines. This host sketch models the pool as a counter, sends 100 data packets with none returned, then one acknowledgement:
#include <stdint.h>
#include <stdio.h>
#define NET_BLOCKS 80u
#define CTRL_RESERVE 8u /* blocks that only control traffic may take */
static uint32_t free_blocks = NET_BLOCKS;
static uint32_t dropped_data;
/* Returns 1 if a block was taken, 0 if the packet must be dropped. */
static int net_take(int is_control)
{
uint32_t floor = is_control ? 0u : CTRL_RESERVE;
if (free_blocks <= floor) {
if (!is_control) dropped_data++; /* tail drop: refuse the newest packet and count it */
return 0;
}
free_blocks--;
return 1;
}
int main(void)
{
int accepted = 0;
for (int i = 0; i < 100; i++) accepted += net_take(0); /* 100 data packets arrive, none are freed */
printf("data accepted=%d dropped=%u free blocks=%u\n", accepted, (unsigned)dropped_data, (unsigned)free_blocks);
printf("ack accepted: %d\n", net_take(1));
return 0;
}
Run in a gcc:14 container (GCC 14.4.0, aarch64, -O2 -Wall -Wextra -fsanitize=address,undefined), it prints:
data accepted=72 dropped=28 free blocks=8
ack accepted: 1
Data traffic stops at 72 blocks (80 minus the 8 reserved), the other 28 packets are dropped and counted, and the acknowledgement still gets one of the reserved blocks. The dropped count is what you would send in telemetry.
What would change the decision. If the network traffic profile were truly bursty and the other subsystems idle during bursts, the network pool and the logging ring could share one larger pool. That saves RAM but lets logging steal buffers from the network, so it needs a quota (a cap on how much of the shared pool the logger may hold). Without measurements showing a real gain, the separate fixed budgets are easier to reason about and to test.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Embedded Developer jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs