Google Embedded Software Engineer Interview Preparation Guide - Junior Level
Google's Embedded SWE interview process for junior-level candidates emphasizes practical embedded systems knowledge and low-level programming proficiency. The interview loop includes an initial recruiter screening, technical phone screen rounds focused on C programming and embedded concepts, and multiple onsite rounds covering embedded systems fundamentals, coding under hardware constraints, system-level problem solving, and behavioral assessment. Unlike standard SWE interviews, embedded roles prioritize bit manipulation, memory optimization, hardware interaction understanding, and driver-level concepts over complex data structures and graph algorithms.
Interview Rounds
Recruiter Screening
What to Expect
Initial conversation with Google recruiter to discuss your background, interest in the embedded systems role, and general qualifications. This is a preliminary assessment round designed to verify basic fit and communication skills. The recruiter will also provide information about the role, team, and interview process.
Tips & Advice
Be clear about your embedded systems experience and genuine interest in hardware-software interaction. Prepare 2-3 concrete examples of embedded projects you've worked on (even if academic). Ask thoughtful questions about the team's technology stack, types of devices/hardware they work with, and real-world challenges they solve. Have a professional summary of your background ready that emphasizes any relevant experience with microcontrollers, IoT, or low-level programming.
Focus Topics
Motivation and Questions
Authentic interest in the specific role and team; thoughtful questions about technology, devices, and technical challenges.
Practice Interview
Study Questions
Communication and Clarity
Ability to explain technical concepts clearly without jargon overload; demonstrates capacity to work effectively with both hardware and software engineers.
Practice Interview
Study Questions
Background and Embedded Experience
Clear articulation of your embedded systems experience, academic or professional projects involving microcontrollers, firmware, or IoT systems.
Practice Interview
Study Questions
Technical Phone Screen - Embedded Fundamentals
What to Expect
First technical interview conducted via phone or video where you'll solve an embedded systems problem combining C programming with hardware concepts. This round assesses your ability to write efficient, correct C code while considering hardware constraints like memory limitations, bit-level operations, and performance. You may be asked to write code or pseudo-code on a shared document. The interviewer will probe your understanding of data types, memory management, and how your code maps to actual hardware behavior.
Tips & Advice
Write clean, efficient C code with proper data types for embedded contexts. Think about memory usage and explain your choices. If using bit manipulation, clearly document what each bit represents. Ask clarifying questions about hardware constraints (e.g., 'Are we working with limited RAM?', 'What is the timing constraint?'). Walk the interviewer through your thought process. Be prepared to optimize code for both speed and memory. Avoid using libraries or abstractions; show you can work at the hardware level. If stuck, discuss your approach with the interviewer rather than staying silent.
Focus Topics
Memory and Resource Constraints
Understanding memory hierarchy, stack vs. heap, register constraints, optimizing for minimal memory footprint, and choosing appropriate data structures for limited-resource environments.
Practice Interview
Study Questions
Arrays and String Handling
Working with arrays, index manipulation, string operations in C (without relying on string.h), buffer management, and avoiding overflow conditions.
Practice Interview
Study Questions
Hardware Interaction Concepts
Basic understanding of how software interacts with hardware: memory-mapped I/O, registers, addresses, and relationship between C code and actual device behavior.
Practice Interview
Study Questions
Bit Manipulation and Bit-Level Operations
Competency with bitwise operators (AND, OR, XOR, shifts), bit masking, bit extraction, setting and clearing specific bits, and understanding binary representations.
Practice Interview
Study Questions
C Programming Fundamentals for Embedded
Solid understanding of C syntax, data types (uint8_t, uint16_t, int32_t), pointers, memory management, and function implementation without relying on standard library abstractions where hardware directly accessed.
Practice Interview
Study Questions
Technical Phone Screen - Driver/Protocol Implementation
What to Expect
Second technical phone interview focusing on driver development or hardware protocol implementation. You may be given a practical scenario such as implementing basic driver functionality, handling interrupts, or communicating with a peripheral (e.g., SPI, I2C, UART). The interviewer assesses your understanding of real-time constraints, interrupt handling, state management, and how software bridges microcontroller capabilities with application requirements. This round is more use-case focused than pure coding.
Tips & Advice
Understand the hardware interface or protocol mentioned in the problem (read datasheets if relevant). Ask about real-time constraints, interrupt priorities, and error conditions. Show understanding of state machines and how interrupt handlers interact with main code flow. Write clear, defensive code that handles edge cases. If a specific protocol or hardware interface is mentioned that you've encountered, discuss your prior experience but don't assume; verify details with the interviewer. Discuss timing implications and potential race conditions. Be prepared to explain how your code handles asynchronous events.
Focus Topics
Register Manipulation and Bit-Banging
Direct hardware register access, setting/clearing bits for device configuration, and understanding memory-mapped I/O. Low-level hardware control without abstraction layers.
Practice Interview
Study Questions
Common Communication Protocols (I2C, SPI, UART)
Familiarity with serial communication protocols, timing requirements, data framing, error handling, and how to implement protocol state machines.
Practice Interview
Study Questions
State Machines and Asynchronous Programming
Designing finite state machines for hardware control, managing state transitions, and writing responsive code that handles asynchronous events efficiently.
Practice Interview
Study Questions
Driver Development Basics
Understanding of device driver architecture, initialization sequences, register configuration, interrupt handling, and state management for hardware peripherals.
Practice Interview
Study Questions
Interrupt Handling and Real-Time Constraints
Concepts of interrupt service routines (ISRs), interrupt priorities, interrupt masking, and timing constraints in real-time systems. Understanding how interrupts interact with main program flow.
Practice Interview
Study Questions
Onsite Round 1 - Embedded Systems Coding
What to Expect
First onsite technical interview focused on embedded systems coding problem. You'll solve a problem that combines C programming with embedded systems constraints. This may involve optimizing code for a specific microcontroller with limited resources, implementing a simple algorithm with hardware awareness, or solving a practical embedded challenge. The interviewer assesses code quality, problem-solving approach, resource optimization, and ability to handle real-time constraints in a collaborative setting.
Tips & Advice
Write your solution on a whiteboard or shared document clearly. Explain your approach before coding. Focus on correctness first, then optimization. Discuss trade-offs between speed and memory explicitly. Consider edge cases and error conditions. Explain how your code would behave on actual hardware. Be ready to refactor based on new constraints (e.g., 'the device now has only 2KB of RAM'). Communicate constantly with the interviewer about your reasoning. If you mention specific hardware in your resume, be ready to discuss it in the context of coding problems.
Focus Topics
Error Handling and Robustness
Handling edge cases, error conditions, and designing code that fails gracefully in resource-constrained environments. Defensive programming for embedded contexts.
Practice Interview
Study Questions
Binary Operations and Low-Level Representation
Understanding how data is represented in binary, performing bit-level operations, working with flags and packed data structures, and manipulating bits for efficiency.
Practice Interview
Study Questions
Communication Between Code and Hardware
Understanding how C code translates to hardware behavior, memory-mapped I/O, register access, and the relationship between software logic and physical device state.
Practice Interview
Study Questions
Code Optimization for Embedded Systems
Techniques for optimizing code for speed and memory: avoiding dynamic allocation, using stack efficiently, understanding compiler optimizations, writing cache-friendly code.
Practice Interview
Study Questions
Algorithm Design for Resource-Constrained Environments
Designing algorithms that minimize memory usage and processing time; selecting appropriate algorithms for embedded contexts where standard solutions may be impractical.
Practice Interview
Study Questions
Onsite Round 2 - System Architecture and Integration
What to Expect
Second onsite technical interview assessing your understanding of system-level embedded architecture. You may be asked to design a simple embedded system component, discuss how different subsystems interact (sensor input, processing, output control), or solve a problem requiring understanding of real-time constraints and system integration. The focus is on practical system thinking rather than complex algorithms—how do components work together, what are the timing implications, how do you ensure reliability and synchronization?
Tips & Advice
Think about the complete system picture: inputs, processing, outputs, timing, and synchronization. Draw diagrams to show system flow and data paths. Discuss potential bottlenecks and failure modes. Be specific about timing requirements and how you'd meet them. Consider both hardware capabilities and software limitations. If discussing an actual embedded system you've worked with, draw its architecture and explain design decisions. Ask about non-functional requirements (latency, power consumption, reliability). Use concrete examples from your experience to illustrate points. Avoid over-engineering; solutions should be practical for a junior-level implementation.
Focus Topics
Power and Energy Efficiency
Understanding power consumption in different modes, low-power design techniques, sleep/wake mechanisms, and trade-offs between performance and power.
Practice Interview
Study Questions
Debugging and Verification in Embedded Systems
Techniques for debugging embedded code, using debuggers and emulators, understanding hardware behavior verification, and testing strategies for systems without easy visibility into internals.
Practice Interview
Study Questions
Sensor Input and Output Control
Reading sensor data with appropriate sampling rates, processing noisy inputs, controlling actuators, understanding analog-to-digital and digital-to-analog conversion at system level.
Practice Interview
Study Questions
Embedded System Architecture Design
Understanding system components (sensors, processors, actuators, communication), data flow between components, architectural patterns suitable for embedded systems, and integration considerations.
Practice Interview
Study Questions
Real-Time Systems and Timing Constraints
Meeting hard and soft real-time deadlines, task scheduling, understanding latency and jitter, synchronizing concurrent activities, and designing systems that respond to time-critical events.
Practice Interview
Study Questions
Onsite Round 3 - Behavioral and Cross-Functional Collaboration
What to Expect
Final onsite round focusing on behavioral competencies and demonstrated ability to work effectively in teams. You'll discuss past projects, how you handled technical challenges, experiences collaborating with hardware engineers, learning from mistakes, and alignment with Google's values. The interviewer assesses communication skills, collaborative approach, ability to handle feedback, growth mindset, and how you think about solving real problems in team contexts. This round often includes discussions about your approach to code quality, testing, and communication.
Tips & Advice
Prepare 3-4 concrete project examples (academic or professional) showing: (1) overcoming a technical challenge in embedded systems, (2) collaborating with hardware engineers or team members, (3) debugging a difficult hardware-software issue, (4) iterating on a design based on constraints. Use the STAR method (Situation, Task, Action, Result) for clear storytelling. Be genuine and specific—avoid generic answers. Discuss what you learned from failures. Show curiosity about how embedded systems work. Emphasize collaborative approach rather than solo achievements. Prepare questions about team dynamics, learning opportunities, and how you'll grow at Google. Listen actively and respond thoughtfully to interviewer's comments.
Focus Topics
Code Quality and Communication
Your approach to writing maintainable code, how you document work, communicating technical decisions to teammates, and valuing code reviews.
Practice Interview
Study Questions
Learning from Failure and Iteration
Examples of mistakes you've made in embedded projects, what you learned, how you changed your approach, and how you've grown as an engineer.
Practice Interview
Study Questions
Debugging and Problem Resolution
Approaches to debugging complex issues, handling ambiguous problems, using tools and techniques effectively, and systematic troubleshooting methodology.
Practice Interview
Study Questions
Collaboration with Hardware Engineers
Experience working with hardware teams, understanding hardware constraints that impact software, communicating across disciplines, and resolving hardware-software integration challenges.
Practice Interview
Study Questions
Project Experience and Technical Problem-Solving
Specific examples of embedded projects you've completed, technical challenges you faced, how you approached problem-solving, and measurable outcomes of your work.
Practice Interview
Study Questions
Frequently Asked Embedded Developer Interview Questions
Explain how multi-master I2C arbitration works and how clock stretching affects master/slave interactions. In firmware, how would you detect arbitration loss and recover from a slave that holds SCL low indefinitely?
Sample Answer
Brief definition / principle
- Multi-master I2C arbitration lets two masters start transfers simultaneously; arbitration is resolved bit-by-bit on SDA while both drive SCL. A master writing '1' but observing SDA low lost arbitration; it must stop immediately and become a receiver/idle.
How arbitration is detected in firmware
- During each transmitted bit the firmware/hardware must sample SDA while driving SDA high (releasing) or low. If you attempted to drive high (let the line float) but read low, another master is pulling it low → arbitration lost.
- Many microcontroller I2C peripherals provide a flag (ARBLST/ARBITRATION_LOST) you should check in the ISR or transaction completion path. If using bit-banged I2C, compare intended bit vs sampled SDA each cycle.
Clock stretching effects
- Slaves can hold SCL low to delay master. Master must respect SCL line state and wait; hardware often stretches by polling SCL before toggling. Excessive stretch may indicate a busy/sluggish slave or hung firmware.
Recovery from a slave holding SCL low indefinitely (firmware steps)
- Implement timeouts: if SCL remains low beyond threshold, consider bus stuck.
- Disable I2C peripheral to avoid fighting hardware.
- Reconfigure SCL/SDA as open-drain GPIOs.
- Generate up to 9 manual SCL pulses while monitoring SDA (release SDA); a stuck slave clocking out a byte will release SDA when done. Code pattern:
// pseudo-C
for (i=0;i<9 && SDA_read()==0;i++) {
drive_SCL_low();
delay_half();
release_SCL(); // let pull-up raise SCL
delay_half();
}
- If SDA still low, drive SDA low then generate a STOP (SCL high then SDA high) to attempt bus reset.
- If unresolved, toggle power or reset the slave (if possible) or re-init bus and report error to higher layer.
Best practices
- Use hardware arbitration flags when available, timeouts for clock-stretching, and a deterministic recovery sequence (9 clocks + STOP) in low-level driver with notifications to the system.
Write a bounds-checked read function (for example a Go safeSliceRead(buf []byte, offset, length int64) ([]byte, error)) that defensively rejects negative values, integer overflow when computing offset+length, and out-of-bounds access, returning an explicit error instead of panicking. Contrast this with an alternative design that returns an Option/optional type instead of an error for an expected-empty case (for example popping from an empty stack) and discuss when each style is preferable.
Sample Answer
Direct answer
Defend a bounds-checked read by explicitly rejecting negative offsets/lengths, checking for integer overflow when computing offset+length (which can wrap around to a small or negative number and defeat a naive bounds check), and rejecting anything past the buffer's actual length, returning an explicit error rather than panicking on any of these conditions.
Structured elaboration
- Negative values: reject immediately; a negative offset or length is never valid for a byte-buffer read.
- Overflow on
offset+length: if both are large positiveint64values, their sum can wrap around to a smaller (or negative) number, which would then incorrectly pass a naiveend > len(buf)check; detect this by checkingend < offset(the sum wrapping pastint64's max means the result becomes smaller than one of its inputs). - Actual bounds: after the overflow check, confirm
enddoesn't exceed the buffer's real length.
Worked example (executed against Go's actual runtime; all assertions passed)
func safeSliceRead(buf []byte, offset int64, length int64) ([]byte, error) {
if offset < 0 || length < 0 {
return nil, fmt.Errorf("negative offset or length: %w", ErrOutOfBounds)
}
end := offset + length
if end < offset { // overflow wrapped around
return nil, fmt.Errorf("offset=%d length=%d: %w", offset, length, ErrOverflow)
}
if end > int64(len(buf)) {
return nil, fmt.Errorf("offset=%d length=%d exceeds buffer len=%d: %w", offset, length, len(buf), ErrOutOfBounds)
}
return buf[offset:end], nil
}
Verified: a normal in-range read returns the correct slice; a negative offset, an out-of-bounds range, and an offset near math.MaxInt64 (which triggers the overflow branch specifically, not just the plain out-of-bounds branch) all correctly return an error instead of panicking or, worse, silently returning wrong data.
A complementary design for an 'expected-empty' case rather than an out-of-range error: C++'s std::optional<int> safe_pop(std::stack<int>& s) returns an empty optional when popping an empty stack, instead of throwing or returning a sentinel like -1 that could collide with a real value; this is preferable when 'nothing there' is a normal, expected outcome (not a caller bug), whereas safeSliceRead's explicit error is right because an out-of-bounds read usually DOES indicate a caller bug worth surfacing loudly.
Trade-offs and pitfalls
The end < offset overflow check works here specifically because Go's integer overflow wraps (rather than panicking or being undefined behavior), which is exactly the property that makes it possible to detect after the fact by checking if the sum became smaller than an input; in a language with undefined behavior on signed overflow (C, in the strict standard sense), this same check is not safe to rely on and you'd need to check BEFORE the addition (e.g., if length > math.MaxInt64 - offset) instead.
Compare C-style manual resource cleanup with RAII in C++. In a codebase where exceptions are disabled, what idiomatic patterns manage resources deterministically, and what would a small RAII wrapper look like?
Sample Answer
Direct answer
C-style cleanup is manual: acquire resources, and on every exit path (success, each early return, each error) release exactly the ones already acquired, usually by jumping to one goto out; label at the end. RAII (resource acquisition is initialisation) ties each resource's release to an object's destructor. A destructor is a function the language calls automatically when an object's lifetime ends, and for a local variable that happens when it leaves its scope (the { } block it was declared in). So the release runs on every path, in reverse order of construction. Turning exceptions off (-fno-exceptions) removes only exception unwinding, which is the compiler-generated cleanup that runs destructors while an exception travels up the call stack; destructors still run on return, break, goto and falling off the end of a scope, so RAII keeps working. What exceptions-off changes is how failure is reported: constructors cannot throw, so you use a factory (a static function that builds and returns the object) that returns an empty or failed object, check it, and keep destructors that cannot fail. The small wrapper below is a move-only owner of a FILE* (move-only means it cannot be copied, only handed over, so there is never a second owner), plus std::unique_ptr with a custom deleter (a small function object telling the pointer how to release this kind of resource) for a malloc buffer.
Structured comparison
C: single-exit goto | C++: RAII objects | |
|---|---|---|
| Who releases | the programmer, at the label | the destructor, by the language |
New early return added later | easy to forget cleanup (leak) | still correct, nothing to update |
| Order | you write reverse order by hand | automatic reverse of construction |
| Ownership transfer | a convention in comments | a move-only type or unique_ptr, enforced by the compiler |
| Partial acquisition | NULL-initialise everything, free what is non-null | an object that failed to acquire is empty, its destructor does nothing |
Patterns that work without exceptions
Patterns 1 to 3 cover most code; 4 to 6 handle particular situations.
- Scope-owned wrapper per resource (the
Fileclass below): constructor or factory acquires, destructor releases, copying is deleted, moving transfers ownership. std::unique_ptr<T, Deleter>for any C handle: a deleter functor callsfree,fclose,close,munmap.- Fallible construction through a factory (fallible: it can fail):
static File open(path, mode)returns an object that is empty on failure; the caller tests it (if (!f) ...). No constructor ever needs to fail. - Explicit status where close can fail: a destructor cannot report an error, so give the class a
close()that returns thefclosestatus for code that cares (writes to disk), and let the destructor be the safety net. The second example below shows it. - Scope guard (a tiny object running a lambda, i.e. a small inline function, in its destructor) for one-off cleanups that do not deserve a class; the second example below shows one. Arenas and pools (one object that owns many allocations and releases them all in one place) fit when many small objects share a lifetime.
- Allocation failure: use the
std::nothrowform ofnew(the version that returns a null pointer instead of throwing) ormalloc, and test the result, since there is no handler to throw to. A function markednoexceptpromises not to throw, which is what the move operations below declare.
Worked example
Compiled with g++ -std=c++17 -O2 -fno-exceptions -Wall -Wextra -fsanitize=address,undefined raii.cpp (GCC 14, Linux; run in a container). The C-style function is compiled as C++ here only so both styles share one file; in a C project it is plain C.
// build: g++ -std=c++17 -fno-exceptions -Wall -Wextra raii.cpp
#include <cstdio>
#include <cstdlib>
#include <memory>
#include <utility>
// 1. C-style: one exit label, cleanup in reverse order of acquisition
int process_c(const char* path) {
int rc = -1;
FILE* f = nullptr;
char* buf = nullptr;
f = std::fopen(path, "r");
if (!f) goto out;
buf = static_cast<char*>(std::malloc(64));
if (!buf) goto out;
if (!std::fgets(buf, 64, f)) goto out;
rc = 0;
out:
std::free(buf); // free(NULL) is a no-op
if (f) std::fclose(f);
return rc;
}
// 2. Small RAII wrapper: owns one FILE*, move-only, no exceptions
class File {
public:
File() = default;
static File open(const char* path, const char* mode) { File f; f.fp_ = std::fopen(path, mode); return f; }
File(const File&) = delete;
File& operator=(const File&) = delete;
File(File&& o) noexcept : fp_(std::exchange(o.fp_, nullptr)) {}
File& operator=(File&& o) noexcept { if (this != &o) { reset(); fp_ = std::exchange(o.fp_, nullptr); } return *this; }
~File() { reset(); }
explicit operator bool() const { return fp_ != nullptr; }
FILE* get() const { return fp_; }
private:
void reset() { if (fp_) { std::fclose(fp_); std::puts(" (File closed)"); fp_ = nullptr; } }
FILE* fp_ = nullptr;
};
// 3. unique_ptr with custom deleter for a C handle
struct FreeDeleter { void operator()(void* p) const { std::free(p); std::puts(" (buffer freed)"); } };
using CBuf = std::unique_ptr<char, FreeDeleter>;
int process_cpp(const char* path) {
File f = File::open(path, "r");
if (!f) return -1; // nothing to clean by hand
CBuf buf(static_cast<char*>(std::malloc(64)));
if (!buf) return -1;
if (!std::fgets(buf.get(), 64, f.get())) return -1; // early return: both destructors still run
std::printf(" read: %s", buf.get());
return 0;
}
int main() {
std::FILE* t = std::fopen("sample.txt", "w"); std::fputs("hello\n", t); std::fclose(t);
std::printf("C style rc=%d\n", process_c("sample.txt"));
std::printf("RAII rc=%d\n", process_cpp("sample.txt"));
std::printf("RAII missing file rc=%d\n", process_cpp("nope.txt"));
}
C style rc=0
read: hello
(buffer freed)
(File closed)
RAII rc=0
RAII missing file rc=-1
Reading it: in process_cpp, f was constructed before buf, and the output shows (buffer freed) before (File closed), so destruction ran in reverse order of construction with no code written for it. In the missing-file call the function returned before buf existed; the empty File destructor printed nothing, because it had nothing to close. No leak report appeared from the sanitizer run. The goto out version works too, but each new resource and each new early exit is a place to make a mistake.
The non-obvious lines of File:
std::exchange(o.fp_, nullptr)storesnullptrino.fp_and returns the old value (cppreference: "Replaces the value ofobjwithnew_valueand returns the old value ofobj"). In the move constructor it takes the handle and leaves the source empty, so only one object will ever close it.File(File&& o) noexceptis the move constructor;noexceptpromises it never throws, which it cannot, since it only copies a pointer.reset()closes the file if there is one and sets the pointer to null, so calling it twice is harmless. The destructor and the move assignment both use it.explicit operator bool()letsif (!f)ask "did the open succeed?";explicitstops aFilefrom silently converting to a number elsewhere.
Scope guard and a checked close
The same approach applies to a one-off cleanup and to a close that can fail. Built the same way as above (g++ -std=c++17 -O2 -fno-exceptions -Wall -Wextra -fsanitize=address,undefined, GCC 14, Linux):
#include <cstdio>
#include <utility>
// Runs a callable when the object leaves scope.
template <class F>
class ScopeGuard {
public:
explicit ScopeGuard(F f) : f_(std::move(f)) {}
ScopeGuard(const ScopeGuard&) = delete;
ScopeGuard& operator=(const ScopeGuard&) = delete;
~ScopeGuard() { f_(); }
private:
F f_;
};
// A file wrapper whose close() reports the fclose result.
class OutFile {
public:
OutFile() = default;
OutFile(const OutFile&) = delete;
OutFile& operator=(const OutFile&) = delete;
OutFile(OutFile&& o) noexcept : fp_(std::exchange(o.fp_, nullptr)) {}
static OutFile open(const char* path) { OutFile f; f.fp_ = std::fopen(path, "w"); return f; }
explicit operator bool() const { return fp_ != nullptr; }
std::FILE* get() const { return fp_; }
int close() { // 0 on success, EOF on failure; safe to call twice
if (!fp_) return 0;
int rc = std::fclose(fp_);
fp_ = nullptr;
return rc;
}
~OutFile() { close(); } // safety net: result ignored here
private:
std::FILE* fp_ = nullptr;
};
int save(const char* path) {
OutFile out = OutFile::open(path);
if (!out) return -1;
ScopeGuard log([] { std::puts(" (save() finished)"); });
std::fputs("data\n", out.get());
return out.close() == 0 ? 0 : -2; // the caller learns whether the flush worked
}
int main() {
std::printf("save rc=%d\n", save("out.txt"));
std::printf("save to missing dir rc=%d\n", save("no_such_dir/out.txt"));
}
(save() finished)
save rc=0
save to missing dir rc=-1
The guard's lambda ran when save returned, without any cleanup call in the body. The checked close() returns the fclose result, so a failed flush becomes the return code -2 instead of being lost. The missing directory made open fail, and the early return -1 ran no guard because it had not been constructed yet.
Pitfalls
- Copying a handle wrapper would double-close the same
FILE*; that is why copy is= deleteand move usesstd::exchangeto null the source. - A destructor that does real work and can fail has nowhere to report it with exceptions off; keep it to the release call and expose
close()for checked shutdown. - Skipped destructors: RAII does not run on
exit()called mid-function,abort(),_exit()or a crash, andlongjmp(the C function that jumps straight back to an earlier saved point, abandoning the frames in between) over C++ frames is undefined behaviour if it skips non-trivial destructors (destructors that do real work). Return error codes up the stack instead. - Two-phase initialisation (default-construct then
init()) recreates the C problem of half-built objects; the factory approach avoids it. - Mixing styles: when a C API requires a callback with a raw pointer, keep the owner object alive in the caller's scope and pass
.get().
Judgement
In a codebase with exceptions disabled, commit to RAII wrappers with factory-style construction and explicit status codes; keep the goto cleanup idiom only in actual C files or in a C library you must match. If a resource is used once in a tiny function, a goto is fine; as soon as there are two resources or more than one exit, wrap them.
You suspect heap corruption in a multitasking RTOS where seemingly unrelated tasks crash. Describe a workflow to identify the corruption source: enabling heap guards, per-task pools, stack/heap canaries, selective watchpoints, link-map analysis, binary search tests, and minimizing instrumentation impact to avoid hiding timing-sensitive bugs.
Sample Answer
Overview / goal
I’d follow a methodical, low-intrusion workflow to localize the heap corruption without masking timing-sensitive bugs.
1. Baseline & reproduce
- Reproduce with deterministic inputs; collect core dump, task list, free/used heap stats, and timestamps.
- Run with hardware trace or SWO to avoid printf-induced timing changes.
2. Passive instrumentation
- Enable lightweight heap guards (red zones) and stack canaries globally but keep logging minimal.
- Turn on malloc/free logging with sampling or ring buffer to avoid timing perturbation.
3. Isolate by partitioning
- Convert global heap into per-task or per-subsystem pools one at a time (or use slab allocators). If corruption disappears when a task uses its own pool, that task is suspect.
4. Selective hardening
- Add bigger canaries and guard pages only around allocations from suspected pools.
- Use compiler options to insert stack canaries for individual tasks.
5. Watchpoints & link-map analysis
- Use the map file to find heap/RO/data ranges; set data-breakpoint/watchpoints on canary/footer addresses or frequently-corrupted buffers.
- If hardware breakpoints limited, set conditional watchpoints in the allocator (check pattern on free).
6. Binary-search tests
- Disable half the tasks, or bisect test cases to find minimal reproducer. Use task-enable masks to preserve timing for remaining tasks.
7. Minimize instrumentation impact
- Prefer hardware tracing, JTAG breakpoints, and sampling over heavy logging.
- Use assertion-based checks that only run in a sampled subset of runs or triggered after suspicious events.
- When needed, replicate the scenario on a simulator/emulator where heavy instrumentation won’t affect timing.
8. Confirm & fix
- Once source found, add deterministic unit tests, fix bounds/ownership errors, add durable per-module pools or hardened allocator, and keep regression tests with stress runs.
This approach balances broad detection, focused isolation, and minimal timing interference so the bug isn’t hidden by the debugger itself.
Describe the common memory types used in embedded systems (SRAM, DRAM, flash, EEPROM). For each type explain volatility, typical access speed, endurance, common uses (stack/heap/data/code), and how an IoT gateway design differs from a constrained MCU sensor node in memory usage.
Sample Answer
SRAM
- Volatility: volatile (loses content when power removed).
- Access speed: very fast, single-cycle on many MCUs.
- Endurance: effectively unlimited (read/write cycles not a concern).
- Common uses: stack, CPU registers, heap, runtime variables, small data buffers.
- Notes: used for deterministic access (real-time tasks).
DRAM
- Volatility: volatile, requires refresh.
- Access speed: slower than SRAM (higher latency), but higher density.
- Endurance: not a limiting factor.
- Common uses: system/main memory in gateways or SoCs where large RAM needed (not typical on small MCUs).
- Notes: needs refresh controller and more power.
Flash (NOR/NAND)
- Volatility: non-volatile.
- Access speed: read fast (NOR supports XIP), write slower, block erase required.
- Endurance: limited (typically 10k–100k erase cycles for NOR, varies).
- Common uses: code storage (firmware image), filesystem (NAND + FTL), large non-volatile data.
- Notes: wear-leveling and erase-block management required.
EEPROM
- Volatility: non-volatile.
- Access speed: slower writes than flash, byte- or page-programmable.
- Endurance: typically 100k–1M cycles depending on type.
- Common uses: small configuration parameters, calibration data, persistent settings.
- Notes: easier byte-level updates than flash, used for infrequent writes.
IoT Gateway vs Constrained MCU Sensor Node
- Gateway: has DRAM and larger flash/NAND + storage (SSD/SD). Uses OS, dynamic memory, file systems, complex networking stacks; memory footprint large, can trade latency for capacity.
- Constrained MCU: relies on SRAM + onboard flash (and sometimes EEPROM). Code often runs XIP or is copied to RAM; tight RAM footprint forces static allocation, stack/heap limits, minimal OS or RTOS; wear-management for flash critical.
- Design implications: gateways focus on throughput and concurrency; sensor nodes prioritize deterministic timing, low power, and careful flash/EEPROM write budgeting.
Find the length of the longest substring without repeating characters. Implement lengthOfLongestSubstring(s) in your preferred language and explain the sliding window approach, why it is O(n), and how you maintain character indices. Example: s = "abcabcbb" -> 3.
Sample Answer
Direct answer
Slide a window over the string while tracking, for each character, the index where it was last seen. When you hit a character already inside the current window, jump the window's left edge to just past that character's previous occurrence, then update the best length seen. Each index is visited a bounded number of times, so the whole scan is O(n).
Structured elaboration
The sliding window invariant
Maintain two indices, left and right, that always bound a substring with no repeated characters. right advances one character at a time. last_seen is a hash map from character to the most recent index where it appeared. When the character at right was last seen at an index that is still inside the current window (last_seen[ch] >= left), move left forward to last_seen[ch] + 1, the smallest new left edge that excludes the earlier occurrence. If the last occurrence is outside the window, or the character has never been seen, the window simply grows.
Why it is O(n)
right moves forward exactly once per character, n times total. left only ever moves forward too, and it never exceeds right, so across the whole scan left advances at most n times as well. Neither pointer ever moves backward. That gives a total of at most 2n pointer movements, so the algorithm is O(n) time, with O(min(n, alphabet size)) space for the hash map, bounded by how many distinct characters can appear in a window at once.
Maintaining character indices correctly
The subtle part is the last_seen[ch] >= left check. Without it, a stale entry from far in the past (before the current window even started) could incorrectly shrink the window, because the map keeps the LAST time a character was seen, which might be from before the window's current left edge if that character has not reappeared since.
Worked example
def length_of_longest_substring(s):
last_seen = {}
left = 0
best = 0
for right, ch in enumerate(s):
if ch in last_seen and last_seen[ch] >= left:
left = last_seen[ch] + 1
last_seen[ch] = right
best = max(best, right - left + 1)
return best
print(length_of_longest_substring("abcabcbb"))
Output:
3
Tracing it by hand: at right=3 the window holds "abc" and the next character is a, last seen at index 0, which is still inside the window (0 >= left=0), so left jumps to 1. The window is now "bca", length 3. The same pattern repeats through the rest of the string, and the best length never exceeds 3, matching the printed result and the question's own stated expectation. The same technique ports directly: a Java solution uses a HashMap<Character, Integer> in place of the dict, and a JavaScript solution uses a Map or a plain object, with identical pointer logic.
Trade-offs and pitfalls
A common wrong turn is checking if ch in last_seen alone and moving left on every repeat ever seen, even ones from before the current window, which can move left backward and silently break correctness. Another is recomputing "is this character in the current window" by scanning the window itself on every step, which turns the algorithm back into O(n^2) or O(n * alphabet size). For very large alphabets (full Unicode rather than ASCII), the map can grow larger, but it is still bounded by the number of distinct characters that actually appear, not by the input length.
Implement a simple fixed-size bump allocator in C for a contiguous memory pool. Provide the API: void pool_init(void* mem, size_t size); void* pool_alloc(size_t n); void pool_reset(void); The allocator does not need to free individual allocations. Ensure proper alignment for typical embedded types.
Sample Answer
What a bump allocator is. A bump (arena) allocator hands out memory by moving one pointer forward through a fixed buffer. Allocation is a compare and an add, so it takes the same few instructions every time (O(1), no searching) and cannot fragment. The cost is that individual blocks cannot be freed; the only way to give memory back is to reset the whole pool at once. That fits firmware that allocates its objects once during initialisation, or a short-lived scratch area that is thrown away at the end of each work cycle.
The implementation. Three details carry the interview weight: align the actual address (not just the offset), check the size against the space remaining before doing any arithmetic on it (so a huge n cannot wrap around), and decide what pool_alloc(0) means. The rounding expression (x + 7) & ~7 rounds x up to the next multiple of 8: adding 7 pushes any value that is not a multiple of 8 past the next multiple, and & ~7 then clears the low three bits (the part that is not a multiple of 8). For example 13 + 7 = 20 and 20 & ~7 = 16, while 16 + 7 = 23 and 23 & ~7 = 16, so a value that is already a multiple of 8 stays where it is. uintptr_t is an unsigned integer type wide enough to hold a pointer, which lets the code do this bit arithmetic on an address.
#include <stdint.h>
#include <stddef.h>
#include <stdio.h>
#define POOL_ALIGN 8u /* covers uint64_t and double on Cortex-M and on the host */
static uint8_t *pool_cur;
static uint8_t *pool_end;
static uint8_t *pool_base;
void pool_init(void *mem, size_t size)
{
uintptr_t start = (uintptr_t)mem;
uintptr_t end = start + size;
start = (start + (POOL_ALIGN - 1u)) & ~(uintptr_t)(POOL_ALIGN - 1u);
if (mem == NULL || start > end) { /* buffer too small to hold even one aligned byte */
pool_base = pool_cur = pool_end = NULL;
return;
}
pool_base = pool_cur = (uint8_t *)start;
pool_end = (uint8_t *)end;
}
void *pool_alloc(size_t n)
{
size_t remaining = (size_t)(pool_end - pool_cur);
size_t rounded;
if (n == 0u || n > remaining) {
return NULL;
}
rounded = (n + (POOL_ALIGN - 1u)) & ~(size_t)(POOL_ALIGN - 1u); /* cannot overflow: n <= remaining */
if (rounded > remaining) {
return NULL;
}
void *p = pool_cur;
pool_cur += rounded;
return p;
}
void pool_reset(void)
{
pool_cur = pool_base;
}
/* ---- test driver ---- */
int main(void)
{
static _Alignas(16) uint8_t raw[70];
pool_init(raw + 1, sizeof raw - 1); /* deliberately misaligned start */
void *a = pool_alloc(1);
void *b = pool_alloc(10);
void *c = pool_alloc(8);
printf("a%%8=%u b%%8=%u c%%8=%u\n", (unsigned)((uintptr_t)a % 8), (unsigned)((uintptr_t)b % 8), (unsigned)((uintptr_t)c % 8));
printf("b-a=%td c-b=%td\n", (uint8_t *)b - (uint8_t *)a, (uint8_t *)c - (uint8_t *)b);
printf("alloc(SIZE_MAX) -> %p\n", pool_alloc((size_t)-1));
printf("alloc(0) -> %p\n", pool_alloc(0));
int filled = 0;
while (pool_alloc(8) != NULL) filled++;
printf("further 8-byte allocs before exhaustion: %d\n", filled);
pool_reset();
printf("after reset, first alloc equals a: %s\n", pool_alloc(1) == a ? "yes" : "no");
return 0;
}
Compiled with gcc -O2 -Wall -Wextra -fsanitize=address,undefined bump.c in a gcc:14 container (GCC 14.4.0, aarch64), the program prints:
a%8=0 b%8=0 c%8=0
b-a=8 c-b=16
alloc(SIZE_MAX) -> (nil)
alloc(0) -> (nil)
further 8-byte allocs before exhaustion: 3
after reset, first alloc equals a: yes
Tracing the output. raw is 70 bytes, and the demo hands over raw + 1, so the pool is 69 bytes starting one byte past a 16-byte boundary. Call that boundary address X: the start is X + 1, and (X + 1 + 7) & ~7 rounds it up to X + 8 (7 bytes lost to alignment), while the end is X + 70, so 70 - 8 = 62 bytes are usable. pool_alloc(1) rounds 1 up to 8 and returns X + 8 (a), leaving 54. pool_alloc(10) rounds 10 up to 16 and returns X + 16 (b), so b - a = 8, leaving 38. pool_alloc(8) returns X + 32 (c), so c - b = 16, leaving 30. Each further 8-byte request takes 8: 30 gives three allocations (24 bytes), leaving 6, and the fourth fails because 8 > 6. That is the 3 in the output. (%td prints a pointer difference, the type ptrdiff_t, in bytes here because the pointers are cast to uint8_t *.)
Why each choice.
- Alignment of the address.
pool_initrounds the start address up to a multiple of 8, because the caller may pass any buffer (the demo deliberately passesraw + 1). Rounding only the running offset would leave every pointer misaligned whenever the buffer itself is. 8 bytes coversuint64_tanddouble; a type with stricter alignment needs a largerPOOL_ALIGN. A misaligned access is a real hazard on microcontrollers: some cores and some instructions raise a fault for it, and even where the hardware tolerates it the access can cost extra cycles. The output showsa,bandcall at multiples of 8, and the 1-byte request consumed 8 bytes (b - a = 8), so rounding costs up to 7 bytes of internal waste per allocation (bytes reserved inside an allocation that the caller never uses). - Overflow-safe size check.
n > remainingis tested first. Only after that isn + 7computed, and it cannot wrap (overflow past the largest valuesize_tcan hold and restart near zero) becausen <= remaining. The naiveif (cur + n > end)test can wrap for a hugenand wrongly succeed: on a 32-bit part withcurat 0x20000100 andnof 0xFFFFFFFF, the sum is 0x1200000FF, which truncates to 0x200000FF, belowend, so the check passes and the caller is given a block that does not exist; the demo'spool_alloc((size_t)-1)returns NULL. pool_alloc(0)returns NULL in this version. The alternative (return a unique valid pointer) is also legitimate; what matters is documenting the choice.uintptr_tfor the alignment arithmetic. Bit masks are applied to integers, not pointers, so the rounding is well defined.pool_resetonly rewinds the cursor. Every pointer handed out earlier now points into memory that will be reused, so the caller must have dropped them all. The demo's last line shows the first allocation after a reset lands on the same address as the first one before it.
Limits to state in the interview. The allocator is not thread safe (it is not correct when several tasks or interrupts call it at the same time): two tasks, or a task and an ISR, can read the same pool_cur and receive the same block. Protect pool_alloc with a critical section (for example masking interrupts briefly) or give each context its own pool. There is also no per-block free, so use it only where lifetimes are grouped. If some objects must be freed individually, a fixed-block pool is the better tool; if every object shares one lifetime, the bump allocator is smaller, faster and impossible to fragment.
What was the biggest technical challenge in that project, and how did you overcome it?
Sample Answer
Direct answer: Pick one real obstacle, not the project's general level of difficulty, and be honest that something didn't work on the first attempt. State what the failure looked like, what you tried, what actually worked, and why.
What "biggest challenge" means to the interviewer
Distinguish ambient difficulty (the project was generally hard) from a specific moment where you were stuck, wrong, or something broke. The question wants the latter: a real obstacle with a resolution arc, not just "the project was hard."
Framework for the answer
- Name the specific obstacle in one sentence (a bug, a wrong initial approach, a constraint discovered late).
- State what you tried first and why it seemed reasonable at the time.
- State why that didn't work, and what new information surfaced.
- State what you changed and why it worked.
- State what you'd do differently to catch it earlier next time.
Common obstacle types
| Type | Example | Resolution pattern |
|---|---|---|
| Technical / design | An approach that worked in testing broke under real conditions | Instrument to find the actual root cause, then isolate the fix to the affected path only |
| Dependency | A team or system you relied on didn't deliver as expected | Renegotiate scope or build a fallback path instead of waiting |
| Knowledge gap | The domain was unfamiliar and the first design missed a real constraint | Bring in a subject-matter reviewer earlier, before the design is finalized |
Worked example (illustrative, no fabricated precision)
Midway through a service migration, the new system passed all pre-launch load tests but started timing out under real production traffic within the first day. The load tests had used synthetic requests with a flat size distribution. Investigating production logs pointed to a long-tail payload size distribution; illustrative assumption for this example: the largest requests ran roughly 50 times the median size, and those large requests were serialized on a single-threaded parser that the flat synthetic test data never exercised. The fix: moved parsing for large payloads onto a separate worker pool bounded by a queue, instead of the shared request-handling thread pool, isolating the slow path without touching the common case. Verified by replaying a sample of real production traffic against the new code path in staging before rollout, rather than trusting the original synthetic load test again.
Trade-offs and pitfalls
- Picking a challenge that wasn't really yours to solve undermines the whole answer once probed.
- Describing only the technical fix without naming what changed in your process afterward misses half the point of the question.
- Avoiding admitting the first approach failed reads as defensive rather than reflective.
- Choosing an obstacle that resolved mostly by luck doesn't showcase reasoning the way a diagnosed-and-fixed obstacle does.
Implement the design for compact nested or hierarchical state machines (HSM) in C to be used on an embedded system with severe RAM constraints. The HSM must support entry/exit actions, shallow history, prioritized events, and dispatching without dynamic allocation. Outline the data structures, table-driven dispatch algorithm, and provide example pseudocode for state transition dispatch.
Sample Answer
Approach (brief)
Design a compact table-driven HSM in C using arrays and enums only, no dynamic allocation. Represent states as small integers, tables for parent, entry/exit handlers, transition targets, guards and prioritized event-action lists. Support shallow history by storing last active child index.
Key data structures
- StateId: uint8_t enum
- EventId: uint8_t enum
- Handler: function pointer type
void (*)(void*) - Tables:
- parent[StateId]
- entry[StateId], exit[StateId]
- history[StateId] (uint8_t, 0 = no history)
- transitions: array of { state, event, guard_idx, action_idx, target_state }
- guards[], actions[] function pointer arrays
Example compact definitions:
typedef uint8_t StateId;
typedef uint8_t EventId;
typedef void (*Action)(void*);
typedef bool (*Guard)(void*);
typedef struct {
StateId src;
EventId evt;
uint8_t guard; // index into guards[], 0xFF = always true
uint8_t action; // index into actions[], 0xFF = none
StateId tgt; // target = same as src means internal
} Transition;
Table-driven dispatch algorithm (steps)
- Build vector of candidate transitions by scanning transitions[] for src = current state or any ancestor (walk parent[]).
- Prioritize by event and transition table ordering (table arranged highest priority first).
- Evaluate guards in ancestor-to-descendant order; pick first true.
- If internal transition (tgt == src) call action only.
- For external transition:
- Compute LCA between current and target using parent[].
- Execute exit handlers from current up to (but not including) LCA; update shallow history for parents.
- Execute transition action.
- Execute entry handlers from LCA down to target (if history exists, descend to recorded child).
- Set current = final entered state.
Pseudocode for dispatch
void dispatch(EventId e, void *ctx){
StateId s = current;
Transition *t = find_transition(s,e,ctx); // scan transitions[] up ancestors
if(!t) return;
if(t->tgt == s){ // internal
if(t->action!=0xFF) actions[t->action](ctx);
return;
}
StateId lca = find_lca(s, t->tgt);
// exit path
for(StateId x = s; x != lca; x = parent[x]){
if(exit[x]) exit[x](ctx);
if(parent[x]!=0xFF) history[parent[x]] = x; // shallow history
}
// action
if(t->action!=0xFF) actions[t->action](ctx);
// entry path stack from lca->tgt
StateId path[DEPTH]; int n=0;
for(StateId y = t->tgt; y != lca; y = parent[y]) path[n++]=y;
while(n--) {
StateId en = path[n];
if(entry[en]) entry[en](ctx);
// follow shallow history
if(history[en]) en = history[en];
}
current = resolve_deepest_active(t->tgt);
}
Complexity & constraints
- Memory: O(#states + #transitions). Tables are packed arrays of uint8_t and small structs.
- CPU: dispatch scans ancestor chain and transitions — acceptable when tables are small; optimize by indexing transitions by event to reduce scan.
Notes / Best practices
- Store tables in flash (const) to save RAM.
- Keep function pointer tables small; use a dispatcher index-to-fn map.
- Precompute parent chains or LCA hints for deep hierarchies if dispatch time critical.
- Use 8-bit types where feasible, align structs to avoid padding.
Describe the criteria you use to decide between applying the smallest hotfix to restore correctness and reverting to the previous stable release. Provide a concrete example where a hotfix is preferable and another where revert is safer. Include risk assessment and customer impact considerations.
Sample Answer
Choosing between a small hotfix and reverting to the previous stable release comes down to which action more reliably and quickly restores correctness with the least additional risk.
Decision criteria
Prefer a hotfix when the root cause is well-understood, the fix is small and isolated (touches little beyond the specific broken behavior), and reverting would also roll back unrelated, wanted changes that shipped in the same release. Prefer a revert when the cause isn't yet fully understood (a revert restores a known-good state without needing to be right about the cause), the "small" fix would actually need to touch several places, or there's any doubt the hotfix itself is fully safe under production conditions.
Two concrete examples
- Hotfix preferable: a null-check is missing on one specific field that a new client type started sending; the fix is one line, well-understood, and reverting would also undo unrelated bug fixes shipped in the same release.
- Revert safer: a new feature interacts with three other systems in ways not yet fully mapped, and a stakeholder needs the feature live today despite the fix genuinely needing more than a day; here, reverting removes both the feature and the risk cleanly, buying time to fix it properly without the pressure of an active production issue.
Risk assessment and customer impact
Weigh: confidence in root cause (low confidence favors revert), blast radius of the broken behavior (wide impact favors whichever restores service fastest), and what else would be lost by reverting (favors hotfix if the release bundled other now-live fixes worth keeping).
Trade-offs and pitfalls
A partial hotfix (fixes some cases but not all) is worse than either a clean revert or a complete fix, since it can create a false sense the issue is resolved while a subset of users remain affected; if a hotfix can't be verified to be complete, defaulting to revert is usually the safer choice.
Want to create your own tailored preparation guide using our deep research?
Get Started for FreeInterview-Ready Courses
Visual-first, interactive, structured learning paths
Browse Embedded Developer jobs
AI-enriched listings across hundreds of company career pages
Explore Jobs