\n\n\n\n.toolbar {\n display: flex;\n flex-wrap: wrap; / lets the row break instead of shrinking the buttons when they grow /\n gap: 4px;\n width: 300px;\n margin: 16px;\n padding: 0;\n}\n.toolbar button {\n box-sizing: border-box;\n width: 28px;\n height: 28px;\n padding: 0;\n border: 1px solid #667;\n border-radius: 4px;\n background: #eef0f6;\n}\n\n/ Primary input is a finger: same markup, bigger boxes. /\n@media (pointer: coarse) {\n .toolbar button {\n width: 44px;\n height: 44px;\n }\n}\n\nThe desktop rules are untouched: 28px buttons with a 4px gap.\nThe coarse block raises width and height to 44px. Because the buttons are real boxes of that size, the whole 44px square is the target and no extra hit-area trick is needed.\nThe toolbar is 300px wide. At 44px with 4px gaps, 6 x 44 + 5 x 4 = 284px fits and 7 x 44 + 6 x 4 = 332px does not, so the eight buttons break into rows of 6 and 2. On a fine pointer, 8 x 28 + 7 x 4 = 252px fits in one row.\nWhere wrapping would make the toolbar too tall (a 14-button toolbar at 44px wraps to three rows of 6, 6 and 2, as the last line of the output shows), the alternatives are horizontal scrolling or moving the least-used actions into an overflow menu. Those change the markup and behaviour, which is the point at which \"no redesign\" stops being possible.\n\nVerifying it\n\nThe test runs in Chromium, Firefox and WebKit, the three browser engines (the rendering cores that lay out the page), driven by Playwright 1.48.2, a library that scripts real browsers; its WebKit build is Playwright's, not Safari, and it runs headless, meaning without a visible window, so it can measure layout and simulate a tap but cannot tell how a real thumb feels on a real phone. A fine pointer is the default context. A touch context is created with hasTouch: true, in which matchMedia('(pointer: coarse)') is true in all three engines. It asserts the sizes, the row counts, that nothing pokes out of the toolbar, and that a real touch tap 40px right and 20px down from the first button's corner (inside a 44px box, outside a 28px one) activates it. Deleting the coarse rule or the flex-wrap declaration makes the assertions fail; both mutants fail the 44px size assertion, the second because a non-wrapping flex row shrinks its items to fit instead of letting them overflow.\n\n// check.mjs\nimport { chromium, firefox, webkit } from 'playwright';\nimport { readFileSync } from 'node:fs';\n\nconst css = readFileSync('toolbar.css', 'utf8');\nconst html = readFileSync('index.html', 'utf8');\nconst variants = {\n shipped: css,\n noCoarseRule: css.replace(/@media \\(pointer: coarse\\) \\{[\\s\\S]*\\n\\}\\n/, ''),\n noWrap: css.replace('flex-wrap: wrap;', 'flex-wrap: nowrap;'),\n};\nconst ok = (c, m) => { if (!c) throw new Error('ASSERT FAILED: ' + m); };\n\nasync function open(browser, variant, touch) {\n const ctx = await browser.newContext({ viewport: { width: 400, height: 400 }, hasTouch: touch });\n const p = await ctx.newPage();\n await p.route('**/toolbar.css', r => r.fulfill({ contentType: 'text/css', body: variants[variant] }));\n await p.route('**/index.html', r => r.fulfill({ contentType: 'text/html', body: html }));\n await p.goto('http://x.test/index.html');\n return p;\n}\nconst layout = p => p.evaluate(() => {\n const bar = document.querySelector('.toolbar');\n const r = [...bar.querySelectorAll('button')].map(b => b.getBoundingClientRect());\n const rows = {};\n r.forEach(b => { rows[Math.round(b.top)] = (rows[Math.round(b.top)] || 0) + 1; });\n return {\n coarse: matchMedia('(pointer: coarse)').matches,\n size: [r[0].width, r[0].height],\n perRow: Object.values(rows),\n overflowsBar: r.some(b => b.right > bar.getBoundingClientRect().right + 0.5),\n minSide: Math.min(...r.map(b => Math.min(b.width, b.height))),\n };\n});\n\nasync function fineTest(b) {\n const p = await open(b, 'shipped', false);\n const m = await layout(p);\n ok(!m.coarse && m.size.join() === '28,28' && m.perRow.join() === '8', 'fine pointer keeps the dense 28px single row');\n return m;\n}\nasync function coarseTest(b, variant) {\n const p = await open(b, variant, true);\n const m = await layout(p);\n ok(m.coarse, 'touch context reports pointer: coarse');\n ok(m.minSide >= 44, 'every button is at least 44 by 44');\n ok(!m.overflowsBar, 'no button pokes out of the toolbar');\n ok(m.perRow.join() === '6,2', 'wraps to rows of 6 and 2');\n // a tap 40px right and 20px down from a button's top-left lands inside its 44px box but outside a 28px one\n const first = await p.evaluate(() => { const r = document.querySelector('.toolbar button').getBoundingClientRect(); return { x: r.x, y: r.y }; });\n await p.touchscreen.tap(first.x + 40, first.y + 20);\n ok((await p.evaluate(() => window.clicks)).join() === 'Bold', 'tap near the edge of the 44px button activates it');\n return m;\n}\n\nfor (const [name, type] of [['chromium', chromium], ['firefox', firefox], ['webkit', webkit]]) {\n const b = await type.launch();\n const f = await fineTest(b);\n const c = await coarseTest(b, 'shipped');\n console.log(name, 'fine:', JSON.stringify({ size: f.size, perRow: f.perRow }), 'coarse:', JSON.stringify({ size: c.size, perRow: c.perRow }));\n const res = {};\n for (const v of ['noCoarseRule', 'noWrap']) { try { await coarseTest(b, v); res[v] = 'STILL PASSES'; } catch (e) { res[v] = 'fails: ' + e.message; } }\n console.log(' without the rule:', JSON.stringify(res));\n // 14 buttons: clone the first button until there are 14, then count buttons per row at 44px\n const p14 = await open(b, 'shipped', true);\n await p14.evaluate(() => { const bar = document.querySelector('.toolbar'); while (bar.children.length < 14) bar.append(bar.firstElementChild.cloneNode(true)); });\n const m14 = await layout(p14);\n ok(m14.perRow.join() === '6,6,2', '14 buttons wrap to rows of 6, 6 and 2');\n console.log(' 14 buttons, coarse pointer, buttons per row:', JSON.stringify(m14.perRow));\n await b.close();\n}\n\nRun in the Playwright 1.48.2 container, it prints the same values for all three engines (repeated runs print the same thing):\n\nchromium fine: {\"size\":[28,28],\"perRow\":[8]} coarse: {\"size\":[44,44],\"perRow\":[6,2]}\n without the rule: {\"noCoarseRule\":\"fails: ASSERT FAILED: every button is at least 44 by 44\",\"noWrap\":\"fails: ASSERT FAILED: every button is at least 44 by 44\"}\n 14 buttons, coarse pointer, buttons per row: [6,6,2]\nfirefox fine: {\"size\":[28,28],\"perRow\":[8]} coarse: {\"size\":[44,44],\"perRow\":[6,2]}\n without the rule: {\"noCoarseRule\":\"fails: ASSERT FAILED: every button is at least 44 by 44\",\"noWrap\":\"fails: ASSERT FAILED: every button is at least 44 by 44\"}\n 14 buttons, coarse pointer, buttons per row: [6,6,2]\nwebkit fine: {\"size\":[28,28],\"perRow\":[8]} coarse: {\"size\":[44,44],\"perRow\":[6,2]}\n without the rule: {\"noCoarseRule\":\"fails: ASSERT FAILED: every button is at least 44 by 44\",\"noWrap\":\"fails: ASSERT FAILED: every button is at least 44 by 44\"}\n 14 buttons, coarse pointer, buttons per row: [6,6,2]\n\nHow to read it: size is the first button's width and height in CSS pixels; perRow lists how many buttons sit on each row, found by grouping the buttons by the rounded top coordinate of their boxes (getBoundingClientRect() returns a box's position and size), so [6,2] means a row of six and a row of two. In layout(), overflowsBar is true if any button's right edge passes the toolbar's right edge and minSide is the smallest side of any button. fineTest opens a normal (mouse-like) page and expects the dense single row; coarseTest opens a touch page and expects the 44px, two-row result plus the tap. without the rule reruns the touch checks against a stylesheet with the coarse block deleted or with flex-wrap switched off, and fails: followed by the assertion message means the checks caught the missing rule (both mutants trip the 44px size assertion first). The last line adds six more buttons (14 in all) and counts rows again.\n\nComplexity and edge cases\n\nThe layout is a single flex pass over eight items. Edge cases to handle:\n\nHybrid devices (laptops or tablets that have both a touchscreen and a mouse or trackpad): pointer reports only the primary input (see the MDN definition above), so a touch laptop whose primary input is the trackpad keeps the dense toolbar; any-pointer is true if any attached input matches, so switch to it if that matters.\nLabels: the buttons carry aria-label so the accessible name does not depend on the glyph; resizing them does not change it.\nPage height: wrapping makes the toolbar taller on touch, which pushes content below it down; account for that in the surrounding layout.\nZoom and user font size: the sizes are in CSS pixels, so they track browser zoom. If the buttons are sized in rem instead, they also track the user's default font size.\n\nRunning the code\n\nmkdir demo && cd demo\necho '{\"type\":\"module\"}' > package.json\nnpm i playwright@1.48.2 && npx playwright install chromium firefox webkit\nsave the three listings above as index.html, toolbar.css and check.mjs, then:\nnode check.mjs"}},{"@type":"Question","name":"Multiple product teams are running experiments that could conflict (UX changes in one flow affecting another). Propose a coordination mechanism to keep experiments aligned with company objectives, avoid user confusion, and preserve statistical validity. Include recommended tooling (experiment registry, feature-flag tracking), processes (registration, review), and governance to resolve conflicts.","acceptedAnswer":{"@type":"Answer","text":"Direct answer\n\nThe coordination mechanism that actually works is a lightweight, mandatory registration step (an experiment registry) paired with automated overlap detection, not a heavyweight approval committee: every team registers an experiment before launch with its affected flows and user segments, a shared feature-flag system (a code-level switch that turns a feature on or off for a subset of users without a new deploy) checks for overlap automatically, and a small governance group only gets involved when a real conflict is flagged, resolving it against pre-agreed priority rules that trace back to company objectives rather than case-by-case politics. Done well, this keeps experiments aligned with company objectives, avoids user confusion from two conflicting changes hitting the same screen, and preserves statistical validity by catching interaction effects before they silently corrupt both results.\n\nStructured elaboration\n\nTooling.\nExperiment registry: a single source of truth recording owner, hypothesis, primary and guardrail metrics (a guardrail metric is one you are not trying to improve but must not let get worse), affected flows or UI components, target segments, and start/end dates for every live experiment.\nFeature-flag platform, wired to the registry so every flag maps back to a registered experiment, with a kill switch for fast rollback.\nOverlap detection: automated checks that flag when two experiments target overlapping user segments or touch the same component, before launch rather than after a confusing result shows up.\n\nProcess.\n1. Registration, a set number of business days before launch, including the affected flows and any known interaction risks.\n2. Automated review: the registry flags segment or component overlap with any other active experiment.\n3. Human review only on a flag: a lightweight weekly sync (design, product, data) resolves real conflicts; if nothing is flagged, no meeting is needed.\n4. Post-launch log: outcomes and any discovered interaction effects get written back to the registry so the next overlapping pair has precedent to work from.\n\nGovernance for resolving conflicts. Pre-agree the priority order before a conflict happens, for example company OKRs first, then user-safety or legal constraints, then UX coherence, then incremental exposure (shrink the holdout on the lower-priority experiment rather than canceling it outright). Give the governance group authority to require orthogonal assignment (a different hash salt per experiment, so which arm a user lands in for one carries no information about which arm they land in for the other, paired with an analysis that actually reads the overlap cell instead of ignoring it) or a staggered rollout when a real conflict is found, and set a clear escalation path with an automatic kill if a guardrail metric moves past an agreed threshold.\n\nWorked example\n\nTeam A tests a new checkout button placement, targeting 50% of traffic. Team B, unaware of Team A's test, independently tests a redesigned cart page that changes the same page's layout, also targeting 50% of traffic. Two very different things can happen from here, and only one of them is a validity problem; telling them apart is most of what the registry is for.\n\nCase 1: the assignments are statistically independent (different hash salts, so which arm a user gets in A tells you nothing about which arm they get in B). The expected overlap, users on the treatment side of both at once, is , so roughly 25% of all traffic sees both changes simultaneously. That overlap on its own does not bias either read: each team is still comparing its treatment against its control over the same mix of the other experiment's arms, so both main effects remain estimable. What it costs is the interaction (the two changes together performing differently than either alone), which gets averaged into both headline numbers instead of being reported, and some power, because the other experiment's effect adds variance to every metric.\n\nCase 2: the assignments are correlated, which looks identical on a dashboard and is the real killer. If both teams bucket users with the same hash of the user id and the same salt, everyone in A's treatment is also in B's treatment, the two changes are perfectly confounded, and neither team can attribute its result to its own change at all. No number inside either experiment reveals this; only something that knows both are live on the same surface does.\n\nThat is what the registry buys, and it is why the choice it unlocks is about assignment and analysis together, not about whether to randomize independently: force the two to be orthogonal (independent salts, verified, so case 2 cannot happen) and analyze the overlap as an explicit 2x2 with a \"both variants\" cell, which turns the hidden interaction from case 1 into something you can read and size; or stagger them so only one runs at a time on the shared surface, which buys a clean single-factor read at the cost of calendar time. Either way the decision is made before launch, rather than discovering the confound (an outside factor that muddles which change actually caused an observed effect) after the fact from a confusing result.\n\nTrade-offs and pitfalls\n\nA registry that requires heavyweight sign-off for every experiment slows delivery enough that teams start quietly bypassing it, which defeats the purpose; the goal is automated detection with human review reserved for actual conflicts, not a gate on every launch. Under-specifying what counts as \"overlap\" (only flagging identical experiments rather than shared components or segments) misses the interaction effects that matter most. And a single centralized governance board can become a bottleneck at scale; a workable version delegates most decisions to the automated rules and only escalates genuinely ambiguous priority conflicts."}},{"@type":"Question","name":"You're designing a component that has to survive real content: variable-length text, user-uploaded images of unpredictable size, and a live data feed that might return nothing or far too much. Walk through how you'd build and stress-test it in Figma so it holds up across breakpoints and platforms instead of breaking the moment the content isn't the tidy placeholder you designed with.","acceptedAnswer":{"@type":"Answer","text":"Direct answer\n\nI'd build the component using Auto Layout (Figma's system for making a frame automatically resize, space, and align its children as content changes, instead of every layer being sized and placed by hand) so its size is driven by whatever content actually fills it, rather than a fixed frame sized to look right with the one tidy example I designed with, then deliberately stress-test it against a spread of realistic and adversarial content, empty, minimal, typical, and maximal, instead of only the happy path, so the breakage shows up in the design file rather than in production.\n\nBuilding for variability\n\nSizing: give every layer the right Auto Layout sizing mode for what it is: \"hug\" for text that should grow or shrink with its content, \"fill\" for anything that should stretch to its container, \"fixed\" only where a size is genuinely constant. The component's overall height or width becomes a function of its actual content, not an assumption baked into the frame.\nText: decide a max-line or truncation rule up front (e.g. a 2-line clamp with an ellipsis) instead of leaving text unconstrained, and deliberately decide what a very short title (one word) should look like too, not just what a very long one does.\nImages: use a fixed-aspect-ratio container with a defined fill/crop behavior, so user-uploaded images of arbitrary source dimensions don't distort the layout or leave visible gaps. Decide and document what \"no image at all\" looks like as its own state (the layout reclaims that space, or shows a placeholder), rather than treating it as a smaller version of \"has an image.\"\nLive data feed: design three states beyond the happy path explicitly: empty (real empty-state guidance, not a blank frame), sparse (one or two items, where a grid can otherwise look broken or lonely), and overflow (far more items than fit, needing pagination, infinite scroll, or a \"show more\" affordance you actually design rather than assume the tool or engineering will handle).\n\nStress-testing method\n\nDon't validate only against your own placeholder copy. Pull or fabricate a genuinely adversarial content set, the longest real title you can find, a small or oddly-cropped real user-uploaded image, an empty feed and an overloaded one, and drop it into the component across every breakpoint and platform frame, checking specifically for text overlap, image distortion, broken row alignment between sibling cards, and inconsistent row heights in a grid. Re-run the same content set per platform, not just per breakpoint width, since comfortable line length and touch-target sizing differ by platform convention even at similar widths.\n\nWorked example\n\nTake a feed card. Content test set: a 1-character title through a genuinely long real headline (~120 characters); images including a square avatar, a wide banner photo, and a \"no image\" case; feed sizes of zero items, one item, exactly enough to fill one row, and far more than fit. Result of the stress test: the long headline without a clamp pushes that card's height noticeably past its row neighbors, breaking alignment, fixed by adding a 2-line clamp. The wide banner photo inside a container built for square avatars either stretches or leaves visible letterboxing, fixed by standardizing the image container's aspect ratio and fill behavior regardless of the source image's shape. The \"no image\" case leaves a blank gray box the same size as a populated one, reading as a broken loading state rather than an intentional one, fixed with a component property that removes the image slot and lets the text stack take that space, used only for the genuine \"no image\" case (a separate, distinct treatment covers \"image still loading\"). A one-item feed on a grid built for three columns leaves two glaring empty cells, fixed by having the grid collapse to a left-aligned row of only the populated cards instead of reserving ghost slots.\n\nTrade-offs and pitfalls\n\nDesigning purely for the absolute worst case (every field maxed out) makes the typical, common case look padded and awkward; the goal is testing the full range, then choosing defaults, like a 2-line clamp, that degrade gracefully rather than breaking at the extreme or looking wrong in the common case.\nTreating \"no image\" as just a smaller version of \"has an image\" instead of its own layout state is the most common source of a card that reads as subtly broken rather than intentionally different.\nEmpty and overflow feed states are often left undesigned because they feel like edge cases, but the choice between pagination, infinite scroll, and a \"show more\" button is a real UX decision. Leaving it undesigned just hands that decision to whoever implements it, with no design intent behind it.\nStress-testing with genuinely messy content matters because generic filler text is uniform in a way that hides exactly the length-variance problems the test exists to catch."}},{"@type":"Question","name":"How do you decide between moderated, unmoderated, and in-person usability testing depending on where you are in the process, say testing a rough prototype early versus validating a near-final flow before launch?","acceptedAnswer":{"@type":"Answer","text":"Direct answer\nMatch the method to where you are in the process. A rough early prototype needs a moderated session, because you need to probe confusion in real time and understand the intent behind a hesitation. A near-final flow before launch needs a larger, less-guided sample (unmoderated, or in-person only if the interaction demands it), because a moderator's presence changes behavior in ways that matter more once you are checking whether people succeed unaided, not diagnosing why they might struggle.\n\nStructured elaboration\n\nMethod versus process stage\n\nMethod | Best process stage | What it tells you | What it costs\n\nModerated (remote or in person) | Early, rough prototype | Rich \"why,\" real-time probing of confusion | Small sample, one participant at a time\nUnmoderated remote | Near-final flow, pre-launch | Larger sample, unaided success or failure at scale | Little to no insight into why something failed\nIn-person specifically | Either stage, when the interaction itself needs it | Everything moderated remote gives you, plus environment control | Most time and logistics cost of the three\n\nWhy the shift happens\nEarly on, the open question is \"where and why do people get confused,\" which needs a human watching and asking follow-ups in the moment. Close to launch, the open question shifts to \"will people get through this without help,\" which a moderator's presence would quietly distort, since people behave differently when someone is coaching or watching closely versus using the product alone.\n\nWhen in-person specifically earns its extra cost\nReach for in-person over remote moderated only when the interaction itself is hard to observe over a screen share, for example something involving a physical device or an environment that matters to the task. Otherwise remote moderated gets you the same real-time probing at lower cost.\n\nThe mechanics of running any one of these methods well (how to structure tasks, avoid leading questions, choose sample size) is a deeper methodology question in its own right; the point at this stage of the process is picking the right method for where you are, not executing it.\n\nWorked example\nA settings flow redesign: at the rough, low-fidelity prototype stage, the team runs five or six moderated sessions, because the goal is to see exactly where people pause and ask why, not just whether they eventually finish. Two weeks before launch, testing the same flow now near-final, the team switches to a larger unmoderated round focused on task completion, because the goal has shifted from diagnosing confusion to confirming the flow holds up without anyone coaching the user through it.\n\nTrade-offs and pitfalls\nUsing unmoderated testing too early wastes a rough prototype's biggest value (real-time probing) on a method built for scale, not diagnosis.\nUsing moderated testing right before launch introduces facilitator effects into a decision that should reflect unaided, real-world use.\nChoosing a method by budget or convenience alone, instead of by what question you are actually answering at that point in the process, is the most common wrong turn here."}},{"@type":"Question","name":"Think back to a time you were stuck on something unfamiliar long enough that it became a problem. How did you work out what was actually blocking you, what did you try, and when did you bring anyone else in?","acceptedAnswer":{"@type":"Answer","text":"Direct answer\n\nWhen I get stuck on something unfamiliar, the first move is to turn \"I'm stuck\" into a precise, testable question: not \"why is this broken\" but \"which of these three things is actually causing it.\" From there I narrow down by testing pieces in isolation, and I bring someone else in on a clock, not a feeling, so I don't burn days flailing but also don't ask before I've done the cheap, obvious checks myself.\n\nStructured elaboration\n\nName the actual blocker. A symptom (\"the report is wrong\") is not a mechanism (\"the join is dropping rows with a null key\"). I spend the first few minutes just narrowing the symptom to something specific enough to test.\nDecompose into testable parts. Instead of staring at the whole system, I split it into pieces I can check independently: does the input look right, does step one behave as expected in isolation, does step two. This is basically bisection: cut the unknown space in half each time instead of guessing at the whole thing.\nOrder of sources, cheapest and most verifiable first. Documentation and existing working examples first (they're free and don't cost anyone else time), then searching for others who hit the same thing, then a specific person, in roughly that order, because each step up costs more of someone else's time and I want to have exhausted the cheap checks first.\nSet a time-box before I start, not after I'm frustrated. I decide up front roughly how long I'll self-serve before escalating, so the decision to ask for help is made by a clock I set with a clear head, not by how annoyed I am two hours in.\nValidate before trusting the fix. Once something looks fixed, I reproduce the original failure once more to confirm I actually understood the cause, not just made a symptom go away, and I check for side effects on the parts I didn't touch.\nHand the finding back. I write down what the actual cause was and how I found it, even briefly, so the next person who hits this doesn't have to repeat the same search from zero.\n\nWorked example\n\nI once inherited a background job that was silently dropping about one in ten messages after a queue migration, with no errors in the logs anywhere. I narrowed the symptom first: not \"messages are lost,\" but \"messages with a specific field are lost,\" which I found by sending a batch of controlled test messages and diffing what came out against what went in. That let me bisect the pipeline: I checked the message right after it entered the queue (present, correct), then right after the first processing stage (some already missing), so the fault was isolated to that one stage. I gave myself until end of day to find the mechanism before pulling in the engineer who'd built the original queue setup. About three hours in, I found it: the new queue silently truncated messages over a certain size, and the dropped ten percent were exactly the ones with a longer optional field. I fixed the truncation limit, then deliberately reran the same failing batch to confirm the fix actually worked rather than just assuming it, and wrote a short note in our team's incident log explaining the cause so nobody else would spend three hours rediscovering it.\n\nTrade-offs and pitfalls\n\nThe two failure modes on either side of this are trying random fixes without narrowing the problem first (which burns time and rarely teaches you anything), and asking for help too early, before doing the cheap checks yourself, which both wastes someone else's time and doesn't build your own ability to do this next time. A fixed time-box protects against both: it stops you from asking too soon out of impatience and from silently struggling too long out of pride. The other real trap is trusting a fix that merely made the symptom disappear once, without reproducing the original failure to confirm you actually understood the cause."}},{"@type":"Question","name":"Explain the pros and cons of handing off fully annotated high-fidelity mockups versus handing off lighter component-based specs with tokens. In which team and product situations would you choose each approach, and how does that choice affect developer velocity and long-term maintainability?","acceptedAnswer":{"@type":"Answer","text":"Direct answer\n\nFully annotated high-fidelity mockups spell out every value directly on the artboard; lighter component-based specs describe intent through named design tokens (shared values like a specific spacing or color, referenced by name rather than a raw number) and existing components, letting the code enforce the actual pixel values. Choose the heavy version for one-off, high-stakes, or unfamiliar-team situations; choose the lighter version once a mature design system exists, because it moves faster and stays correct as the system evolves, while heavy mockups tend to go stale the moment the underlying values change.\n\nStructured elaboration\n\nFully annotated mockups | Lighter token-based specs\n\nPros | Unambiguous; no guessing required; works well with a new or junior team; self-contained | Fast to produce; updates automatically when a token changes; forces genuine reuse\nCons | Slow to produce and to update; the numbers were typed once and frozen, so they drift the moment a token changes; can encourage one-off values instead of reuse | Assumes a mature, well-documented system and a team fluent in it; can leave real gaps for a team without one\nBest for | A team without a design system yet, an unfamiliar or critical flow like checkout, an external vendor handoff, compliance-sensitive UI | A mature product with an established system, strong design-engineering trust, fast-iterating features\n\nWorked example\n\nThe same card component, spec'd two ways.\n\nHeavy mockup: \"Padding: 24px top and bottom, 16px left and right. Corner radius 8px. Shadow: 0px horizontal offset, 2px vertical offset, 4px blur, black at 8% opacity.\" Every number hand-typed on the artboard. Note that the heavy approach only delivers the \"unambiguous\" column of the table above if every value is genuinely a number: a phrase like \"a soft shadow with light blur\" on an annotated mockup reintroduces exactly the guessing the heavy approach was chosen to eliminate, while costing all of its maintenance overhead.\n\nLighter spec: \"Padding uses the large spacing token vertically and the medium spacing token horizontally; corners use the medium radius token; the shadow uses the small elevation token.\" The two padding values have to stay two tokens. Collapsing them into a single \"large spacing token for padding\" is the most common way a token spec silently loses information the heavy mockup carried, and it is worse than the heavy version, because the engineer implements one uniform padding, it looks deliberate, and nothing in the spec records that the asymmetry was ever intended.\n\nNow suppose the design system's large-spacing token changes value across the whole product, for a slightly denser layout. The token-based spec updates automatically everywhere it's referenced, because the token itself changed. The fully annotated mockup for this one card still shows its original hand-typed number, and nobody knows to revisit it unless someone remembers that specific artboard exists, so it silently drifts out of sync with the rest of the product.\n\nOn velocity specifically: for a single, one-off screen the two are close enough that velocity doesn't decide anything. The heavy version of the card above hand-types seven values (two paddings, a radius, and four shadow parameters) where the lighter one looks up four token names, and that gap is seconds of typing, not a difference in how the team works. The heavy approach starts costing more only once there are many components and many updates over time, which is the long-term maintainability axis, not the day-one speed of producing the first spec. Any count quoted in this comparison is worth re-checking against the example it describes, because a tidy symmetric-sounding count flatters the heavy spec's day-one cost and argues the question on the one axis where the two approaches are nearly tied anyway.\n\nTrade-offs and pitfalls\n\nChoosing the lighter approach before a real token system and shared library exist just recreates the ambiguity of having no spec at all. Sticking with the heavy approach forever, even once a design system matures, means the system's real benefit, cheap propagation of change, never actually gets used. And the two aren't mutually exclusive: a mature system can still add a fully annotated spec for the one truly novel, high-stakes screen, like a compliance disclosure flow, while using lightweight token specs everywhere else. The check that makes the lighter approach safe is a lossless-ness test: read the token spec and ask whether every distinction the heavy version would have made, asymmetric padding, a different radius on one corner, a state-specific elevation, still survives as a named token. Wherever it doesn't, the answer is to add the missing token reference, not to fall back to hand-typed numbers, because a token spec that quietly drops a distinction is the one failure mode that neither approach's pros column protects against."}}]}

Amazon UI Designer (Mid-Level) Interview Preparation Guide

UI Designer
Amazon
Mid Level
7 rounds
Updated 6/23/2026

Amazon's UI Designer interview process for mid-level candidates typically involves a combination of portfolio review, design problem-solving, system design thinking, technical collaboration assessments, and behavioral evaluation. The process is designed to assess visual design skills, design systems knowledge, prototyping ability, technical communication, and cultural alignment with Amazon's leadership principles.

Interview Rounds

1

Recruiter Screening

2

Design Case Study Phone Interview

3

Onsite Round 1: Portfolio and Visual Design Fundamentals

4

Onsite Round 2: Design Systems and Scalable Design

5

Onsite Round 3: Technical Design and Developer Collaboration

6

Onsite Round 4: Product Sense and Cross-Functional Impact

7

Onsite Round 5: Behavioral and Culture Fit with Hiring Manager

Frequently Asked UI Designer Interview Questions

Visual Design Fundamentals: Typography, Color, and BrandEasyTechnical
66 practiced

An e-commerce home goods brand needs a clear photography direction. Walk through the choices you'd make around lighting, composition, staging, and color treatment, how hero shots should differ from plain product shots, and how those choices reinforce the brand's personality while keeping the page scannable.

Design Critique, Iteration, and Decision RationaleHardTechnical
24 practiced

Engineering leads are reluctant to give up a sprint for usability fixes. Walk through how you'd persuade them using evidence and trade-offs: how you'd quantify the ROI, what artifacts you'd bring, and how you'd address their concerns about scope, risk, and technical debt.

Design Systems and Component LibrariesEasyTechnical
42 practiced

Your Button component ships in several sizes, tones, and interaction states, and different teams keep inventing their own names for the variants in both Figma and code. Design a naming convention that covers the component itself, its variants, and its file/folder layout. Walk through what your naming pattern actually looks like end to end, and explain how it improves discoverability and consistency across design and code.

Accessibility and Inclusive DesignMediumTechnical
77 practiced

Plan a usability test session focused on onboarding for participants with vision, motor, or cognitive impairments. Include recruitment criteria, accessibility accommodations, consent and ethics, and task scenarios.

Frontend Fundamentals: HTML, CSS, and Responsive StylingEasyTechnical
116 practiced

What size should touch targets be on mobile, and why? How would you keep a dense desktop toolbar usable on touch screens without redesigning it?

Design Metrics and Impact MeasurementHardSystem Design
20 practiced

Multiple product teams are running experiments that could conflict (UX changes in one flow affecting another). Propose a coordination mechanism to keep experiments aligned with company objectives, avoid user confusion, and preserve statistical validity. Include recommended tooling (experiment registry, feature-flag tracking), processes (registration, review), and governance to resolve conflicts.

Design Tools and AI-Assisted WorkflowsHardTechnical
34 practiced

You're designing a component that has to survive real content: variable-length text, user-uploaded images of unpredictable size, and a live data feed that might return nothing or far too much. Walk through how you'd build and stress-test it in Figma so it holds up across breakpoints and platforms instead of breaking the moment the content isn't the tidy placeholder you designed with.

Design Thinking and the End-to-End Design ProcessMediumTechnical
42 practiced

How do you decide between moderated, unmoderated, and in-person usability testing depending on where you are in the process, say testing a rough prototype early versus validating a near-final flow before launch?

Growth Mindset and Learning AgilityMediumBehavioral
40 practiced

Think back to a time you were stuck on something unfamiliar long enough that it became a problem. How did you work out what was actually blocking you, what did you try, and when did you bring anyone else in?

Design Handoff and Developer CollaborationMediumTechnical
61 practiced

Explain the pros and cons of handing off fully annotated high-fidelity mockups versus handing off lighter component-based specs with tokens. In which team and product situations would you choose each approach, and how does that choice affect developer velocity and long-term maintainability?

Want to create your own tailored preparation guide using our deep research?

Get Started for Free

Interview-Ready Courses

Visual-first, interactive, structured learning paths

Browse UI Designer jobs

AI-enriched listings across hundreds of company career pages

Explore Jobs