Cross-Browser and Cross-Platform Testing Questions
Verifying that a product behaves consistently across browsers, browser versions, operating systems, devices, and screen sizes. Covers building a browser and device coverage matrix, prioritizing platforms by real usage telemetry, deciding what to automate versus leave to manual exploratory checks, risk-based sampling between PR, nightly, and pre-release runs with escalation to the full matrix, wiring cross-browser suites into CI with parallel and selective runs, cloud grids versus self-hosted labs (Selenium Grid in Docker, on-prem device labs) versus emulators and simulators, real devices and native mobile apps, sizing infrastructure for large numbers of parallel sessions, running one suite across Chromium, Firefox, and WebKit (Playwright projects, driver factories driven by environment config), headless versus headed differences, feature detection versus user-agent sniffing, browser-specific defects (Safari timing and event ordering, cookie policy and third-party embeds, font and DPI rendering, input and event differences) and how to debug them on physical devices, and OS-level differences (filesystem, encoding, endianness, numeric behavior). Excludes framework and Page Object design, flaky-test management, locator and wait mechanics, and accessibility or visual-diff tooling, which are covered elsewhere.
Your UI tests are green on Chrome, but several fail on Firefox and Safari. What kinds of rendering, input and event differences between browsers usually cause this, and what quick mitigation would you apply to each?
Sample Answer
Direct answer
When a suite is green on Chrome and red on Firefox and Safari, the cause is usually one of four things: the test hard-codes an assumption that only Chromium satisfies (an exact pixel value, a default font, a focus side effect), the engines really do render or handle input differently, the app uses a feature one engine lacks, or the environment differs. Triage each failure into one of those buckets before touching code, because the mitigation differs: loosen or rewrite the test, normalise the styling, add a fallback in the app, or pin the environment.
Differences that cause failures, and the quick mitigation for each
| Kind | What differs | Quick mitigation |
|---|---|---|
| Rendering: default styles | Each engine ships its own default stylesheet for form controls and body text | Set font: inherit on form controls and an explicit font stack on the page; self-host web fonts so all engines load the same file |
| Rendering: sub-pixel layout | Engines round fractional sizes differently, so an exact === on a measured size fails | Assert with a tolerance, or assert on structure (which row an item landed in) instead of raw pixels |
| Input: typing into form fields | Engines disagree on what a field accepts and what its value reads while text is invalid | Drive input the way the user does, then assert on the validated result or the error message, not on the raw value string |
| Input: native widgets | Date pickers, select popups and file choosers are drawn by the browser, not the page | Set values through the API (fill, selectOption, setInputFiles) rather than clicking inside native UI |
| Events: focus side effects | Whether a click moves focus onto a button varies by browser and OS | Do not assert on focus after a click; focus the element explicitly when the test needs it |
| Events: timing | Animations, transitions and layout settle at slightly different times | Wait for a condition (an element state, a response) instead of a fixed delay |
Evidence from a run
Run in the Playwright 1.48.2 image (Chromium 130.0, Firefox 131.0, WebKit 18.0, on Linux), a probe page printed:
- Sub-pixel width. A sub-pixel size is a fractional CSS pixel such as 33.33px; each engine stores layout sizes on its own fine grid, so the leftover fraction lands differently. Three
flex: 1items in a 100px container reported widths of33.328125for each item in Chromium and WebKit, and33.33332824707031,33.333343505859375,33.33332824707031in Firefox. 100px cannot be split into three equal whole parts, and the engines resolve the remainder differently: Chromium and WebKit report 33.328125, which is exactly 2133/64 (a whole number of 1/64-pixel steps, so three items total 99.984375), while Firefox reports numbers near 33.3333. A test assertingwidth === 33.328125passes on two engines and fails on Firefox.expect(width).toBeCloseTo(33.33, 1)(within 0.05) passes on all three. - Number input.
12e3is a valid number: theeintroduces an exponent, so it means 12 x 10^3 = 12000.validity.badInputis true when the user has typed something the browser cannot convert to a number, and an invalid number field reads as an emptyvalue(""). After clicking an<input type="number">and typing12e3abc, Chromium reportedvalue"12e3"withbadInputfalse (in the same run, typing onlyabcleft Chromium's value""withbadInputfalse, so Chromium discards the letters that cannot be part of a number as they are typed), while Firefox and WebKit reportedvalue""withbadInputtrue (they keep the text in the box, so it is invalid). That is why""and"12e3"differ: one means no usable number, the other a number the engine accepted. A test that reads.valueafter typing sees different strings. Mitigation: type only valid digits when the test is about something else, and test invalid input in a dedicated case that asserts the visible validation message. - Event order on a button click.
pointerdownis the press of a mouse, finger or pen,mousedownis the older mouse-only press event, andfocusfires when the element becomes the target of keyboard input. All three engines loggedpointerdown, mousedown, focus, pointerup, mouseup, click, so ordering of mouse events is not the usual culprit. - Focus on click. MDN documents that whether clicking a button focuses it varies by browser and OS, and that Safari does not focus it, by design. Playwright's Linux WebKit focused the button in the run above, which shows why a green WebKit project does not prove macOS Safari behaviour. A test that asserts
toBeFocused()after a click is a bad test; a keyboard-accessibility test should press Tab or callfocus()itself.
Triage order
- Open the failure's trace (Playwright's recorded timeline of actions, screenshots and page snapshots, opened with
npx playwright show-trace) or screenshot on the failing browser and read the assertion: is it an exact number, a string from a field, or a missing element? - Re-run that single test on that browser alone (
--project=firefoxruns only the Playwright config project named firefox, one browser setup) to rule out interference from parallel tests. - Classify: test assumption (loosen or rewrite), real cross-engine difference (normalise in CSS or code), missing feature (feature detection plus fallback), or environment (fonts, locale, time zone).
- Fix at the cheapest level. Normalising CSS fixes real users as well as the test, so prefer it when the difference is visible to users.
Pitfalls
- Do not skip a test on the failing browser to make CI green without a ticket; the skip hides a possible real bug.
- Playwright WebKit is not iOS Safari. Mobile-only failures need a real device or a cloud device.
Login works in Chrome, but in our automated runs an embedded third-party widget and the sign-in flow break on Safari and Firefox. How can differences in cookie policy between browsers cause this, how would you write automated tests that validate cookie behaviour across browsers, and how would you configure test environments to cover both legacy and modern policies?
Sample Answer
Direct answer
Browsers disagree on when a cookie set by an embedded (third-party) page is allowed to come back. Chrome's default does not block third-party cookies, but a cookie only travels in a cross-site request if it was set with SameSite=None; Secure; a cookie with no SameSite is treated as Lax and held back. So "Chrome accepts third-party cookies" and "a cookie is missing in Chromium" are both true: the missing cookies lack the required attributes. Safari's tracking prevention (Intelligent Tracking Prevention, ITP, Safari's built-in anti-tracking rules) and Firefox's Total Cookie Protection (a separate cookie jar for each top-level site) restrict or partition third-party cookies by default, even when the attributes are right. So a widget that "just works" in Chrome can lose its session in Safari and Firefox. Test it by encoding a per-browser expectation table and running it in every engine, with browser profiles that match the real privacy defaults, and confirm Safari on real Safari because Playwright's WebKit is not Safari.
Terms: a third-party cookie is one set or sent by a site other than the page in the address bar (for example a widget in an <iframe>). Site means scheme plus registrable domain (the part of a host name that can be registered, such as example.com or example.co.uk). SameSite is a cookie attribute controlling cross-site sending. Partitioned storage (CHIPS, Cookies Having Independent Partitioned State) gives each top-level site its own copy of an embed's cookie. The top-level page is the one in the address bar; an embed is a page inside it. A top-level redirect sends the whole tab to the other site's address, so that site is first-party for the visit. Single sign-on is one login that works on several sites. Loopback is the address 127.0.0.1, the machine itself. The Storage Access API lets an embedded frame ask the browser for its normal, unpartitioned cookies, usually after a user gesture.
What each browser does (documented)
| Browser | Documented default (MDN, third-party cookies guide) |
|---|---|
| Chrome | Does not block third-party cookies by default, only in Incognito or when the user sets blocking |
| Firefox | Total Cookie Protection when Enhanced Tracking Protection is on (the default): a separate cookie jar per site |
| Safari | Tracking prevention (ITP) gives similar third-party protections by default; per-frame access only through the Storage Access API |
Also from MDN: SameSite=None must be paired with Secure; Lax cookies are not sent for fetch(), subresources or <iframe> navigations; when SameSite is omitted, some browsers (Chromium) default to Lax and the rest vary, so always set it explicitly.
A cross-browser test you can run
Site A and site B embed a widget from a third host. The widget's response sets a cookie; the widget then calls its own server, which reports whether the cookie came back. Four attribute variants, five browser profiles. The hostnames are mapped to loopback with --add-host, and an in-process HTTPS server with a throwaway certificate makes Secure real.
// cookies.mjs
import fs from 'node:fs';
import https from 'node:https';
import { execFileSync } from 'node:child_process';
import { chromium, firefox, webkit } from 'playwright';
// Self-signed certificate so Secure cookies can be tested over real HTTPS.
execFileSync('openssl', ['req', '-x509', '-newkey', 'rsa:2048', '-nodes', '-keyout', 'key.pem', '-out', 'cert.pem',
'-subj', '/CN=localhost', '-days', '1', '-addext', 'subjectAltName=DNS:app-a.test,DNS:app-b.test,DNS:widget.test'], { stdio: 'ignore' });
const tls = { key: fs.readFileSync('key.pem'), cert: fs.readFileSync('cert.pem') };
const VARIANTS = {
'None; Secure': 'sid=1; SameSite=None; Secure; HttpOnly',
'None (no Secure)': 'sid=1; SameSite=None',
'no SameSite': 'sid=1',
'None; Secure; Partitioned': 'sid=1; SameSite=None; Secure; Partitioned',
};
const WIDGET = 'https://widget.test:4100'; // the third party
const SITE_A = 'https://app-a.test:4100'; // first embedding site
const SITE_B = 'https://app-b.test:4100'; // a different embedding site
// One server answers for all three hosts. Embedding pages carry the widget in an iframe.
https.createServer(tls, (req, res) => {
const u = new URL(req.url, 'https://x');
const host = req.headers.host.split(':')[0];
if (u.pathname === '/embed') {
res.setHeader('content-type', 'text/html');
res.end(`<iframe src="${WIDGET}/widget?v=${u.searchParams.get('v')}&set=${u.searchParams.get('set')}"></iframe>`);
} else if (u.pathname === '/widget') {
if (u.searchParams.get('set') === '1') res.setHeader('set-cookie', VARIANTS[u.searchParams.get('v')]);
res.setHeader('content-type', 'text/html');
res.end(`<script>fetch('/whoami', { credentials: 'include' }).then(r => r.text()).then(t => { window.result = t; });</script>`);
} else {
res.end(/(^|;\s*)sid=1/.test(req.headers.cookie || '') ? 'sent' : 'missing');
}
}).listen(4100);
const PROFILES = {
'chromium': async () => { const b = await chromium.launch(); return [await b.newContext({ ignoreHTTPSErrors: true }), () => b.close()]; },
'chromium+phaseout': async () => { const b = await chromium.launch({ args: ['--test-third-party-cookie-phaseout'] }); return [await b.newContext({ ignoreHTTPSErrors: true }), () => b.close()]; },
'chromium+legacy': async () => { const b = await chromium.launch({ args: ['--disable-features=SameSiteByDefaultCookies,CookiesWithoutSameSiteMustBeSecure'] }); return [await b.newContext({ ignoreHTTPSErrors: true }), () => b.close()]; },
'firefox': async () => { const b = await firefox.launch(); return [await b.newContext({ ignoreHTTPSErrors: true }), () => b.close()]; },
'firefox+TCP pref': async () => { const b = await firefox.launch({ firefoxUserPrefs: { 'network.cookie.cookieBehavior': 5 } }); return [await b.newContext({ ignoreHTTPSErrors: true }), () => b.close()]; },
'webkit': async () => { const b = await webkit.launch(); return [await b.newContext({ ignoreHTTPSErrors: true }), () => b.close()]; },
};
async function widgetSees(context, site, v, set) {
const page = await context.newPage();
await page.goto(`${site}/embed?v=${encodeURIComponent(v)}&set=${set}`);
const frame = page.frames().find((f) => f.url().includes('/widget'));
await frame.waitForFunction(() => window.result !== undefined);
const out = await frame.evaluate(() => window.result);
await page.close();
return out;
}
const names = Object.keys(PROFILES);
const row = (label, cells) => console.log(label.padEnd(28) + cells.map((c) => c.padEnd(20)).join(''));
console.log('A. cookie set and read inside one embedding site');
row('cookie attributes', names);
for (const v of Object.keys(VARIANTS)) {
const cells = [];
for (const n of names) { const [ctx, close] = await PROFILES[n](); cells.push(await widgetSees(ctx, SITE_A, v, 1)); await close(); }
row(v, cells);
}
console.log('\nB. cookie set under site A, then read when the widget is embedded on site B');
row('cookie attributes', names);
for (const v of ['None; Secure', 'None; Secure; Partitioned']) {
const cells = [];
for (const n of names) {
const [ctx, close] = await PROFILES[n]();
await widgetSees(ctx, SITE_A, v, 1);
cells.push(await widgetSees(ctx, SITE_B, v, 0));
await close();
}
row(v, cells);
}
process.exit(0);
Run in the Playwright 1.48.2 container (Chromium 130, Firefox 131, WebKit 18.0), it printed the same table on each of 10 runs:
A. cookie set and read inside one embedding site
cookie attributes chromium chromium+phaseout chromium+legacy firefox firefox+TCP pref webkit
None; Secure sent sent sent sent sent missing
None (no Secure) missing missing missing missing missing missing
no SameSite missing missing missing sent sent missing
None; Secure; Partitioned sent sent sent sent sent missing
B. cookie set under site A, then read when the widget is embedded on site B
cookie attributes chromium chromium+phaseout chromium+legacy firefox firefox+TCP pref webkit
None; Secure sent sent sent sent missing missing
None; Secure; Partitioned missing missing missing sent missing missing
How to read the tables: each row is a Set-Cookie header the widget sends, each column is a browser profile, and each cell is what the widget's own server received when the widget called it. sent means the cookie came back, so the widget can recognise the signed-in user. missing means the server saw no cookie; in a real product that is the user who looks signed out inside the embed, or a sign-in loop that starts over. Table A sets and reads the cookie under one embedding site. Table B sets it under site A, then loads the widget under site B (the single-sign-on case).
How the script produces them:
app-a.test,app-b.testandwidget.testare three different sites. The--add-host name:127.0.0.1options in the run command point all three at loopback, so one Node server on port 4100 can play every role.Securecookies are only stored and sent over HTTPS, soopensslcreates a throwaway self-signed certificate valid for the three names (subjectAltName), andignoreHTTPSErrors: truelets each browser accept it.- The server has three routes.
/embedreturns a page containing an iframe for the widget./widgetsets the cookie (whenset=1) and runs afetchto/whoamiwithcredentials: 'include'./whoamianswerssentormissingfrom the request'sCookieheader. That fetch is cross-site because the page in the address bar belongs to a different site from the widget. widgetSeesopens the embedding page, finds the widget frame and waits forwindow.resultto exist, so it waits for a condition instead of sleeping.- Two Chromium profiles pass command-line switches;
firefox+TCP prefsetsnetwork.cookie.cookieBehaviorto 5, which Mozilla's cookie service defines asBEHAVIOR_REJECT_TRACKER_AND_PARTITION_FOREIGN(block known trackers, partition other third-party cookies).
What the tables show:
- Chrome-family behaviour matches the docs.
None; Secureis returned;NonewithoutSecureis not (every engine here refused it); a missingSameSiteis not returned in the cross-site iframe, as theLaxdefault predicts. - The
no SameSiterow differs by engine. Firefox returns the cookie (sent); Chromium and this WebKit build do not (missing). A widget that forgetsSameSitetherefore looks fine in Firefox and breaks for Chrome users. A suite that ran only in Firefox would pass; one that ran only in Chromium would fail but would say nothing about how Firefox or Safari treat the same cookie. Running every engine catches the bug each one alone would hide. - Partitioned (CHIPS) cookies isolate per embedding site. In Chromium and in Firefox with the pref from finding 4, the cookie set under site A is not visible under site B (table B,
missing); a plainNone; Securecookie is shared in Chromium (sent), which is exactly what privacy-restricted browsers stop. - Playwright's bundled Firefox does not match retail Firefox. By default the cross-site cookie crossed from A to B (
sent). Launching with the prefnetwork.cookie.cookieBehaviorset to 5 (the value that turns on Total Cookie Protection's partitioning) producesmissing, so the CI profile must set it. - This WebKit build returned no third-party cookie in any variant. That fits Safari's documented default, but it is a Playwright WebKit build, not Safari, so use it as an early alarm and verify on real Safari (macOS and iPhone) before concluding.
- Neither the documented Chrome flag
--test-third-party-cookie-phaseoutnor the legacy-behaviour feature switches changed anything in this Chromium build (thechromium+phaseoutandchromium+legacycolumns equal plainchromium). The block-third-party lane for Chrome therefore needs branded Chrome or Chrome for Testing configured through the vendor's documented setting, not this bundled build.
Turning it into automated tests
- Contract test on the header (no browser): the widget's
Set-Cookiemust containSameSite=None; Secure, andPartitionedif you adopt CHIPS. It fails the build in seconds if someone changes the cookie helper. - Expectation matrix per engine: keep the table above as data (
variant x profile -> sent|missing) and assert it in one test per Playwright project. A browser upgrade that changes policy then fails a named cell instead of a vague login timeout. - End-to-end sign-in test: log in inside the embedded widget, reload the host page, and assert the widget still shows the signed-in state; run with and without third-party cookies. Failures here tell you the session did not survive, which is the production symptom.
- Real Safari lane: the same flow on real Safari (a Mac runner or a device cloud with real iPhones), because ITP details such as the seven-day storage cap are not reproduced by Playwright's WebKit.
Environments for legacy and modern policy
| Lane | Represents | How |
|---|---|---|
| Chromium default | Chrome today (third-party allowed, Lax default) | Playwright Chromium |
| Chrome with blocking | Third-party cookies blocked | branded Chrome with the setting or flag Google documents |
| Firefox default | An older, permissive profile (no Lax default, no partitioning in this build) | Playwright Firefox |
| Firefox with Total Cookie Protection | Retail Firefox | firefoxUserPrefs: {'network.cookie.cookieBehavior': 5} |
| WebKit plus real Safari | Safari tracking prevention | Playwright WebKit as the fast check, real Safari to decide |
| Old browser versions | Pre-SameSite-default behaviour | a cloud grid that offers those versions |
Fixing the product
Ask one question: must a user who signed in through the widget on site A also be signed in on site B? If no, use a partitioned cookie, the smallest change. If yes, use the Storage Access API, or avoid third-party state altogether. Commit to one:
- Session only matters inside each host site:
SameSite=None; Secure; Partitioned. Each embedder gets its own cookie and table B's isolation is the intended behaviour. It is not a universal fix: in table A the Playwright WebKit column returned no cookie for this variant either, so confirm on real Safari, and use the Storage Access API for any browser where the partitioned cookie does not come back. - One login shared across many sites: the Storage Access API.
document.requestStorageAccess()must be called from a user gesture in a secure context; browsers may prompt (Safari prompts for every embed that has not received access, Firefox after a threshold), the grant is stored per top-level and embedded site pair, and an iframe with thesandboxattribute (which removes powers unless tokens give them back) needs theallow-storage-access-by-user-activationtoken. - Avoid the problem when you can: serve the widget from a subdomain of the host site (same site, so it is first-party) or sign in through a top-level redirect or a popup window and hand the widget a token.
Running the code
# save cookies.mjs and a package.json containing {"type":"module"} in one folder, then:
docker run --rm --ulimit core=0 --add-host app-a.test:127.0.0.1 --add-host app-b.test:127.0.0.1 \
--add-host widget.test:127.0.0.1 -v "$PWD":/w -w /w mcr.microsoft.com/playwright:v1.48.2-jammy sh -c \
'npm i --silent playwright@1.48.2 && timeout 300 node cookies.mjs'
A UI test passes locally on Chrome and Firefox but fails intermittently in CI on macOS Safari with a timing-related assertion. Outline a practical debugging plan: which artifacts you would collect, how you would run the failing scenario remotely or with remote devtools, and what Safari-specific diagnostics you would ask developers for.
Sample Answer
Direct answer
Work from the cheapest, most informative check to the most expensive one. First measure how often it fails and under what conditions (artifacts). Second, try to reproduce the failure on the WebKit engine in a Playwright run you control, with tracing on; if it reproduces there, the trace usually names the cause. Third, only if it does not reproduce, reproduce on real Safari on a Mac (a CI machine you can reach, or a cloud grid session) and attach Safari's Web Inspector. Fourth, hand developers a precise list of Safari-specific questions. Do not add a longer sleep or a retry before you know which of the causes below you have, because that hides the defect rather than fixing it.
Two terms. WebKit is the browser engine Safari is built on. Playwright (a test framework that drives real browsers) ships its own WebKit build, taken from WebKit's main branch, and its docs state it does not work with the branded Safari (the real Safari application that Apple ships, as opposed to a build made by Playwright). So a Playwright WebKit run is a close proxy for Safari's engine, and a CI job that drives real Safari on macOS (for example through Selenium) is a different thing. Name which one your CI uses before you start, because it decides which of the steps below you can run.
Step 1: Collect artifacts from the failing CI runs
Before changing anything, gather for several failing and several passing runs. Start with the versions, the failure rate and conditions, and the screenshot or video, because they say first whether the cause is a version change, a timing pattern or a visible page state; collect the others when those point at them:
| Artifact | What it tells you |
|---|---|
Exact OS, Safari and driver versions (the driver is the program that turns test commands into browser actions, such as Apple's safaridriver) | Whether failures cluster on one build (a Safari or macOS update changes results) |
| Failure rate and when it fails | Always on the first test after browser start (a cold start: caches empty, the machine still busy loading), or random (a race)? |
| Screenshot and video at the failure | Whether the page was still loading, animating, or showing a different state |
| The test's step log with timestamps and the assertion's expected versus actual value | How long the page had to reach the state and what it showed instead |
| Network log (HAR, a recorded list of requests with timings) from the app or a proxy | Whether a request was slow, failed, or returned in a different order than in Chrome |
| Browser console errors and the app's own client logs | A JavaScript error that only Safari throws |
document.visibilityState and window focus logged at the failing step | Whether the page was hidden or in the background; browsers may throttle timers and animation frames in hidden pages |
| Machine load (CPU, memory) on the runner | Whether a slow shared Mac stretched an animation or fetch past the assertion's timeout |
Compare the Safari failures against the Chrome and Firefox runs of the same commit: if only Safari is slow, ask what Safari does differently at that step.
Step 2: Reproduce on WebKit with a trace
On a laptop or a Linux container, run the single test repeatedly on the WebKit project with tracing forced on:
npx playwright test checkout.spec.js --project=webkit --repeat-each=50 --trace on --workers=1
A trace holds the step log, network and console for each step, plus DOM snapshots you can inspect (Playwright's trace viewer documentation lists these). Open the failing run's trace and read the snapshot at the failing assertion. Typical results:
- The element exists but is mid-transition or not yet laid out: a timing assumption on an animation, a web font load or lazy-loaded image. Wait on the real condition (the element reaching its final state or the request completing), not on time.
- A response arrives after the assertion in WebKit but before it in Chromium: a race in the app or an over-tight timeout. Fix the app race if the user could see it.
- Nothing reproduces after 50 repeats in Playwright WebKit: the cause is outside the engine (real Safari behaviour, macOS machine, driver). Go to step 3.
Step 3: Reproduce on real Safari and attach remote devtools
- On a Mac you control: in Safari's Settings, Advanced pane, enable "Show features for web developers", then use the Develop menu (the Safari menu-bar item that appears once that setting is on; it offers Show Web Inspector, or right-click and Inspect Element), as WebKit documents. Then use the Network and Timelines tabs to find the slow request or the layout work that ran late.
- On a remote CI Mac or a cloud grid: most cloud providers expose a live session or screen sharing for a failing job; reproduce the failing scenario by hand in that session with Web Inspector open. If the cloud session is real Safari on a Mac, the same Develop menu applies.
- For an iPhone or iPad failure: turn on Settings, Safari, Advanced, Web Inspector on the device, connect it to a Mac by cable (or set up wireless debugging in Xcode), and the device appears in Safari's Develop menu. iOS Simulator sessions are always inspectable that way, per WebKit's documentation.
- Record a Timeline during the failing step and compare the same recording from Chrome.
Step 4: What to ask developers for (Safari-specific)
- Whether the code under test uses an API that behaves differently or is missing in Safari: check the feature on MDN or caniuse before trusting it (for example
requestIdleCallbackis absent from Playwright's WebKit build, per a probe run in the Playwright 1.48.2 container). - Whether timers, animations or
requestAnimationFrameare used to drive the state the test asserts on, since hidden or background pages may be throttled. - Whether a Safari-specific storage or cookie rule affects the flow (third-party or cross-site cookies, storage persistence), and whether it appears only in Safari's Network and Storage tabs.
- A build with source maps (files that map minified production code back to the original source, so the Inspector shows readable code) and an easy way to enable verbose application logging.
- Whether they can reproduce with Safari Technology Preview (Apple's early-access Safari, which tracks newer WebKit) to tell a known fixed engine bug from an app bug.
Fix and verify
Replace time-based waits with a wait on the condition the user sees; fix or document the Safari-specific behaviour; then prove it. Run the test 100 times on WebKit (and on real Safari in the grid if that is where it failed) and require zero failures before closing the ticket. Keep the trace from the original failing run attached to the ticket.
Pitfalls
- Blaming "Safari is flaky" and adding retries: this hides real defects that users on Safari see.
- Debugging only in Playwright WebKit when CI runs real Safari: the machine, the driver and branded Safari differ from the Playwright build.
- Raising a global timeout for one slow step instead of finding what the step waits on.
Running the code
The Step 2 command assumes a suite that already has a Playwright project named webkit. A minimal setup that makes it run, inside mcr.microsoft.com/playwright:v1.48.2-jammy: a package.json containing { "private": true, "type": "module" }, npm i @playwright/test@1.48.2, your spec file (here checkout.spec.js), and this playwright.config.js:
// playwright.config.js
import { defineConfig, devices } from '@playwright/test';
export default defineConfig({
testDir: '.',
projects: [
{ name: 'chromium', use: { ...devices['Desktop Chrome'] } },
{ name: 'firefox', use: { ...devices['Desktop Firefox'] } },
{ name: 'webkit', use: { ...devices['Desktop Safari'] } },
],
});
With a trivial passing spec, the Step 2 command ran 50 repeats on WebKit and wrote one trace.zip per run under test-results/. Without a project named webkit, --project=webkit exits with an error.
Write a Playwright configuration that runs one suite across Chromium, Firefox and WebKit in parallel and produces a single consolidated JSON report. Explain how you keep browser versions identical across CI and developer machines.
Sample Answer
Direct answer
Declare three projects (Chromium, Firefox, WebKit) in one playwright.config.ts, turn on fullyParallel, and add the json reporter with a fixed outputFile. The JSON reporter writes one file that contains every project's results, so there is nothing to stitch together for a single run. Browser versions stay identical because Playwright ties browser builds to its own version: pin @playwright/test to an exact version with a lockfile, and run both CI and developer machines from the same Playwright container image tag.
The configuration
// playwright.config.ts
import { defineConfig, devices } from '@playwright/test';
export default defineConfig({
testDir: './tests',
fullyParallel: true,
workers: process.env.CI ? 3 : undefined,
retries: process.env.CI ? 1 : 0,
reporter: [
['list'],
// One JSON file holds every project's results (chromium, firefox, webkit).
['json', { outputFile: 'reports/results.json' }],
],
projects: [
{ name: 'chromium', use: { ...devices['Desktop Chrome'] } },
{ name: 'firefox', use: { ...devices['Desktop Firefox'] } },
{ name: 'webkit', use: { ...devices['Desktop Safari'] } },
],
});
projectsare named configurations; each one runs every test file once with its own settings.devices['Desktop Chrome'],devices['Desktop Firefox']anddevices['Desktop Safari']are Playwright's built-in device descriptors: ready-made bundles of user agent, viewport size and related settings for those engines, spread intouse.fullyParallel: trueruns tests inside a file in parallel too, not only separate files. Tests from all three projects share the same worker pool;workerscaps the pool (the number of parallel worker processes) at 3 on CI.undefinedlets Playwright choose; the Playwright API reference says the default is half the number of logical CPU cores, so a 10-core laptop runs 5 workers.['json', { outputFile: 'reports/results.json' }]writes the report to a fixed path, with each test entry carrying itsprojectName. A reporter is the part of the test runner that turns results into output; thelistreporter prints one line per test in the terminal, whilejsonwrites the machine-readable file.retries: 1on CI re-runs a failed test once; the report records every attempt in the test'sresultsarray, which is why the summariser reads the last one.
A test file and a summariser
// tests/basics.spec.ts
import { test, expect } from '@playwright/test';
test('flex row wraps and keeps each card at its basis', async ({ page }) => {
await page.setContent(`
<style>body{margin:0}</style>
<div id="row" style="display:flex;flex-wrap:wrap;width:500px">
<div class="c" style="flex:0 0 200px;height:40px"></div>
<div class="c" style="flex:0 0 200px;height:40px"></div>
<div class="c" style="flex:0 0 200px;height:40px"></div>
</div>`);
// $$eval runs the function inside the page on every element matching '.c'
const tops = await page.$$eval('.c', (els) => els.map((e) => e.getBoundingClientRect().top));
// 200 + 200 = 400 fits in the 500px row; a third 200px card would need 600, so it wraps
// onto a second line that starts 40px (one card height) down
expect(tops).toEqual([0, 0, 40]);
});
test('feature detection, not user-agent sniffing', async ({ page, browserName }) => {
await page.setContent('<p>x</p>');
const has = await page.evaluate(() => CSS.supports('selector(:has(a))'));
expect(has).toBe(true);
test.info().annotations.push({ type: 'engine', description: browserName });
});
test('reports the browser build it ran on', async ({ page, browser }) => {
await page.setContent('<p>x</p>');
test.info().annotations.push({ type: 'browser-version', description: browser.version() });
});
The flex rule flex:0 0 200px is shorthand for flex-grow: 0; flex-shrink: 0; flex-basis: 200px: the card never grows or shrinks and stays 200px wide. With flex-wrap: wrap the first two cards share the first line (offsets 0 and 0) and the third starts the second line at the 40px height of the first line, which gives [0, 0, 40] in every engine.
// summarize.mjs
import { readFileSync } from 'node:fs';
const report = JSON.parse(readFileSync('reports/results.json', 'utf8'));
const rows = [];
const versions = new Map();
const walk = (suite) => {
for (const spec of suite.specs ?? []) {
for (const t of spec.tests) {
rows.push([t.projectName, spec.title.slice(0, 44), t.results.at(-1).status]);
const v = t.annotations.find((a) => a.type === 'browser-version');
if (v) versions.set(t.projectName, v.description);
}
}
(suite.suites ?? []).forEach(walk);
};
report.suites.forEach(walk);
rows.sort().forEach((r) => console.log(r.join(' | ')));
for (const [p, v] of [...versions].sort()) console.log(`${p} build: ${v}`);
console.log('stats:', JSON.stringify({ expected: report.stats.expected, unexpected: report.stats.unexpected }));
The third test records the browser build in the test's annotations. An annotation is a { type, description } pair a test can attach to itself while it runs; the JSON reporter copies annotations into the report, which makes them a convenient way to carry the browser version into the file, so each report states exactly which builds produced it.
Reading summarize.mjs
The JSON report is an object with config, suites, errors and stats. Each top-level suite is one test file; a file's specs array holds one entry per test(...) call, and each spec has a tests array with one entry per project, carrying projectName, annotations and results (one element per attempt, each with a status). A describe block would appear as a nested suite inside suites, so the script defines walk, which handles one suite and calls itself on every nested suite; report.suites.forEach(walk) starts it on the files. Per test, it pushes [projectName, shortened title, results.at(-1).status] (the last attempt decides the outcome after a retry), and if the test carries a browser-version annotation it remembers the description in the versions Map keyed by project. The rows.sort() call orders lines by project name, then title, and report.stats.expected and unexpected are the counts of tests that ended as expected and unexpectedly.
Result
Run in the Playwright 1.48.2 image (Linux), npx playwright test printed 9 passed, and node summarize.mjs printed:
chromium | feature detection, not user-agent sniffing | passed
chromium | flex row wraps and keeps each card at its ba | passed
chromium | reports the browser build it ran on | passed
firefox | feature detection, not user-agent sniffing | passed
firefox | flex row wraps and keeps each card at its ba | passed
firefox | reports the browser build it ran on | passed
webkit | feature detection, not user-agent sniffing | passed
webkit | flex row wraps and keeps each card at its ba | passed
webkit | reports the browser build it ran on | passed
chromium build: 130.0.6723.31
firefox build: 131.0
webkit build: 18.0
stats: {"expected":9,"unexpected":0}
The suite was run 11 more times with the JSON reporter and all exited 0, so these tests are not timing-sensitive. One consolidated file covers all three projects (9 results, 3 tests x 3 engines).
Keeping browser versions identical on CI and developer machines
- Pin Playwright exactly. Use
"@playwright/test": "1.48.2"with no caret, commit the lockfile, and install withnpm ci. Playwright's documentation states that each Playwright version needs specific browser binaries, so pinning the package pins the browser builds (Chromium 130.0.6723.31, Firefox 131.0 and WebKit 18.0 for 1.48.2, as printed above). - Run from one image. Use
mcr.microsoft.com/playwright:v1.48.2-jammyas the CI job container and as the developer's local runner (docker run --rm ...). That also fixes the operating system, fonts and system libraries, which change rendering as much as browser versions do. - Where Docker is not used, run
npx playwright install --with-depsafter every Playwright upgrade, because each Playwright version needs specific browser binaries and an upgrade needs a re-install. - Upgrade deliberately. Bump the package, the lockfile and the image tag in one pull request, and read the
build:lines in the report to confirm what ran. - Do not mix in branded browsers silently. A project with
channel: 'chrome'or'msedge'uses the Chrome or Edge installed on the machine, so its version follows that machine. Playwright's own Chromium also runs ahead of the branded browsers.
Pitfalls
- Sharded runs write one JSON file per shard. Sharding splits the suite across several CI jobs, each running a slice. For sharding, use the blob reporter (a reporter that saves a raw, mergeable result bundle per shard) on each shard and combine the bundles with
npx playwright merge-reports, as described in the Playwright sharding docs; the singleresults.jsonabove is for an unsharded run. - WebKit here is not iOS Safari. It is a Playwright build of WebKit on Linux; it does not exercise iOS touch, keyboard or storage behaviour.
- Exact pixel assertions differ per engine. The layout test asserts on which row each box lands in, which is stable across engines, not on fractional sizes.
Running the code
# directory: package.json (type: module, @playwright/test 1.48.2), playwright.config.ts,
# tests/basics.spec.ts, summarize.mjs
docker run --rm --ulimit core=0 -v "$PWD":/src:ro \
mcr.microsoft.com/playwright:v1.48.2-jammy sh -c \
'cp -r /src /w && cd /w && npm i @playwright/test@1.48.2 && \
npx playwright test && node summarize.mjs'
Your app has shipped a bug where the order in which scheduled callbacks run differs between Safari and Chrome, and users on one engine saw a broken flow. You want an automated check that guards this across browsers and, where relevant, Node. How would you design it so results are deterministic, what would you assert, and how would you stop timing variability from making the check flaky?
Sample Answer
Direct answer
Do not compare absolute timings, and do not assert an order the specifications leave open. Instead, record every callback into a log of labelled events, wait until all expected labels have arrived (a settle barrier driven by a counter, not a sleep), and assert relative order by index in that log. Split the rules into two groups: guarantees every engine must meet (synchronous code first, microtasks before any timer, equal-delay timers in scheduling order, a microtask queued inside a timer callback running before the next timer) and per-engine expectations for orderings the standards leave open (kept as data in a config object). Run the same page in Playwright's Chromium, Firefox and WebKit projects, repeat each test ten times, and keep the check in CI knowing that Playwright's WebKit build is not real Safari.
Why the orders differ
JavaScript runs on one thread with a microtask queue (promise callbacks, queueMicrotask) that drains completely after each running piece of code, and several task queues (timers, message events, rendering callbacks such as requestAnimationFrame) from which the browser's event loop picks the next task. The HTML specification fixes the microtask rules and the order of timers with equal delays, but it does not require any particular order between different task sources, such as a MessageChannel message versus a setTimeout(fn, 0) timer, or where a requestAnimationFrame callback lands relative to timers. Engines choose their own scheduling there. A flow that accidentally depends on such an order works in one engine and breaks in another, which is the reported bug.
Terms the probe uses. A task source is a category of scheduled work (timers, messages, rendering) that the browser keeps in its own queue. MessageChannel creates two linked ports; posting on one port fires a message event on the other as a separate task. requestAnimationFrame asks the browser to call a function just before the next screen repaint. process.nextTick and setImmediate exist only in Node: nextTick runs a callback as soon as the current operation finishes, before the event loop moves on, and setImmediate runs it in the loop's check phase, after the poll phase that handles I/O. A settle barrier is a counter that starts at the number of callbacks scheduled and resolves a promise when the last one has fired.
A five-line example shows the basic order:
// tiny.mjs
console.log('A: synchronous');
setTimeout(() => console.log('D: timer, delay 0'), 0);
Promise.resolve().then(() => console.log('C: promise callback'));
console.log('B: synchronous');
Run in Node 20.18, it prints A, B, C, D in that order. The two synchronous lines finish first (A, B). When the running code ends, the microtask queue drains, so the promise callback runs (C). Only then does the event loop take the next task, the timer (D), although its delay is 0. The HTML specification gives browsers the same rule: microtasks run when the call stack empties, before the next task. The probe below repeats this with more kinds of callback and checks which parts of the order are guaranteed.
What to log and what to assert
The probe schedules the same work in every environment and logs a label when each callback fires:
// probe.mjs
// Runs in a browser page (via page.evaluate) and in Node. Self-contained on purpose.
export function runProbe() {
return new Promise((resolve) => {
const log = [];
let pending = 0;
const track = (label, schedule) => {
pending++;
schedule(() => {
log.push(label);
if (--pending === 0) resolve(log); // settle barrier: every scheduled callback has fired
});
};
log.push('sync-start');
track('micro-promise', (cb) => Promise.resolve().then(cb));
track('micro-queue', (cb) => queueMicrotask(cb));
track('timeout-a0', (cb) => setTimeout(cb, 0));
track('timeout-b0', (cb) => setTimeout(cb, 0));
track('timeout-c10', (cb) => setTimeout(cb, 10));
// Microtask queued from inside a timer callback: must run before the next timer.
track('timeout-d5', (cb) => setTimeout(() => { Promise.resolve().then(() => log.push('micro-in-d5')); cb(); }, 5));
track('timeout-e5', (cb) => setTimeout(cb, 5));
if (typeof MessageChannel !== 'undefined') {
track('message', (cb) => { const ch = new MessageChannel(); ch.port1.onmessage = () => { ch.port1.close(); cb(); }; ch.port2.postMessage(0); });
}
if (typeof requestAnimationFrame === 'function') track('raf', (cb) => requestAnimationFrame(() => cb()));
if (typeof setImmediate === 'function') track('immediate', (cb) => setImmediate(cb));
if (typeof process !== 'undefined' && process.nextTick) track('nexttick', (cb) => process.nextTick(cb));
log.push('sync-end');
});
}
How track(label, schedule) works: schedule is a function that arranges for its argument to be called later ((cb) => setTimeout(cb, 0)). track adds 1 to pending, then passes schedule a wrapper that logs the label, subtracts 1, and resolves the whole promise when pending reaches 0. The test therefore waits exactly as long as the slowest callback and no longer. Labels that a primitive cannot produce (no MessageChannel in some hosts, no setImmediate in browsers) are simply not scheduled, so one probe file runs everywhere.
The rules live in a second file. invariantFailures returns a list of broken rules, so a failure message names the exact rule; ENGINE_RULES holds the allowed message-versus-timer relation per engine:
// invariants.mjs
// Order rules that every engine must satisfy, plus per-engine rules kept as data.
const at = (log, label) => log.indexOf(label);
// Engines may legitimately disagree about message-channel vs zero-delay timer order.
export const ENGINE_RULES = {
chromium: { messageVsTimeout: ['after'] },
firefox: { messageVsTimeout: ['after'] },
webkit: { messageVsTimeout: ['before', 'after'] },
};
export function invariantFailures(log) {
const fails = [];
const need = (ok, msg) => { if (!ok) fails.push(msg); };
const macro = ['timeout-a0', 'timeout-b0', 'timeout-d5', 'timeout-e5', 'timeout-c10'];
need(at(log, 'sync-end') === at(log, 'sync-start') + 1, 'synchronous code must finish before any callback');
for (const m of ['micro-promise', 'micro-queue']) {
need(at(log, m) > at(log, 'sync-end'), `${m} must run after the synchronous code`);
for (const t of macro) need(at(log, m) < at(log, t), `${m} must run before ${t}`);
}
need(at(log, 'micro-promise') < at(log, 'micro-queue'), 'microtasks run in the order they were queued');
need(at(log, 'timeout-a0') < at(log, 'timeout-b0'), 'equal-delay timers run in scheduling order');
need(at(log, 'timeout-d5') < at(log, 'micro-in-d5') && at(log, 'micro-in-d5') < at(log, 'timeout-e5'),
'a microtask queued inside a timer callback runs before the next equal-delay timer');
return fails;
}
export function messageRelation(log) {
return at(log, 'message') < at(log, 'timeout-a0') ? 'before' : 'after';
}
Why these assertions and not others: each guaranteed rule compares two labels by their position in the log, so machine speed cannot change the verdict. The two 5 ms timers (timeout-d5, timeout-e5) share a delay, so their relative order is fixed by scheduling order; the 10 ms timer is deliberately not compared with them, because two timers with different delays can swap when the page stalls between scheduling them. requestAnimationFrame is only required to come after the synchronous code: its slot among timers varies between runs even inside one engine.
Making it deterministic and non-flaky
- Settle barrier, not a sleep. The probe resolves only when its counter reaches zero, so the test never reads the log early and never waits a fixed time. A hung callback fails by the test timeout rather than silently passing.
- Relative order by index. Compare positions in the log, never durations or timestamps.
- Open orders become data. Where engines legitimately differ, the allowed relations sit in
ENGINE_RULES. If WebKit is seen producing both orders across repeated runs, as it is here, its entry allows both; if an engine is stable, pin its single value so a change in either direction is noticed. - Repeat. Ten repeats per engine surface an order that is only usually true. A rule that fails in one repeat out of ten is a real defect or a rule that was never guaranteed, and either way it should not stay in the suite as written.
- Isolate the page. The page is created fresh with
setContent, with no network and no app code, so the check fails only because of scheduling.
The Playwright project runs the same spec in three engines with ten repeats each:
// playwright.config.mjs
import { defineConfig, devices } from '@playwright/test';
export default defineConfig({
testDir: '.',
testMatch: 'order.spec.mjs',
repeatEach: 10, // repeat each test 10 times in every project
workers: 1, // one page at a time keeps timer pressure comparable
timeout: 15_000,
reporter: 'line',
projects: [
{ name: 'chromium', use: { ...devices['Desktop Chrome'] } },
{ name: 'firefox', use: { ...devices['Desktop Firefox'] } },
{ name: 'webkit', use: { ...devices['Desktop Safari'] } },
],
});
// order.spec.mjs
import { test, expect } from '@playwright/test';
import { runProbe } from './probe.mjs';
import { ENGINE_RULES, invariantFailures, messageRelation } from './invariants.mjs';
test('scheduling order keeps the guaranteed rules', async ({ page, browserName }) => {
await page.setContent('<!doctype html><title>order</title>');
const log = await page.evaluate(runProbe); // resolves only when every callback has fired
expect(invariantFailures(log), log.join(' > ')).toEqual([]);
// rAF only has to come after the synchronous code; its slot among timers is not asserted.
expect(log.indexOf('raf')).toBeGreaterThan(log.indexOf('sync-end'));
// Engine-specific expectation, read from data rather than hard-coded in the test.
expect(ENGINE_RULES[browserName].messageVsTimeout).toContain(messageRelation(log));
});
The same probe runs in Node, where it also proves the rule set can fail, by feeding it a doctored log in which a timer jumped ahead of a microtask:
// node-run.mjs
import { runProbe } from './probe.mjs';
import { invariantFailures } from './invariants.mjs';
const log = await runProbe();
console.log(log.join(' > '));
console.log('failures:', invariantFailures(log));
// Teeth check: a doctored log where a timer jumped ahead of a microtask must be reported.
const bad = log.filter((l) => l !== 'timeout-a0');
bad.splice(bad.indexOf('micro-promise'), 0, 'timeout-a0');
console.log('doctored log failures:', invariantFailures(bad));
The last four lines of node-run.mjs exist because a rule set that has never failed has not been shown to work. They edit the real log so a timer jumps ahead of a microtask and require invariantFailures to report it.
Worked result
Run in the Playwright 1.48.2 image (Chromium 130.0, Firefox 131.0, WebKit 18.0, Node 20.18, probe loaded as an ES module), node node-run.mjs printed:
sync-start > sync-end > micro-promise > micro-queue > nexttick > immediate > timeout-a0 > timeout-b0 > message > timeout-d5 > micro-in-d5 > timeout-e5 > timeout-c10
failures: []
doctored log failures: [
'micro-promise must run before timeout-a0',
'micro-queue must run before timeout-a0'
]
npx playwright test reported 30 passed (3 engines x 10 repeats) on each of three consecutive invocations. In separate exploratory loops on the same page, the observed logs showed what the config encodes: Chromium and Firefox put message after timeout-a0 in every observed run, WebKit sometimes put message first, and raf appeared at different slots from run to run in all three engines, which is why it has no position rule. Node-only labels (nexttick, immediate) are logged but not asserted; they are absent from the browser rules and their placement relative to promise callbacks is an environment detail the guard does not depend on.
Reading the Node log label by label:
sync-start,sync-end: the probe function runs to its end first; everything else was only queued.micro-promise,micro-queue: the microtask queue drains in the order the callbacks were queued, before any timer.nexttickcomes after the promise callbacks here, which can look backwards becausenextTickis described as running first. Node's documentation explains it: in a CommonJS file the next-tick queue runs before promise callbacks, but an ES module is processed as part of the microtask queue, so Node is already draining microtasks and the promise callbacks go first.probe.mjsis an ES module. The same three lines show both orders:
// ticks.cjs (the same three lines are saved as ticks.mjs)
Promise.resolve().then(() => console.log('promise'));
process.nextTick(() => console.log('nextTick'));
console.log('sync');
Run in Node 20.18, ticks.cjs prints sync, nextTick, promise, and ticks.mjs prints sync, promise, nextTick.
4. immediate precedes timeout-a0. Node's event-loop guide says that when setTimeout(fn, 0) and setImmediate are both called from the main module, the order is non-deterministic and depends on the performance of the process. That is why the probe records immediate and asserts nothing about it.
5. timeout-a0, timeout-b0: equal delays run in scheduling order, a guaranteed rule.
6. message lands between timeout-b0 and timeout-d5. Node delivers message ports through its own machinery, so this slot is a Node detail. The specifications do not order a message task against timers, which is why the browser projects check it against ENGINE_RULES and not against a fixed position.
7. timeout-d5, micro-in-d5, timeout-e5: the d5 callback queued a promise callback, and the microtask queue drains straight after that callback, before the next timer is taken.
8. timeout-c10 is last only because nothing stalled; it is not compared with the 5 ms timers.
If you keep only two rules from this, keep these: synchronous code finishes before any callback, and every queued microtask runs before the next timer or message task. The other assertions follow from the same two ideas.
What the guard does not prove
Playwright's WebKit is a build of the WebKit engine running on Linux, not Safari on iOS or macOS, so it can miss behaviour that comes from Safari's own scheduling or from iOS itself. Keep the production bug as a regression case here (it catches engine-level scheduling), and add a small real-Safari check on a cloud device or a macOS runner for the specific flow that broke, before relying on this alone. Also fix the app: the durable remedy is that the flow stops depending on an order the specifications leave open, for example by chaining promises or awaiting an explicit completion signal instead of racing a timer against a message.
Running the code
# put the files above in an empty folder (ticks.cjs and ticks.mjs hold the same three lines)
echo '{"type":"module"}' > package.json
docker run --rm --ulimit core=0 -v "$PWD":/w -w /w mcr.microsoft.com/playwright:v1.48.2-jammy \
sh -c 'timeout 200 sh -c "npm i --no-audit --no-fund @playwright/test@1.48.2 && node tiny.mjs && node ticks.cjs && node ticks.mjs && node node-run.mjs && npx playwright test"'
Unlock Full Question Bank
Get access to all 15 Cross-Browser and Cross-Platform Testing interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.