Payment and Transaction Processing Systems Questions
Designing systems that move money correctly: idempotent payment flows, exactly-once semantics, reconciliation, ledgers, double-entry accounting, and fraud-detection architecture. Covers handling retries and partial failures without double-charging, and the consistency guarantees payments demand. Also covers protecting cardholder data through tokenization and PCI scope reduction, reconciling processor webhooks, and merchant and partner payouts. A high-stakes specialization of distributed transactions.
Design the architecture for scoring a payment for fraud risk inside the authorization path, where the entire decision (rules plus model) must complete within a strict latency budget of roughly 100ms. Cover how you'd decide fail-open versus fail-closed when the scoring service is slow or unavailable, how chargeback outcomes feed back into the system over time, and how you'd safely roll back a scoring change that starts degrading approval rates.
Sample Answer
Direct answer
Treat the 100 ms as a hard deadline that the scoring service owns end to end: every input the decision needs is precomputed and read from memory, rules and the model run in parallel against a fixed per-stage budget, and when the deadline hits the service returns a decision from a cheap local fallback instead of waiting. Whether that fallback leans "approve" (fail-open) or "decline/step up" (fail-closed) is a per-segment business policy computed from expected loss, not a global switch; for most card traffic it is fail-open with guardrails. Chargebacks (a cardholder's bank forcibly reversing a payment after a dispute) come back weeks later as labels (the fraud or not-fraud outcome attached after the fact to a past decision) that flow into offline analysis and model retraining (updating the model with those new examples), never into the live request. And every rule set and model ships as an immutable, versioned bundle behind shadow mode and a canary, so rollback is a pointer flip triggered automatically by an approval-rate guardrail.
Requirements I am designing to
- Deadline: p99 (the 99th-percentile latency, the time 99 of 100 requests beat) of about 100 ms for the full fraud decision, measured at the authorization service that calls us. The issuer round trip is outside this budget.
- Decision: approve, decline, or step-up (send the customer through 3-D Secure, the issuer's authentication challenge), plus a reason code and the bundle version that produced it.
- Availability: the payment path must not go down because fraud scoring did. A fraud outage that stops all checkouts is itself a revenue incident.
- Traffic: I will use 2,000 TPS (transactions per second) at peak for the arithmetic below.
Architecture
flowchart LR
AUTH[Authorization service] -->|deadline 100ms| GW[Scoring API]
GW --> FS[(In-memory feature store)]
GW --> RULES[Rule evaluator in process]
GW --> MODEL[Model server]
GW --> FB[Local fallback policy]
GW -->|async| LOG[(Decision log)]
STREAM[Payment and dispute events] --> AGG[Streaming aggregates]
AGG --> FS
LOG --> LABEL[Label join with chargebacks]
LABEL --> TRAIN[Offline retrain and rule review]
TRAIN --> REG[(Versioned bundle registry)]
REG --> GW
The critical rule: nothing on the request path does work that could be done earlier. Velocity counts ("attempts on this card in the last 10 minutes"), device reputation, customer history and merchant risk tiers are maintained by a streaming job (a program that continuously updates its output as new events arrive, instead of running on a fixed schedule) and written into a low-latency key-value store (an in-memory store such as Redis, or an in-process cache for slow-changing data). The request path only does point reads by key. No joins, no calls to the ledger (the system of record for completed payments and balances), no third-party lookups without their own hard timeout.
Budget allocation
| Stage | Budget | Notes |
|---|---|---|
| Network into and out of the scoring service | 10 ms | Same region, persistent connections |
| Parse and enrich from the request payload | 5 ms | BIN (bank identification number, the card's first digits) lookup from memory |
| Feature fetch | 25 ms hard timeout | All keys fetched in parallel in one batched round trip |
| Rules and model | 30 ms hard timeout | Run in parallel once features arrive; rules are cheap, the model is the long pole |
| Combine and respond | 5 ms | Decision log written asynchronously |
| Total on the critical path | 75 ms | 10 + 5 + 25 + 30 + 5 |
| Headroom | 25 ms | Garbage-collection pauses (brief stalls while the language runtime reclaims memory), a retry to a second feature replica, queueing at peak |
Why the headroom matters: tail latencies compound. If a request makes k independent parallel calls and each is slower than its own p99 1% of the time, the chance at least one is slow is
P(at least one slow)=1−0.99kFor k=5 that is 1 − 0.951 = 4.9%, so a request with five fan-out calls hits a p99-slow dependency about once in 20 requests. The design answers this by batching feature reads into one round trip, hedging (sending a duplicate read to a second replica if the first has not answered after a short delay), and treating each stage's budget as a timeout that falls through to defaults rather than waiting.
Deadline propagation
The authorization service sends the absolute deadline with the request. Each stage computes remaining time and gives up when it runs out: a missing feature becomes an explicit "unknown" value the rules and model were trained to handle, and a model timeout means the decision is made by rules plus fallback. The scoring API always answers before the deadline; it never lets the caller's own timeout be the thing that fires.
Fail-open versus fail-closed
Fail-open means approving (letting the issuer decide) when we cannot score; fail-closed means declining or stepping up. The right answer is the one with the lower expected cost for that segment, so compute it.
Illustrative inputs (not measurements): at the 2,000 TPS peak stated above, a 10-minute scoring outage is 2,000 × 600 = 1,200,000 payments; average ticket 60 USD; baseline fraud rate on this traffic 0.3%.
- Fail-open loss ≈ 1,200,000 × 0.003 × 60 = 216,000 USD of fraud, plus chargeback fees and some dispute-rate damage.
- Fail-closed loss ≈ 1,200,000 × 0.997 × 60 = 71,784,000 USD of good sales declined, plus customers who do not come back.
For ordinary card traffic, fail-open wins by more than two orders of magnitude (216,000 USD versus 71,784,000 USD, about 330 times less). But "open" should never mean "unguarded". The fallback policy is a small, fast rule set compiled into the scoring service itself, with no external dependencies:
- hard caps on amount per transaction and per card during degraded mode;
- a denylist of known-bad cards, devices and BIN ranges held in memory;
- step-up to 3DS instead of approve for transactions above a threshold, which shifts fraud liability toward the issuer where 3DS applies;
- fail-closed for segments where the arithmetic flips: gift cards, crypto purchases, high-value electronics, or a merchant currently under active card-testing attack (an attacker running many small, automated authorizations to find which stolen card numbers still work), where fraud rates are high and goods are instantly resellable.
The policy is data (per merchant category and amount band), is reviewed by the risk team, and every degraded-mode decision is tagged so it can be reviewed and its losses measured afterwards.
The chargeback feedback loop
Chargebacks (the cardholder's bank forcibly reversing a payment after a dispute) are the ground truth for fraud, but they arrive late: card networks typically let cardholders dispute within about 120 days, and most disputes arrive well inside that. The architecture accepts that delay instead of fighting it:
- Decision log. Every decision is written with the payment ID, the features actually used (as read, including which were missing), the bundle version, and the outcome. This is what makes later analysis honest: you train on what the system saw, not on today's recomputed values.
- Label join. Dispute events from the PSP (payment service provider) and settlement reports are joined back to the decision log by payment ID. Early fraud warnings (issuer fraud reports such as Visa TC40 and Mastercard SAFE data, which often precede or replace a formal dispute) give an earlier, noisier label.
- Label maturity (waiting long enough after a payment for its true fraud-or-not outcome to be known before trusting it as a training label). A payment only counts as "not fraud" once it is old enough that most disputes would have arrived. Retraining on last week's payments treats not-yet-disputed fraud as good traffic.
- Fast path for rules. Rules react in hours (an analyst writes a rule for a new pattern and ships it through the same canary); the model retrains on a slower cadence. Chargeback outcomes also feed the streaming aggregates (the same continuously updated feature values the streaming job maintains) directly: a card with a fresh fraud dispute goes onto the denylist immediately.
- Approved-only bias. You only observe chargebacks on payments you approved. Holding back a tiny random slice of would-be declines for review, or routing them to 3DS instead of declining, is how you learn whether the rules are rejecting good customers.
Safe rollout and rollback
A bundle is an immutable, versioned artifact: rule set, model, feature schema, thresholds and fallback policy, together. The scoring service holds a pointer to the active bundle and can hold a second one.
- Shadow: the candidate bundle scores live traffic alongside the active one; only the active decision is enforced. Compare decision distributions by segment.
- Canary: the candidate enforces on a small, randomly assigned slice (say 5%) of traffic, assigned by hashing the payment ID so it is stable.
- Guardrails: automatic rollback when approval rate, decline rate, step-up rate or issuer decline rate for the canary diverges from control beyond a threshold.
- Rollback: repoint to the previous bundle, which every node already has loaded. No deploy, no restart.
Set the threshold from the noise, not from intuition. At the 2,000 TPS peak, a 5% canary over 5 minutes sees 2,000 × 0.05 × 300 = 30,000 payments. With a 90% approval rate the standard error (roughly, how much this percentage would wobble from sample to sample by chance alone, computed as sqrt(p(1-p)/n)) of the canary's approval rate is
300000.9×0.1≈0.00173that is about 0.17 percentage points (0.18 pp for the difference against the much larger control group, which adds a little of its own, smaller sampling noise). A three-standard-error threshold is then about 3 × 0.18 ≈ 0.54 pp. A real 2 pp drop sits (2 − 0.54) / 0.18 ≈ 8 standard errors above that threshold, so it is caught with near certainty inside one 5-minute window; only a drop within about a percentage point of the threshold would need a longer window or a bigger canary to tell from noise, which is exactly the trade-off to state out loud.
Approval rate alone is not enough: a bundle that approves more is also dangerous. Guard both directions, and keep a delayed guardrail on early fraud warnings and dispute rate for the weeks after full rollout, since that is where a too-permissive change shows up.
Trade-offs and pitfalls
- Synchronous feature computation ("just query the last 30 days of transactions") is the most common way a 100 ms budget becomes 800 ms at peak.
- Global fail-closed looks safe and is usually the most expensive choice; global fail-open without caps is exploited within hours once attackers notice.
- Training-serving skew: features computed one way in the stream and another way offline produce a model that scores well in evaluation and poorly in production. Logging the served features avoids it.
- Rollback that requires a deploy is not a rollback under incident conditions; keep the previous bundle hot.
- Timeouts that are not tested: run chaos tests (deliberately injected failures, used to check the system degrades the way it was designed to) that stall the model server (the service that runs the fraud model and returns a score) and the feature store (the low-latency key-value store described above), and assert that the fallback answers within the deadline.
Architect a fraud rule engine for a merchant platform that supports rule authoring, staging, and fast evaluation in the live payment path. Describe how rules are stored, compiled/deployed, evaluated under high throughput, how to rollback problematic rules, and how to avoid performance impacts on the authorization latency budget.
Sample Answer
Direct answer
Treat rules as versioned data compiled into an immutable ruleset artifact, not as code in the payment service and not as rows the evaluator queries at request time. Analysts author rules in a restricted expression language through a UI; every change creates a new ruleset version that is validated, backtested against historical payments, run in shadow (evaluated and logged but not enforced) and then promoted. Evaluator instances load the compiled ruleset into memory and swap to a new version atomically, so evaluation is pure in-memory work over features that were precomputed before the request arrived. Rollback is repointing to the previous version, which every instance already holds, and individual rules have a kill switch.
Requirements
- Analysts (not engineers) add or change a rule and have it live within minutes, safely.
- Evaluation fits inside a small slice of the authorization latency budget: say 5 ms of a roughly 100 ms fraud decision, at thousands of TPS (transactions per second).
- Every decision is explainable: which rules fired, in which ruleset version.
- A bad rule can be undone in seconds without a deploy.
Architecture
flowchart LR
UI[Rule authoring UI] --> VAL[Validate and lint]
VAL --> STORE[(Rule store, versioned)]
STORE --> BT[Backtest on history]
BT --> COMP[Compile ruleset vN]
COMP --> REG[(Artifact registry)]
REG --> EV[Evaluators in the payment path]
FEAT[(Precomputed features)] --> EV
EV --> LOG[(Decision log with version)]
LOG --> DASH[Rule hit rates and shadow diffs]
Storage
- A rule has an ID, owner, description, a condition in the rule language, an action (
allow,review,block,step_upto 3-D Secure authentication), a priority, a scope (all merchants, one merchant, one category), and a lifecycle state:draft,shadow,active,disabled. - A ruleset version is an immutable snapshot: the exact list of rules and states at publish time, with who approved it. Rules are never edited in place; a change produces a new version. The audit trail comes for free.
- The store is a normal relational database. It is on the authoring path, not the evaluation path.
The rule language
Give analysts a restricted expression language, not a general programming language: field references from a fixed catalogue, comparison and set operators, boolean logic, and a small set of functions. Google's CEL (Common Expression Language) is an off-the-shelf example designed for exactly this: safe, side-effect-free, fast to evaluate. The restriction is the point: a rule cannot call the network, loop, or read arbitrary data, so its cost is bounded and it cannot break the evaluator.
Compile and deploy
- Validate: every field exists in the feature catalogue, operators suit the field types, thresholds are in range. Reject anything that needs a feature not available within the latency budget.
- Lint (an automated check for likely mistakes, without actually running anything): flag rules that would block a large share of traffic, duplicate another rule, or can never fire.
- Backtest: replay the last N days of logged decisions and features through the new version; report how many decisions change, on which segments, and how many known-fraud and known-good payments each changed rule catches.
- Compile: turn the ruleset into an in-memory structure: parsed expressions or generated predicates (small functions that test one condition and return true or false), rules indexed by scope so a payment only evaluates rules relevant to its merchant and category.
- Publish: write the artifact to a registry; evaluators poll or are notified, load it in the background, and swap an atomic pointer (flip a single reference to the new version in one indivisible step, so no request ever sees a half-loaded ruleset). Requests in flight finish on the old version; new ones use the new one.
Evaluation under high throughput
- Features are precomputed. Velocity counts ("attempts on this card in 10 minutes"), customer age and device reputation are maintained by a streaming job and read by key in one batched lookup before rules run. A rule can only reference catalogue features, so no rule can add a database query to the request.
- Cost is bounded by construction: evaluation is O(rules in scope × predicates per rule) (its cost scales with just two numbers: how many rules apply, times how many conditions each one checks), with no I/O. Indexing by scope keeps the "rules in scope" term small even when the total rule count grows into the thousands.
- Horizontal scale: evaluators are stateless apart from the loaded ruleset, so throughput scales by adding instances.
- A per-evaluation deadline still exists: if something unexpected (a pathological rule, a slow feature read) exceeds the budget, the evaluator returns a decision from the rules evaluated so far plus a conservative default, and records that it did so.
Staging and rollback
- Shadow first: a new or changed rule runs in shadow for a period, and a dashboard shows how many live decisions it would have changed.
- Promote with an approval step (a second person for rules that block).
- Rollback of a whole version: repoint to the previous version, a config change measured in seconds.
- Kill switch per rule: disable one rule without republishing everything, for the common case where a single new rule is misbehaving.
- Automatic guardrails: alert or auto-disable when a rule's hit rate jumps far above its shadow-period baseline, which catches a mistyped threshold (500 instead of 50,000) within minutes.
Runnable sketch
A minimal version of the compile, evaluate and shadow steps: rules stored as data, validated against an allowed field and operator list at compile time (no eval, Python's built-in for running an arbitrary string as code, which would let a "rule" read secrets, call the network or do anything else the process can do), each rule compiled into a closure (a small function that remembers the specific field, operator and value it was built from, so it can later be called with just a transaction) rather than re-parsed every time, and a candidate version compared against the live one on the same traffic.
import operator
# Rules as authored and stored (data, not code). Version 42 of the ruleset.
# Amounts are in cents, the ledger's minor unit, so 50000 below means 500.00 in the merchant's currency.
RULESET_V42 = [
{"id": "R1", "when": [["amount", ">", 50000], ["country_mismatch", "==", True]], "action": "review"},
{"id": "R2", "when": [["card_attempts_10m", ">=", 5]], "action": "block"},
{"id": "R3", "when": [["email_domain", "in", ["mailinator.com", "tempmail.io"]]], "action": "review"},
]
# Candidate v43 changes R2's threshold from 5 to 3; it runs in shadow first.
RULESET_V43 = [dict(r) for r in RULESET_V42]
RULESET_V43[1] = {"id": "R2", "when": [["card_attempts_10m", ">=", 3]], "action": "block"}
OPS = {">": operator.gt, ">=": operator.ge, "==": operator.eq, "in": lambda a, b: a in b}
ALLOWED_FIELDS = {"amount", "country_mismatch", "card_attempts_10m", "email_domain"}
SEVERITY = {"allow": 0, "review": 1, "block": 2}
def compile_ruleset(rules):
"""Validate once at deploy time, then return one closure per rule. No eval()."""
compiled = []
for r in rules:
preds = []
for field, op, value in r["when"]:
if field not in ALLOWED_FIELDS or op not in OPS:
raise ValueError(f"{r['id']}: unknown field or operator")
f = OPS[op]
# field=field, f=f, value=value bind each closure to THIS loop iteration's
# values; without the defaults, every closure would share the same loop
# variables and all of them would end up testing the LAST rule parsed.
preds.append(lambda tx, field=field, f=f, value=value: f(tx[field], value))
compiled.append((r["id"], r["action"], preds))
# Most severe rules first (a production engine could stop at the first block).
compiled.sort(key=lambda c: -SEVERITY[c[1]])
return compiled
def evaluate(compiled, tx):
hits = [rid for rid, action, preds in compiled if all(p(tx) for p in preds)]
action = max((a for rid, a, _ in compiled if rid in hits), key=SEVERITY.get, default="allow")
return action, hits
live, shadow = compile_ruleset(RULESET_V42), compile_ruleset(RULESET_V43)
traffic = [
{"amount": 1200, "country_mismatch": False, "card_attempts_10m": 1, "email_domain": "gmail.com"},
{"amount": 60000, "country_mismatch": True, "card_attempts_10m": 1, "email_domain": "gmail.com"},
{"amount": 900, "country_mismatch": False, "card_attempts_10m": 4, "email_domain": "gmail.com"},
{"amount": 700, "country_mismatch": False, "card_attempts_10m": 3, "email_domain": "tempmail.io"},
{"amount": 300, "country_mismatch": False, "card_attempts_10m": 6, "email_domain": "gmail.com"},
]
diffs = 0
for i, tx in enumerate(traffic):
a_live, h_live = evaluate(live, tx)
a_shadow, _ = evaluate(shadow, tx) # logged, never enforced
diffs += a_live != a_shadow
print(f"tx{i}: live={a_live:<6} hits={h_live} shadow={a_shadow}")
print(f"shadow would change {diffs} of {len(traffic)} decisions")
try:
compile_ruleset([{"id": "BAD", "when": [["__import__", "==", 1]], "action": "block"}])
except ValueError as e:
print("rejected at deploy:", e)
Output:
tx0: live=allow hits=[] shadow=allow
tx1: live=review hits=['R1'] shadow=review
tx2: live=allow hits=[] shadow=block
tx3: live=review hits=['R3'] shadow=block
tx4: live=block hits=['R2'] shadow=block
shadow would change 2 of 5 decisions
rejected at deploy: BAD: unknown field or operator
The shadow comparison shows exactly what the backtest and shadow stages are for: lowering R2's threshold from 5 to 3 attempts would turn one allow and one review into blocks. Whether that is good depends on whether those payments turn out to be fraud, which is what the decision log and later chargeback data answer. The last line shows the validator refusing a rule that references a field outside the catalogue before it can reach production.
Trade-offs and pitfalls
- Rules that query the database at evaluation time are the usual reason a rule engine blows the latency budget; the feature catalogue is the enforcement point.
- Editing active rules in place loses the audit trail and makes rollback impossible; always version.
- Rule sprawl: hundreds of overlapping rules nobody owns. Track hit rates, give every rule an owner and a review date, retire rules that never fire.
- Rules versus models: rules are transparent and fast to change, models generalise better. Most platforms run both, with rules for known patterns, hard policy and fast response to a new attack.
A merchant wants your recommendation on how to capture card data while keeping PCI scope down: route it through a PSP-hosted redirect, embed an iframe/hosted-field widget, or capture via direct API-level tokenization inside their own page. Walk through the PCI-scope, UX, and integration-complexity trade-offs of each, note the typical failure modes, and make a recommendation.
Sample Answer
Direct answer
For most merchants, and specifically for a medium-sized SaaS company that wants customers to adopt its product quickly, recommend embedded hosted fields (the payment service provider's, or PSP's, card inputs rendered in iframes inside the merchant's own checkout page): they keep the merchant in the smallest PCI self-assessment (SAQ A) like a redirect does, while keeping the checkout on-brand and in-page. Use a full redirect when speed of integration matters above everything else, and direct API tokenization in the merchant's own page only when there is a feature need that hosted fields cannot meet, because it moves the merchant into a much larger compliance scope.
The terms
- PCI DSS (Payment Card Industry Data Security Standard): the card networks' security rules for anyone who stores, processes or transmits card data. Smaller merchants prove compliance with a SAQ (self-assessment questionnaire); which SAQ applies depends on how card data touches their systems. The tiers differ hugely in size: SAQ A is a short questionnaire of a few dozen questions, SAQ A-EP runs to roughly 150 to 190 questions covering nearly the full standard, and SAQ D covers all 12 PCI DSS requirements (commonly cited as 300-plus questions across 90-plus pages) plus ongoing obligations SAQ A never asks for, such as quarterly external vulnerability scans.
- PSP (payment service provider): the company that processes cards for the merchant, for example Stripe, Adyen or Braintree.
- Tokenization: exchanging the card number for a token (a random reference) that the merchant can store and charge without ever holding the card number.
- Separate origin: the browser's same-origin security policy, which stops the parent page's JavaScript from reading anything inside an iframe served from a different domain (here, the PSP's). This is why a hosted-fields iframe keeps card data out of the merchant's own code even if that code is compromised.
The three options
| PSP-hosted redirect | Iframe / hosted fields | Direct API tokenization in the merchant's page | |
|---|---|---|---|
| What happens | Shopper leaves to the PSP's payment page, returns afterwards | Merchant page, but card inputs are iframes served by the PSP | Merchant's own form fields; merchant JavaScript sends the card to the PSP's API (or, worse, to the merchant's server) |
| Card data touches merchant | Never | Never (the iframe is a separate origin the page cannot read) | Merchant's page code handles it; server too if posted there |
| Typical SAQ | SAQ A | SAQ A | SAQ A-EP if the browser posts straight to the PSP; SAQ D if it passes through the merchant's servers |
| Branding and UX | PSP's look, limited styling, visible hand-off | Merchant's layout; field styling via PSP options | Complete control |
| Integration effort | Lowest: configure and handle the return and webhook | Low to medium: JavaScript SDK (software development kit: the PSP's browser library), styling, error states | Highest: own form, validation, security controls, audits |
About SAQ A: since 31 March 2025 the SAQ A version in force no longer lists the payment-page script requirements 6.4.3 (keeping and justifying an inventory of every script that runs on the payment page) and 11.6.1 (a mechanism that detects and alerts on unauthorized changes to that page); instead the merchant must confirm, as an eligibility condition, that its site is not susceptible to attacks from scripts that could affect its e-commerce systems. That means even an iframe merchant still has to keep its own page's scripts under control; it is just assessed through eligibility rather than those two controls. SAQ A-EP merchants must meet those script requirements in full.
Typical failure modes
Redirect
- Return-URL failures: shopper closes the tab after paying, so the merchant never sees the redirect back. Fix: rely on the server-side webhook to confirm the order, and treat the return as a hint.
- Trust drop at the hand-off: shoppers abandon when they land on an unfamiliar domain.
- Session loss across the round trip (cart emptied, logged out) if checkout state lives only in the browser.
Hosted fields
- The iframe script fails to load (ad blockers, strict Content Security Policy headers (a browser header that restricts which domains a page is allowed to load scripts and iframes from) that do not allow the PSP's domain, PSP incident): show a clear error and a redirect fallback.
- Styling and accessibility limits: the iframe controls focus and screen-reader labels, so test with keyboard and screen readers.
- Page-level script compromise: an attacker who injects script into the parent page can overlay a fake form on top of the iframe. This is why the SAQ A eligibility condition about scripts matters.
Direct API tokenization
- Skimming attacks (malicious JavaScript injected into the page reads the card fields directly); the merchant is responsible for detecting script changes.
- Scope creep: a developer adds logging on the checkout page or posts through the backend "for convenience" and the merchant is suddenly in SAQ D scope.
- Higher engineering and audit cost for every change to the checkout.
Worked example: the SaaS recommendation
A SaaS company with about 200 employees sells monthly plans at 49 to 499 USD and wants self-serve signup to convert quickly. Its constraints: a branded, in-app upgrade flow (sending users to another domain during signup hurts trust in a product they are just evaluating), a small security team, and no appetite for a large PCI programme.
- Redirect: fastest to ship (days), SAQ A, but the off-brand hop sits in the middle of the signup funnel.
- Direct API tokenization: full design control, but SAQ A-EP means payment-page script inventory, integrity checks and change detection, more audit questions, and ongoing engineering. Nothing the SaaS needs requires it.
- Hosted fields: recommended. SAQ A, the upgrade form looks native to the product, and the integration is a JavaScript SDK plus the same webhook handling a redirect needs. It fits the "fast adoption" goal because it removes the domain hop from signup without adding compliance work. Plan a redirect fallback for when the SDK fails to load.
What would change the recommendation: if the company needs a fully custom card form experience that the PSP's fields cannot style, or wants to route across several PSPs (in which case a vault or orchestration provider (a PSP-agnostic layer that tokenizes the card once and can route the resulting token to more than one downstream PSP), whose own hosted fields preserve SAQ A); if it sells mostly via invoices, the PSP's hosted invoice page (a redirect) is simpler still.
Trade-offs and pitfalls
- "We use tokens, so we are out of scope" is wrong if the card number passed through your page code or servers before becoming a token.
- Branding differences between hosted fields and native fields are smaller than teams expect; the conversion loss usually comes from the redirect's domain change, not from input styling.
- Whatever the option, confirm orders from server-side webhooks (HTTP callbacks the PSP sends directly to your server when a payment's status changes, bypassing the browser), never from what the browser reports back.
Outline the main PCI DSS control areas relevant to a cloud-based payments platform, and describe the practical architecture-level approaches a team can use to reduce PCI scope for a merchant. Explain the trade-offs and the residual compliance responsibilities that remain after scope reduction.
Sample Answer
Direct answer
PCI DSS (Payment Card Industry Data Security Standard, currently v4.0.1) has 12 requirements grouped under six goals: secure networks, protect account data, manage vulnerabilities, control access, monitor and test, and maintain a security policy. The cheapest control is not needing most of them: design the payment flow so card numbers go straight from the customer's browser or device to the payment provider and never touch the merchant's systems. Scope reduction shrinks the audit, but it never removes all responsibility: the merchant still owns the page that loads the payment form, its vendors, its policies and its incident response.
Key terms
- Cardholder data: the PAN (primary account number, the long card number), plus cardholder name, expiry date and service code when stored with it.
- SAD (sensitive authentication data): the CVV security code, full track or chip data, and PINs. It must never be stored after authorization.
- CDE (cardholder data environment): every system that stores, processes or transmits cardholder data, plus systems connected to them. PCI scope is the CDE.
- PSP (payment service provider): the company that processes card payments on the merchant's behalf.
- SAQ (self-assessment questionnaire): the form smaller merchants complete to validate compliance. Different SAQs apply to different payment designs, and the shortest ones apply when the merchant never handles card data.
The main PCI DSS control areas
| Goal | Requirements (v4.0.1) | What it means in a cloud payments platform |
|---|---|---|
| Build and maintain a secure network and systems | 1. Network security controls. 2. Secure configurations | Security groups and network policies (security groups are cloud firewall rules attached to a resource, controlling exactly which network traffic may reach it) isolating the CDE; hardened images (server images with unnecessary services and default accounts stripped out before use, to shrink what an attacker could exploit); no default credentials |
| Protect account data | 3. Protect stored account data. 4. Strong cryptography over open networks | Store as little as possible; encrypt stored PANs with keys held in a KMS (key management service, a system dedicated to creating, storing and controlling access to encryption keys) or HSM (hardware security module); TLS (Transport Layer Security, the encryption protocol behind HTTPS) everywhere |
| Maintain a vulnerability management program | 5. Protect against malicious software. 6. Develop and maintain secure systems and software | Patching, secure coding, code review, managing scripts on payment pages |
| Implement strong access control | 7. Restrict access by business need to know. 8. Identify users and authenticate access. 9. Restrict physical access | Least-privilege IAM (identity and access management) roles, MFA (multi-factor authentication) for all CDE access, physical security inherited from the cloud provider |
| Regularly monitor and test networks | 10. Log and monitor all access. 11. Test security regularly | Centralised, tamper-evident logs; vulnerability scans; penetration tests (authorized, simulated attacks meant to find exploitable weaknesses before a real attacker does); change detection (alerting when a file or configuration that is supposed to stay static gets modified, which can signal a compromise) |
| Maintain an information security policy | 12. Support information security with policies and programs | Risk assessment, vendor management, training, incident response plan |
In a cloud deployment, the provider covers part of this (physical security, hypervisor (the virtualization layer that isolates one customer's virtual machines from another's on the same physical hardware), some managed services) and publishes a responsibility matrix saying which requirements it meets, which you meet, and which are shared. You still own your configuration of everything the provider hands you.
Architecture approaches that reduce scope
Ordered from least to most merchant scope:
- Redirect to a PSP-hosted payment page. The customer leaves the merchant site to pay on the PSP's page, then returns. The merchant never sees card data.
- Embedded iframe or hosted fields. (an iframe, inline frame, is a small embedded webpage inside the merchant's page that is actually served by, and under the control of, a different origin, here the PSP) The card inputs are rendered inside frames served by the PSP, so the PAN goes from the browser to the PSP even though the page looks like the merchant's. The merchant receives a token (a random substitute for the card number).
- PSP JavaScript that posts directly to the PSP from a merchant-served page, or client-side encryption of the card in the browser with the PSP's public key. The merchant's servers never receive a readable PAN, but the merchant's page itself can be tampered with to steal it, so more controls apply than in options 1 and 2.
- Tokenization for storage and recurring billing. After the first payment, the merchant stores only the PSP's token and charges it later, so no PAN is at rest in merchant systems.
- P2PE (point-to-point encryption) for card-present terminals: a validated terminal encrypts the card at swipe or tap and only the provider can decrypt it.
- Network segmentation for anything that must remain in scope: isolate the CDE into its own account and network so the rest of the estate is out of scope, and prove the isolation with testing.
Worked example
An online retailer accepts 400,000 card payments a year on its website. That number matters for what follows: under the card networks' volume-based merchant levels, a merchant processing roughly 20,000 to 1,000,000 e-commerce transactions a year falls in the tier that validates PCI compliance by completing a Self-Assessment Questionnaire (SAQ), not the highest-volume tier (Level 1, over roughly six million transactions a year), which faces a mandatory annual on-site audit by a Qualified Security Assessor (QSA). At 400,000 payments, this retailer is squarely in SAQ territory, which is exactly why the SAQ A eligibility question below is the one that decides its validation path.
- Design A: its checkout page posts the card form to its own API, which calls the PSP. Its web servers, API servers, load balancers, logs and the admins of all of them are in the CDE. It must satisfy essentially all 12 requirements across that estate.
- Design B: it switches to PSP hosted fields and stores only tokens. Its servers never see a PAN. Its validation shrinks to a short questionnaire (for fully outsourced e-commerce this is SAQ A). What remains is protecting the checkout page that embeds the PSP frames, managing the PSP as a vendor, and its policies and incident response.
The SAQ A revision published in January 2025 (effective 31 March 2025) removed three payment-page requirements from SAQ A's own checklist (6.4.3: keep an inventory of every script that runs on the payment page, justify why each one is needed, and verify none has been tampered with; 11.6.1: run a change- and tamper-detection mechanism, checked at least weekly, that alerts on unauthorized changes to the payment page's security-relevant HTTP headers and script contents as the browser actually receives them; and 12.3.1: the targeted risk analysis that had backed how often 6.4.3 and 11.6.1 needed to run, no longer needed on SAQ A once those two are gone) and replaced them with two eligibility criteria: every payment-page element must be sourced only from a PCI DSS-compliant third-party service provider, and the merchant must confirm its site is not susceptible to attacks from scripts that could affect its e-commerce systems. In practice that still means controlling which scripts run on the checkout page and vetting who those scripts come from. The underlying requirements 6.4.3, 11.6.1 and 12.3.1 still exist in PCI DSS itself and still apply to merchants on SAQ D or a full on-site assessment; only SAQ A's own paperwork changed for the merchants eligible for it.
Trade-offs
| Approach | Scope | Cost to the merchant |
|---|---|---|
| Redirect | Smallest | Less control over look and flow; the extra hop can cost conversion |
| Iframe / hosted fields | Very small | Styling limits; tied to that PSP's components |
| Direct post or client-side encryption | Moderate | Merchant owns page integrity and more controls |
| Own vault | Large | Full control and PSP portability; full audit burden |
Tokens also create PSP lock-in: a token from one provider does not work at another, which matters if you want a backup processor.
Residual responsibilities after scope reduction
- The page that hosts the payment form. A compromised checkout page can swap the PSP frame for a fake one. Script control, a content security policy (a browser header listing which script sources may run) and change monitoring on that page stay the merchant's job.
- Vendor management. Keep the PSP's attestation of compliance (a formal document, typically signed by a PCI-approved assessor, stating that the PSP's own systems were found compliant) on file, know which requirements it covers, and review it yearly.
- Policies, training, incident response. A breach plan and staff awareness are required regardless of design.
- Anything that drifts back into scope. A support agent who takes a card number over the phone and types it into a CRM (customer relationship management system, the software used to track customer interactions) pulls that CRM and its users back in. Scans for PAN patterns in logs, tickets and data warehouses catch this.
- Annual validation. Scope reduction lowers the effort; it does not remove the obligation to validate and to keep the evidence.
Design a public payments API for a client that must handle 1 billion requests per day across multiple regions, comply with PCI-DSS, provide sub-100ms median latency for non-card operations, and support partner integrations. Cover the architecture components, security controls, data-residency strategy, and how you'd version and evolve the API without breaking partner integrations.
Sample Answer
Direct answer
Build it as a set of regional cells behind a global edge (a worldwide network of entry points close to each partner, which forwards each request into the right region), with each partner account pinned to a home region inside its data-residency jurisdiction, and with card data confined to a small, separately secured card vault so that almost every request (reads, refunds, payouts, webhooks) never touches raw card numbers. The sub-100ms median target is met by answering non-card calls entirely inside the home region from caches and a local primary database. The API evolves through date-based versions pinned per partner, with the server translating between versions, so partners upgrade on their own schedule and nothing they depend on changes underneath them.
Requirements and load
First, a few terms. PCI DSS (Payment Card Industry Data Security Standard) is the card industry's security standard; its current version is v4.0.1. The CDE (cardholder data environment) is every system that stores, processes or transmits card numbers, and every system connected to those; all of it is audited against PCI DSS. A PAN (primary account number) is the long card number. Data residency means a customer's personal and payment data must be stored and processed inside a named jurisdiction (for example, EU data stays in the EU).
Turn "1 billion requests per day" into a rate:
86,400 s109 requests11,574×3 (assumed peak-to-average ratio)≈11,574 RPS average≈34,722 RPS at peak(RPS is requests per second.) The 3x peak factor is an assumption to state out loud and replace with measured traffic shape. Assume 5% of requests carry card data (creating a payment method or a card payment). That is 50 million card operations per day, about 579 RPS on average. This split is the whole architecture: 95% of the traffic can be served by systems outside the CDE, and the CDE can stay small.
Assume three jurisdictions (US, EU, and a third such as India), each with two regions so a regional failure can fail over without data leaving the jurisdiction. If peak traffic splits evenly across jurisdictions, each carries about 11,574 RPS at peak, and each region in a pair must be sized to absorb its partner's full load: about 11,574 RPS per region, not half of it. Residency is what forces that doubling. You cannot fail EU traffic over to a US region.
Architecture components
flowchart LR
P[Partner client] --> E[Global edge: TLS, WAF, geo routing]
E --> G[Regional API gateway]
G --> A[Auth and rate limiter]
G --> I[Idempotency store]
G --> S[Payments services]
S --> D[(Regional primary DB)]
S --> V[Card vault in CDE]
V --> H[HSM or KMS]
S --> Q[Event bus]
Q --> W[Webhook dispatcher]
Q --> R[Reporting store]
| Component | Responsibility | Why it is there |
|---|---|---|
| Global edge | TLS termination (the point where the partner's encrypted connection is decrypted before the request continues inside the network) near the partner, WAF (web application firewall, which blocks malicious request patterns), absorption of DDoS (distributed denial-of-service) floods, routing to the account's home region | Cuts round-trip time and keeps attack traffic off the regional cells |
| Regional API gateway | API-key or OAuth token validation (OAuth is the standard for delegated access tokens), per-partner rate limits, request schema validation, version translation | One place to enforce the contract before business logic runs |
| Idempotency store | Records each write's idempotency key (a client-chosen unique ID for one intended operation) and its stored response | Partners retry on timeouts; without this, retries create duplicate charges |
| Payments services | Payment state machine, refunds, payouts, partner configuration | Business logic, outside the CDE |
| Card vault | Accepts PANs, returns tokens, talks to card networks and processors | The only place PANs exist, so the audited footprint is small |
| HSM / KMS | A hardware security module or a key management service holds encryption keys; software never sees raw key material | Keys that protect card data never sit next to the data |
| Event bus and webhook dispatcher | Publishes state changes; delivers signed webhooks to partners with retries | Partners integrate asynchronously instead of polling |
| Reporting store | Read-optimised copy for dashboards and exports | Keeps heavy analytical reads off the transactional database |
Routing to the home region. Every account ID and object ID embeds its home-region code (for example acct_eu1_...). The edge reads the partner's credential, looks up the home region from a small replicated routing table (non-personal data only), and forwards there. Object IDs carrying the region mean any request for an existing payment routes correctly without a global lookup.
Cells. A cell is a complete, independent copy of the service stack. A region is the geographic and legal boundary (where the data is allowed to live); a cell is a smaller blast-radius boundary inside it, and a region normally runs several of them. Partners are spread across a region's independent cells (each serving a subset of partners). A bad deployment or a runaway partner then affects one cell, not the region.
Meeting sub-100ms median for non-card operations
A median (p50, the 50th percentile) under 100ms is a budget you allocate, then verify with load tests. An allocation that fits:
| Step | Budget |
|---|---|
| Partner to edge, TLS reuse on a warm connection | 20 ms |
| Edge to home region | 25 ms |
| Gateway: auth (cached key lookup), rate limit, validation | 5 ms |
| Idempotency check and record (writes only) | 5 ms |
| Service logic plus one or two primary-database queries by primary key | 20 ms |
| Headroom | 25 ms |
These are allocations, not measurements. What makes them achievable: no cross-region calls on the request path, API-key verification cached in the gateway, every read by indexed key, and no call to external processors for non-card operations. Card operations are explicitly exempt from the 100ms target because they wait on card networks and issuers you do not control; publish a separate objective for them.
Security controls
- Scope reduction first. Partners collect cards through hosted fields (payment form fields served by an iframe from the vault's own origin, so the card number is typed directly into the vault, never the partner's page) or client SDKs that send the PAN straight to the vault, which returns a token. The partner's servers and your main services only ever see tokens, so they are outside the CDE.
- Card vault hardening. A separate network segment and separate cloud account, no human interactive access by default, envelope encryption (each record encrypted with a data key, and data keys encrypted by a master key held in the HSM or KMS), and key rotation.
- Partner authentication. Secret API keys scoped per environment (test vs live) and per permission set; OAuth for partner platforms acting on behalf of merchants; optional mTLS (mutual TLS, where the client also presents a certificate) for high-value partners.
- Service-to-service. mTLS between internal services and authorization checks on every call, so a compromised service cannot freely call the vault.
- Abuse controls. Per-partner and per-endpoint rate limits, anomaly detection on card-testing patterns (many small authorizations with different cards), and fraud scoring on payments.
- Webhooks. Signed with an HMAC (a keyed hash) over the timestamp and body so partners can verify origin and reject replays.
- Audit. Append-only audit logs of administrative actions and key usage, shipped to storage the platform's own operators cannot edit.
Data-residency strategy
- The partner chooses a home jurisdiction at onboarding; all personal and payment data for that account lives in that jurisdiction's two regions.
- The global control plane (the one piece of the platform allowed to be non-regional; it maps account IDs to home regions and holds nothing else) stores only non-personal routing metadata (account ID to region), never names, emails or card data.
- Replication and backups stay within the jurisdiction's region pair.
- Cross-border needs (a global partner selling in both EU and US) are handled as separate accounts per jurisdiction linked by a parent organisation, not by moving data.
- Analytics that must be global use aggregated or anonymised extracts produced inside each jurisdiction.
Versioning and evolution
- Additive changes ship to everyone. New optional request fields, new response fields, new endpoints and new enum values are non-breaking. Partners are told up front to ignore unknown fields and handle unknown enum values; that expectation is part of the contract.
- Breaking changes get a new dated version (for example
2026-09-01). Each partner account is pinned to the version it first integrated against. A request can override with a version header, which lets a partner test the next version one call at a time. - Translation layer. Internally there is one current model. The gateway runs a chain of small transforms that downgrade responses (and upgrade requests) between adjacent versions. Old versions cost one transform each rather than a forked codebase.
- Webhooks follow the account's pinned version, so an event payload does not change shape under a partner.
- Safety net. Consumer contract tests (automated checks that a real partner request shape still gets back the response that partner's code expects) for the most common partner request shapes run in CI; per-version traffic metrics show who still uses what; deprecation is announced with a long notice period and only retired when traffic reaches zero or the partner has been contacted.
Worked example
A partner in Germany integrated in 2025, pinned to version 2025-03-01. In 2026 the platform renames amount to a nested amount: {value, currency} object. That is breaking, so it ships as 2026-09-01. The German partner's GET /v1/payments/pay_eu1_8f2... routes by the eu1 prefix to Frankfurt, reads one row from the regional primary, and the gateway's downgrade transform flattens the nested object back into the old shape. No card data is read, no processor is called, and the partner sees exactly the response it saw last year. If Frankfurt fails, the request is served from the EU partner region, never from the US.
Trade-offs and pitfalls
- Residency doubles capacity cost. Each jurisdiction needs two full regions; that is the price of failing over without moving data. A single global active-active database (one dataset that every region can write to directly, with every write visible everywhere) would be cheaper and simpler, and illegal for many customers.
- Version transforms accumulate. Each old version adds a transform to maintain and test. Put a lifetime on versions and invest in migration tooling, or the chain becomes the slowest part of the gateway.
- Common wrong turn: putting the whole platform in PCI scope because one service touches PANs. That makes every deploy an audit event and slows everything.
- Common wrong turn: one global idempotency store. Keys must be checked in the home region with strong consistency (every read sees the latest write immediately, never a stale replica's answer), or two regions can both accept the same retried payment.
- Latency pitfalls: synchronous calls to the fraud engine or to a global service on the non-card path quietly break the median; move them off the request path or make them regional.
Unlock Full Question Bank
Get access to all 7 Payment and Transaction Processing Systems interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.