Payment and Transaction Processing Systems Questions
Designing systems that move money correctly: idempotent payment flows, exactly-once semantics, reconciliation, ledgers, double-entry accounting, and fraud-detection architecture. Covers handling retries and partial failures without double-charging, and the consistency guarantees payments demand. Also covers protecting cardholder data through tokenization and PCI scope reduction, reconciling processor webhooks, and merchant and partner payouts. A high-stakes specialization of distributed transactions.
Design the architecture for scoring a payment for fraud risk inside the authorization path, where the entire decision (rules plus model) must complete within a strict latency budget of roughly 100ms. Cover how you'd decide fail-open versus fail-closed when the scoring service is slow or unavailable, how chargeback outcomes feed back into the system over time, and how you'd safely roll back a scoring change that starts degrading approval rates.
Sample Answer
Direct answer
Treat the 100 ms as a hard deadline that the scoring service owns end to end: every input the decision needs is precomputed and read from memory, rules and the model run in parallel against a fixed per-stage budget, and when the deadline hits the service returns a decision from a cheap local fallback instead of waiting. Whether that fallback leans "approve" (fail-open) or "decline/step up" (fail-closed) is a per-segment business policy computed from expected loss, not a global switch; for most card traffic it is fail-open with guardrails. Chargebacks (a cardholder's bank forcibly reversing a payment after a dispute) come back weeks later as labels (the fraud or not-fraud outcome attached after the fact to a past decision) that flow into offline analysis and model retraining (updating the model with those new examples), never into the live request. And every rule set and model ships as an immutable, versioned bundle behind shadow mode and a canary, so rollback is a pointer flip triggered automatically by an approval-rate guardrail.
Requirements I am designing to
- Deadline: p99 (the 99th-percentile latency, the time 99 of 100 requests beat) of about 100 ms for the full fraud decision, measured at the authorization service that calls us. The issuer round trip is outside this budget.
- Decision: approve, decline, or step-up (send the customer through 3-D Secure, the issuer's authentication challenge), plus a reason code and the bundle version that produced it.
- Availability: the payment path must not go down because fraud scoring did. A fraud outage that stops all checkouts is itself a revenue incident.
- Traffic: I will use 2,000 TPS (transactions per second) at peak for the arithmetic below.
Architecture
flowchart LR
AUTH[Authorization service] -->|deadline 100ms| GW[Scoring API]
GW --> FS[(In-memory feature store)]
GW --> RULES[Rule evaluator in process]
GW --> MODEL[Model server]
GW --> FB[Local fallback policy]
GW -->|async| LOG[(Decision log)]
STREAM[Payment and dispute events] --> AGG[Streaming aggregates]
AGG --> FS
LOG --> LABEL[Label join with chargebacks]
LABEL --> TRAIN[Offline retrain and rule review]
TRAIN --> REG[(Versioned bundle registry)]
REG --> GW
The critical rule: nothing on the request path does work that could be done earlier. Velocity counts ("attempts on this card in the last 10 minutes"), device reputation, customer history and merchant risk tiers are maintained by a streaming job (a program that continuously updates its output as new events arrive, instead of running on a fixed schedule) and written into a low-latency key-value store (an in-memory store such as Redis, or an in-process cache for slow-changing data). The request path only does point reads by key. No joins, no calls to the ledger (the system of record for completed payments and balances), no third-party lookups without their own hard timeout.
Budget allocation
| Stage | Budget | Notes |
|---|---|---|
| Network into and out of the scoring service | 10 ms | Same region, persistent connections |
| Parse and enrich from the request payload | 5 ms | BIN (bank identification number, the card's first digits) lookup from memory |
| Feature fetch | 25 ms hard timeout | All keys fetched in parallel in one batched round trip |
| Rules and model | 30 ms hard timeout | Run in parallel once features arrive; rules are cheap, the model is the long pole |
| Combine and respond | 5 ms | Decision log written asynchronously |
| Total on the critical path | 75 ms | 10 + 5 + 25 + 30 + 5 |
| Headroom | 25 ms | Garbage-collection pauses (brief stalls while the language runtime reclaims memory), a retry to a second feature replica, queueing at peak |
Why the headroom matters: tail latencies compound. If a request makes k independent parallel calls and each is slower than its own p99 1% of the time, the chance at least one is slow is
P(at least one slow)=1−0.99kFor k=5 that is 1 − 0.951 = 4.9%, so a request with five fan-out calls hits a p99-slow dependency about once in 20 requests. The design answers this by batching feature reads into one round trip, hedging (sending a duplicate read to a second replica if the first has not answered after a short delay), and treating each stage's budget as a timeout that falls through to defaults rather than waiting.
Deadline propagation
The authorization service sends the absolute deadline with the request. Each stage computes remaining time and gives up when it runs out: a missing feature becomes an explicit "unknown" value the rules and model were trained to handle, and a model timeout means the decision is made by rules plus fallback. The scoring API always answers before the deadline; it never lets the caller's own timeout be the thing that fires.
Fail-open versus fail-closed
Fail-open means approving (letting the issuer decide) when we cannot score; fail-closed means declining or stepping up. The right answer is the one with the lower expected cost for that segment, so compute it.
Illustrative inputs (not measurements): at the 2,000 TPS peak stated above, a 10-minute scoring outage is 2,000 × 600 = 1,200,000 payments; average ticket 60 USD; baseline fraud rate on this traffic 0.3%.
- Fail-open loss ≈ 1,200,000 × 0.003 × 60 = 216,000 USD of fraud, plus chargeback fees and some dispute-rate damage.
- Fail-closed loss ≈ 1,200,000 × 0.997 × 60 = 71,784,000 USD of good sales declined, plus customers who do not come back.
For ordinary card traffic, fail-open wins by more than two orders of magnitude (216,000 USD versus 71,784,000 USD, about 330 times less). But "open" should never mean "unguarded". The fallback policy is a small, fast rule set compiled into the scoring service itself, with no external dependencies:
- hard caps on amount per transaction and per card during degraded mode;
- a denylist of known-bad cards, devices and BIN ranges held in memory;
- step-up to 3DS instead of approve for transactions above a threshold, which shifts fraud liability toward the issuer where 3DS applies;
- fail-closed for segments where the arithmetic flips: gift cards, crypto purchases, high-value electronics, or a merchant currently under active card-testing attack (an attacker running many small, automated authorizations to find which stolen card numbers still work), where fraud rates are high and goods are instantly resellable.
The policy is data (per merchant category and amount band), is reviewed by the risk team, and every degraded-mode decision is tagged so it can be reviewed and its losses measured afterwards.
The chargeback feedback loop
Chargebacks (the cardholder's bank forcibly reversing a payment after a dispute) are the ground truth for fraud, but they arrive late: card networks typically let cardholders dispute within about 120 days, and most disputes arrive well inside that. The architecture accepts that delay instead of fighting it:
- Decision log. Every decision is written with the payment ID, the features actually used (as read, including which were missing), the bundle version, and the outcome. This is what makes later analysis honest: you train on what the system saw, not on today's recomputed values.
- Label join. Dispute events from the PSP (payment service provider) and settlement reports are joined back to the decision log by payment ID. Early fraud warnings (issuer fraud reports such as Visa TC40 and Mastercard SAFE data, which often precede or replace a formal dispute) give an earlier, noisier label.
- Label maturity (waiting long enough after a payment for its true fraud-or-not outcome to be known before trusting it as a training label). A payment only counts as "not fraud" once it is old enough that most disputes would have arrived. Retraining on last week's payments treats not-yet-disputed fraud as good traffic.
- Fast path for rules. Rules react in hours (an analyst writes a rule for a new pattern and ships it through the same canary); the model retrains on a slower cadence. Chargeback outcomes also feed the streaming aggregates (the same continuously updated feature values the streaming job maintains) directly: a card with a fresh fraud dispute goes onto the denylist immediately.
- Approved-only bias. You only observe chargebacks on payments you approved. Holding back a tiny random slice of would-be declines for review, or routing them to 3DS instead of declining, is how you learn whether the rules are rejecting good customers.
Safe rollout and rollback
A bundle is an immutable, versioned artifact: rule set, model, feature schema, thresholds and fallback policy, together. The scoring service holds a pointer to the active bundle and can hold a second one.
- Shadow: the candidate bundle scores live traffic alongside the active one; only the active decision is enforced. Compare decision distributions by segment.
- Canary: the candidate enforces on a small, randomly assigned slice (say 5%) of traffic, assigned by hashing the payment ID so it is stable.
- Guardrails: automatic rollback when approval rate, decline rate, step-up rate or issuer decline rate for the canary diverges from control beyond a threshold.
- Rollback: repoint to the previous bundle, which every node already has loaded. No deploy, no restart.
Set the threshold from the noise, not from intuition. At the 2,000 TPS peak, a 5% canary over 5 minutes sees 2,000 × 0.05 × 300 = 30,000 payments. With a 90% approval rate the standard error (roughly, how much this percentage would wobble from sample to sample by chance alone, computed as sqrt(p(1-p)/n)) of the canary's approval rate is
300000.9×0.1≈0.00173that is about 0.17 percentage points (0.18 pp for the difference against the much larger control group, which adds a little of its own, smaller sampling noise). A three-standard-error threshold is then about 3 × 0.18 ≈ 0.54 pp. A real 2 pp drop sits (2 − 0.54) / 0.18 ≈ 8 standard errors above that threshold, so it is caught with near certainty inside one 5-minute window; only a drop within about a percentage point of the threshold would need a longer window or a bigger canary to tell from noise, which is exactly the trade-off to state out loud.
Approval rate alone is not enough: a bundle that approves more is also dangerous. Guard both directions, and keep a delayed guardrail on early fraud warnings and dispute rate for the weeks after full rollout, since that is where a too-permissive change shows up.
Trade-offs and pitfalls
- Synchronous feature computation ("just query the last 30 days of transactions") is the most common way a 100 ms budget becomes 800 ms at peak.
- Global fail-closed looks safe and is usually the most expensive choice; global fail-open without caps is exploited within hours once attackers notice.
- Training-serving skew: features computed one way in the stream and another way offline produce a model that scores well in evaluation and poorly in production. Logging the served features avoids it.
- Rollback that requires a deploy is not a rollback under incident conditions; keep the previous bundle hot.
- Timeouts that are not tested: run chaos tests (deliberately injected failures, used to check the system degrades the way it was designed to) that stall the model server (the service that runs the fraud model and returns a score) and the feature store (the low-latency key-value store described above), and assert that the fallback answers within the deadline.
Provide a recommended microservice decomposition for a merchant payments platform. List core services (for example gateway adapter, authorization service, ledger/settlement service, fraud service, reconciler, notifications), describe responsibilities, typical APIs between them, and how you would handle ownership of critical data like transaction state and merchant configuration.
Sample Answer
Direct answer
Split the platform along ownership of money-critical data: one service owns the lifecycle of each payment, one owns balances (the ledger), one owns merchant configuration, and each external processor sits behind its own adapter. The synchronous path that a buyer waits on (API gateway, payment service, fraud check, processor adapter) stays short, and everything else, ledger postings, notifications, reconciliation (matching internal records against what the processor and the bank report), and reporting, reacts to events. The rule that keeps it correct: every critical piece of data has exactly one service that writes it, and every other service reads it through that service's API or through events it publishes.
Key terms
- PSP (payment service provider) / processor: the external company that sends authorizations to the card networks and settles funds.
- Authorization: the bank approves and holds funds. Capture: the merchant claims them. Settlement: the money actually moves, net of fees.
- Ledger: an append-only record of money movements using double-entry bookkeeping (every movement is a debit on one account and an equal credit on another, so the books always sum to zero).
- Reconciliation: matching internal records against processor and bank reports.
- Outbox: a table written in the same database transaction as a state change, from which events are published, so a state change and its event can never disagree.
- Idempotency key: a client-chosen unique ID attached to a request so that if the same request is retried (for example after a timeout), the server can recognize the repeat and return the original result instead of creating a second payment.
- Event bus: the shared channel a service publishes domain events to (such as
payment.captured) so other services can subscribe and react, without the publisher calling any of them directly. - Debit and credit: in double-entry bookkeeping every ledger line is a debit or a credit. A debit increases an asset or expense account and decreases a liability or revenue account; a credit does the reverse. A receivable is an asset, money owed TO the platform (for example by the PSP, once it has taken a card payment but not yet settled it); debiting it records that the platform is now owed more. A payable is a liability, money the platform owes someone else (a seller's pending balance, a merchant's fee credit); crediting it records that the platform now owes more.
Proposed services
flowchart LR
API[API gateway] --> PAY[Payment service]
PAY --> AUTH[Authorization service]
AUTH --> FR[Fraud service]
AUTH --> AD[Processor adapters]
AD --> EXT[External PSPs]
PAY --> BUS[Event bus]
BUS --> LED[Ledger and settlement]
BUS --> NOT[Notifications]
REC[Reconciler] --> LED
REC --> AD
CFG[Merchant config] -.-> PAY
| Service | Responsibilities | Owns (sole writer) | Typical API |
|---|---|---|---|
| API gateway | Authentication of merchants, rate limits, idempotency-key handling, request validation | Idempotency records | Public REST: POST /payments, POST /payments/{id}/capture, POST /refunds |
| Payment service | The payment state machine: created, authorized, captured, partially refunded, refunded, failed, disputed; enforces legal transitions | Payment and refund state | Internal: CreatePayment, Capture, Refund over gRPC (a binary remote-procedure-call protocol); publishes payment.authorized, payment.captured, refund.succeeded |
| Authorization service | Decides how to authorize: calls fraud, picks the processor (routing), handles 3-D Secure (cardholder authentication) | Routing decisions and authorization attempts | Authorize(payment_id, amount, token, merchant) returns approved, declined or unknown |
| Fraud service | Scores risk from device, velocity (how rapidly this device or card is attempting payments, a common fraud signal) and history; returns allow, review or block within a strict time budget | Risk scores and rules | Score(payment context), plus async feedback events for chargebacks (a cardholder disputing a charge through their bank, which pulls the funds back from the platform) |
| Processor (gateway) adapters | One gateway adapter per external PSP: translate the internal model to that PSP's API, map decline codes (the processor's classification of why an attempt failed, such as insufficient funds or a stolen card, which routing logic needs in order to decide whether to retry elsewhere), handle PSP-specific retries and status queries | Raw PSP request and response log | Authorize, Capture, Refund, GetStatus, and a webhook receiver (an HTTP endpoint the PSP calls to push status updates to us, instead of us polling for them) that emits normalised events |
| Ledger and settlement | Double-entry postings for every money movement; merchant balances; payout calculation and scheduling | Journal entries, balances, payouts | Consumes payment events; GetBalance(merchant); publishes payout.created |
| Reconciler | Ingests PSP settlement files and bank statements; matches them to ledger and payment records; raises breaks (unmatched or mismatched items) | Reconciliation results and break cases | Batch jobs plus ListBreaks for finance tooling |
| Notifications | Merchant webhooks (signed, retried), customer receipts, email | Delivery attempts | Consumes events; ReplayWebhook(event_id) |
| Merchant configuration | Merchant profile, enabled payment methods, fees, payout schedule, API keys, risk settings | Merchant config, versioned | GetMerchantConfig(id) returning a version number; publishes merchant.config_changed |
| Token vault (often the PSP's) | Stores raw card numbers and issues a token in their place (a token is a random reference that stands in for the card number, so other services never need to handle the real number) | Card data | Tokenize (turn a card number into a token), restricted Detokenize (reverse a token back to the real card number) |
Ownership of critical data
- Transaction state is owned by the payment service only. The adapter reports what the PSP said; the payment service decides the state transition. This avoids two services disagreeing about whether a payment is captured.
- Balances are owned by the ledger only. Nobody else computes "how much does this merchant have". The ledger derives it from postings, and reporting reads a copy.
- Merchant configuration is owned by the config service, but read on the hot path from a local cache. Every payment records the config version it used, so a fee change at 12:00 cannot make a 11:59 payment's fee ambiguous later.
- Events carry facts, not commands to change someone else's data. The ledger posts entries because it heard
payment.captured, not because the payment service wrote into the ledger's tables. Each consumer de-duplicates by event ID, because delivery is at-least-once (an event may arrive more than once). - No shared databases. Each service's store is private; a shared table is where two writers eventually corrupt each other.
Worked example: one card payment
- Merchant calls
POST /paymentsfor 80.00 EUR with idempotency keyord-5531. The gateway records the key. - The payment service creates
pay_901in statecreated, reads merchant config version 14 from cache. - The authorization service gets a fraud score (allow), routes to PSP A, and the adapter returns approved. The payment service moves
pay_901toauthorizedand writespayment.authorizedto its outbox in the same transaction. - The merchant ships and calls capture; the state becomes
capturedandpayment.capturedis published. - The ledger posts: debit "PSP A receivable" 80.00, credit "merchant pending balance" 78.40, credit "platform fee revenue" 1.60 (a 2% fee from config version 14). The three lines net to zero.
- Two days later the reconciler matches PSP A's settlement line for
pay_901against the receivable and clears it.
Trade-offs and pitfalls
- Separate authorization service or not. At small scale, fold authorization into the payment service; split it out when routing across several PSPs and 3-D Secure logic grows enough to deserve its own team and deploy cadence.
- Too many synchronous hops on the buyer's path add latency and failure points. Keep it to gateway, payment, fraud and adapter; everything else is asynchronous.
- Fraud outage policy must be explicit: fail open (accept with lower limits) or fail closed (decline). Most platforms fail open for low amounts and closed for high ones.
- A distributed transaction across services (making the payment update and the ledger posting commit together) is the wrong tool; use the outbox plus idempotent consumers and let reconciliation catch the rare gap.
- Over-splitting (a separate refunds service, a separate captures service) scatters one state machine across services, which is exactly where money bugs come from.
Your primary payment processor experiences an extended outage. Design a disaster recovery plan that ensures merchants can continue accepting payments (or degrade gracefully), eventual settlement and reconciliation occur correctly, fees and mapping are handled for backup processors, and customers and regulators are informed as required.
Sample Answer
Direct answer
Plan for three things in order: keep taking payments by routing new authorizations to a pre-integrated backup processor (and degrading payment methods that cannot move), never lose or double-count money by freezing every in-flight transaction with the primary as "unknown" until the primary can be asked what happened, and settle and reconcile per processor afterwards, because each processor pays out, charges fees and reports in its own way. Communication runs on a pre-agreed clock: merchants and customers through status pages and API signals within minutes, regulators and card-scheme partners according to their incident-reporting obligations.
Key terms
- Processor / PSP (payment service provider): the company that sends card authorizations to the card networks and later pays out the money.
- Authorization: the issuing bank approves and places a hold on funds. Capture: the merchant claims the held funds. Settlement: the processor actually transfers the money, usually a day or more later, minus fees.
- Reconciliation: matching what your own records say should have happened against what processors and banks report did happen.
- Circuit breaker: logic that stops sending traffic to a failing dependency once errors cross a threshold, then probes it periodically.
- MID (merchant ID): the identifier a processor or acquirer assigns to a merchant account; the backup needs its own.
What must exist before the outage
A DR (disaster recovery) plan for payments is mostly work done in advance:
- A second processor, integrated and warm (fully set up and already carrying a slice of live traffic, not sitting idle). Signed contract, underwriting completed (the processor's own review and approval of your business before it will process for you), MIDs issued for every merchant you intend to fail over, and a small share of live traffic (for example 2 to 5%) routed there permanently, so you know it works today, not in theory.
- Portable card credentials. Tokens issued by the primary processor do not work at the backup. Stored cards (subscriptions, one-click checkout) can only fail over if you hold network tokens (issued by the card networks, usable across processors) or run your own vault that can send card data to either processor. Without that, only customers typing a card fresh can pay through the backup.
- A mapping layer. A table per processor translating your internal payment methods, currencies, merchant category codes (MCCs, the standard 4-digit codes that classify what kind of business a merchant is), statement descriptors (the text that shows up on the cardholder's statement), decline codes and 3-D Secure (the card-holder authentication step) settings into that processor's vocabulary. Divergence here (for example a merchant category code the backup does not recognise, or a decline code mapped to the wrong retry behavior) is what turns a failover into a flood of declines.
- A routing service with per-processor health, a circuit breaker, and a manual override switch.
During the outage
flowchart TD
A[New payment] --> B{Primary healthy?}
B -- yes --> C[Primary processor]
B -- no --> D{Card usable at backup?}
D -- yes --> E[Backup processor]
D -- no --> F[Offer other method or retry later]
G[In-flight at primary when it failed] --> H[State: unknown, no retry elsewhere]
H --> I[Resolve via primary status query or settlement report]
- Route new authorizations to the backup once the breaker opens. Keep the idempotency key (the client's unique ID for this payment attempt) as your internal reference so each processor call is traceable to one payment.
- Do not replay in-flight payments at the backup. A request that timed out at the primary may have been approved there. Retrying it at the backup risks charging the customer twice. Mark it
unknown, tell the merchant the outcome is pending, and resolve it once the primary answers status queries or its settlement file arrives. - Captures, refunds and voids follow the original processor. An authorization made at the primary can only be captured or refunded through the primary. Queue these operations durably and drain them when the primary returns; if an authorization would expire first (online card authorizations are typically valid for about 7 days), decide per merchant whether to re-authorize at the backup.
- Degrade gracefully. Recurring charges on non-portable tokens are deferred and retried later rather than failed; merchants see a clear "processing delayed" status rather than hard declines. For card-present terminals, offline "store and forward" approvals (the terminal approves the sale itself on the spot, with no live authorization call, then submits it for real authorization once connectivity returns) under a low per-transaction limit are possible but the merchant carries the risk of later declines; enable it only where merchants opted in.
- Protect the backup. Ramp traffic gradually and watch approval rates; a sudden jump in volume can trigger the backup's own fraud controls.
After the outage: settlement and reconciliation
- Reconcile per processor. Each processor's settlement report is matched against your ledger (your internal, append-only record of money movements) for the transactions it handled. The routing decision is recorded on every payment, so each payment belongs to exactly one processor's reconciliation.
- Resolve every
unknown. For each, the primary either reports it approved (keep it, capture as normal) or never received or declined it (mark failed, and if the customer retried and succeeded at the backup, there is no double charge to unwind). Any real duplicate found is refunded proactively. - Different payout timing. The backup may pay out a day later, to a different bank account, with different fee deductions. Merchants' payout projections and your treasury forecast (the internal projection of cash moving in and out of the business) must show this explicitly.
- Chargebacks and disputes (a customer disputing a charge through their bank) for backup-processed payments arrive from the backup for months afterwards; the dispute system must know to look there.
Fees and who pays them
Fees differ between processors, so decide in advance whether the platform absorbs the difference or passes it on. Illustrative numbers (assumed contract rates, not any real vendor's price list):
- A merchant processes 2,000,000 USD per day, spread evenly, at an average ticket of 50 USD. A 6-hour outage moves 6/24 of that, 500,000 USD, or 10,000 payments, to the backup.
- Primary rate 2.2% + 0.10 USD: 11,000 + 1,000 = 12,000 USD.
- Backup rate 2.9% + 0.30 USD: 14,500 + 3,000 = 17,500 USD.
- Extra cost of the outage: 5,500 USD for this merchant.
Most platforms absorb this for short outages (it is small against lost sales and merchant trust) and write the rule into their merchant terms so it is not negotiated during an incident.
Communication
| Audience | When | What |
|---|---|---|
| Merchants | Within minutes of routing change | Status page, API status field on affected payments, dashboard banner: what works, what is delayed, whether payouts shift |
| Customers | Via merchants | Clear "payment pending" messages, never "failed" for an unknown outcome |
| Backup processor | Before ramping | Warn them of volume so their risk controls do not block it |
| Card schemes (another name for the card networks, such as Visa or Mastercard) and acquiring partners (the banks or institutions that hold your merchant accounts with those networks) | Per contract | Incident notices where agreements require them |
| Regulators | Per jurisdiction | Where the platform is a regulated payment institution (licensed by a financial regulator to move customers' money, not just a software vendor), major operational incidents are reportable within fixed deadlines measured in hours to days (for example, under the EU's DORA, the Digital Operational Resilience Act, for financial entities). Legal and compliance own the filing; engineering supplies the timeline and impact numbers |
Trade-offs and pitfalls
- Cost of warmth. Keeping a second processor live costs contract minimums and engineering on two integrations, but a cold backup is almost always broken when needed.
- Automatic vs manual failover. Automatic failover reacts in seconds but can flap (switch back and forth rapidly as health checks briefly pass, then fail, then pass again) on a partial outage and send traffic to the backup that the primary would have approved. A common compromise: automatic for hard failures (connection refused, sustained 5xx), human decision for soft degradation (elevated latency, falling approval rate).
- The dangerous mistake is converting timeouts into failures and retrying them elsewhere. That produces the double charges that customers remember long after the outage.
- Approval rates will drop on the backup (new MIDs, no transaction history, different issuer relationships). Plan and tell merchants; do not treat it as a new incident.
For a payment that requires authorization, settlement, ledger update, and downstream fulfillment, analyze trade-offs between using two-phase commit/distributed transactions (XA) versus sagas with compensating transactions. Propose a robust saga-based architecture that minimizes inconsistent states and handles failures such as partial refunds and late-arriving chargebacks.
Sample Answer
Direct answer
Two-phase commit (2PC; XA is the specific standard, dating to X/Open, that lets databases and message brokers act as coordinated participants in a 2PC transaction) cannot span the systems a payment touches: the card network and the PSP (payment service provider) will never join your transaction coordinator, and a capture that has reached a bank cannot be "rolled back", only reversed with a new money movement. So the real choice is how to build a saga: a sequence of local transactions, each committed on its own, where every step has a compensating action (a new transaction that semantically undoes it, such as a void or refund) instead of a rollback. Use an orchestrated saga with a durable state machine per payment, idempotent steps, a transactional outbox, and an append-only double-entry ledger; model refunds and chargebacks as new sagas that post new entries, never as edits to old ones.
2PC/XA versus sagas for this flow
Two-phase commit: a coordinator asks every participant "can you commit?" (phase 1, participants lock their rows and vote), then tells all of them "commit" or "abort" (phase 2).
| Concern | 2PC / XA | Saga |
|---|---|---|
| Atomicity | All-or-nothing across participants that support it | Each step atomic locally; the whole is eventually consistent, with intermediate states visible |
| External parties (PSP, card network, warehouse system) | Cannot participate: no XA interface over HTTP APIs | Natural fit: each call is a step with a known compensation |
| Locks | Held across the whole protocol, including network waits | Only for each local transaction |
| Coordinator failure | Participants that voted "yes" are blocked holding locks until it returns | Orchestrator resumes from its persisted state |
| Availability | Every participant must be up for any payment to commit | Steps can wait and retry independently |
| Time horizon | Milliseconds to seconds | Can span days (capture at shipping, chargebacks months later) |
2PC is still reasonable inside one database or between two databases you own (say, the payments table and the ledger in the same Postgres cluster (one PostgreSQL database installation): just use one local transaction). It fails for the cross-company, multi-day shape of a payment.
The saga
stateDiagram-v2
[*] --> Authorized: authorize
Authorized --> Captured: capture
Authorized --> Voided: fulfilment fails / timeout
Captured --> Booked: ledger posting
Booked --> Fulfilled: fulfilment ok
Booked --> Settled: PSP deposits funds (async, hours to days later)
Booked --> Refunding: fulfilment fails
Refunding --> Refunded: refund settled
Fulfilled --> PartiallyRefunded: partial refund saga
Fulfilled --> Disputed: chargeback arrives
PartiallyRefunded --> Disputed: chargeback arrives
Steps and compensations:
| Step | Local transaction | Compensation |
|---|---|---|
| 1. Authorize | PSP call; store auth ID | Void the authorization (releases the hold; no money moved) |
| 2. Capture | PSP capture call (can be at shipping time) | Refund (a new money movement back to the card) |
| 3. Ledger update | Post balanced entries | Post reversing entries; never delete |
| 4. Fulfilment | Reserve and ship | Cancel shipment or trigger return |
| 5. Settlement | PSP deposits the captured funds into the platform's bank (asynchronous, hours to days after capture) | No compensation of its own: a refund or chargeback that arrives after settlement is recovered from the merchant's payable balance instead, as the worked example below shows |
Ordering to minimize inconsistent states, reconciled with the table and diagram above. The state machine and step table run Authorize, Capture, Ledger update, then Fulfilment, and that is the order this design recommends: put the cheapest-to-compensate steps first and the irreversible ones last. Concretely, split Fulfilment into its two halves and run the free half early: reserve inventory (the first half of step 4) before capture, so that if stock is gone you only void the authorization (free, step 1's compensation) instead of refunding a completed capture (fees lost, customer annoyed); only once the reservation confirms stock does the saga proceed to Capture, the Ledger update, and finally Fulfilment's second half, shipping. An alternative some businesses use instead, capture on ship (Authorize, reserve inventory, ship, then Capture, then Ledger update), delays charging the customer longer but risks the card expiring or being declined between order and ship, and it changes Capture's compensation from a refund to a void, since nothing was captured until shipping. This design's table and diagram assume the first pattern, capture at order time; capture-on-ship would reorder the Capture and Fulfilment rows and change Capture's compensation accordingly.
Making each step safe
- Durable orchestrator. One state row per payment (
state,version,next_action,deadline), updated with a compare-and-set (an atomic "update this row only if its version still equals the value I last read" operation) onversion, so two orchestrator workers cannot both advance it. A workflow engine such as Temporal (durable execution: the workflow's history is persisted, so code resumes after a crash) is a good fit; a state table plus a job scheduler also works. - Idempotency keys on every external call (a value attached to a request so that repeating it with the same key returns the original result instead of repeating its effect), derived deterministically:
payment_id + step + attempt_group(attempt_groupidentifies one logical attempt at a step: the orchestrator's own automatic retries of that same attempt, for example three network retries of one capture call, reuse the same attempt_group and so share one idempotency key and collapse into one PSP-side effect, while a fresh attempt started for a different reason, such as retrying capture after fixing a different bug, gets a new attempt_group and therefore a new key). A retried capture with the same key returns the original result from the PSP instead of capturing twice. - Transactional outbox. When a service updates its own table, it writes the event it needs to publish into an
outboxtable in the same local transaction; a relay publishes from the outbox. This removes the "committed but never announced" gap (the dual-write problem). - Unknown outcomes are a state. A PSP call that times out is
capture_unknown, not failed. The orchestrator queries the PSP by idempotency key before deciding to retry or compensate. Compensating on a timeout is how you refund money that was never captured, or capture twice. - Compensations must themselves be idempotent and retried until they succeed; a refund that fails is escalated to a human queue, never silently dropped.
Partial refunds
A refund is a new saga against the payment, not an edit. Rules the refund saga enforces:
- Refundable amount = captured minus refunds already succeeded or in flight. Checked under the payment row lock so two concurrent partial refunds cannot exceed the total.
- Each refund has its own ID and idempotency key and its own states (
requested → submitted → succeeded | failed). - Ledger entries post when the PSP confirms, and reverse the right accounts: merchant revenue, and a platform fee if your policy returns it.
Late-arriving chargebacks
A chargeback is a reversal forced by the cardholder's bank; it can arrive weeks or months after the sale, after refunds, after the merchant was paid out. Handle it as a separate dispute saga keyed by the original payment:
- Receive the dispute notification (webhook or dispute file); idempotent on the network's dispute ID.
- Post ledger entries immediately: move the disputed amount (plus the dispute fee) from the merchant's balance to a
disputes_heldaccount. If the merchant's balance is already paid out, it goes negative and is recovered from future sales or a reserve. - Check for overlap with refunds: a chargeback for money you already refunded is contestable ("credit already processed"), so submit that evidence automatically.
- Evidence deadline timer; on win, reverse the held entries; on loss, move them to a loss account.
Worked example
A 100.00 order. Each line is one balanced transaction: it debits one account and credits another by the same amount. "Merchant payable" is what the platform owes the merchant (a credit increases it, a debit reduces it); "PSP receivable" is what the PSP owes the platform, and it is cleared by an explicit Settlement entry once the PSP actually deposits the funds, which is what funds the Bank account below.
| Event | Debit | Credit | Amount |
|---|---|---|---|
| Capture | PSP receivable | Merchant payable | 100.00 |
| Settlement: the PSP deposits the captured funds into the platform's bank | Bank | PSP receivable | 100.00 |
| Partial refund of 30.00 (damaged item) | Merchant payable | PSP receivable | 30.00 |
| Merchant paid out | Merchant payable | Bank | 70.00 |
| Chargeback for the full 100.00 plus a 15.00 dispute fee (illustrative), eight weeks later | Disputes held | PSP receivable | 115.00 |
| Recover from merchant (the chargeback plus the fee; policy: the merchant bears network dispute fees) | Merchant payable | Disputes held | 115.00 |
After the chargeback the merchant payable balance is 100 − 30 − 70 − 115 = −115.00: the merchant owes the platform 115.00 (the 100.00 disputed sale plus the 15.00 dispute fee they bear per policy). The system automatically submits evidence that 30.00 was already refunded; if the issuer accepts it, 30.00 comes back (debit PSP receivable, credit Merchant payable 30.00), leaving the merchant at −85.00, which exactly matches the 70.00 they actually received for a sale the cardholder successfully disputed, plus the 15.00 dispute fee. Every intermediate state is explainable from the entries, and nothing was edited.
A note on Settlement and the account balances, since two of them go negative and neither is an error: once Settlement moves the captured funds into Bank, PSP receivable returns to zero if nothing else happens. A refund or a chargeback that happens AFTER settlement, as both do here, pushes PSP receivable negative, because that money has already left PSP receivable and landed in Bank; the negative balance is what the PSP will net out of its next settlement to the platform (or, absent one in time, an amount the platform must wire back). Bank itself never goes negative in this example: Settlement funds it with 100.00 before any of that money is paid out (70.00) or clawed back, which is exactly why Settlement has to be an explicit step even though the question only names it in passing. Without it, the "Merchant paid out" line would debit Bank with no deposit ever having credited it.
Trade-offs and pitfalls
- Sagas expose intermediate states. A customer may see "charged" before "order confirmed". Design the UI and support tooling around explicit states rather than pretending to atomicity.
- Choreography versus orchestration. Choreography (each service reacts to events, no central brain) looks decoupled but scatters the payment's logic across services; for money, one orchestrator you can query per payment wins.
- Semantic, not exact, undo. A refund is not the inverse of a capture: fees may not come back, currency rates may have moved. Model those differences as explicit ledger lines.
- Reconciliation is the safety net. A daily job compares the PSP's settlement report with your ledger by payment ID; every mismatch is a saga that got stuck or a step your system never saw.
Design a public payments API for a client that must handle 1 billion requests per day across multiple regions, comply with PCI-DSS, provide sub-100ms median latency for non-card operations, and support partner integrations. Cover the architecture components, security controls, data-residency strategy, and how you'd version and evolve the API without breaking partner integrations.
Sample Answer
Direct answer
Build it as a set of regional cells behind a global edge (a worldwide network of entry points close to each partner, which forwards each request into the right region), with each partner account pinned to a home region inside its data-residency jurisdiction, and with card data confined to a small, separately secured card vault so that almost every request (reads, refunds, payouts, webhooks) never touches raw card numbers. The sub-100ms median target is met by answering non-card calls entirely inside the home region from caches and a local primary database. The API evolves through date-based versions pinned per partner, with the server translating between versions, so partners upgrade on their own schedule and nothing they depend on changes underneath them.
Requirements and load
First, a few terms. PCI DSS (Payment Card Industry Data Security Standard) is the card industry's security standard; its current version is v4.0.1. The CDE (cardholder data environment) is every system that stores, processes or transmits card numbers, and every system connected to those; all of it is audited against PCI DSS. A PAN (primary account number) is the long card number. Data residency means a customer's personal and payment data must be stored and processed inside a named jurisdiction (for example, EU data stays in the EU).
Turn "1 billion requests per day" into a rate:
86,400 s109 requests11,574×3 (assumed peak-to-average ratio)≈11,574 RPS average≈34,722 RPS at peak(RPS is requests per second.) The 3x peak factor is an assumption to state out loud and replace with measured traffic shape. Assume 5% of requests carry card data (creating a payment method or a card payment). That is 50 million card operations per day, about 579 RPS on average. This split is the whole architecture: 95% of the traffic can be served by systems outside the CDE, and the CDE can stay small.
Assume three jurisdictions (US, EU, and a third such as India), each with two regions so a regional failure can fail over without data leaving the jurisdiction. If peak traffic splits evenly across jurisdictions, each carries about 11,574 RPS at peak, and each region in a pair must be sized to absorb its partner's full load: about 11,574 RPS per region, not half of it. Residency is what forces that doubling. You cannot fail EU traffic over to a US region.
Architecture components
flowchart LR
P[Partner client] --> E[Global edge: TLS, WAF, geo routing]
E --> G[Regional API gateway]
G --> A[Auth and rate limiter]
G --> I[Idempotency store]
G --> S[Payments services]
S --> D[(Regional primary DB)]
S --> V[Card vault in CDE]
V --> H[HSM or KMS]
S --> Q[Event bus]
Q --> W[Webhook dispatcher]
Q --> R[Reporting store]
| Component | Responsibility | Why it is there |
|---|---|---|
| Global edge | TLS termination (the point where the partner's encrypted connection is decrypted before the request continues inside the network) near the partner, WAF (web application firewall, which blocks malicious request patterns), absorption of DDoS (distributed denial-of-service) floods, routing to the account's home region | Cuts round-trip time and keeps attack traffic off the regional cells |
| Regional API gateway | API-key or OAuth token validation (OAuth is the standard for delegated access tokens), per-partner rate limits, request schema validation, version translation | One place to enforce the contract before business logic runs |
| Idempotency store | Records each write's idempotency key (a client-chosen unique ID for one intended operation) and its stored response | Partners retry on timeouts; without this, retries create duplicate charges |
| Payments services | Payment state machine, refunds, payouts, partner configuration | Business logic, outside the CDE |
| Card vault | Accepts PANs, returns tokens, talks to card networks and processors | The only place PANs exist, so the audited footprint is small |
| HSM / KMS | A hardware security module or a key management service holds encryption keys; software never sees raw key material | Keys that protect card data never sit next to the data |
| Event bus and webhook dispatcher | Publishes state changes; delivers signed webhooks to partners with retries | Partners integrate asynchronously instead of polling |
| Reporting store | Read-optimised copy for dashboards and exports | Keeps heavy analytical reads off the transactional database |
Routing to the home region. Every account ID and object ID embeds its home-region code (for example acct_eu1_...). The edge reads the partner's credential, looks up the home region from a small replicated routing table (non-personal data only), and forwards there. Object IDs carrying the region mean any request for an existing payment routes correctly without a global lookup.
Cells. A cell is a complete, independent copy of the service stack. A region is the geographic and legal boundary (where the data is allowed to live); a cell is a smaller blast-radius boundary inside it, and a region normally runs several of them. Partners are spread across a region's independent cells (each serving a subset of partners). A bad deployment or a runaway partner then affects one cell, not the region.
Meeting sub-100ms median for non-card operations
A median (p50, the 50th percentile) under 100ms is a budget you allocate, then verify with load tests. An allocation that fits:
| Step | Budget |
|---|---|
| Partner to edge, TLS reuse on a warm connection | 20 ms |
| Edge to home region | 25 ms |
| Gateway: auth (cached key lookup), rate limit, validation | 5 ms |
| Idempotency check and record (writes only) | 5 ms |
| Service logic plus one or two primary-database queries by primary key | 20 ms |
| Headroom | 25 ms |
These are allocations, not measurements. What makes them achievable: no cross-region calls on the request path, API-key verification cached in the gateway, every read by indexed key, and no call to external processors for non-card operations. Card operations are explicitly exempt from the 100ms target because they wait on card networks and issuers you do not control; publish a separate objective for them.
Security controls
- Scope reduction first. Partners collect cards through hosted fields (payment form fields served by an iframe from the vault's own origin, so the card number is typed directly into the vault, never the partner's page) or client SDKs that send the PAN straight to the vault, which returns a token. The partner's servers and your main services only ever see tokens, so they are outside the CDE.
- Card vault hardening. A separate network segment and separate cloud account, no human interactive access by default, envelope encryption (each record encrypted with a data key, and data keys encrypted by a master key held in the HSM or KMS), and key rotation.
- Partner authentication. Secret API keys scoped per environment (test vs live) and per permission set; OAuth for partner platforms acting on behalf of merchants; optional mTLS (mutual TLS, where the client also presents a certificate) for high-value partners.
- Service-to-service. mTLS between internal services and authorization checks on every call, so a compromised service cannot freely call the vault.
- Abuse controls. Per-partner and per-endpoint rate limits, anomaly detection on card-testing patterns (many small authorizations with different cards), and fraud scoring on payments.
- Webhooks. Signed with an HMAC (a keyed hash) over the timestamp and body so partners can verify origin and reject replays.
- Audit. Append-only audit logs of administrative actions and key usage, shipped to storage the platform's own operators cannot edit.
Data-residency strategy
- The partner chooses a home jurisdiction at onboarding; all personal and payment data for that account lives in that jurisdiction's two regions.
- The global control plane (the one piece of the platform allowed to be non-regional; it maps account IDs to home regions and holds nothing else) stores only non-personal routing metadata (account ID to region), never names, emails or card data.
- Replication and backups stay within the jurisdiction's region pair.
- Cross-border needs (a global partner selling in both EU and US) are handled as separate accounts per jurisdiction linked by a parent organisation, not by moving data.
- Analytics that must be global use aggregated or anonymised extracts produced inside each jurisdiction.
Versioning and evolution
- Additive changes ship to everyone. New optional request fields, new response fields, new endpoints and new enum values are non-breaking. Partners are told up front to ignore unknown fields and handle unknown enum values; that expectation is part of the contract.
- Breaking changes get a new dated version (for example
2026-09-01). Each partner account is pinned to the version it first integrated against. A request can override with a version header, which lets a partner test the next version one call at a time. - Translation layer. Internally there is one current model. The gateway runs a chain of small transforms that downgrade responses (and upgrade requests) between adjacent versions. Old versions cost one transform each rather than a forked codebase.
- Webhooks follow the account's pinned version, so an event payload does not change shape under a partner.
- Safety net. Consumer contract tests (automated checks that a real partner request shape still gets back the response that partner's code expects) for the most common partner request shapes run in CI; per-version traffic metrics show who still uses what; deprecation is announced with a long notice period and only retired when traffic reaches zero or the partner has been contacted.
Worked example
A partner in Germany integrated in 2025, pinned to version 2025-03-01. In 2026 the platform renames amount to a nested amount: {value, currency} object. That is breaking, so it ships as 2026-09-01. The German partner's GET /v1/payments/pay_eu1_8f2... routes by the eu1 prefix to Frankfurt, reads one row from the regional primary, and the gateway's downgrade transform flattens the nested object back into the old shape. No card data is read, no processor is called, and the partner sees exactly the response it saw last year. If Frankfurt fails, the request is served from the EU partner region, never from the US.
Trade-offs and pitfalls
- Residency doubles capacity cost. Each jurisdiction needs two full regions; that is the price of failing over without moving data. A single global active-active database (one dataset that every region can write to directly, with every write visible everywhere) would be cheaper and simpler, and illegal for many customers.
- Version transforms accumulate. Each old version adds a transform to maintain and test. Put a lifetime on versions and invest in migration tooling, or the chain becomes the slowest part of the gateway.
- Common wrong turn: putting the whole platform in PCI scope because one service touches PANs. That makes every deploy an audit event and slows everything.
- Common wrong turn: one global idempotency store. Keys must be checked in the home region with strong consistency (every read sees the latest write immediately, never a stale replica's answer), or two regions can both accept the same retried payment.
- Latency pitfalls: synchronous calls to the fraud engine or to a global service on the non-card path quietly break the median; move them off the request path or make them regional.
Unlock Full Question Bank
Get access to all 15 Payment and Transaction Processing Systems interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.