Payment and Transaction Processing Systems Questions
Designing systems that move money correctly: idempotent payment flows, exactly-once semantics, reconciliation, ledgers, double-entry accounting, and fraud-detection architecture. Covers handling retries and partial failures without double-charging, and the consistency guarantees payments demand. Also covers protecting cardholder data through tokenization and PCI scope reduction, reconciling processor webhooks, and merchant and partner payouts. A high-stakes specialization of distributed transactions.
Propose a practical idempotency-key scheme for HTTP payment requests: what should uniquely identify a request, what should the key be scoped to, and how would you choose TTLs for a short-lived authorization versus a longer-running capture? Explain how you'd expire or archive keys safely in a distributed environment.
Sample Answer
Direct answer
A request is identified by who sent it plus the key they chose: the authenticated merchant or API credential, and a client-generated random key (a UUID, universally unique identifier) for one intended operation. Alongside it, store a fingerprint of the request (method, path and a hash of the canonical body (the request body serialized in one fixed, deterministic form, for example JSON with keys sorted and no extra whitespace, so two payloads that mean the same thing always hash the same)) so the same key with a different body is rejected rather than silently replayed. TTLs follow the operation, not a single global number: a key must outlive the client's retry window and the time the operation can remain unfinished, which is short for an authorization and longer for a capture. Keys expire only once they are in a completed state, and expiry is enforced by comparing timestamps at read time, never by trusting a background deleter to have run.
Key terms
- Idempotency key: a unique value the client sends with a request (for example in an
Idempotency-Keyheader) so that repeats of that request produce one effect and the same response. - Authorization: the bank approves and holds funds. Capture: the merchant claims held funds. An authorization expires if not captured in time.
- TTL (time-to-live): how long a record is kept before it may be removed.
- Fingerprint: a hash of the request's meaningful content, used to detect a reused key on a different request.
- Lease: a time-limited claim on a record, so a crashed worker's claim eventually lapses.
What uniquely identifies a request
| Part | Why it is needed |
|---|---|
| Scope: authenticated merchant (or API key) ID | Keys are only unique per client. Scoping stops one client's key from ever colliding with, or reading the stored response of, another's |
| The idempotency key | Identifies one intended operation: "charge order 5531", not "this HTTP attempt" |
| Fingerprint: method, path, hash of the canonical body | Detects the same key reused for a different amount, currency or card, which is a client bug to reject loudly |
The storage key is (merchant_id, idem_key); the fingerprint is a stored attribute compared on every repeat. Do not include the fingerprint in the key itself, or a buggy client that changes the amount on retry would create a second charge instead of an error.
Client-generated or server-generated keys
- Client-generated (default). Only the client knows that two HTTP attempts are the same intention, so only it can supply the key before the first attempt. Tell clients to use a random UUID, not a sequential number or anything containing personal data.
- Server-generated, two-step. For clients you do not trust to generate keys properly (a mobile app with flaky storage), first
POST /payment_intentscreates an object and returns its ID; thenPOST /payment_intents/{id}/confirmis naturally idempotent because the object's state machine only allows one confirmation. The first step still needs protection against duplicates, but a duplicate there creates an unused object, not a charge.
Defending against malicious or replayed keys
- Keys are identifiers, not secrets. A key seen by another party must not grant anything. Because lookups are always scoped to the authenticated credential, another client replaying your key gets a fresh, independent request under their own scope, and can never read your stored response.
- Fingerprint mismatch returns an error (for example HTTP 422) rather than the stored response, so a key cannot be used to fetch a response for a different request.
- Validate the format and length (Stripe, for example, accepts keys up to 255 characters) and rate-limit key creation per client, so an attacker cannot fill the store with junk keys.
- Never let a key cross merchants in a platform setting: a partner acting for merchant A and merchant B must have keys scoped by the merchant it is acting for.
Handling a key that is still in flight
The first request atomically inserts (merchant_id, idem_key, fingerprint, state=in_progress, lease_until). The unique constraint is the lock: a concurrent duplicate fails the insert, sees in_progress, and gets HTTP 409 ("still processing, retry shortly") instead of starting a second charge.
If the worker crashes, the lease lapses. The next retry takes over the record, but it must not simply re-run the charge: the crashed worker may have reached the processor (the external payment processor, such as a card network's gateway, that actually authorizes or captures the funds). It first queries the processor by the same reference (send the idempotency key, or a value derived from it, as your reference to the processor) and only charges if the processor has no record.
Runnable demonstration
This uses SQLite and threads to show the in-flight, replay, mismatch and cross-merchant cases, then 20 concurrent retries of one new key.
import hashlib, json, sqlite3, threading, os
DB = "/tmp/idem-demo.db"
if os.path.exists(DB):
os.remove(DB)
c = sqlite3.connect(DB)
c.execute("""CREATE TABLE idem (
merchant_id TEXT, idem_key TEXT, fingerprint TEXT NOT NULL,
state TEXT NOT NULL, response TEXT,
PRIMARY KEY (merchant_id, idem_key))""")
c.commit(); c.close()
charges_executed = []
exec_lock = threading.Lock()
def fingerprint(method, path, body):
canon = json.dumps(body, sort_keys=True, separators=(",", ":"))
return hashlib.sha256(f"{method} {path} {canon}".encode()).hexdigest()
def handle(merchant_id, key, method, path, body, gate=None):
fp = fingerprint(method, path, body)
conn = sqlite3.connect(DB, timeout=30, isolation_level=None)
try:
try: # claim the key atomically: the primary key is the lock
conn.execute("INSERT INTO idem VALUES (?,?,?,'in_progress',NULL)",
(merchant_id, key, fp))
except sqlite3.IntegrityError:
row_fp, state, resp = conn.execute(
"SELECT fingerprint, state, response FROM idem "
"WHERE merchant_id=? AND idem_key=?", (merchant_id, key)).fetchone()
if row_fp != fp:
return 422, "key reused with a different request"
if state == "in_progress":
return 409, "original request still in flight, retry later"
return 200, "replayed: " + resp
if gate: gate.wait() # hold the first request open (demo only)
with exec_lock:
charges_executed.append((merchant_id, key))
resp = json.dumps({"charge": f"ch_{len(charges_executed)}", "amount": body["amount"]})
conn.execute("UPDATE idem SET state='completed', response=? "
"WHERE merchant_id=? AND idem_key=?", (resp, merchant_id, key))
return 201, resp
finally:
conn.close()
body = {"amount": 5000, "currency": "USD", "payment_method": "pm_abc"}
gate = threading.Event()
out = {}
t = threading.Thread(target=lambda: out.setdefault("first", handle("m_1", "k-123", "POST", "/v1/charges", body, gate)))
t.start()
import time
while True: # wait until the first request has claimed the key
r = sqlite3.connect(DB).execute("SELECT state FROM idem").fetchone()
if r: break
time.sleep(0.01)
print("retry while first is in flight:", handle("m_1", "k-123", "POST", "/v1/charges", body))
gate.set(); t.join()
print("first request:", out["first"])
print("retry after completion: ", handle("m_1", "k-123", "POST", "/v1/charges", body))
print("same key, amount changed: ", handle("m_1", "k-123", "POST", "/v1/charges", dict(body, amount=9000)))
print("same key, different merchant: ", handle("m_2", "k-123", "POST", "/v1/charges", body))
# 20 concurrent retries of one brand-new key: exactly one charge may run.
threads = [threading.Thread(target=handle, args=("m_1", "k-race", "POST", "/v1/charges", body))
for _ in range(20)]
for th in threads: th.start()
for th in threads: th.join()
print("charges executed for k-race:", charges_executed.count(("m_1", "k-race")))
Output:
retry while first is in flight: (409, 'original request still in flight, retry later')
first request: (201, '{"charge": "ch_1", "amount": 5000}')
retry after completion: (200, 'replayed: {"charge": "ch_1", "amount": 5000}')
same key, amount changed: (422, 'key reused with a different request')
same key, different merchant: (201, '{"charge": "ch_2", "amount": 5000}')
charges executed for k-race: 1
The last line is the property that matters: 20 simultaneous retries produce exactly one charge, because only one insert of the key can succeed.
Runnable demonstration: the lease and crash takeover
The demo above shows the fast paths (in-flight, replay, mismatch, cross-merchant, and the concurrent race). This one adds the lease_until column the prose above described and actually traces the two crash cases it talks about: a worker that dies before ever reaching the processor, and the more dangerous case, a worker that dies right after the processor accepted the charge but before the key is marked completed.
import hashlib, json, sqlite3, threading, os, time
DB = "/tmp/idem-lease-demo.db"
if os.path.exists(DB):
os.remove(DB)
c = sqlite3.connect(DB)
c.execute("""CREATE TABLE idem (
merchant_id TEXT, idem_key TEXT, fingerprint TEXT NOT NULL,
state TEXT NOT NULL, response TEXT, lease_until REAL NOT NULL,
PRIMARY KEY (merchant_id, idem_key))""")
c.commit(); c.close()
processor_charges = {} # stands in for asking the real processor "did you already see this reference?"
charges_executed = []
exec_lock = threading.Lock()
def fingerprint(method, path, body):
canon = json.dumps(body, sort_keys=True, separators=(",", ":"))
return hashlib.sha256(f"{method} {path} {canon}".encode()).hexdigest()
def handle(merchant_id, key, method, path, body, lease_seconds=30, crash_after_claim=False,
crash_after_processor_charge=False):
fp = fingerprint(method, path, body)
conn = sqlite3.connect(DB, timeout=30, isolation_level=None)
now = time.time()
try:
try:
conn.execute("INSERT INTO idem VALUES (?,?,?,'in_progress',NULL,?)",
(merchant_id, key, fp, now + lease_seconds))
except sqlite3.IntegrityError:
row_fp, state, resp, lease_until = conn.execute(
"SELECT fingerprint, state, response, lease_until FROM idem "
"WHERE merchant_id=? AND idem_key=?", (merchant_id, key)).fetchone()
if row_fp != fp:
return 422, "key reused with a different request"
if state == "completed":
return 200, "replayed: " + resp
if lease_until > now: # the claimant might just be slow, not crashed
return 409, f"original request still in flight (lease held for {lease_until - now:.1f}s more), retry later"
cur = conn.execute( # lease expired: try to take over, atomically
"UPDATE idem SET lease_until=? WHERE merchant_id=? AND idem_key=? "
"AND state='in_progress' AND lease_until<=?",
(now + lease_seconds, merchant_id, key, now))
if cur.rowcount == 0:
return 409, "another worker already took over this expired lease, retry later"
existing = processor_charges.get((merchant_id, key))
if existing is not None: # the crashed worker DID reach the processor: never re-charge
conn.execute("UPDATE idem SET state='completed', response=? "
"WHERE merchant_id=? AND idem_key=?", (existing, merchant_id, key))
return 201, "took over crashed worker's key, found existing processor charge, finalized: " + existing
with exec_lock: # the crashed worker never reached the processor: safe to charge
charges_executed.append((merchant_id, key))
resp = json.dumps({"charge": f"ch_{len(charges_executed)}", "amount": body["amount"]})
processor_charges[(merchant_id, key)] = resp
conn.execute("UPDATE idem SET state='completed', response=? "
"WHERE merchant_id=? AND idem_key=?", (resp, merchant_id, key))
return 201, "took over crashed worker's key, processor had no record, charged: " + resp
if crash_after_claim:
return None, "SIMULATED CRASH before charging or completing"
with exec_lock:
charges_executed.append((merchant_id, key))
resp = json.dumps({"charge": f"ch_{len(charges_executed)}", "amount": body["amount"]})
if crash_after_processor_charge:
processor_charges[(merchant_id, key)] = resp
return None, "SIMULATED CRASH after the processor accepted the charge, before marking completed"
processor_charges[(merchant_id, key)] = resp
conn.execute("UPDATE idem SET state='completed', response=? "
"WHERE merchant_id=? AND idem_key=?", (resp, merchant_id, key))
return 201, resp
finally:
conn.close()
body = {"amount": 5000, "currency": "USD", "payment_method": "pm_abc"}
print("=== Case A: crash before any processor call ===")
print("worker 1 (crashes):", handle("m_1", "k-crash-a", "POST", "/v1/charges", body, lease_seconds=1, crash_after_claim=True))
print("retry immediately (lease not yet expired):", handle("m_1", "k-crash-a", "POST", "/v1/charges", body, lease_seconds=1))
time.sleep(1.1)
print("retry after lease expiry (takes over, charges):", handle("m_1", "k-crash-a", "POST", "/v1/charges", body, lease_seconds=30))
print("retry after completion (replays):", handle("m_1", "k-crash-a", "POST", "/v1/charges", body))
print()
print("=== Case B: crash after processor accepted the charge, before marking completed ===")
print("worker 1 (crashes after processor charge):", handle("m_1", "k-crash-b", "POST", "/v1/charges", body, lease_seconds=1, crash_after_processor_charge=True))
time.sleep(1.1)
print("retry after lease expiry (takes over, finds existing processor charge, does NOT re-charge):",
handle("m_1", "k-crash-b", "POST", "/v1/charges", body, lease_seconds=30))
print()
print("total charges actually executed against the simulated processor:", len(charges_executed))
Output:
=== Case A: crash before any processor call ===
worker 1 (crashes): (None, 'SIMULATED CRASH before charging or completing')
retry immediately (lease not yet expired): (409, 'original request still in flight (lease held for 1.0s more), retry later')
retry after lease expiry (takes over, charges): (201, 'took over crashed worker\'s key, processor had no record, charged: {"charge": "ch_1", "amount": 5000}')
retry after completion (replays): (200, 'replayed: {"charge": "ch_1", "amount": 5000}')
=== Case B: crash after processor accepted the charge, before marking completed ===
worker 1 (crashes after processor charge): (None, 'SIMULATED CRASH after the processor accepted the charge, before marking completed')
retry after lease expiry (takes over, finds existing processor charge, does NOT re-charge): (201, 'took over crashed worker\'s key, found existing processor charge, finalized: {"charge": "ch_2", "amount": 5000}')
total charges actually executed against the simulated processor: 2
Case A is the ordinary crash: a retry while the lease is still valid gets 409 (nobody assumes a crash just because a request is slow), and only once the lease actually expires does a retry take over and charge, exactly once. Case B is the one the direct-answer prose calls out: processor_charges stands in for asking the real processor by reference before charging, and because the takeover finds that record, it finalizes the existing charge instead of calling the processor a second time. Skip that processor lookup and Case B would double-charge the customer.
Choosing TTLs
A key must be kept at least as long as the longest of: the client's retry window, the time the operation can stay unfinished, and a safety margin for clock and queue delays.
| Operation | What bounds it | Suggested TTL for completed keys |
|---|---|---|
| Authorization (short-lived) | Clients retry for minutes to hours after a timeout; the authorization itself completes in seconds | 24 hours. Stripe documents the same shape: keys may be removed once at least 24 hours old, and a reused key after pruning is treated as a new request |
| Capture (longer-running) | Capture can happen days after authorization (online card authorizations are typically valid for about 7 days), and some captures complete asynchronously, staying pending for hours | Key TTL of about 8 days (authorization validity plus a margin), and rely on the payment's state machine as a second guard: a payment can only move from authorized to captured once, so even after the key expires a duplicate capture is rejected by state |
The general rule: idempotency keys protect against retries; the domain state machine protects against duplicates forever. Short TTLs are safe only where state guards exist.
Expiring and archiving safely in a distributed environment
- Expire only completed records. A record in
in_progressis never removed by TTL; it is resolved (by processor status query) or escalated. - Check expiry at read time. Stores with built-in TTL deletion (for example DynamoDB's TTL feature (Amazon's managed NoSQL database) or Redis key expiry (Redis is an in-memory key-value data store)) may delete late. Store
expires_at, and treat a record past it as expired regardless of whether it still exists, so behaviour does not depend on the deleter's timing. - Use server time from one source. Compute
expires_aton the server that owns the record, not from client clocks. - One home per key. In multi-region setups (the service runs in more than one geographic data-center region, usually for latency or resilience), route each merchant's writes to its home region and keep the idempotency store strongly consistent there (every read reflects the latest write immediately; no lagging copy of the data can answer a lookup with a stale "not found"). Two regions each checking their own replica (a copy of the data kept in sync, sometimes with a delay) can both accept the same retry if that consistency guarantee is missing.
- Archive before delete. Move expired records (key, fingerprint, outcome reference, timestamps, never full card data) to cheap storage for dispute and audit investigations, then delete from the hot store in small batches to avoid load spikes.
- Store the idempotency record in the same database transaction as the business change where possible, so "charge recorded" and "key marked completed" cannot disagree.
Trade-offs and pitfalls
- Storing full responses makes replays exact but costs storage; store the response body for small responses and a reference to the created object for large ones.
- Caching error responses: replaying a stored 500 error forever blocks a legitimate retry. Store results only once the operation has actually started (validation failures should not consume the key), and let retryable failures be retried.
- Keys in logs are fine; keys built from email addresses or card data are not.
- Wrong turn: deriving the key from the body hash on the server. Two genuinely separate purchases of the same item for the same amount would be merged into one.
Describe, step-by-step, the end-to-end flow for a typical online card payment (card-not-present) from customer checkout through authorization, capture, settlement, and reconciliation. Highlight which components commonly run in the merchant, payment gateway, processor, and card network, and where latency, failure, retries, and fraud checks typically occur.
Sample Answer
Direct answer
An online card payment is two separate journeys. The authorization is a real-time question ("will the cardholder's bank stand behind this amount?") that travels merchant → payment gateway → processor/acquirer → card network → issuing bank and back while the customer waits. The money movement happens later and in batches: the merchant captures (confirms the amount it wants), captured transactions are cleared through the network, funds are settled from the issuer to the merchant's bank a day or more later, and the merchant reconciles what arrived against what it expected. Latency and fraud checks concentrate on the synchronous authorization leg; failures that matter most are timeouts where you do not know whether the bank said yes.
The cast (who runs what)
| Party | What it is | What runs there |
|---|---|---|
| Merchant | The shop | Checkout UI, order service, its own fraud rules, the order and payment records, reconciliation |
| Payment gateway | The front door the merchant integrates with (often bundled with the processor as a PSP, payment service provider, e.g. Stripe or Adyen) | Card-field collection and tokenization (swapping the card number for a reference so the merchant never stores it), API, idempotency, routing, often a fraud-scoring product |
| Processor / acquirer | The acquirer is the merchant's bank for card acceptance; the processor is the technical platform that talks to the networks on its behalf | Formatting authorization messages, batching captures for clearing, paying the merchant out |
| Card network | Visa, Mastercard and similar | Routing authorizations to the right issuer, running clearing and settlement between banks, network-level fraud and dispute rules |
| Issuer | The cardholder's bank | Approve/decline decision, available-credit check, its own fraud model, placing the hold, and later the customer's statement |
Card-not-present (CNP) means the card is not physically at a terminal (online, in-app, phone), so there is no chip to prove the card is real. That is why CNP leans on extra checks such as the CVC (the three- or four-digit card verification code), address verification and 3-D Secure (3DS: a step where the issuer authenticates the cardholder, for example with a push to the banking app; completing it successfully also shifts fraud liability from the merchant to the issuer, since the issuer vouched for the cardholder).
sequenceDiagram
participant C as Customer
participant M as Merchant
participant G as Gateway/PSP
participant A as Processor/Acquirer
participant N as Card network
participant I as Issuer
C->>M: Pay 80.00
M->>G: Create payment (token, amount, idempotency key)
G->>A: Authorization request
A->>N: Route by card number range
N->>I: Authorize?
I-->>N: Approved, hold placed
N-->>A: Approved
A-->>G: Approved
G-->>M: Authorized
M-->>C: Order confirmed
M->>G: Capture 80.00 (later)
A->>N: Clearing batch
N->>A: Net settlement funds
A->>M: Payout net of fees
Step by step
- Checkout and tokenization (merchant + gateway). The card number is typed into fields served by the gateway, not the merchant's own form, and comes back as a token. This keeps most of the merchant's systems out of scope for PCI DSS (the Payment Card Industry Data Security Standard that governs anyone storing or handling card numbers).
- Create the payment (merchant → gateway). The merchant sends amount, currency, token and an idempotency key (a unique ID for this attempt so a retried request cannot create a second charge). Merchant-side fraud rules usually run here, before the bank is ever asked.
- Authentication, if needed (issuer). For risky transactions, or where regulation requires strong customer authentication (for example much of Europe), the issuer runs 3DS. This is the step with human-scale latency, because the customer may have to approve in an app.
- Authorization (gateway → processor → network → issuer and back). The issuer checks the account, available funds, CVC/address results and its own fraud score, then approves or declines. An approval places a hold: the customer's available credit drops, but no money has moved yet.
- Capture (merchant → gateway). The merchant confirms the final amount, immediately for digital goods or at shipment for physical goods. Authorizations expire: for online card payments Stripe documents a typical 7-day validity window for customer-initiated transactions (Visa merchant-initiated ones are shorter). An expired authorization cannot be captured.
- Clearing (processor → network → issuer). Captured transactions are sent in batches; the issuer converts the hold into a posted charge on the statement.
- Settlement (issuer → network → acquirer → merchant). Banks settle net positions through the network, and the acquirer pays the merchant the captured amount minus fees (interchange, the per-transaction fee paid to the card-issuing bank, set by the network; network fees; the acquirer/PSP markup), usually one to a few business days later depending on the contract.
- Reconciliation (merchant). The merchant matches the settlement report and the bank deposit against its own ledger: every captured payment should appear once, at the expected net amount, and anything else (missing items, fee differences, chargebacks, meaning a cardholder's bank forcibly reversing a payment after a dispute) becomes a break (an unexplained mismatch between what was expected and what arrived) to investigate.
Where latency, failure, retries and fraud checks sit
| Stage | Latency | Typical failures | Retry rule | Fraud control |
|---|---|---|---|---|
| Tokenize/create | Low | Validation errors | Safe with idempotency key | Merchant rules, velocity checks (limits on how many attempts, or how much value, one card, device or account can push through in a short window) |
| 3DS | Seconds to minutes (human) | Customer abandons | Customer-driven | Issuer authentication |
| Authorization | The bulk of machine latency: several network hops | Timeout (outcome unknown), soft decline (e.g. issuer unavailable), hard decline (e.g. stolen card) | Retry a timeout with the SAME idempotency key or query status first; do not blindly retry hard declines | PSP risk score, issuer model, network checks |
| Capture | Low, not customer-facing | Auth expired, amount above authorized | Retry with idempotency key; re-authorize if expired | Rarely |
| Clearing/settlement | Hours to days, batch | Missing or partial items | Handled by reconciliation, not API retries | n/a |
| Post-settlement | Weeks to months | Refunds, chargebacks | Separate flows | Dispute evidence |
The most important failure in the list is the authorization timeout: the request may have reached the issuer and been approved even though the merchant saw an error. Treating that as a decline and letting the customer click Pay again is how customers get charged twice. The correct handling is to mark the payment as unknown, retry with the same idempotency key (the gateway returns the original result) or query the gateway for the payment's status.
Worked example
A customer buys an 80.00 USD jacket on Monday.
- Monday 10:00: authorization for 80.00 approved; the customer's banking app shows an 80.00 pending item.
- Tuesday: the jacket ships; the merchant captures 80.00.
- Tuesday night: the capture goes out in the processor's clearing batch.
- Thursday: payout arrives. If the merchant's all-in pricing were 2.9% + 0.30 USD (an example rate, not a quote), the fee is 80.00 × 0.029 = 2.32, plus 0.30, so 2.62; the deposit line is 80.00 − 2.62 = 77.38 USD.
- Thursday reconciliation: the merchant's ledger expected 80.00 gross; the report shows 80.00 gross, 2.62 fee, 77.38 net. Matched. Had the merchant reconciled the bank deposit against gross revenue it would have seen a false 2.62 shortfall on every order.
Trade-offs and pitfalls
- Authorization is not payment. An approved authorization can still expire uncaptured, and captured money can still come back as a refund or chargeback months later. Fulfilment logic should key off the state it actually needs.
- Unknown is a real state. Systems that model only success/failure convert timeouts into duplicate charges or into lost sales.
- Soft vs hard declines. Retrying a "do not honour" (the generic issuer decline code that gives no specific reason) or stolen-card decline over and over looks like card testing (an attacker running many small charges to find out which stolen card numbers still work) to issuers and networks; retry only declines that are explicitly transient.
- Gross vs net. Reconciliation must model fees, currency conversion and timing, or it produces noise that ops teams learn to ignore.
- Capture timing is a business decision (capture now vs at shipment), with consequences for refunds, disputes and expiry; it should be chosen deliberately, not left at an SDK default.
Explain what idempotency means in payment processing. Provide concrete examples of problems that arise without idempotency (for example duplicate charges and duplicate fulfillment) and outline high-level strategies to implement idempotency across stateless, horizontally scaled services.
Sample Answer
Direct answer
An operation is idempotent if doing it twice has the same effect as doing it once. In payments that property has to be engineered, because "charge this card 50 USD" is naturally not idempotent: sending it twice charges twice. Duplicates are unavoidable (networks drop responses, clients retry, users double-click, message queues redeliver), so the system must recognise a repeat of the same intent and return the first result instead of acting again. In a horizontally scaled service that recognition cannot live in any one server's memory; it has to live in shared, durable state, usually a unique idempotency key stored alongside the payment and passed on to the processor.
Where duplicates come from
| Source | What happens | Why it is not a bug you can simply fix |
|---|---|---|
| Lost response | Server charged the card, the response died on a flaky mobile network, the app retries | The client cannot tell "never arrived" from "arrived, answer lost" |
| Impatient user | Double-click on Pay, or refresh on the confirmation page | Humans retry when things look slow |
| Load balancer or gateway retry | A proxy retries a request that timed out upstream | Retries are how distributed systems survive transient faults |
| Message redelivery | A queue delivers "payment succeeded" twice | Most queues are at-least-once (every message arrives, some more than once), because exactly-once delivery across a network is not achievable in general |
| Webhook replay | The PSP (payment service provider) resends an event because your endpoint was slow to acknowledge | Same at-least-once contract |
Problems without idempotency
- Duplicate charges. Two charge requests, two captures, one angry customer, a refund, and quite possibly a dispute.
- Duplicate fulfilment. A "payment succeeded" event processed twice ships two parcels or issues two gift-card codes. This one is often worse than a double charge because the goods are gone.
- Duplicate refunds or payouts. Money leaves the platform twice and must be clawed back.
- Corrupted ledger (the running record of every debit and credit). The same transaction posted twice means balances, revenue reports and reconciliation (matching two independent records of the same money to confirm they agree) are all wrong, and the error is discovered days later.
Strategies for stateless, horizontally scaled services
"Stateless" means any request can land on any of, say, 40 identical instances. So the memory of "I have seen this before" must be in a shared store, and the check-and-act must be atomic.
- Client-generated idempotency key. The client creates a random unique ID (for example a UUID, a universally unique identifier drawn from 122 random bits) once per payment attempt, stores it, and sends the same value on every retry. The server treats the key as the identity of the intent.
- Unique constraint, not check-then-insert. "SELECT to see if the key exists, then INSERT" lets two instances both see "not found" at the same moment. A database unique index on the key (scoped per merchant or per user) makes exactly one insert win.
- Store and replay the result. When the first request finishes, save its status code and body against the key. A retry gets the saved response, byte for byte, instead of a fresh execution.
- Fingerprint the request. Store a hash of the request body with the key. The same key with a different amount is a client bug, and should be rejected rather than silently replayed.
- Propagate the key downstream. When calling the processor, pass a key derived from yours. Then even if your service crashes between calling the processor and recording the result, the re-driven call cannot charge twice.
- Idempotent consumers for events. A service consuming "payment succeeded" records the event ID in a processed-events table in the same database transaction as its side effect. A redelivered event hits the unique constraint and is skipped.
- Make state transitions conditional. "Mark order shipped where status = 'paid'" is naturally idempotent: the second attempt matches zero rows. Modelling a payment as a state machine with legal transitions gets you a lot of idempotency for free.
Worked example
Maya taps Pay for a 50.00 USD order on a train.
Without a key: request 1 reaches the server, the card is charged, the response is lost in a tunnel. The app shows "Something went wrong", Maya taps Pay again, request 2 charges another 50.00. Her statement shows 100.00; support refunds 50.00; she may also have filed a dispute with her bank.
With a key: the app generated key 7f3c... when the order was created. Request 1 inserts the key, charges the card through the processor using the same key, stores 201 {payment_id: pay_881}. The retry with 7f3c... hits the unique constraint, finds the stored response and returns pay_881. One charge, and Maya sees "Paid".
And if the retry arrived while request 1 was still talking to the processor, it would find the key in an "in progress" state and get a "try again shortly" answer rather than starting a second charge.
Trade-offs and pitfalls
- Idempotency is not exactly-once. The request may still arrive and be checked several times; idempotency makes the effect happen once. People call this "effectively once".
- A key generated on each retry is useless. The key must be created once per intent and persisted on the client before the first attempt.
- Keys expire. Stores keep keys for a bounded window (Stripe documents that keys can be pruned once they are at least 24 hours old), so a retry after that window is treated as new.
- Scope matters. A key unique per merchant is fine; a key unique only per server instance is not.
- Do not cache failures you should retry. Decide deliberately whether a stored error is replayed or whether a retry is allowed to try again; validation errors and "conflict, still in progress" are usually not stored.
That is every published Payment and Transaction Processing Systems question for Full-Stack Developer so far. Browse the other topics in this category, or practice this one interactively.