Direct answer
Use opaque tokens at the edge and short-lived, audience-scoped JWTs inside the mesh (the network of proxies that carries and secures service-to-service traffic across the estate), minted by a central token service through token exchange at each trust hop. The edge gateway turns the user's external token into an internal token that carries only who the user is and which service it is meant for, not every permission they hold. Downstream services authenticate the calling service with its workload identity (for example mTLS, mutual TLS where both sides present certificates), validate the internal token locally, and ask a policy decision point for fine-grained decisions, with a local sidecar cache (a small helper process deployed alongside each service instance, sharing its network so it can intercept and answer calls without a remote hop) to keep that off the latency critical path. Revocation is handled by keeping internal tokens to a few minutes and pushing revocation events for the rare cases that cannot wait.
Vocabulary first
- JWT (JSON Web Token): a signed, self-contained token. Any service holding the issuer's public key can verify it without a network call. Downside: once issued it is valid until it expires.
- Opaque token: a random string meaningful only to the issuer. Verifying it means calling the issuer (introspection, RFC 7662). Revocation is instant, but every check is a network call.
- Token exchange (RFC 8693): an OAuth 2.0 grant where a service presents a token it holds and receives a new token for a different audience, typically with narrower scope.
- PEP / PDP / PAP / PIP: the policy enforcement point (code that blocks or allows the request), policy decision point (evaluates the policy), policy administration point (where policies are authored and versioned) and policy information point (the source of attributes such as team membership).
Architecture
mermaid
flowchart LR
U[Client] -->|opaque token| GW[Edge gateway]
GW -->|exchange| TS[Token service]
TS -->|internal JWT aud=orders| GW
GW -->|JWT + mTLS| O[Orders service]
O -->|exchange for aud=payments| TS
O -->|new JWT + mTLS| P[Payments service]
O -.->|check| SC1[Policy sidecar]
P -.->|check| SC2[Policy sidecar]
PAP[Policy admin and store] -->|bundles| SC1
PAP -->|bundles| SC2
1. Token formats: commit per boundary
| Boundary | Format | Why |
|---|
| Internet client to edge | Opaque access token (or a JWT validated only at the edge) | Clients are untrusted and tokens leak (logs, browser storage); opaque tokens can be revoked instantly and reveal nothing |
| Edge to internal services | Short-lived JWT (2 to 5 minutes), signed by the internal token service | Hundreds of services each verify locally with a cached public key, so there is no per-hop call to the issuer |
| Service to service | Workload identity (mTLS certificate) plus the propagated user JWT when acting on a user's behalf | mTLS proves which service is calling; the JWT proves on whose behalf |
Introspecting an opaque token on every internal hop does not scale: 20,000 requests per second fanning out through an average of 6 internal hops is 120,000 introspection calls per second against one service, which becomes the platform's single point of failure.
2. Token exchange for internal calls
Forwarding the user's original token unchanged through the call chain is the common anti-pattern: every downstream service receives a token that is valid at every other service (a replayable bearer credential: whoever holds the bytes can use it, with no extra proof that they are the party it was issued to), and a compromised low-value service can replay it against a high-value one.
Instead, each hop exchanges:
- Orders receives a JWT with
aud: orders.
- To call Payments it presents that token plus its own workload identity to the token service and requests
aud: payments.
- The token service checks policy ("may orders call payments on behalf of a user?") and returns a JWT with
aud: payments, the same sub (the user), and an act (actor) claim recording that orders is the intermediary. RFC 8693 defines act for exactly this delegation chain.
Payments rejects any token whose aud is not payments, so a token stolen from Orders is useless against Payments. Cost: one extra call per cross-service hop, which you cache: the exchanged token is reusable for its lifetime for the same user and audience.
3. Refresh and revocation
- Internal tokens are not refreshed; they are re-minted. They live 2 to 5 minutes, so the worst-case window in which a revoked user keeps access internally is that lifetime.
- Refresh tokens exist only between the client and the authorization server, are rotated on every use (a reused refresh token signals theft and revokes the whole family), and never enter the mesh.
- Urgent revocation (account takeover, fired employee): the auth server publishes a revocation event (
sub or session ID plus a timestamp) on a stream; sidecars and gateways keep a small in-memory deny set and reject any token for that subject issued before the revocation time. The set stays small because entries can be dropped once every token issued before them has expired.
4. Minimizing token bloat
Putting a user's permissions in the token is the fastest way to break the platform. Measured by encoding real JWTs (ES256: an elliptic-curve signing algorithm producing a compact, fixed 64-byte signature, versus RS256, an RSA signature that runs 256 bytes for a 2048-bit key; compact JSON): jwt_len builds the same three parts a real token has (a header naming the algorithm and key, a claim set, and a fixed-length signature), base64url-encodes each, and adds them up, once with a bare claim set and once with a growing list of permission strings appended to it, to measure how token size scales with how many permissions ride inside it.
python
import base64, json
def b64url(b):
return base64.urlsafe_b64encode(b).rstrip(b"=")
def jwt_len(claims):
header = {"alg": "ES256", "kid": "2026-09-a", "typ": "JWT"}
h = b64url(json.dumps(header, separators=(",", ":")).encode())
p = b64url(json.dumps(claims, separators=(",", ":")).encode())
sig = b64url(bytes(64)) # an ES256 signature is always 64 bytes (r || s)
return len(h) + 1 + len(p) + 1 + len(sig)
base = {"iss": "https://auth.internal", "sub": "user:48213", "aud": "orders-api",
"exp": 1790000000, "iat": 1789999100, "jti": "8f14e45fceea167a"}
print("minimal claims:", jwt_len(base), "bytes")
for n in (50, 200, 500):
fat = dict(base, perms=[f"tenant-{i:04d}:orders:read" for i in range(n)])
print(f"{n} permission strings:", jwt_len(fat), "bytes")
text
minimal claims: 319 bytes
50 permission strings: 2066 bytes
200 permission strings: 7266 bytes
500 permission strings: 17666 bytes
Each permission string like tenant-0004:orders:read is 23 raw characters; as one element of a JSON array that is 26 bytes (23 plus 2 quote bytes plus a comma), and base64url encoding then expands every 3 raw bytes into 4 encoded characters, a 4/3 blow-up, so each permission costs about 26 x 4/3 ~= 34.7 bytes in the final token. That matches the numbers above exactly: 200 permissions minus 50 is 7266 - 2066 = 5200 bytes over 150 permissions, 5200 / 150 ~= 34.7 bytes each. nginx's default buffer for one large request header line is 8 KB, so a user with a few hundred permissions starts getting unexplained 4xx errors at the proxy, and a 7 KB header repeated across 6 hops is 42 KB of overhead per request. The fix is structural, not compression: keep identity claims (sub, aud, exp, iat, jti, tenant, maybe a coarse role) in the token and resolve permissions at decision time from the PIP. Using ES256 rather than RS256 also helps: its signature is 64 bytes versus 256 bytes for a 2048-bit RSA key.
5. Fine-grained authorization downstream
Three options, and where each fits:
| Pattern | What it is | Use it when |
|---|
| Authorization middleware in each service | A shared library checks scopes and simple rules in-process | Coarse checks: "does this token have orders:write?" |
| Centralized PDP (policy-as-a-service) | Services call one authorization service per decision | Relationship-heavy rules ("can this user edit this document?") where the data is large and central |
| Distributed PDP: policy sidecar with pushed bundles | Policies authored centrally (PAP), compiled into bundles, pushed to a sidecar next to each service (Open Policy Agent, OPA: a widely used open-source policy engine that evaluates authorization rules written in its own policy language, is the common one) | The default for hundreds of services: sub-millisecond local decisions, no central runtime dependency |
Recommendation: middleware for coarse scope checks, a sidecar PDP for most fine-grained rules, and a central relationship service only for resource-level sharing graphs. Every decision is logged with the policy version that produced it so an auditor can reconstruct why access was granted.
Trade-offs and pitfalls
- JWTs trade revocation for scale. The honest statement of the design is "internal access can outlive revocation by at most the internal token lifetime", and that lifetime is the knob.
- Token exchange adds a dependency. If the token service is down, cross-service calls fail. Run it per region, cache exchanged tokens, and treat it with the same SLO (service-level objective) as the edge.
- Do not let services mint their own internal tokens. Only the token service holds signing keys; services only verify.
- Policy drift: sidecars running stale bundles make inconsistent decisions. Export the bundle version each sidecar is running and alert when any lags the published version beyond a threshold.
- Common wrong turn: using the same token for user identity and service identity. Keep them separate: the certificate says which service, the token says which user.