Zero Trust, Segmentation, and Service-to-Service Security Questions
Designing network and service-communication trust models where no implicit trust is granted by network location. Covers zero-trust access, microsegmentation and identity-aware perimeters, least-privilege network access, lateral-movement prevention, and segmenting environments to contain blast radius, together with securing service-to-service communication in distributed and microservices architectures: mutual authentication between services, service mesh security, multi-tenancy isolation, east-west traffic, and the security implications of scale and geographic distribution. The architectural trust-boundary pattern and its enforcement across decomposed, high-scale systems, distinct from device-level firewall configuration.
Compare JSON Web Tokens and opaque tokens for service authentication: local verification versus introspection, how each is revoked, size and transport considerations, and when you'd prefer one over the other. What common mistakes should a reviewer look for when a service validates a JWT (algorithm confusion, missing audience or expiry checks)?
Sample Answer
Direct answer: a JSON Web Token (JWT) is self-contained and verified locally with a signature check, fast, no network call needed, but hard to revoke early. An opaque token is a meaningless random string that must be looked up (introspected) against the issuing server on every use, slower and an added dependency, but instantly revocable by deleting it server-side.
Local verification versus introspection: a JWT's verifier checks its cryptographic signature against the issuer's public key locally, no network call per request. An opaque token carries no embedded meaning, so the verifier must call the issuing server's introspection endpoint to ask whether it is still valid, on every use.
Revocation: a JWT is hard to revoke before its natural expiry, since any verifier holding the issuer's public key validates it independently, options are keeping expiry very short, or maintaining a deny-list every verifier must also check, which erodes the "no network call" advantage. An opaque token is trivially revocable, delete or invalidate the record in the issuer's store and every future introspection call immediately reports it invalid.
Size and transport: a JWT is bigger, a header, a payload of claims, and a signature, all base64-encoded, and this adds up if it rides along on many internal calls. An opaque token is small, just an identifier, cheap to pass around.
When to prefer each: prefer JWTs when many verifiers need to check tokens fast and independently, without a hot-path dependency on a central auth server, and when the token's lifetime can be kept short enough that limited revocability is not a real risk. Prefer opaque tokens with introspection when you need to instantly kill a specific session or credential, a compromised account, a departed employee's access, or when you do not want token contents exposed to anyone who can read the token in transit or in logs.
Common JWT validation mistakes a reviewer should look for:
- Algorithm confusion: trusting the
algfield the token itself claims instead of pinning the expected algorithm, this is how a server ends up accepting a token signed withnone, or gets tricked into verifying a forged token using its OWN public key as if it were a shared secret. - Missing audience (
aud) check: without verifying the token was issued FOR this specific service, a valid token minted for one internal service can be replayed against a different service that trusts the same issuer. - Missing or ignored expiry (
exp) check: a technically expired token is still accepted because the timestamp claim is never checked. - Missing or mismatched issuer (
iss) check when multiple issuers are trusted: a token from a lower-trust issuer gets accepted where only a specific higher-trust issuer should be.
Worked example: a reviewer should specifically confirm the code hardcodes an expected algorithm, for example explicitly requiring RS256 rather than reading alg from the token, explicitly compares the aud claim against an expected value, and explicitly checks exp, rather than trusting a library's defaults without confirming what those defaults enforce. Library defaults on this have changed across versions and ecosystems, so "we used a JWT library" is not itself something you can verify without reading the actual validation call.
Trade-offs & pitfalls: teams sometimes treat "JWT" and "secure" as synonyms, an improperly validated JWT is not safer than an opaque token, it is often worse, because it looks cryptographically strong while actually accepting forged tokens. Long-lived JWTs adopted because "introspection was too slow" reintroduce the exact revocation problem opaque tokens exist to solve.
Design a fine-grained authorization model for a large number of microservices: evaluate RBAC, ABAC, and a hybrid approach, and decide where policy evaluation should happen (central decision point, sidecar, or in-process library). Give a concrete example of one authorization rule your design would enforce.
Sample Answer
Direct answer: use role-based access control (RBAC) for the coarse, stable question of whether one service may talk to another at all, layer attribute-based access control (ABAC) on top for the finer, dynamic question of under exactly what conditions, and put the actual decision-making close to the data path, typically at the sidecar or a shared policy library, reserving a fully centralized decision point for cases where the extra network hop is an acceptable cost for stronger central auditability.
Role-based access control (RBAC): permissions attached to roles, roles assigned to services or identities. Simple to reason about and audit, "the billing-service role can call the accounts-service role," but does not scale well to fine-grained, resource-level rules, expressing "billing-service can read balances but only for accounts in its own tenant" starts requiring an explosion of narrow roles if forced entirely into RBAC.
Attribute-based access control (ABAC): the decision is evaluated against attributes of the caller, the resource, and the environment at request time, for example "the caller's tenant ID must match the requested resource's tenant ID," which handles that fine-grained, dynamic case naturally, but policies become harder to read, test, and audit as they accumulate, and evaluating richer attribute-based rules costs more per request than a role lookup.
A practical hybrid: use RBAC as the coarse first-pass gate, does this service identity have any business calling that service at all, then apply ABAC-style conditions on top for the fine-grained per-resource decision. This keeps the coarse, auditable structure RBAC is good at while still expressing the dynamic per-resource rules pure RBAC cannot.
Where policy evaluation should happen:
- A central policy decision point (a shared decision service, for example built on Open Policy Agent, OPA, an open source policy engine): one source of truth and one place to audit every decision, at the cost of a network round trip added to every call, and that shared service becoming a scaling bottleneck and a single point of failure as call volume grows.
- A sidecar (the same OPA engine, or a mesh's built-in authorization enforcement, running alongside each service): keeps the round trip local rather than crossing the network, and scales horizontally with the number of service instances, at the cost of needing a reliable mechanism to distribute policy updates to every sidecar, with inevitable lag before an update reaches everywhere.
- An in-process library embedded directly in application code: the lowest possible latency, no hop at all, but every team's language and runtime needs its own working integration, and a policy change requires each service to reload or redeploy, reintroducing exactly the "every team has to get this right independently" risk a shared enforcement layer exists to avoid.
Worked example, a concrete rule this design would enforce: billing-service may call GET /accounts/{accountId}/balance on accounts-service only when the caller's cryptographic service identity matches billing-service's expected identity, the coarse RBAC-style gate, AND the accountId in the request path belongs to a tenant billing-service's own request context is currently scoped to, the finer ABAC-style condition layered on top. The first half is a stable, auditable role-to-role permission; the second half depends on request-time data and would need constant new roles to express in pure RBAC, but is a single readable condition in the hybrid model.
Trade-offs & pitfalls: a common failure is letting ABAC-style conditions accumulate ad hoc, with no shared vocabulary for what attributes mean across services, until nobody can confidently describe the FULL authorization policy from reading it. Choosing a fully centralized decision point purely because it is easiest to audit, without accounting for the added latency and new single point of failure, is also a common early mistake that becomes expensive to unwind once call volume grows past what that central service was sized for.
Write a policy-as-code snippet (Open Policy Agent / Rego, or an equivalent policy language of your choice) that authorizes a service-to-service request only when: the caller's JWT audience claim matches the target service, the caller's role is on that service's access list, and the caller's device posture score meets a minimum bar. Explain what each clause is protecting against.
Sample Answer
Use Open Policy Agent (OPA), a general-purpose policy engine, with its policy language Rego to compute a single allow decision as the AND of three independent checks against the incoming request: the caller's JSON Web Token (JWT, a compact, signed way of encoding claims about who someone is), an access-control list, and a device posture score. Each clause enforces a distinct part of the zero trust story: who is calling, what they're allowed to do, and whether the thing presenting that identity is healthy enough to be trusted with it.
Approach
Model the request as input (the caller's JWT claims, the target service name, and the device posture) and keep the access lists and posture thresholds in a separate data document, so a security team can update who's allowed without touching the policy logic itself.
Code (Rego v1)
package service.authz
import rego.v1
default allow := false
allow if {
input.jwt.aud == input.target_service
input.jwt.role in data.access_lists[input.target_service]
input.device.posture_score >= data.posture_thresholds[input.target_service]
}
Control data (data.json):
{
"access_lists": {
"billing-service": ["payments-writer", "payments-admin"]
},
"posture_thresholds": {
"billing-service": 70
}
}
A request that should be allowed (input.json):
{
"jwt": {"aud": "billing-service", "role": "payments-writer"},
"target_service": "billing-service",
"device": {"posture_score": 82}
}
Run it:
opa eval -d policy.rego -d data.json -i input.json "data.service.authz.allow" --format pretty
Output: true
Change only posture_score to 55 in the input (below the 70 threshold) and re-run the same command: the output flips to false, and the same happens if jwt.aud is set to a different service than target_service.
Key points: what each clause protects against
input.jwt.aud == input.target_service: protects against a token issued for one service being replayed against a different one. Without an audience check, a token that's valid but meant for the inventory service could be presented to the billing service and pass identity verification even though it was never intended for that call.input.jwt.role in data.access_lists[...]: protects against a caller who is correctly authenticated but not authorized for this specific service. Identity (who you are) is kept separate from entitlement (what you're allowed to do), which is the least-privilege half of the model.- the posture check: protects against a compromised or non-compliant device or workload using otherwise-valid credentials. Even a correctly-scoped, correctly-authenticated caller shouldn't get access if the thing presenting that identity fails a health check. This is the "continuous verification" piece of zero trust: trust is re-evaluated against the current state of the caller, not granted once and assumed to hold.
Complexity and edge cases
Evaluation is a constant number of map lookups per request against the loaded data document, so the cost that actually scales is the size of data and how often it's refreshed, not the policy logic itself. Edge cases worth testing: a missing key anywhere in the lookup chain should resolve to undefined and therefore deny, matching the default allow := false; a request with posture_score entirely absent should fail closed rather than being treated as passing; and a role that exists on some OTHER service's access list but not this one's should still deny, since the lookup is keyed strictly by target_service.
Trade-offs and pitfalls
Hardcoded, hand-maintained access lists don't scale as the number of services grows; production setups usually generate data from a service catalog or a CI pipeline rather than editing it by hand. If this policy runs behind a centralized Policy Decision Point (PDP, the component that computes authorization decisions) rather than as a local OPA sidecar, add a timeout and an explicit fail-closed or short-TTL-cached behavior for when the PDP is unreachable, since otherwise an identity or posture system outage silently becomes an outage of all service traffic.
Compare the main approaches for authenticating one service to another: mutual TLS via a service mesh with workload identities (e.g. SPIFFE/SPIRE), signed JWTs issued by a central authority, and cloud-native IAM roles or client certificates. What's the trust model, key-rotation story, and operational overhead for each, and how would you decide between them for a multi-team, multi-namespace environment?
Sample Answer
Direct answer: the three approaches sit on a spectrum from infrastructure-managed identity (mutual TLS via a mesh) to application-issued identity (signed tokens) to cloud-managed identity (native identity and access management, IAM, roles), and for a multi-team, multi-namespace environment the right choice usually is not picking one, it is using mesh-based mutual TLS for baseline service-to-service trust, then layering the other two in for specific needs like cross-cloud calls or user-context propagation.
Mutual TLS via service mesh with workload identities (for example SPIFFE/SPIRE, a standard and implementation for issuing cryptographic workload identities): trust is rooted in a shared certificate authority (CA) every workload and the mesh agree on, and each workload gets a short-lived certificate tied to its identity. Rotation is automatic and frequent, minutes to hours, handled transparently by sidecar proxies. Operational overhead is highest to stand up (control plane, CA infrastructure, sidecar injection) but lowest ongoing per call, since verification happens at the network layer with no app code changes.
Signed JSON Web Tokens (JWTs) from a central authority: a central issuer signs tokens, and any verifier trusting the issuer's public key validates a token locally, no callback needed. The issuer rotates its own signing key on a schedule, published via a key-set endpoint verifiers refresh; the tokens themselves are simply re-issued rather than "rotated." Operational overhead is low to start, but every service must correctly validate audience, expiry, and issuer itself, a per-service correctness burden instead of a centralized one.
Cloud-native IAM roles or client certificates: the cloud platform attests that a given piece of compute is who it claims to be, based on the platform's own control plane, and hands out short-lived credentials on request, fully managed rotation, no certificate files for engineers to touch. Overhead is lowest for workloads staying inside one cloud provider, but this does not natively extend across clouds or to on-premises workloads, and ties your identity model to that provider's semantics.
Deciding for a multi-team, multi-namespace environment: if most traffic is service-to-service inside one platform, mesh-based mutual TLS gives every team the same baseline identity and encryption without each team writing its own auth code, which matters more as team count grows. Add signed JWTs where a human user's identity or fine-grained claims need to travel through a call chain. Reach for native cloud IAM roles specifically at the boundary where a workload calls a CLOUD service, a storage bucket, a managed queue, not another one of your own services, since that is the boundary the cloud provider's identity model is actually built for.
Worked example (a cloud-IAM branch in practice): a nightly batch job running as a Kubernetes pod needs to read from a cloud storage bucket. Instead of embedding a static access key, the pod's service account is federated to a cloud IAM role, for example AWS's IAM Roles for Service Accounts (IRSA), and at startup the job exchanges a short-lived, automatically-rotated token for temporary credentials from the cloud's Security Token Service (STS), typically valid for under an hour. If the pod is compromised, the exposed credential expires within that window, rather than being a static key that works indefinitely until someone manually revokes it.
Trade-offs & pitfalls: mixing all three without a clear rule for which is used when is itself a common failure, teams pick whichever they found a tutorial for, and you end up auditing three inconsistent identity systems instead of one. Cloud-native IAM roles are the easiest of the three to reach for and the easiest to over-scope, since there is no mesh-level review step forcing a second set of eyes on the grant.
How does continuous authentication and authorization differ from a one-time login? What signals (behavioral, location, device posture) should trigger re-authentication or an adaptive change in access, and how do you avoid re-prompting the user so often that they get fatigued?
Sample Answer
Direct answer
A one-time login checks identity once, at the start of a session, then trusts that session until it expires or is explicitly logged out. Continuous authentication and authorization keep re-evaluating trust throughout the session as new signals arrive, so access can be tightened, challenged, or revoked mid-session if something changes, not only at the door.
Structured elaboration
Signals that should trigger re-authentication or an adaptive change in access:
- Behavioral: a user performing actions well outside their normal pattern, such as bulk-querying data they never normally touch, or an unusually rapid sequence of requests that looks automated.
- Location: a login or request originating from a network location or geography inconsistent with the user's recent activity, sometimes described as impossible travel, an active session in one place and a new request appearing to originate from somewhere else shortly after.
- Device posture: the device's security state changing mid-session, disk encryption disabled, an outdated patch level, malware detection triggering, or the device no longer matching the enrolled, managed device that started the session.
Not every signal should trigger the same response. A well-designed system grades severity: a low-confidence anomaly might just be logged; a moderate anomaly might trigger a step-up challenge, such as an additional multi-factor authentication (MFA) prompt, scoped to the specific sensitive action being attempted rather than the whole session; a high-confidence anomaly, a clearly failed device-posture check or a hijacking signal, should terminate the session and force full re-authentication.
Avoiding re-prompt fatigue:
- Scope step-up challenges to the specific risky action, not the whole session, so a user browsing normal, low-sensitivity resources is never interrupted.
- Use risk-based thresholds rather than fixed intervals. Prompting every user every fifteen minutes regardless of behavior trains people to reflexively approve prompts, a habit attackers exploit through prompt bombing; prompting only when a meaningful signal changes keeps prompts rare enough to be taken seriously.
- Prefer passive signals, device posture, network reputation, behavioral baselining, over active prompts wherever possible, since passive checks add friction to the system rather than to the user, escalating to an active prompt only when those passive signals actually indicate elevated risk.
Worked example
A user logs in from a managed laptop on the corporate network at 9am, low risk, no prompts needed for normal work. At 2pm, the same session starts downloading a far larger volume of customer records than that user has ever accessed in one sitting, a behavioral anomaly, while the device's posture check reports its disk encryption was disabled ten minutes earlier, a device-posture anomaly. Both signals firing together push the computed risk well past a step-up threshold into a high-risk band, so the system does not just ask for MFA, it suspends the download and forces full re-authentication plus a device compliance check before any further access to customer data, while the user's earlier, low-risk browsing that morning was never interrupted at all.
Trade-offs and pitfalls
The common wrong turn is treating every anomaly as equally important and re-prompting for anything unusual, which produces fatigue and trains users to click through prompts without reading them. The fix is grading signals by confidence and severity and scoping the response to the specific action at risk, not the entire session.
Unlock Full Question Bank
Get access to all 15 Zero Trust, Segmentation, and Service-to-Service Security interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.