Microsoft Azure Services and Architecture Questions
Microsoft Azure's core service catalog and architectural patterns: Virtual Machines, managed Kubernetes (AKS), App Service and Azure Functions, Storage accounts and managed disks, Azure SQL and Cosmos DB, VNets with hybrid connectivity and global load balancing, Microsoft Entra ID and RBAC, and Key Vault secrets and encryption. Covers Azure service selection, infrastructure as code (ARM, Bicep, Terraform), observability with Azure Monitor and Kusto queries, cost governance and Azure Policy, the Azure Well-Architected design principles, and hybrid management via Azure Arc, common in enterprise Azure estates. For provider-agnostic trade-offs, see the cross-cloud entries.
Intermittent 502 Bad Gateway responses are observed from an Application Gateway fronting a backend web API. Create a troubleshooting runbook: which Application Gateway metrics and logs to check, how to validate backend health probes, probe path and host header issues, timeout and keep-alive settings, and SSL termination mismatches that can cause 502s.
Sample Answer
Direct answer
Treat intermittent 502 Bad Gateway responses as a symptom with at least four distinct root causes behind Application Gateway: backend health probe failures, request timeout or keep-alive mismatches, Transport Layer Security (TLS) or Server Name Indication (SNI) mismatches, and backend connection exhaustion. Work through them in that order, correlating each candidate cause against the exact timestamps of the 502s, rather than guessing at a single fix.
Structured elaboration
Step 1: scope and reproduce. Record the exact times, client IPs, and request paths for the failures, then reproduce from a virtual machine inside the same virtual network (VNet) as the Application Gateway, using the same host header and path a real client would send, so you are testing the same path the gateway tests.
Step 2: check metrics and logs. Pull Application Gateway's FailedRequests and BackendHealth (healthy versus unhealthy host count) metrics, and overlay them against the 502 timestamps; a spike in unhealthy host count that lines up with the 502s points at probe or backend-health issues, while a flat healthy count during the 502 spike points elsewhere (timeouts or TLS). Pull access logs for the affected requests to see the returned status code and which backend pool member served (or failed to serve) each one, and pull backend-side application logs for the same timestamps.
Step 3: validate health probes. Confirm the probe's configured protocol, path, host header, and unhealthy threshold, then manually send the identical request:
curl -v -H "Host: api.example.com" http://10.0.1.5/healthz
If the probe path redirects, requires authentication, or returns anything outside the expected success range, the gateway will intermittently mark the backend unhealthy exactly when that condition is hit, which produces 502s that look random but are not.
Step 4: host header and path issues. If the backend only serves virtual-hosted responses for a specific host header, confirm the probe's host-header override matches; a probe hitting the backend's IP directly without the right host header can get a default, non-2xx response even though the same request through the real listener would succeed.
Step 5: timeouts and keep-alive. Application Gateway's backend HTTP settings have a request timeout, defaulting to 20 seconds (the configurable range is 1-86,400 seconds for a private backend and 1-240 seconds for an external backend); if the backend occasionally takes longer than the configured value, the gateway returns a 502 or 504 rather than waiting indefinitely. Separately, if the backend does not support HTTP keep-alive or drops idle connections faster than the gateway expects, a request can be sent on a connection the backend already closed, producing an intermittent 502 that correlates with connection reuse rather than backend load:
curl --max-time 35 -v https://api.example.com/slow-endpoint
Step 6: TLS and SNI mismatches. If Application Gateway terminates TLS and re-encrypts to the backend (end-to-end TLS), confirm the backend's certificate is valid for the SNI the gateway sends, and that the backend is not trying to redirect HTTP to HTTPS on a connection the gateway already re-encrypted:
openssl s_client -connect 10.0.1.5:443 -servername api.example.com
A handshake failure or certificate mismatch here explains 502s that cluster around certificate rotation events or backend deployments that changed the certificate.
Step 7: correlate and remediate. Line up the 502 timestamps against probe transitions, backend deploys or restarts, and any autoscale events; fix the specific mismatch found (probe path/host, timeout value, keep-alive configuration, or certificate/SNI), not all of them speculatively.
Worked example
Suppose the 502s cluster in short bursts every few minutes, each burst lasting under 30 seconds. The BackendHealth metric shows brief dips in healthy host count at the exact same timestamps, ruling out a pure timeout or TLS issue (those would not typically flip the health status) and pointing at the probe. Manually running the curl command above against the probe's exact path and host header during a normal period returns 200 quickly, but retrying it in a loop reveals it occasionally takes 4 to 5 seconds, longer than the probe's configured 3-second timeout. The fix is either to raise the probe timeout to comfortably exceed the backend's occasional slow response, or to investigate why a supposedly lightweight health endpoint is intermittently slow, treating the slow probe response as the real signal rather than only patching the timeout.
Trade-offs & pitfalls
The most common mistake is fixing the first plausible cause found (usually raising the timeout) without confirming it actually correlates with the failure timestamps, which can mask a real backend performance problem behind a longer timeout instead of fixing it. The second is testing from outside the VNet, which introduces its own network variables and does not reproduce what the gateway itself experiences on its private path to the backend; always reproduce from inside the same network boundary the gateway operates in.
Compare TCP, HTTP, and HTTPS health probe behaviors on Azure Load Balancer and Application Gateway. Explain how misconfigurations (wrong probe path, host header, response codes) can cause healthy backend instances to be marked unhealthy and cause scale set instance replacement or traffic blackholing.
Sample Answer
Direct answer
A TCP probe only confirms the port accepts a connection; it says nothing about the application behind it. An HTTP or HTTPS probe actually exercises the application by requesting a specific path, but what counts as a "healthy" response is not identical across the two services this question compares: Azure Load Balancer's HTTP/HTTPS probe accepts only an exact HTTP 200 as healthy (a 403, 404, or 500 all count as failure), while Application Gateway's default probe is more permissive and treats any response in the 200 to 399 range as healthy (a custom probe on Application Gateway can further widen or narrow that with an explicit status-code match). Either way, a probe can be misconfigured to fail against a perfectly healthy backend, which then gets pulled from rotation, or in a virtual machine scale set (VMSS), gets replaced outright, even though nothing is actually wrong with it.
Structured elaboration
How the probe types differ. A TCP probe on Azure Load Balancer performs only a three-way handshake against an IP and port; it cannot see an HTTP error, a redirect, or a slow application response, so it will report "healthy" even when the application itself is failing every real request. An HTTP or HTTPS probe (available on both Load Balancer and Application Gateway) sends an actual request to a configured path and host header and evaluates the response code, but the two services do not agree on what response code counts as success: Load Balancer requires exactly HTTP 200, while Application Gateway accepts the whole 200-399 range by default. Application Gateway, being a Layer 7 (application-layer) proxy with multiple listeners and backend pools, additionally has to match the probe's host header and path to what each specific backend expects, since one gateway can front several different virtual hosts.
How each misconfiguration causes a false failure.
| Misconfiguration | What happens | Result |
|---|---|---|
| Wrong probe path (probes /health, app only serves /ready) | Application returns 404 to a path it never registered | Instance marked unhealthy despite being fully functional on its real endpoint |
| Missing or wrong host header / Server Name Indication (SNI) | Virtual-hosted or certificate-bound backend returns 403, 404, or aborts the TLS handshake entirely | Probe fails even though the same request with the correct host header would succeed |
| Unexpected response code (redirect or authentication challenge) | App returns a 302 redirect to a login page because the probe path requires authentication | Non-2xx/3xx-boundary response counts as a failure even though the backend is up |
| TCP probe hiding an HTTP-layer problem | Port accepts the handshake, but the application layer is broken | Backend stays "healthy" while every real request fails, the opposite failure mode |
Operational impact. When enough instances are falsely marked unhealthy, the load balancer or gateway simply stops sending them traffic, which looks like the service going dark even though the instances are fine; this is traffic blackholing. On a VMSS, sustained probe failures can trigger the platform's own health-based instance repair, replacing instances that were never actually broken, which burns through deployment time and obscures the real, configuration-level root cause.
Worked example
Suppose an Application Gateway probe is configured with path /health and no host header override, while the backend application only responds to requests carrying the host header api.example.com and returns a 404 for any other host. Trace it step by step: the probe connects to the backend's private IP, sends a GET for /health with either no host header or the gateway's own default, the backend's routing layer does not recognize that host, returns 404, the gateway logs that as a probe failure, and after the configured unhealthy threshold (say 3 consecutive failures), removes the backend from rotation. Meanwhile a real user request through the gateway's listener carries the correct host header and would have succeeded. The fix is not to guess: reproduce the exact probe request with curl using the same path and host header the gateway is configured to send, confirm it also returns 404, and only then you have proven the probe configuration, not the application, is the defect.
Trade-offs & pitfalls
The recurring pitfall is trusting the load balancer's health dashboard as ground truth about the application; it is ground truth about the probe configuration matching the application, which is a different claim. A second, subtler pitfall is setting unhealthy thresholds too aggressively low to "detect failures fast," which then trips on ordinary transient load spikes and starts a churn cycle of instance replacement that never lets the application stabilize. Expose a dedicated, unauthenticated, lightweight health endpoint that matches exactly what the probe is configured to call, and treat any probe failure as a signal to reproduce the exact probe request before touching the fleet.
Design a secure Azure network architecture to expose internal microservices only to authenticated internal clients, using Private Link, Azure Firewall with WAF, API Management, and service endpoints. Describe the ingress/egress flows, DNS and private-zone setup, policy enforcement, certificate handling, and the residual risks of the approach.
Sample Answer
Direct answer
Put API Management (APIM) in internal (VNet [Virtual Network, Azure's private network boundary]-injected) mode as the only entry point authenticated internal clients ever talk to, expose each microservice to APIM through a Private Link connection (not a service endpoint, which restricts by network identity but still uses a public IP path; Private Link gives the backend a private IP and removes it from the public internet's routable surface entirely), and route anything that needs internet egress (image pulls, third-party API calls) through Azure Firewall so it's inspected and logged centrally instead of each service managing its own outbound rules.
Design
- Ingress: internal clients reach
APIMover a private VNet address (APIM deployed in internal mode has no public data-plane endpoint at all). APIM enforces authentication (validate a Microsoft Entra ID token viavalidate-jwtpolicy) and authorization (scopes/roles in the token) before a request is allowed anywhere near a backend. - APIM to microservice: each backend (App Service, AKS ingress, or a container app) sits behind its own Private Endpoint, so from APIM's perspective the backend resolves to a private IP inside the VNet, never crossing the public internet even encrypted. This is the actual "only authenticated internal clients" boundary: a microservice with a Private Endpoint and no public network access enabled simply has no public attack surface to find.
- Egress and inspection: any outbound call a microservice needs to make (a third-party API, a public package registry) is routed through Azure Firewall via user-defined routes, which gives you one place to apply FQDN-based egress allow-lists and one place logs land, instead of each service's NSG doing it inconsistently.
- DNS: Private Link requires a private DNS zone (e.g.,
privatelink.azurewebsites.netfor App Service) linked to the VNet, so internal name resolution for the backend's hostname returns the private IP instead of the public one. This is the step most commonly missed: without the private DNS zone linked, clients still resolve the public FQDN and either fail or, worse, quietly succeed over the public path if public access wasn't actually disabled on the backend. - Certificate handling: APIM terminates client TLS with a certificate from Key Vault (via managed identity, same pattern as Application Gateway); the APIM-to-backend hop can use the backend's own certificate or an internal CA-issued one, since that hop never leaves the VNet.
- Policy enforcement: rate limiting, IP filtering (if a subset of internal ranges should be further restricted), and request/response transformation all live as APIM policies, giving one consistent enforcement point rather than duplicating auth logic in every microservice.
graph LR
Client[Authenticated internal client] -->|private VNet address, JWT| APIM[API Management, internal mode]
APIM -->|Private Endpoint, private DNS| SvcA[Microservice A]
APIM -->|Private Endpoint, private DNS| SvcB[Microservice B]
SvcA -->|UDR forces egress| FW[Azure Firewall]
SvcB -->|UDR forces egress| FW
FW -->|allow-listed FQDNs only| Internet((Internet))
Worked example: what breaks if a step is skipped
A concrete failure mode this design has to guard against: a team enables a Private Endpoint on a microservice's App Service, but forgets to link the corresponding private DNS zone to APIM's VNet. Clients (and APIM) resolving the App Service's hostname still get the public IP, because DNS resolution never changed, only the network path did. If the App Service's public network access wasn't also explicitly disabled, traffic silently continues to flow over the public endpoint, defeating the entire Private Link setup while every dashboard shows the private endpoint as "provisioned" and healthy. The concrete fix and verification step: after wiring Private Link, explicitly set the backend's public network access to disabled (not just "add a private endpoint," which by itself doesn't remove the public one), and confirm from a VM inside the VNet that nslookup on the backend's FQDN returns the private IP, not the public one.
Residual risks
- A compromised identity inside the VNet still has the same network path as a legitimate client. Private Link and APIM authentication stop external attackers and stop anything without a valid VNet presence, but they don't substitute for authorization logic; a compromised internal service with an over-scoped token can still call anything APIM allows that scope to reach. Least-privilege token scopes per client remain necessary on top of the network boundary.
- DNS misconfiguration is a silent failure mode, not a loud one, as shown above; it needs an explicit verification step in the deployment pipeline, not just a "the resource exists" check.
- Azure Firewall is a single logical egress point, which is good for control but means its availability and throughput now gate every microservice's outbound calls; size and monitor it accordingly rather than treating it as a fire-and-forget appliance.
- APIM's internal mode requires the client to already be on the VNet (or connected via VPN/ExpressRoute/peering), so this design does not, by itself, solve "authenticated users on the public internet need access"; that's a different, additional problem (typically APIM's external mode plus a separate auth layer) layered on top if it's ever needed.
Trade-offs and pitfalls
The main trade-off against a simpler "just use NSGs and service endpoints" design is operational: Private Link plus private DNS zones plus internal-mode APIM is more moving parts to deploy and verify correctly, and the DNS failure mode above is exactly the kind of thing that passes a cursory review. The payoff is that the microservices genuinely have no public IP-based attack surface at all, versus service endpoints which restrict by source but still route over Azure's public backbone and still require the backend to have a public IP to protect in the first place.
For a three-tier application (public web, private app, private DB) in Azure, design NSG rules and Azure Firewall placement. Include how to restrict management access using Azure Bastion, capture and analyze flow logs, and control egress to external services while enabling required PaaS access with minimal blast radius.
Sample Answer
Direct answer
Isolate the three tiers with per-subnet Network Security Groups (NSGs) that only allow the next tier's specific traffic, route egress (outbound traffic leaving the virtual network) through a centrally placed Azure Firewall for auditable, allow-listed outbound access, and remove every management port from the public internet entirely by using Azure Bastion for administrative access instead.
NSG rules per tier
Building on a web / app / db subnet layout:
| Subnet | Inbound allow | Outbound allow | Everything else |
|---|---|---|---|
| Web | Port 443 from the internet (or narrower, from a gateway's published range) | To App subnet, app port only | Denied |
| App | App port, from Web subnet only | To DB subnet, database port; to firewall for any required outbound | Denied |
| DB | Database port, from App subnet only | Explicitly denied to the internet | Denied |
Azure Firewall placement for egress control
Place Azure Firewall in its own dedicated subnet (AzureFirewallSubnet, sized at least /26, a block of 64 IP addresses, Azure's documented minimum for this subnet) and route each subnet's default outbound traffic through it via a User-Defined Route (UDR) with the firewall's private IP as next hop, rather than letting subnets reach the internet directly. This lets Domain Name System (DNS)-name-based (fully qualified domain name, FQDN) filtering restrict exactly which external endpoints each tier is allowed to call, for example, letting the app tier reach only a named payment gateway and a specific Azure SQL server's FQDN, and denying everything else outbound by default. Relying on Azure's older implicit default outbound access instead of an explicit egress path (a firewall or NAT Gateway, which translates private IP addresses to a shared public one for outbound calls) is no longer the recommended pattern for new deployments and should not be the design's actual egress mechanism.
Restricting management access with Azure Bastion
Deploy Azure Bastion into its own dedicated subnet (AzureBastionSubnet, sized at least /26), and remove any public IP address from every virtual machine (VM) across all three tiers. Administrators connect to Bastion over Transport Layer Security (TLS) on port 443 through the Azure portal or Bastion's native client, and Bastion makes the private hop to the target VM's private IP from there. No VM in any tier ever needs an open Remote Desktop Protocol (RDP) or Secure Shell (SSH) port reachable from the internet, closing the single most common way a three-tier deployment gets compromised in the first place: an exposed management port.
Flow logs and monitoring
Enable NSG flow logs (via Network Watcher) on all three subnets' NSGs, sent to a Log Analytics workspace, and enable Azure Firewall's own diagnostic logs (network rule log, application rule log, and threat intelligence log) if the firewall is in place. Layer Traffic Analytics on top of the raw flow logs to visualize east-west and north-south flows at a glance and catch a tier doing something it shouldn't, for example, the database tier suddenly attempting an outbound connection, which the NSG should already block, but which the logs should also surface as an anomaly worth investigating.
Minimal blast radius while still enabling required PaaS access
flowchart LR
Internet((Internet)) -->|443 only| Web[Web subnet]
Web -->|app port| App[App subnet]
App -->|db port| DB[DB subnet]
App -->|Private Link| PaaS[(Storage / Key Vault / SQL)]
Web -. egress via UDR .-> FW[Azure Firewall: FQDN allow-list]
App -. egress via UDR .-> FW
Admin[Administrator] --> Bastion[Azure Bastion] --> Web
Bastion --> App
Bastion --> DB
Rather than opening broad outbound internet access "just in case" a tier needs a platform-as-a-service (PaaS) dependency (Storage, Key Vault, Azure SQL), use Private Link (Private Endpoints: a private IP address inside your VNet that fronts the PaaS service) for those specific dependencies so a tier never needs general internet egress for its normal operation at all. That reserves the firewall's FQDN-filtered egress path for genuinely external, non-Azure calls the application needs (a third-party Application Programming Interface, API, for example), which both shrinks the blast radius of a compromised instance and keeps the firewall's allow-list small enough to actually audit, instead of a wildcard rule nobody can meaningfully review.
A multi-tenant extension
If this three-tier application is itself multi-tenant, the network defaults above don't automatically get stronger for that reason: if tenants must not reach each other's data even from a compromised app instance, either isolate at the database level (separate databases or schemas per tenant) or, for the highest-sensitivity tenants, dedicate a separate app-subnet-and-NSG pair per tenant rather than relying on rules that only see "the app subnet" as one undifferentiated source.
Trade-offs and pitfalls
- Leaving Azure's default outbound internet access as the actual egress mechanism, rather than an explicit firewall or NAT Gateway path, undermines the entire point of centralizing egress control; verify the UDR is actually forcing traffic through the firewall.
- Skipping Bastion in favor of a jump-box VM with its own public IP just reintroduces the exposed-management-port problem one hop later.
- An allow-list that grows to include broad service tags "to be safe" instead of specific FQDNs defeats the minimal-blast-radius goal; keep it as narrow as the application's actual dependencies.
Design hybrid connectivity between an on-prem datacenter and Azure across two regions with stringent latency (<50ms) and high availability requirements. Compare ExpressRoute (private circuits) vs VPN Gateway (IPsec) for this scenario, explain BGP peering and route advertisement, failover strategies, and how to avoid asymmetric routing or single points of failure.
Sample Answer
Direct answer
Make ExpressRoute (Microsoft's private, non-internet circuit into Azure) the primary path for its predictable sub-50ms latency and service-level agreement (SLA), with a Site-to-Site VPN Gateway (an IPsec-encrypted tunnel over the public internet) as an active backup, both duplicated across two regions and two physically diverse circuits so no single carrier, router, or Azure gateway is a single point of failure. Use Border Gateway Protocol (BGP, the routing protocol that lets networks exchange reachability information) on every link, with route preference tuned so ExpressRoute always wins when it is healthy and the VPN takes over automatically, symmetrically, when it is not.
Structured elaboration
Why ExpressRoute over VPN Gateway here. ExpressRoute is a dedicated private circuit with predictable latency and a formal SLA, which is what a sub-50ms, high-availability requirement actually needs. A Site-to-Site VPN runs over the public internet (even if encrypted end to end), so its latency and jitter vary with internet conditions and it carries no comparable SLA. VPN's advantages are that it provisions in hours, not weeks, and costs far less, which is exactly why it is the right role for a backup path, not the primary.
Topology. Provision two ExpressRoute circuits terminating at physically diverse peering locations (different buildings, ideally different carriers), each connected into a regional Azure Virtual WAN hub or ExpressRoute gateway. Deploy active-active VPN gateways in both regions as the failover path. On-premises, terminate on at least two edge routers, each with its own physical uplink, both speaking BGP.
BGP peering and route advertisement. Every peering (ExpressRoute private peering and the VPN's BGP session) advertises prefixes in both directions: on-premises advertises its internal subnets to Azure, and Azure advertises the connected VNet prefixes back. Control which path wins with standard BGP attributes: a higher local preference or a shorter advertised AS-path on the ExpressRoute session makes routers prefer it over the VPN under normal conditions; when ExpressRoute fails, BGP simply stops hearing those routes and the VPN's routes become the best path automatically, with no manual re-routing.
Avoiding asymmetric routing and single points of failure. Asymmetric routing (traffic leaving over one path and returning over a different one) causes stateful devices like firewalls to drop return traffic they never saw the initial packet for. Prevent it by keeping route preference consistent on both ends: if on-premises prefers ExpressRoute outbound, Azure must also prefer the matching ExpressRoute path inbound, using the same BGP attributes on both sides rather than one side using static routes and the other using BGP. Redundant circuits on physically diverse paths, redundant gateways in active-active mode, and Bidirectional Forwarding Detection (BFD, a fast link-failure detection protocol that reacts in milliseconds instead of BGP's default tens of seconds) close the remaining single-point-of-failure gaps.
flowchart TD
OP1[On-prem router A] -- ExpressRoute circuit 1 --> AZ1[Azure region 1 hub]
OP2[On-prem router B] -- ExpressRoute circuit 2 --> AZ2[Azure region 2 hub]
OP1 -. VPN backup .-> AZ2
OP2 -. VPN backup .-> AZ1
AZ1 <--> AZ2
Validation before go-live. Integrate this into the hub-and-spoke design already in place for the workload's VNets: the ExpressRoute and VPN gateways terminate in the hub, and spoke VNets reach on-premises only through hub peering, so a single set of gateways serves every spoke. Before declaring the link production-ready, run a controlled failover test: withdraw the ExpressRoute route deliberately (not by physically cutting a cable) and confirm the VPN path takes over within the BFD detection window, traffic resumes without a duplicate or dropped-session storm, and route tables on both ends converge to the expected state.
Worked example
A workload like an SAP application landscape (used here only as an illustrative example of a genuinely latency-sensitive enterprise workload, not a prescription) typically needs application-tier-to-database-tier round trips well under the latency budget of a single user transaction. If the physical distance between the on-premises datacenter and the nearer Azure region is roughly 500 km, the one-way propagation delay in fiber (light travels at roughly two-thirds the speed of light in glass, so about 200,000 km per second) is 500 divided by 200,000 seconds, or 2.5 milliseconds, and a round trip is 5 milliseconds. That leaves roughly 45 of the 50ms budget for switching, queuing, and application-level processing at each hop, which is exactly why the design targets a private, low-jitter circuit like ExpressRoute instead of the public internet's much less predictable delay: the physics is not the constraint, the variance is.
Trade-offs & pitfalls
The pitfall this scenario is built to catch is treating BGP local preference as something you tune once and forget: if the two ends of the connection are configured by different teams (network engineering on-premises, cloud engineering in Azure) and they drift out of sync, asymmetric routing reappears silently and only shows up as intermittent, hard-to-reproduce connection resets. The second pitfall is skipping the failover rehearsal because the design "should" work; BGP convergence behavior under a real link failure is exactly the kind of thing that differs from the whiteboard, so test it before it is load-bearing.
Unlock Full Question Bank
Get access to all 17 Microsoft Azure Services and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.