Microsoft Azure Services and Architecture Questions
Microsoft Azure's core service catalog and architectural patterns: Virtual Machines, managed Kubernetes (AKS), App Service and Azure Functions, Storage accounts and managed disks, Azure SQL and Cosmos DB, VNets with hybrid connectivity and global load balancing, Microsoft Entra ID and RBAC, and Key Vault secrets and encryption. Covers Azure service selection, infrastructure as code (ARM, Bicep, Terraform), observability with Azure Monitor and Kusto queries, cost governance and Azure Policy, the Azure Well-Architected design principles, and hybrid management via Azure Arc, common in enterprise Azure estates. For provider-agnostic trade-offs, see the cross-cloud entries.
Explain how Azure Arc can be used to manage Kubernetes clusters and SQL Servers running on-prem or in other clouds. Design a governance and deployment model using Arc-enabled servers and Arc-enabled Kubernetes with GitOps, and explain limitations, operational overhead, and typical enterprise use-cases.
Sample Answer
Direct answer
Azure Arc projects non-Azure infrastructure (on-prem servers, other clouds' VMs, and Kubernetes clusters anywhere) into Azure Resource Manager as first-class resources, so the same governance tools that already work on native Azure resources (Azure Policy, RBAC or role-based access control, tagging, Azure Monitor) also work on that outside infrastructure, without moving the workload itself into Azure. Arc-enabled servers register a machine as an Azure resource via a lightweight agent; Arc-enabled Kubernetes registers an existing cluster (any CNCF-conformant distribution, anywhere) the same way, and the recommended way to actually deploy and manage configuration on both is GitOps (a Flux-based extension that continuously reconciles the cluster or machine's state against a Git repository) rather than one-off imperative changes.
Structured elaboration
Arc-enabled servers. Install the Azure Connected Machine agent on a Windows or Linux server anywhere (on-prem, another cloud) and it registers as a Microsoft.HybridCompute/machines resource in a resource group of your choice. Once registered, that server can receive Azure Policy guest configuration assignments (audit or remediate OS-level settings), be onboarded to Microsoft Defender for Cloud, have Azure Monitor's agent installed and centrally managed the same way as a native Azure VM, and be governed by the same tag and RBAC model as everything else in the subscription. It does not become an Azure VM: compute still runs where it always ran, and Arc adds a management and governance plane on top, not a migration.
Arc-enabled Kubernetes. An existing cluster (on-prem via VMware/Hyper-V/bare metal, or a cluster in another public cloud) gets connected via an agent that establishes an outbound-only connection to Azure (no inbound firewall holes needed, since the cluster initiates the connection), after which the cluster appears as a Microsoft.Kubernetes/connectedClusters resource. From there, the GitOps extension (Flux v2) can be configured per cluster to continuously pull manifests from one or more Git repositories and reconcile the cluster's actual state to match, which is the recommended deployment model specifically because it gives drift detection for free: Flux does not just apply the desired state once, it continuously watches for and corrects manual changes made directly against the cluster, surfacing a compliance signal back to Azure Policy if a cluster has drifted from its assigned GitOps configuration.
Arc-enabled SQL Server (distinct from Arc-enabled Kubernetes above, and worth naming specifically). Where Arc-enabled Kubernetes brings GitOps to clusters, "SQL Server enabled by Azure Arc" is a different Arc capability aimed at existing SQL Server instances running anywhere (on-prem, another cloud, an edge site): it layers the Azure Connected Machine agent plus a separate Azure extension for SQL Server on top of an already Arc-enabled machine, giving a single Azure-side inventory of every SQL Server instance and database (version, edition, core count, host OS, queryable centrally via Azure Resource Graph), a best-practices configuration assessment, Microsoft Defender for Cloud's vulnerability scanning and threat alerts extended to those instances, Microsoft Entra ID authentication for the SQL Server instance itself (SQL Server 2022+), governance integration with Microsoft Purview, and, notably, an optional pay-as-you-go licensing model billed through Azure instead of a traditional perpetual license, which suits instances with variable compute demand (scaled down or shut off nights/weekends) or a short expected lifetime. This is governed the same way as the Kubernetes side: centrally via Azure Policy and RBAC on the Arc resource, not by logging into each SQL Server box individually, which is the whole point for an enterprise with dozens or hundreds of scattered SQL Server instances it currently has to inventory and patch by hand.
Designing the governance model. A workable structure: one Git repository (or a clearly organized mono-repo with per-cluster/per-environment folders) as the single source of truth for what should be running on each Arc-enabled Kubernetes cluster, an Azure Policy initiative assigned at the subscription or management-group level that requires every Arc-enabled cluster to have a GitOps configuration pointed at an approved repository (catching a cluster that was connected to Arc but never actually put under GitOps management), and RBAC on the Arc resource itself (not just on the underlying cluster) so who can view or modify the Arc-managed configuration in Azure is controlled centrally even though the cluster's own Kubernetes RBAC is separate and still applies to direct cluster access.
Limitations and operational overhead. Arc does not virtualize or abstract away the underlying infrastructure's operational reality: patching the underlying OS on an Arc-enabled server, upgrading the Kubernetes version on an Arc-enabled cluster, and managing the underlying hypervisor or bare-metal hardware all remain exactly the customer's responsibility; Arc adds a governance and visibility layer, not a managed-service layer, on top of infrastructure Azure does not operate. Network connectivity from the managed infrastructure back to Azure's control plane endpoints must be reliable (outbound HTTPS, specific required endpoints allowlisted through any corporate proxy or firewall); a server or cluster that loses that connectivity for an extended period shows as disconnected/stale in Azure and stops receiving policy evaluation or configuration updates until connectivity is restored, though the workload itself keeps running locally unaffected. There is also a real agent and extension footprint (CPU/memory overhead per node for the Connected Machine agent or the Arc Kubernetes agents) that should be accounted for in capacity planning on resource-constrained on-prem hardware, and a genuine skills requirement: a team needs to be fluent in both their existing on-prem/other-cloud operational tooling and Azure's governance tooling, which is a real, ongoing cross-training cost, not a one-time setup cost.
Typical enterprise use cases. Centralizing security posture and compliance reporting (Microsoft Defender for Cloud, Azure Policy compliance dashboards) across a genuinely hybrid or multi-cloud estate from one pane of glass, rather than maintaining separate governance tooling per environment; standardizing Kubernetes application deployment via GitOps across on-prem, edge, and multi-cloud clusters from the same pipeline used for AKS; and providing a migration on-ramp where infrastructure is governed consistently with Azure-native resources well before (or instead of) an actual lift-and-shift migration, letting an organization improve its security and compliance posture on infrastructure it has no near-term plan to physically move.
Positioning this to a worried infrastructure manager. The two most common (and reasonable) worries are "is this the first step toward being forced to migrate everything to Azure" and "what does this agent actually do to my servers." On the first: Arc explicitly does not require migration; the workload keeps running exactly where it is, and Arc's value (unified governance, security posture, one policy and RBAC model) is realized without ever moving a workload, which is worth stating plainly rather than letting it be assumed as a migration play in disguise. On the second: the agent's footprint, required outbound endpoints, and exact permissions it needs should be reviewed concretely (not hand-waved) against the manager's own change-control and security-review process, the same way any new agent or piece of infrastructure software would be, since "it's from Microsoft" is not itself sufficient answer to a legitimate security-review question about what an agent with configuration-management privileges is capable of doing to a production server.
Worked example
A retailer with 40 on-prem VMware clusters running Kubernetes for in-store point-of-sale systems, plus a growing AKS footprint for its e-commerce platform, wants one deployment pipeline and one compliance view across both. Arc-enabling the on-prem clusters and standardizing all of them (on-prem and AKS alike) on the same Flux-based GitOps configuration, pulling from the same Git repository structure with per-store-region overlays, means a single Azure Policy initiative can report compliance ("is this cluster running the currently approved application manifest version") across the entire hybrid fleet from one dashboard, and a security patch to the shared base manifest propagates to every connected cluster through the same Git-merge-and-wait-for-reconciliation process, rather than a separate manual rollout process for the on-prem fleet.
Trade-offs and pitfalls
The most common early mistake is connecting infrastructure to Arc for the visibility and compliance dashboard value without ever actually assigning GitOps or Policy guest configuration to it, which leaves the connected resource visible in Azure but not actually governed by anything, a false sense of control. The second is underestimating the network-reliability requirement: an environment with unreliable or heavily-filtered outbound internet access (common in some industrial/OT or classified on-prem environments) will see Arc resources flapping between connected and disconnected, which degrades the very policy-evaluation and drift-detection value Arc is meant to provide, and should be scoped and tested against the actual network environment before a wide rollout, not assumed to just work because it worked in a lab.
Explain the difference between Azure Active Directory (authentication/identity) and Azure RBAC (authorization). How would you design role assignments and administrative boundaries for a multi-team environment with dev, staging, and prod subscriptions to enforce least privilege and separation of duties?
Sample Answer
Direct answer
Microsoft Entra ID (the current name for what was Azure Active Directory) answers "who are you": it authenticates a user, service principal, or managed identity and issues a token. Azure RBAC (role-based access control) answers "what can you do": it evaluates a role assignment, a security principal plus a role definition plus a scope, against every Azure Resource Manager request after Entra ID has already authenticated it. For a multi-team environment, put dev, staging, and prod in separate subscriptions and grant each team roles scoped only to the subscription their work touches, never at the tenant level by default.
Structured elaboration
Authentication (Entra ID). Covers identity lifecycle (users, groups, app registrations and service principals, managed identities), sign-in security (multi-factor authentication, Conditional Access policies that can require a compliant device or a specific network location before a token is issued), and directory-level roles such as Global Administrator, which manage the directory itself, not Azure resources.
Authorization (Azure RBAC). A role assignment binds a security principal, always sourced from Entra ID, to a role definition (a list of allowed actions, such as Microsoft.Compute/virtualMachines/start/action) at a scope: management group, subscription, resource group, or a single resource. RBAC is additive: a principal's effective permissions are the union of every assignment applying at or above its own scope, so a broad assignment at a management group is inherited by everything beneath it. That inheritance is exactly why scope discipline is the design problem, not the role names themselves.
Design for dev, staging, and prod. Put each environment in its own subscription, not just a resource group, since a subscription is the strongest RBAC, cost-reporting, and Azure Policy isolation boundary Azure offers, then group the three under a shared management group so tenant-wide guardrails apply consistently. Assign the built-in Contributor role to the engineering team at the dev subscription; a narrower custom role (deploy without delete) at staging; and, at prod, no standing human access beyond Reader, with real changes flowing only through a pipeline's service principal holding the minimum RBAC it needs. Use Entra ID groups, never individual users, as the RBAC principal, so onboarding and offboarding become a group-membership change instead of a per-subscription audit.
Separation of duties, concretely. The person who approves a pipeline's production deployment stage should not also hold standing Contributor on the prod subscription. The person who manages Entra ID directory roles, who could in principle grant themselves any Azure RBAC role by first granting themselves ownership, is a distinct, separately audited function from day-to-day Azure administration.
Worked example
A team of 8 engineers and 2 leads works across dev, staging, and prod subscriptions under one management group. Dev: the "eng-dev" Entra ID group gets Contributor at the dev subscription, so all 8 engineers create and delete resources freely. Staging: the same group gets a custom role granting most write actions but no delete actions, so promotion happens but destructive mistakes are harder to make by accident. Prod: no standing human access at all; only the CI/CD pipeline's service principal holds Contributor scoped to prod, and the 2 leads hold Reader plus eligibility for a time-boxed, approval-gated Owner role activation (Privileged Identity Management) for genuine emergencies. This gives a clean, mechanical answer to "who could have made this prod change on this date": either the pipeline's identity, which is logged and code-reviewed, or a specific, timestamped emergency activation.
Trade-offs and pitfalls
Putting all three environments in one subscription and trying to enforce the boundary with resource groups and RBAC alone works at small scale, but subscription-level quotas, cost reporting, and the blast radius of a single subscription-wide network or policy mistake all leak across environments anyway. Assigning roles to individual users instead of groups creates an offboarding gap: a departed employee's access lingers until someone remembers to remove it, because it is not visible in one place. Confusing Owner, which includes the ability to grant others any role, with Contributor, which cannot manage access at all, is a common mistake when writing a role for someone who should deploy but never re-permission the environment.
A high-throughput OLTP workload in Azure SQL Database is showing increased latency during peak hours. Walk through diagnostic and optimization steps: using Query Store and Query Performance Insights, analyzing wait stats, index tuning and missing index analysis, parameter sniffing mitigation, partitioning strategies, in-memory OLTP, and when to consider horizontal scaling or Hyperscale.
Sample Answer
Direct answer
Diagnose before tuning: use Query Store and Query Performance Insight to find which specific queries regressed and when, check wait statistics to confirm what resource the database is actually waiting on, then apply the fix that matches the diagnosis (an index, a parameter-sniffing mitigation, partitioning, or In-Memory OLTP, OLTP meaning Online Transaction Processing) rather than reaching for the first plausible-sounding technique. Only after the workload has been tuned as far as a single database reasonably goes should horizontal scaling or Hyperscale (an Azure SQL service tier that separates compute, storage, and log so you can scale each independently and add read replicas) enter the conversation.
Diagnostic steps, in order
- Query Store and Query Performance Insight: Query Store retains historical execution plans and runtime statistics per query, and Query Performance Insight (built on top of it in the Azure portal) surfaces the top resource-consuming and longest-running queries over the peak window. Start here to identify which queries regressed, not just that "the database is slow."
- Wait statistics (
sys.dm_os_wait_statsand Query Store's per-query wait breakdown): confirm what the engine is actually blocked on, lock contention, disk I/O, CPU, or memory grants, before assuming the fix is an index. Tuning an index when the real wait is lock contention wastes effort and doesn't move the needle. - Missing-index analysis (
sys.dm_db_missing_index_details,_groups,_group_stats): these dynamic management views (DMVs) surface indexes the optimizer would have used had they existed, ranked by estimated impact. Cross-check candidates against Query Store's top offenders before creating anything, since the missing-index DMVs can suggest indexes that overlap with existing ones or that only help a rare query. - Parameter sniffing check: if Query Store shows the same query plan-cached statement performing very differently across executions, and wait stats don't point at a hardware bottleneck, parameter sniffing is the likely cause: the optimizer cached a plan built for one parameter value's data distribution and it performs badly for a different value. Mitigations, cheapest first:
OPTION (RECOMPILE)on the specific statement if it's infrequent enough to tolerate replanning cost each time; Query Store forced plans to pin a known-good plan;OPTIMIZE FOR UNKNOWNto plan for an average case; and, on database compatibility level 160, Parameter Sensitive Plan (PSP) optimization, an intelligent query-processing feature that automatically maintains multiple cached plans for one parameterized statement when the data distribution is genuinely non-uniform, which is the current, lowest-effort fix when available. - Partitioning: for a large, date-ordered table (order history, event logs), range-partitioning by date lets old partitions be maintained (index rebuilds, archiving) independently of the hot, recent partition, and lets the optimizer skip irrelevant partitions entirely for date-scoped queries. This addresses maintenance-window and query-scan cost, not raw peak-hour throughput.
- In-Memory OLTP: memory-optimized tables and natively compiled stored procedures remove locking and latching almost entirely for the tables converted, which is the right tool specifically when wait statistics point at lock or latch contention on a small number of very hot tables, not a general-purpose speed-up for the whole database.
- Horizontal scaling or Hyperscale: once single-database tuning is exhausted, add a free built-in read-scale-out replica to offload reporting-style reads from the primary, or move to the Hyperscale service tier, which separates compute, storage, and log so you can add up to 30 independently-scaled named read replicas and scale compute up or down without a size-of-data operation, precisely for workloads that have outgrown vertical tuning on a single compute node.
Worked example
Query Performance Insight shows one specific stored procedure spiking in duration only during the 9am login rush. Query Store's plan history shows two different execution plans for that procedure over the past week, with the slow plan doing a table scan and the fast plan doing an index seek, a signature of parameter sniffing rather than a missing index (the index exists; the wrong plan just isn't using it well for some parameter values). Forcing the known-good plan via Query Store gives an immediate fix; migrating the database to compatibility level 160 to get Parameter Sensitive Plan optimization is the durable fix, since it lets the engine keep multiple plans instead of relying on one to fit every login pattern.
Trade-offs and pitfalls
- Creating an index suggested by the missing-index DMVs without checking Query Store's actual top offenders can add write overhead (every index has an insert/update cost) for a query that wasn't actually the peak-hour bottleneck.
OPTION (RECOMPILE)fixes parameter sniffing but adds CPU cost to compile a fresh plan on every execution; it's a reasonable stopgap for a low-frequency statement, a poor permanent fix for a hot one.- Partitioning helps scan cost and maintenance windows, but it does not, by itself, relieve lock contention on hot rows within the current partition; don't reach for it as a fix for contention symptoms.
- Jumping straight to Hyperscale or horizontal scaling before ruling out an index, plan, or contention problem trades a real fix for a bigger, more expensive database that still has the same underlying issue.
A sudden spike in 5xx errors correlates with an Azure scheduled platform maintenance window. Describe how you would confirm whether platform maintenance caused the issue, engage Microsoft support effectively, and design mitigations to reduce exposure in future maintenance events (e.g., using zones, maintenance control options, multi-region strategies).
Sample Answer
Direct answer
Do not accept the correlation at face value. Confirm causation first with Azure's own maintenance records (Service Health's planned-maintenance notifications, Resource Health's history for the specific resource, and, for infrastructure-as-a-service virtual machines, the Scheduled Events feed), because "correlates with a maintenance window" and "was caused by the maintenance window" are different claims, and only one of them is worth escalating to Microsoft support on. Once confirmed, engage support with the specific tracking ID and timestamps rather than a general ticket, and treat the incident as evidence that your own architecture, not Microsoft's schedule, is the thing you actually control going forward.
Structured elaboration
Confirming the cause
Azure Service Health, in the portal, has a "Planned maintenance" blade scoped to your subscription; it lists scheduled maintenance events with a tracking ID, the affected resource types, and a time window. Cross-reference your 5xx timestamps against this list first.
Resource Health for the specific affected resource (a virtual machine, an App Service, a database) keeps a history of availability state changes and categorizes each one, including a PlannedMaintenance reason code distinct from an unplanned interruption or a user-initiated change. This is the most direct evidence available, because it is scoped to the exact resource that produced the 5xx responses, not just "something happened in the region."
For infrastructure-as-a-service virtual machines, the Instance Metadata Service (IMDS, a service reachable only from inside the VM at a fixed address) exposes a Scheduled Events feed that records upcoming Freeze, Reboot, Redeploy, or Terminate events with an event ID and a time window, before they happen. If the application was not already polling this feed, you cannot get advance warning after the fact, but Resource Health's retrospective record still lets you confirm whether a Scheduled Event lines up with the 5xx window.
Check the Activity Log for the resource group as a secondary corroborating source, and resist the temptation to stop at "the timing lines up": a deployment, a dependent service's own incident, or a capacity issue elsewhere can produce a 5xx spike that happens to coincide with, but was not caused by, platform maintenance.
Engaging Microsoft support effectively
Open the support request directly from the Service Health "Planned maintenance" or "Health history" blade for the specific event, which pre-fills the tracking ID; a ticket that references a concrete tracking ID and exact timestamps gets routed and triaged faster than a ticket that says "we saw errors."
State the business impact in the ticket (error rate, duration, affected customers) and explicitly ask for a Post Incident Review, Microsoft's term for a written root-cause report, not just confirmation that maintenance occurred. Confirmation alone will not explain why your specific workload was impacted more than expected.
If response time matters, this is where your support plan tier matters: a plan with faster initial response times gets you an engineer sooner, which is worth knowing before the next incident, not during this one.
Reducing exposure to future maintenance events
Availability Zones (AZs): an AZ is a physically separate datacenter facility within a region with independent power and networking. Spreading a tier across three zones means a maintenance operation scoped to one zone, or to the underlying rack or cluster, does not take out the whole tier at once. This is the highest-leverage fix for the class of maintenance that affects a subset of the region's infrastructure.
Maintenance control options: be precise about scope here, because it is easy to overclaim. Azure's Maintenance Configurations feature lets you pick your own maintenance window, but its host scope, the one covering platform-level updates, only applies to isolated virtual machines, isolated scale sets, and dedicated hosts, a premium, single-tenant compute option. A regular multi-tenant VM size does not get to defer Microsoft's platform maintenance schedule. What a regular VM does get is the Scheduled Events feed described above, which gives a warning window, as little as a few seconds for a Freeze event, longer for others, to drain connections or checkpoint state before the maintenance action hits: a real mitigation even without control over the timing itself.
Multi-region strategies: an active-active or active-passive deployment behind Azure Front Door or Traffic Manager, with health probes tuned tighter than the defaults, routes traffic away from a region experiencing a maintenance-related blip. This only helps if the probe interval and unhealthy-threshold are tight enough to catch a short-duration event. Traffic Manager's own defaults, a 30-second probing interval with a tolerated-failures count of 3, mark an endpoint Degraded after roughly 90 to 120 seconds of continuous failure by Microsoft's own documented timeline, which is comfortably inside an 8-minute window, so the default probe cadence is not actually the risk here. The real risk with Traffic Manager specifically is DNS caching: once an endpoint is marked Degraded, clients and recursive resolvers that already cached the old DNS answer keep sending traffic to the unhealthy region until that cached answer's TTL expires, so a profile left at a long TTL can keep sending a meaningful share of traffic to the bad region for several minutes after Traffic Manager itself has already detected the problem. A probe interval loose enough to be the bottleneck on its own would need to be much slower than the default, for example around a 90-second interval with a five-consecutive-failure threshold, which pushes detection out into roughly the 6-to-8-minute range and does start to compete with an 8-minute window. Front Door, which fails over at its edge routing layer rather than through DNS TTL expiry, does not carry this specific caching risk.
Worked example
Say Resource Health confirms the outage window was 02:14 to 02:22 UTC (8 minutes) and, during that window, the error rate on the affected service was 40% instead of its normal baseline of under 1%. To size the business impact for the support ticket and for your own service level agreement (SLA, the contractual or internal target for acceptable downtime or error rate) tracking: 8 minutes out of a 30-day month (30 x 24 x 60 = 43,200 minutes) is 8 / 43,200 = 0.0185% of the month, which by itself sounds small. But at a 40% error rate on real user traffic during that window, if the service normally handles 500 requests/minute, that is roughly 500 x 8 x 0.40 = 1,600 failed requests concentrated in 8 minutes, which is the number worth putting in the support ticket and the postmortem, not the diluted monthly percentage. This is also the number that tells you whether adding Availability Zones is worth the added cost: if this pattern repeats monthly, 1,600 failed requests/month is a very different business case than a one-time event.
Trade-offs & pitfalls
Spreading a tier across Availability Zones adds cross-zone latency, typically low but non-zero, and cost, since redundant capacity now runs in zones you did not previously need for pure scale reasons. Justify it with the impact math above, not with "zones are best practice."
Isolated VMs and dedicated hosts, the only compute options that get real control over host-scope maintenance timing, carry a real price premium. Reach for them when the workload's sensitivity to any freeze or reboot genuinely justifies it (gaming, media streaming, financial transactions are the textbook cases), not as a default reaction to one incident.
A multi-region failover strategy only protects you if the health probe is tuned for the failure you actually saw; a default, loosely-tuned probe configuration gives the appearance of resilience without the substance.
The most common wrong turn here is treating "Microsoft caused it" as the end of the investigation. Even a confirmed platform-maintenance root cause is a finding about your architecture's exposure, not an excuse: the Post Incident Review will not change Microsoft's schedule quickly, but redundancy design is fully in your control today.
Explain Azure Managed Identities (system-assigned vs user-assigned). Show how you'd use a managed identity to allow an AKS workload to access Azure Key Vault without storing credentials, and outline operational concerns such as lifecycle, principal rotation, and RBAC scoping.
Sample Answer
Direct answer
A system-assigned managed identity is created and destroyed with a single Azure resource; a user-assigned managed identity is its own standalone object you create once and attach to many resources. For an Azure Kubernetes Service (AKS) workload calling Key Vault, use a user-assigned identity through Microsoft Entra Workload ID (the current name for what was called Azure AD Workload Identity federation) so pods get short-lived tokens with no credential ever stored in a secret or a config file.
Structured elaboration
System-assigned versus user-assigned. A system-assigned identity is the simplest option for a one-to-one relationship: create a virtual machine, it gets an identity automatically, delete the virtual machine, the identity disappears with it. A user-assigned identity is created independently in Microsoft Entra ID (Azure's cloud identity and directory service, the current name for Azure Active Directory), has its own lifecycle, and can be attached to many resources at once, which matters here because you want the same identity to survive an AKS node pool being replaced or scaled, and you may want several pods across the cluster to share one identity's permissions rather than provisioning one identity per pod.
Wiring it to AKS and Key Vault, step by step.
- Create a user-assigned managed identity.
- Grant it least-privilege access on the specific Key Vault, using Azure role-based access control (RBAC) with a scoped role like Key Vault Secrets User (get and list only, not set or delete), rather than a broader role or a subscription-wide scope.
- Enable Microsoft Entra Workload ID on the AKS cluster and create a Kubernetes ServiceAccount annotated with the managed identity's client ID:
apiVersion: v1
kind: ServiceAccount
metadata:
name: kv-client-sa
annotations:
azure.workload.identity/client-id: "<user-assigned-client-id>"
- Deploy the pod using that ServiceAccount. The application code uses the Azure SDK's default credential chain, which automatically exchanges a Kubernetes-issued, short-lived token for an Azure access token through the federated identity trust, with no secret stored anywhere in the cluster or the container image.
Operational concerns.
- Lifecycle. Favor user-assigned identities for anything that must survive infrastructure churn (node pool upgrades, pod restarts); document which application owns which identity so an identity is never orphaned when the app it belonged to is decommissioned.
- Principal rotation. The underlying credential the managed identity uses is rotated automatically by the platform; what you actually manage is the federated trust between the Kubernetes ServiceAccount and the identity, so if you ever replace the identity itself, you must update the ServiceAccount annotation and redeploy the affected pods.
- RBAC scoping. Grant only the specific permission needed (secrets get/list, not full key or certificate management) and use a separate identity per trust boundary (per environment, per application) so a compromise of one workload's token cannot reach another workload's secrets.
- Auditing. Enable Key Vault diagnostic logs and Microsoft Entra sign-in logs so every token exchange and every secret access is attributable to a specific identity and, transitively, a specific pod.
Worked example
Trace a single request end to end: a pod using the kv-client-sa ServiceAccount calls the Azure SDK to fetch a database connection string. The SDK's credential chain detects it is running under workload identity federation, presents the pod's Kubernetes-issued, short-lived service account token to Microsoft Entra ID, which validates the federated trust configured between that specific ServiceAccount and the user-assigned managed identity, and issues a short-lived Azure access token scoped to whatever the identity can do. The application uses that token to call Key Vault's get-secret operation; Key Vault's RBAC check confirms the identity has the Key Vault Secrets User role and returns the value. No password, connection string, or service principal secret was ever written to a Kubernetes secret, an environment variable at rest, or a config map; the only long-lived object in the whole chain is the federated trust configuration itself.
Trade-offs & pitfalls
The most common mistake is defaulting to a system-assigned identity for convenience and then discovering it disappears the moment the node pool is recreated during an upgrade, silently breaking every pod that depended on it. The second is over-scoping the RBAC grant "to save time," handing the workload identity broader Key Vault permissions than get/list; that turns a single compromised pod into a path to delete or overwrite every secret in the vault, not just read the one it needed.
Unlock Full Question Bank
Get access to all Microsoft Azure Services and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.