Microsoft Azure Services and Architecture Questions
Microsoft Azure's core service catalog and architectural patterns: Virtual Machines, managed Kubernetes (AKS), App Service and Azure Functions, Storage accounts and managed disks, Azure SQL and Cosmos DB, VNets with hybrid connectivity and global load balancing, Microsoft Entra ID and RBAC, and Key Vault secrets and encryption. Covers Azure service selection, infrastructure as code (ARM, Bicep, Terraform), observability with Azure Monitor and Kusto queries, cost governance and Azure Policy, the Azure Well-Architected design principles, and hybrid management via Azure Arc, common in enterprise Azure estates. For provider-agnostic trade-offs, see the cross-cloud entries.
An enterprise wants to migrate hundreds of VMs and databases to Azure. Compare lift-and-shift (IaaS) to re-platforming (PaaS) and refactoring. Discuss migration effort, downtime, cost, operational benefits, and propose a phased migration roadmap with pilot, migration waves, validation, and rollback plans. Also cover tools (Azure Migrate, Database Migration Service) you would use.
Sample Answer
Direct answer
For an enterprise moving hundreds of virtual machines (VMs) and databases, run all three migration strategies in parallel against different slices of the estate rather than picking one: lift-and-shift (infrastructure-as-a-service, IaaS, meaning you keep managing the operating system) the low-risk majority for a fast win, re-platform (platform-as-a-service, PaaS, where Microsoft manages the OS) the workloads that would clearly benefit from managed scaling and patching, and reserve refactoring for the small number of strategic applications where a redesign pays for itself. Sequence it as a pilot, then waves, with validation and rollback built into every wave, not bolted on at the end.
Structured elaboration
| Strategy | Migration effort | Downtime | Cost pattern | Operational benefit |
|---|---|---|---|---|
| Lift-and-shift (IaaS) | Low: replicate the VM image and cut over | Low to moderate, shrinkable with replication tools | Lower upfront effort, but the team keeps paying the same operational cost (patching, backup, HA configuration) it paid on-premises | Fast time-to-cloud, minimal application change |
| Re-platform (PaaS) | Medium: adapt the app to a managed service (Azure App Service, Azure SQL Managed Instance) | Moderate, minimized with a blue-green cutover (running old and new side by side and switching traffic) | Lower ongoing operations cost, better built-in scaling and patching | Reduced management burden, built-in high availability |
| Refactor | High: redesign into microservices or serverless components | Larger per component, but rolled out incrementally | Higher upfront engineering cost, best long-run agility and cost efficiency at scale | Cloud-native resilience and faster feature delivery |
Tools. Use Azure Migrate for discovery, dependency mapping between VMs (so you do not split an application across migration waves and break it), sizing recommendations, and orchestrating the lift-and-shift itself. Use Azure Database Migration Service (DMS) for the database moves, whether homogeneous (SQL Server to SQL Server) or heterogeneous (for example, Oracle to an Azure-native database). Use Azure Site Recovery for VM-level replication and a tested failback path. Use Azure Cost Management and Azure Advisor after each wave to right-size what you just moved instead of leaving it at its on-premises sizing forever.
Phased roadmap.
- Pilot (4 to 8 weeks). Migrate 10 to 20 non-critical VMs plus one representative database using Azure Migrate. The goal is to validate the tooling and the rollback procedure under real conditions, not to prove the whole migration works.
- Wave planning. Categorize every remaining workload as rehost, re-platform, refactor, or retire (a common "four Rs" framework), grouped by dependency graph so a wave never splits two systems that talk to each other synchronously.
- Migration waves. Each wave runs pre-migration compatibility checks, replication, a scheduled cutover window, and post-cutover validation scripts before the next wave starts.
- Validation. Functional tests, performance comparison against the on-premises baseline, security review, and a named stakeholder sign-off gate the wave's completion.
- Rollback plan. Keep the source systems in a stopped-but-restorable state (VM snapshots, database backups) for a defined retention window per wave, with a documented, rehearsed runbook for reverting a wave, not just a verbal understanding of "we can always go back."
Worked example
Consider an enterprise with 300 VMs and 40 databases. The dependency map (built from Azure Migrate's discovery data) shows a cluster of 60 VMs and 5 databases behind one customer-facing order system, and 240 largely independent VMs running departmental tools. Sequence: pilot on 15 of the departmental VMs with no cross-dependencies. Wave 2 through 5 lift-and-shift the remaining departmental VMs in batches of roughly 55 to 60, since they have no coupling risk. The order system cluster becomes its own wave, re-platformed together (all 60 VMs and 5 databases) because splitting it across waves would leave part of the system talking to Azure and part still on-premises mid-cutover, adding exactly the kind of latency and consistency risk the roadmap exists to avoid.
Trade-offs & pitfalls
The dominant failure mode is treating migration strategy as an enterprise-wide policy decision ("we are a re-platform shop") instead of a per-workload one; that produces either needless re-engineering of simple utility VMs or missed savings on workloads that would have benefited from PaaS. The second is skipping the pilot's real purpose: a pilot exists to rehearse rollback and validation, not just to prove migration is technically possible, so pick a pilot workload that is genuinely representative of a later wave's risk, not the easiest thing to move.
Propose a cost-optimized architecture for a batch data processing pipeline that ingests 10 TB/day with a 6-hour SLA. Compare using Azure Batch (spot VMs), AKS with scale-to-zero nodes, Databricks with spot workers, and serverless orchestration. Discuss throughput, checkpointing, handling spot preemptions, and operational overhead for each choice.
Sample Answer
Direct answer
For 10 TB/day with a 6-hour SLA, Azure Batch with Spot VMs is the strongest default if the workload can checkpoint and tolerate node preemption (lowest cost, most operational overhead to build the checkpointing yourself); Databricks with a mostly-Spot job cluster is the strongest default if the workload is genuinely Spark-shaped and the team wants Delta Lake's transactional guarantees (Delta Lake is a transactional table format built on top of a data lake) with less custom infrastructure to build; AKS (Azure Kubernetes Service) with scale-to-zero fits best when this batch job is one of several workload types already living on a shared Kubernetes platform and the team wants one operational model instead of a dedicated batch service; serverless orchestration (Azure Functions/Durable Functions or Logic Apps fanning out to one of the above compute options) is the wrong primary compute layer for 10 TB of data processing itself, but is a strong choice for the orchestration and glue around whichever compute layer does the heavy lifting.
Structured elaboration
Sizing the requirement first. 10 TB/day inside a 6-hour processing window means an average sustained throughput of 10 x 10^12 bytes / (6 x 3600 seconds) = 10,000,000,000,000 / 21,600 ≈ 463,000,000 bytes/second ≈ 463 MB/s ≈ 3.7 Gbps sustained, before accounting for any headroom for slow starts, retries, or uneven daily volume (a real pipeline should size for its P95 daily volume within the SLA window, not its average, since an SLA measured against the average day is not really a 6-hour SLA on the days that matter). This number matters because it rules out anything that can only reach a fraction of that sustained rate regardless of how the compute layer is chosen.
Azure Batch with Spot VMs. Best throughput-per-dollar of the four options for embarrassingly parallel or checkpoint-friendly workloads (partition the 10 TB into independent chunks, process each on a pool of Spot VMs). Checkpointing is entirely the team's responsibility: each work item must be small enough, and progress tracked durably enough (a completion marker per chunk in Blob Storage or a database row), that a preempted node's in-flight chunk can be picked up by another node without reprocessing the whole thing or double-counting a partial result. Handling preemption means subscribing to Batch's node preemption notification and requeuing the affected tasks; Batch does this requeueing for you at the task level if the pool is configured correctly, but the task itself must be idempotent (safe to run again) for that requeue to be safe. Operational overhead is the highest of the four here: no managed job scheduler UI beyond Batch's own, no built-in data-lineage or table format, so the team is building more of the plumbing.
AKS with scale-to-zero node pools. Good fit if the team already runs other workloads on this AKS cluster and wants one platform rather than a separate batch-specific service. A user node pool's cluster autoscaler minimum can be set to 0 (system node pools cannot go to zero, since the control plane (the managed Kubernetes components that run the cluster itself) has its own critical pods that need somewhere to run), so the pool costs nothing when idle and scales up only when the daily job starts, using KEDA (Kubernetes Event-Driven Autoscaling, a component that scales workloads based on external event metrics) or a scheduled CronJob to trigger the scale-up. Spot node pools are supported for AKS user node pools too, combining scale-to-zero with Spot pricing. Preemption handling here is Kubernetes-native (a SIGTERM and a grace period before eviction, and a Pod Disruption Budget to control how many can go at once), which a team already fluent in Kubernetes will find more familiar than Batch's task-preemption model, at the cost of needing genuine Kubernetes operational maturity (node pool sizing, PDBs, resource requests/limits tuned correctly) to avoid the scheduler thrashing under a sudden 10 TB/day burst.
Databricks with Spot workers. Best fit when the transformation logic is naturally expressed in Spark and the team wants Delta Lake's ACID transactions and schema enforcement instead of hand-rolling idempotency and checkpoint tracking. Databricks Job Clusters support a mix of on-demand and Spot workers (a small on-demand "driver-adjacent" core plus a larger Spot-heavy worker pool), and Spark's own task-level retry and lineage tracking absorbs a fair amount of the preemption-handling burden that Batch or raw AKS would otherwise put entirely on the team. Operational overhead is the lowest of the compute-heavy three here for a Spark-shaped workload, at a real dollar premium over raw Batch/AKS for the same underlying VM cost (Databricks' own per-DBU charge (DBU = Databricks Unit, Databricks' own compute-billing unit) sits on top of the VM cost), which needs to be weighed against the engineering time saved.
Serverless orchestration. Azure Functions/Durable Functions have per-invocation duration and payload-size limits that make them a poor fit for moving or transforming 10 TB directly; use them instead as the layer that kicks off the batch job on a schedule, fans out work items with retry semantics, and orchestrates dependent steps (validate inputs, trigger the Batch/AKS/Databricks job, wait for completion, trigger the next stage), while the actual data movement and transformation happens in one of the three compute options above.
Worked example: throughput and checkpoint granularity
Given the ≈463 MB/s sustained requirement, if the pipeline partitions the day's 10 TB into fixed-size chunks of, say, 2 GB each, that is 10,000 GB / 2 GB = 5,000 chunks to process within the 6-hour window. At an assumed per-node sustained throughput of roughly 100 MB/s per worker (a conservative, illustrative figure for a general-purpose VM doing real transformation work, not just a network copy), reaching the required 463 MB/s aggregate needs on the order of 5 concurrent workers running for the full window in the ideal case, but Spot preemption and uneven chunk processing time mean the pool should be sized with real headroom (commonly 30-50% above the theoretical minimum) so a wave of simultaneous preemptions does not, by itself, blow the 6-hour SLA. This is exactly the kind of number a design answer should compute rather than assert: it turns "add enough Spot VMs" into a concrete target the on-call team can check the pool against during an actual run.
Trade-offs and pitfalls
The single biggest risk across all Spot-based options is correlated preemption: if the pipeline's compute is concentrated in one VM size/region/zone, a regional capacity crunch can evict a large fraction of the pool simultaneously, which is very different from losing one node at a time; diversify across a few VM sizes/families in the pool (Batch and AKS both support this) so a capacity shortage in one SKU does not take down the whole job at once. A second common pitfall is treating "6-hour SLA" as license to start the job right at the edge of a 6-hour window every night with no margin; build in a genuine safety margin (start with enough time that even a bad-but-plausible night, with above-average preemption or above-average data volume, still finishes inside the SLA) rather than a plan that only works on an average day.
You're designing compute for a latency-sensitive transactional service. Describe the selection process for Azure VM sizes including considerations for vCPU, memory ratio, local/ephemeral disk availability, managed disk IOPS and throughput, network bandwidth, and cost. Explain how you'd validate sizing with benchmarks and telemetry.
Sample Answer
Direct answer
Size compute by working backward from the P99 latency budget (the 99th-percentile response time target, meaning 99 out of 100 requests must finish faster than this) for the slowest step in the critical path, usually a disk write or a network round trip, not raw CPU, which typically points to a general-purpose or compute-optimized SKU paired with Premium SSD v2 or Ultra Disk rather than the cheapest SKU that merely has "enough" vCPU and RAM on paper. Then validate that choice with an actual load test against production-representative traffic before committing to it in capacity planning, since a synthetic single-request benchmark systematically misses the tail-latency effects that matter most for a transactional workload.
Structured elaboration
vCPU and memory ratio. A transactional service is usually not CPU-bound the way a batch job is, so start from a general-purpose D-series (roughly 1:4 vCPU-to-GiB) rather than compute-optimized F-series (roughly 1:2) unless profiling shows the service is genuinely CPU-saturated at its target request rate. Oversizing vCPU for a service actually gated by I/O or network wastes budget without moving the metric you are trying to improve.
Local and ephemeral disk. Relevant only for genuinely disposable state, a request-scoped cache or a scratch path. A transactional service's real data belongs on durable managed disk, so this is a secondary sizing decision here, not the primary one it is for a memory-heavy analytics workload.
Managed disk IOPS and throughput. For the database or log volume backing the transactional path, Premium SSD v2 is the current default recommendation over the older Premium SSD (v1), since it lets you provision IOPS (input/output operations per second, how many read or write operations the disk can handle each second) and throughput independently of capacity, so a small volume can still get high IOPS without over-provisioning size just to unlock performance, generally at a lower cost per unit of performance than v1. Move to Ultra Disk only if the sustained requirement exceeds Premium SSD v2's ceiling, roughly 80,000 IOPS and up to somewhere in the 1,200 to 2,000 MB/s range depending on configuration, figures that should be checked against current Azure documentation before being written into a specific capacity plan, or if the workload needs its provisioned performance changed dynamically without a disk swap.
Network bandwidth. Every Azure VM size has a fixed network bandwidth cap tied to its size, with larger SKUs getting proportionally more. For a low-latency service, accelerated networking, which bypasses the host's software network stack, should be enabled by default on any supported SKU, since it measurably reduces both latency and CPU overhead spent on network processing, close to a free win rather than a trade-off.
Cost and pricing model, including bursting. For a genuinely steady, always-on transactional service, standard general-purpose SKUs on a Reserved Instance or Savings Plan are the right pricing model. For a service with a low steady baseline and occasional bursts, common for a smaller or newly ramping transactional workload, B-series burstable VMs accumulate CPU credits during idle periods and spend them during a burst, which can be materially cheaper than provisioning a non-burstable SKU sized for peak, but only if the real burst pattern stays within what the accumulated credit balance can cover. A service with frequent, sustained bursts exhausts its credits and falls back to throttled baseline performance mid-burst, exactly the failure mode a latency-sensitive service cannot tolerate, so B-series should be chosen only after checking the real burst profile against the credit-accumulation math for that specific size, not assumed by default because it is cheaper on paper.
Contrast with a CPU-bound batch job. A CPU-bound nightly batch job is the mirror image of this exercise: it benefits from F-series' higher vCPU-to-memory ratio, tolerates much higher single-request latency, and can often run on Spot VMs (spare Azure compute capacity sold at a steep discount that Azure can reclaim with little notice), since a delayed or restarted batch run rarely breaches an SLA (Service Level Agreement, a contractual performance or uptime guarantee) the way a delayed transactional request does. Naming this contrast explicitly matters for correctly scoping which sizing principles apply to which workload shape.
Validating sizing with benchmarks and telemetry. Run a load test replaying production-representative request shapes and concurrency, not a single-request synthetic benchmark, which misses queueing and contention effects that only appear under concurrent load, against a candidate SKU, capturing P50, P95, and P99 latency (the median, 95th-percentile, and 99th-percentile response times) alongside the VM's own CPU, memory, and disk-queue-length metrics from Azure Monitor during the test. If P99 latency is acceptable but disk queue length is climbing toward the provisioned IOPS ceiling, that is a leading indicator the current disk configuration will become the bottleneck under future growth, well before the SLA actually breaks in production.
Worked example
Two candidate configurations for a transactional API with a 50 ms P99 latency budget: Config A is a 4 vCPU, 16 GiB SKU with a Premium SSD v2 disk provisioned for 5,000 IOPS. Config B is a smaller, cheaper 2 vCPU, 8 GiB SKU with the same 5,000 IOPS disk. A load test replaying the production request mix at expected peak concurrency shows Config A at P99 = 38 ms with CPU peaking at 55 percent, and Config B at P99 = 61 ms with CPU peaking at 91 percent, breaching the 50 ms budget under load despite having "enough" memory on paper. The telemetry, not the spec sheet, is what disqualifies Config B: its vCPU count, not its disk or memory, is the bottleneck at this concurrency level, which the load test surfaces and a single-request synthetic benchmark would likely have missed, since a single request never contends for CPU with itself.
Trade-offs and pitfalls
Choosing a SKU purely from vCPU and RAM numbers on a pricing page, without ever load-testing under realistic concurrency, is exactly how Config B above would have shipped. Defaulting to B-series burstable VMs for cost savings on a workload whose actual traffic pattern turns out to be sustained rather than bursty leads to credit exhaustion and a latency cliff in production. And treating accelerated networking as optional or "something to turn on later," when it is a same-cost setting that measurably helps the exact metric, tail latency, this whole sizing exercise is protecting, is a needless gap.
Describe how to use Azure Policy to enforce resource-level constraints such as requiring an enforced 'cost-center' tag, permitted VM SKUs, and encryption at rest. Provide an example policy effect (Audit, Deny, DeployIfNotExists) for tagging and explain how to remediate existing non-compliant resources without breaking production.
Sample Answer
Direct answer
Azure Policy enforces resource-level rules declaratively: a policy definition states a condition and an effect, and an assignment applies it at a scope (management group, subscription, or resource group). For the three examples in the question: a required cost-center tag uses Deny (or Modify/DeployIfNotExists if you want to auto-remediate rather than just block) at creation time; permitted VM SKUs use Deny against any SKU outside an allowed list; encryption at rest is usually already enforced by the platform for many resource types, but where it is optional (some storage or database configurations), Audit or Deny verifies it is turned on. Remediating existing non-compliant resources without breaking production means using Audit first to find the population, then a scoped, reviewed remediation task (for tags) or a planned, tested change (for anything that risks an outage) rather than a blanket auto-fix.
Structured elaboration
The three effects asked about, and what each actually does.
- Audit: allows the operation to proceed but flags the resource as non-compliant in Policy's compliance dashboard. Zero risk of breaking anything, since nothing is blocked or changed; the trade-off is that non-compliant resources keep getting created until someone acts on the audit findings.
- Deny: blocks the create/update operation outright if it violates the condition. Effective for stopping new non-compliance immediately, but has zero effect on resources that already exist and already violate the rule; a
Denypolicy assigned today does not retroactively fix yesterday's resources. - DeployIfNotExists (DINE): after a resource is created or updated, checks whether a related resource or configuration exists (a diagnostic setting, an extension, a specific configuration) and if not, deploys it using a specified template, via a managed identity attached to the policy assignment. This is how you get "every VM automatically gets a monitoring agent" or "every storage account automatically gets a diagnostic setting pointed at the central workspace" without anyone manually configuring it per resource. (A close relative,
Modify, directly changes properties of the resource itself, such as adding or updating a tag, rather than deploying a related resource.)
Example: enforcing the cost-center tag. A Deny policy with a condition like field 'tags[cost-center]' exists 'false' blocks any new resource created without that tag, which is simple and effective for new resources but has the sharp edge that automation, IaC pipelines, or a team's existing deployment templates that do not set this tag will start failing outright the moment the policy is assigned, potentially in the middle of an unrelated urgent deployment. A more commonly recommended pattern for tags specifically is a Modify effect with DeployIfNotExists-style remediation: instead of blocking the create, automatically add a default cost-center value (e.g. unassigned) to any resource missing the tag, paired with a separate Audit policy or a Workbook that surfaces every resource still tagged unassigned for a human to correct later. This keeps deployments from breaking while still guaranteeing the tag is always present in some form.
Remediating existing non-compliant resources without breaking production. For a tag fix, this is low-risk: a Policy remediation task (built on the Modify or DeployIfNotExists effect) applies to the existing non-compliant population directly and adding or correcting a tag has no runtime impact on the resource. For something with real blast radius (how much could break, and how far the damage could spread, if the change goes wrong), such as retroactively enforcing a VM SKU restriction on VMs that are already running a now-disallowed SKU, do not remediate automatically: Deny going forward stops the bleeding, and existing non-compliant VMs go into a planned, scheduled migration (resize during a maintenance window, tested against the specific workload) rather than an automated remediation task that could resize or replace a running production VM without warning. The general rule: remediate automatically when the fix is purely metadata or purely additive (a missing diagnostic setting, a missing tag); remediate manually and deliberately when the fix changes the resource's actual runtime configuration or availability.
Permitted VM SKUs and encryption at rest. A SKU allowlist is typically a built-in policy (Allowed virtual machine SKUs) parameterized with your list, assigned with Deny. For encryption at rest, note that many Azure storage/database services encrypt at rest by default with platform-managed keys and cannot be turned off, so a policy here is more often about enforcing customer-managed keys where that's a compliance requirement, or catching a specific resource type where encryption is genuinely optional; check what the specific resource type's actual default is before writing a policy to enforce something that may already be non-optional.
Worked example
A subscription-scoped Deny policy definition targeting Microsoft.Compute/virtualMachines and Microsoft.Compute/virtualMachineScaleSets, condition not (field 'tags[cost-center]' exists 'true'), effect Deny, assigned to the subscription with an exemption for a specifically named "sandbox" resource group where short-lived experimentation is allowed to skip tagging. Rollout sequence to avoid breaking production: assign the same policy as Audit first for two weeks, review the compliance report to find every pipeline/team that would have been blocked, fix those pipelines' templates to include the tag, then flip the assignment to Deny. This order (Audit before Deny) is the single most reliable way to introduce a blocking policy without an unplanned outage caused by the policy itself.
Trade-offs and pitfalls
Deny policies assigned without an Audit-first rollout are the most common source of "why did my deployment suddenly start failing" incidents in a governed environment, especially when the policy is assigned at a management-group scope covering many subscriptions a central platform team does not operate day-to-day. DeployIfNotExists remediation runs under a system-assigned managed identity with specific role assignments granted at policy-assignment time; forgetting to grant that identity the roles its remediation template actually needs is a common, silent failure mode where the policy shows resources as "non-compliant, remediation in progress" indefinitely because the identity cannot actually perform the deployment. Finally, policies stack: a resource can be subject to policies assigned at the management group, the subscription, and the resource group simultaneously, and the most restrictive Deny anywhere in that chain wins, which means testing a new policy's effect at the exact scope and resource types it will actually apply to, not just in isolation, matters before a wide rollout.
Write a Terraform (HCL) snippet that creates a user-assigned Managed Identity in Azure, assigns it the 'Reader' built-in role on a specific resource group, and configures the remote backend to use Azure Storage with locking enabled. Assume azurerm provider; show the key resource blocks and backend configuration.
Sample Answer
Direct answer
Below is Terraform HCL that creates a user-assigned Managed Identity (an Azure-managed credential an app or automation can use without storing a secret), grants it the built-in Reader role on a specific resource group, and configures the azurerm remote backend with locking enabled. It validates cleanly (terraform init -backend=false && terraform validate against provider azurerm ~> 3.110, resolved to 3.117.1).
Structured elaboration: approach
A user-assigned identity is a standalone Azure resource (unlike a system-assigned identity, which is born and dies with the resource it is attached to), so it is declared directly with azurerm_user_assigned_identity. The role assignment is a separate resource, azurerm_role_assignment, scoped to the resource group's ARM ID with role_definition_name = "Reader" (Terraform resolves the built-in role's GUID from the name). Locking on the azurerm backend is not a Terraform-side setting you toggle: it is inherent to how the backend works. Terraform takes out a blob lease on the state blob for the duration of every plan/apply, so a second concurrent run against the same state key blocks until the lease is released; this is why the backend configuration below matters (the resource group, storage account, and blob key that anchor that shared, leased state file), not a separate locking = true argument, which does not exist for this backend.
terraform {
required_version = ">= 1.5.0"
required_providers {
azurerm = {
source = "hashicorp/azurerm"
version = "~> 3.110"
}
}
backend "azurerm" {
resource_group_name = "rg-tfstate"
storage_account_name = "sttfstateprod001"
container_name = "tfstate"
key = "networking/prod.terraform.tfstate"
use_azuread_auth = true
}
}
provider "azurerm" {
features {}
}
variable "resource_group_name" {
type = string
description = "Existing resource group the identity is granted Reader on."
}
data "azurerm_resource_group" "target" {
name = var.resource_group_name
}
resource "azurerm_user_assigned_identity" "this" {
name = "id-reader-${var.resource_group_name}"
resource_group_name = data.azurerm_resource_group.target.name
location = data.azurerm_resource_group.target.location
}
resource "azurerm_role_assignment" "reader" {
scope = data.azurerm_resource_group.target.id
role_definition_name = "Reader"
principal_id = azurerm_user_assigned_identity.this.principal_id
}
output "identity_id" {
value = azurerm_user_assigned_identity.this.id
}
output "identity_client_id" {
value = azurerm_user_assigned_identity.this.client_id
}
output "identity_principal_id" {
value = azurerm_user_assigned_identity.this.principal_id
}
Verified: terraform validate returns Success! The configuration is valid. use_azuread_auth was checked against HashiCorp's current azurerm backend documentation: it is a real, current attribute that authenticates the backend's data-plane calls (reading/writing the state blob) via Microsoft Entra ID (Azure AD) instead of a storage account access key, which HashiCorp's docs now recommend over key-based auth for new configurations.
Worked example: what "locking" actually protects against
If two pipeline runs (say, a scheduled drift-check and a developer's manual terraform apply) target this same backend key at the same moment, the second one to reach init/plan fails fast with a state blob is already locked error identifying the lock ID and who holds it, rather than both writing to state concurrently and corrupting it. This is why the key in the backend block should be unique per environment (dev.terraform.tfstate vs prod.terraform.tfstate): sharing one key across environments would mean a dev apply and a prod apply contend for the same lock even though they touch unrelated infrastructure.
Trade-offs and pitfalls
Key points. Using role_definition_name = "Reader" rather than hardcoding the role's GUID keeps the code readable and is resolved consistently by the provider; principal_id (the identity's Azure AD object ID), not client_id, is what an azurerm_role_assignment needs, and mixing those two up is a common, confusing-to-debug mistake because both are GUIDs and Terraform will not catch the swap at plan time, only Azure will reject it (or worse, silently assign the role to the wrong nonexistent-looking principal in some cases) at apply time.
Edge cases. Role assignments in Azure are eventually consistent: immediately after terraform apply finishes, a check of the identity's effective permissions can briefly show Reader as not yet in effect (typically within seconds, occasionally longer), which is a common source of "the pipeline says it worked but the next step got a 403" flakiness; a short retry/backoff on the very next dependent operation is cheaper than debugging a phantom failure. If var.resource_group_name does not exist, data.azurerm_resource_group.target fails at plan time with a clear "resource group not found" error rather than creating one, which is intentional here (the exercise assumes an existing resource group) but worth calling out since some data sources fail far less clearly.
Pitfall to avoid. Never grant this identity a broader role "just in case" because it is convenient during testing: an over-scoped Reader-turned-Contributor identity attached to automation is a standing privilege-escalation path if that automation is ever compromised, and least-privilege for identities used by pipelines matters more, not less, than for human accounts, because nobody reviews a service identity's access as often as they review a person's.
Unlock Full Question Bank
Get access to all Microsoft Azure Services and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.