Microsoft Azure Services and Architecture Questions
Microsoft Azure's core service catalog and architectural patterns: Virtual Machines, managed Kubernetes (AKS), App Service and Azure Functions, Storage accounts and managed disks, Azure SQL and Cosmos DB, VNets with hybrid connectivity and global load balancing, Microsoft Entra ID and RBAC, and Key Vault secrets and encryption. Covers Azure service selection, infrastructure as code (ARM, Bicep, Terraform), observability with Azure Monitor and Kusto queries, cost governance and Azure Policy, the Azure Well-Architected design principles, and hybrid management via Azure Arc, common in enterprise Azure estates. For provider-agnostic trade-offs, see the cross-cloud entries.
You inherited a large monolithic Terraform codebase that manages hundreds of Azure resources across multiple subscriptions. Propose a migration plan to modularize the code into reusable modules, implement remote state with locking, adopt a module registry, and integrate with CI/CD. Describe steps to avoid accidental resource recreation during refactoring and how to stage the migration with minimal disruption.
Sample Answer
Direct answer
Modularize by extracting behind a stable interface, not by refactoring in place: write each new module so its planned resource addresses match what Terraform already tracks in state for that resource, then use terraform state mv to relocate the existing resources into the new module's state paths without touching the real infrastructure at all. The single rule that prevents accidental recreation is checking terraform plan shows zero create/destroy actions after every state move, only address changes; if a plan ever shows a resource being destroyed and recreated during this migration, stop, because that's real infrastructure being torn down, not a refactor.
Migration plan
- Stop the bleeding first: remote state with locking, before touching modularization. If the monolith isn't already on remote state (Azure Storage backend with a lease-based lock), migrate to that first, in isolation, as its own change. This is the highest-leverage single step, since it's what makes safe concurrent work (multiple engineers, or a later CI pipeline) possible at all, and it carries zero resource-recreation risk since it only moves state storage location, not resource addresses.
- Establish a module boundary by domain, not by resource type. Group resources the way the organization actually reasons about them (e.g., "the networking module," "the AKS module," "the shared Key Vault module"), matching subscription/resource-group boundaries where they already exist, rather than an abstract "one module per resource type" scheme that doesn't map to how anyone will actually use or reason about it later.
- For each module, in order of lowest risk first:
- Write the new module with resource blocks that, when planned, will produce the exact same resource addresses and arguments Terraform already has in state (same names, same
for_each/countkeys if used), which requires reading the current state (terraform state list,terraform show) as the source of truth, not the old monolithic.tffiles, since those can have drifted from what's actually deployed. - Run
terraform state mv 'module_root.resource_type.name' 'module.new_module_name.resource_type.name'for each resource (or-dry-runfirst, always) to relocate it in state. - Immediately run
terraform planand confirm it shows no changes (or only genuinely intended changes, kept as a separate, explicit step, never bundled into the same plan as a state move). Any create/destroy pair here means the new module's resource definition doesn't actually match what's in state, and needs fixing before proceeding, not overriding with-targetor forcing through. - Only after the plan is clean, commit the new module and the updated root configuration together, so the repository's history reflects "this was purely a state reorganization" as one atomic, reviewable change.
- Write the new module with resource blocks that, when planned, will produce the exact same resource addresses and arguments Terraform already has in state (same names, same
- Adopt a module registry once a module is used more than once, not before; a private registry (Terraform Cloud/Enterprise's registry, or a Git-tag-based source reference as a lighter-weight interim step) matters once you're versioning a module independently of the root configuration that consumes it, which is premature for a module still being extracted for the first time.
- Integrate with CI/CD incrementally, module by module, running
planon every PR against the affected module's state, and requiring a human review of the plan output (not just an automated apply) for the first several migrations, until the team has confidence the state-move process is being applied correctly and consistently.
Worked example: avoiding accidental recreation
The concrete failure this guards against: an engineer extracts a virtual network resource into a new networking module but writes the new resource block with a slightly different for_each key expression than the original (say, keying by list index in the old code versus a named map key in the new module). Terraform sees this as a different resource identity entirely, and terraform plan will show it as destroy-then-create, not a rename, because from Terraform's perspective the address changed. The moment plan shows that, before ever running apply, is where the mistake is caught: the fix is either to match the original for_each/count scheme exactly in the new module (preferred, since changing key schemes on live infrastructure is a separate, deliberate decision to make later, not a side effect of a refactor), or to add explicit moved blocks (Terraform's declarative alternative to imperative state mv, which documents the old-to-new address mapping directly in code and is reviewable in a pull request) if the key scheme genuinely needs to change as part of this work.
Trade-offs and pitfalls
The most damaging mistake is treating "the code looks cleaner now" as sufficient validation and skipping the zero-diff plan check before applying; Terraform will happily destroy and recreate a resource if the address or a force-new argument changed, often with no warning beyond the plan output itself, which is exactly why that plan output has to be read carefully every single time, not skimmed. The second common mistake is doing the full migration in one large PR across every module at once, which makes a mistake in one module hard to isolate and revert independently from the rest; staging the migration module by module, each with its own clean-plan checkpoint, keeps blast radius small and keeps rollback (reverting one module's state moves) actually feasible.
Design hybrid connectivity between an on-prem datacenter and Azure across two regions with stringent latency (<50ms) and high availability requirements. Compare ExpressRoute (private circuits) vs VPN Gateway (IPsec) for this scenario, explain BGP peering and route advertisement, failover strategies, and how to avoid asymmetric routing or single points of failure.
Sample Answer
Direct answer
Make ExpressRoute (Microsoft's private, non-internet circuit into Azure) the primary path for its predictable sub-50ms latency and service-level agreement (SLA), with a Site-to-Site VPN Gateway (an IPsec-encrypted tunnel over the public internet) as an active backup, both duplicated across two regions and two physically diverse circuits so no single carrier, router, or Azure gateway is a single point of failure. Use Border Gateway Protocol (BGP, the routing protocol that lets networks exchange reachability information) on every link, with route preference tuned so ExpressRoute always wins when it is healthy and the VPN takes over automatically, symmetrically, when it is not.
Structured elaboration
Why ExpressRoute over VPN Gateway here. ExpressRoute is a dedicated private circuit with predictable latency and a formal SLA, which is what a sub-50ms, high-availability requirement actually needs. A Site-to-Site VPN runs over the public internet (even if encrypted end to end), so its latency and jitter vary with internet conditions and it carries no comparable SLA. VPN's advantages are that it provisions in hours, not weeks, and costs far less, which is exactly why it is the right role for a backup path, not the primary.
Topology. Provision two ExpressRoute circuits terminating at physically diverse peering locations (different buildings, ideally different carriers), each connected into a regional Azure Virtual WAN hub or ExpressRoute gateway. Deploy active-active VPN gateways in both regions as the failover path. On-premises, terminate on at least two edge routers, each with its own physical uplink, both speaking BGP.
BGP peering and route advertisement. Every peering (ExpressRoute private peering and the VPN's BGP session) advertises prefixes in both directions: on-premises advertises its internal subnets to Azure, and Azure advertises the connected VNet prefixes back. Control which path wins with standard BGP attributes: a higher local preference or a shorter advertised AS-path on the ExpressRoute session makes routers prefer it over the VPN under normal conditions; when ExpressRoute fails, BGP simply stops hearing those routes and the VPN's routes become the best path automatically, with no manual re-routing.
Avoiding asymmetric routing and single points of failure. Asymmetric routing (traffic leaving over one path and returning over a different one) causes stateful devices like firewalls to drop return traffic they never saw the initial packet for. Prevent it by keeping route preference consistent on both ends: if on-premises prefers ExpressRoute outbound, Azure must also prefer the matching ExpressRoute path inbound, using the same BGP attributes on both sides rather than one side using static routes and the other using BGP. Redundant circuits on physically diverse paths, redundant gateways in active-active mode, and Bidirectional Forwarding Detection (BFD, a fast link-failure detection protocol that reacts in milliseconds instead of BGP's default tens of seconds) close the remaining single-point-of-failure gaps.
flowchart TD
OP1[On-prem router A] -- ExpressRoute circuit 1 --> AZ1[Azure region 1 hub]
OP2[On-prem router B] -- ExpressRoute circuit 2 --> AZ2[Azure region 2 hub]
OP1 -. VPN backup .-> AZ2
OP2 -. VPN backup .-> AZ1
AZ1 <--> AZ2
Validation before go-live. Integrate this into the hub-and-spoke design already in place for the workload's VNets: the ExpressRoute and VPN gateways terminate in the hub, and spoke VNets reach on-premises only through hub peering, so a single set of gateways serves every spoke. Before declaring the link production-ready, run a controlled failover test: withdraw the ExpressRoute route deliberately (not by physically cutting a cable) and confirm the VPN path takes over within the BFD detection window, traffic resumes without a duplicate or dropped-session storm, and route tables on both ends converge to the expected state.
Worked example
A workload like an SAP application landscape (used here only as an illustrative example of a genuinely latency-sensitive enterprise workload, not a prescription) typically needs application-tier-to-database-tier round trips well under the latency budget of a single user transaction. If the physical distance between the on-premises datacenter and the nearer Azure region is roughly 500 km, the one-way propagation delay in fiber (light travels at roughly two-thirds the speed of light in glass, so about 200,000 km per second) is 500 divided by 200,000 seconds, or 2.5 milliseconds, and a round trip is 5 milliseconds. That leaves roughly 45 of the 50ms budget for switching, queuing, and application-level processing at each hop, which is exactly why the design targets a private, low-jitter circuit like ExpressRoute instead of the public internet's much less predictable delay: the physics is not the constraint, the variance is.
Trade-offs & pitfalls
The pitfall this scenario is built to catch is treating BGP local preference as something you tune once and forget: if the two ends of the connection are configured by different teams (network engineering on-premises, cloud engineering in Azure) and they drift out of sync, asymmetric routing reappears silently and only shows up as intermittent, hard-to-reproduce connection resets. The second pitfall is skipping the failover rehearsal because the design "should" work; BGP convergence behavior under a real link failure is exactly the kind of thing that differs from the whiteboard, so test it before it is load-bearing.
Design a secure key rotation and secret lifecycle for a high-compliance environment (e.g., PCI-DSS) using Azure Key Vault, Managed HSM, Azure AD, and automation. Include rotation frequency, emergency rotation plans, secret versioning, access revocation, auditing, role separation, and rollback mechanisms to meet audit requirements.
Sample Answer
Direct answer
Split the design into two independent problems that only merge at policy level: which vault tier holds which secret (Key Vault for general secrets and shared-tenant HSM-backed keys, Managed HSM for keys that must never leave a single-tenant, customer-controlled cryptographic boundary), and what governs when a key or secret changes (a routine rotation schedule tied to a defined cryptoperiod, versus an emergency path triggered by suspected compromise). PCI-DSS (Payment Card Industry Data Security Standard) does not mandate a universal rotation calendar; it mandates that you define and document a cryptoperiod per key based on its strength, algorithm, and exposure, and that you can prove role separation and an audit trail for every key-management action. Key Vault's built-in role-based access control (RBAC), version history, and Managed HSM's customer-held security domain give you exactly that natively, if you design around them instead of bolting audit on afterward.
Structured elaboration
Vault tiering
| Asset | Where it lives | Why |
|---|---|---|
| Cardholder-data-scope encryption keys (transparent data encryption root keys, tokenization keys) | Managed HSM | Single-tenant, FIPS (Federal Information Processing Standard) 140-3 Level 3 validated hardware, with a customer-controlled security domain that even Microsoft cannot access; its local RBAC model lets designated HSM administrators lock out even subscription- or management-group-level administrators, the strongest form of role separation Azure offers |
| Application secrets, connection strings, API keys | Key Vault (standard or premium tier) | Not cryptographic key material in the HSM sense; needs versioning and rotation-policy support, not a dedicated single-tenant HSM |
| TLS certificates | Key Vault Certificates | Built-in integration with managed certificate authorities for automated renewal ahead of expiry |
Rotation frequency and cryptoperiod
PCI-DSS Requirement 3.7 is the one that requires documented key-management processes and procedures covering the full key lifecycle, generation, distribution, storage, cryptoperiod-based changes, retirement, and custodian acknowledgment, with sub-requirement 3.7.6 specifically covering manual cleartext key operations and requiring split knowledge and dual control wherever a human ever touches key material in the clear. Requirement 3.6 is the companion requirement covering how the keys themselves must be secured, for example stored encrypted, inside a secure cryptographic device, or as full-length key components, with cleartext key access restricted to the fewest custodians necessary. Neither requirement hands you a fixed number of days; you set the cryptoperiod based on key strength, algorithm, and exposure, following NIST (National Institute of Standards and Technology) Special Publication 800-57 guidance, and you write that reasoning down.
In practice, for a PCI-DSS environment: rotate data-encryption keys (DEKs) on a defined cadence, commonly quarterly to annually, using envelope encryption, meaning the DEK is wrapped by a key-encryption key (KEK) held in the vault. Rotating the KEK does not require re-encrypting the underlying cardholder data, only re-wrapping the DEK, which is what makes frequent KEK rotation operationally cheap. Rotate application secrets and API keys more aggressively (30 to 90 days is a common baseline) since they are lower-value, higher-exposure targets and cheaper to rotate. Key Vault supports an automated rotation policy natively for keys, setting a rotation time and emitting an Event Grid notification, rather than requiring you to build a scheduler.
Certificates should auto-renew at a fixed percentage of their remaining lifetime (Key Vault's integrated-CA renewal typically triggers well before expiry) so rotation is never a manual fire drill.
Emergency rotation
Trigger conditions: confirmed or suspected key compromise, an offboarding event for someone who held cleartext access, or an audit finding. The first action is always to disable the affected key or secret version, not delete it: disabling stops new use immediately while preserving the version, so anything that legitimately still needs to decrypt old data with the old key version can still do so under access control, whereas deletion, even with soft-delete and purge protection, is a recovery operation, not a revocation one.
Whether the emergency rotation is seamless depends on a design decision made long before the emergency: applications that reference a secret by its versionless URI pick up a new version automatically on the next fetch, while applications pinned to a specific version ID require a redeploy or config change to move. Anything in the PCI-DSS cardholder data environment (CDE, the network segment that stores, processes, or transmits cardholder data) that might need emergency rotation should default to versionless references specifically so an emergency does not turn into an outage.
For any manual, cleartext key-management step, which in practice comes up specifically during a Managed HSM security-domain download or restore (the encrypted backup artifact that lets you recover the HSM in a disaster), PCI-DSS 3.7.6's split-knowledge and dual-control requirement maps directly onto Managed HSM's own quorum design: the security domain is protected using Shamir's Secret Sharing among a minimum of 3, and up to 10, customer-held RSA key pairs, so no single custodian can reconstruct it alone. Use that native feature as your compliance control rather than inventing a separate manual process, since it is the same requirement Managed HSM was built to satisfy.
Versioning, revocation, and rollback
Both Key Vault and Managed HSM keep an automatic, immutable version history for every key and secret; this is what "rollback" actually means in this context, since you cannot literally restore a value, only re-point consumers at, or re-enable, a prior version. Design the emergency-rotation runbook to include the rollback path explicitly: if a new key version breaks a downstream integration, the fastest safe recovery is re-enabling the previous version, not disabling the vault or the application.
Access revocation is a role-assignment change, not a secret change, whenever possible: prefer Managed Identities (an Entra ID, Microsoft's cloud identity and access management service formerly called Azure AD, identity tied to an Azure resource itself, with no long-lived credential for an application to hold or leak) over service principals with client secrets for anything talking to the vault, because there is then no long-lived secret to rotate or revoke on the application side at all, only a role assignment to remove.
Role separation in practice: a small, break-glass-only Key Vault or HSM Administrator role that manages the vault and its access policy, a Crypto Officer role that can create and rotate keys, and a Crypto User role, assigned to applications via managed identity, that can only use keys (encrypt, decrypt, wrap, unwrap) and can never export or manage them. Auditing every operation under these roles through Azure Monitor and a Log Analytics workspace, retained per your PCI-DSS retention requirement, is what turns "we have RBAC" into "we can prove who did what, when" for an assessor.
flowchart TD
A[Compromise suspected: alert, audit finding, or offboarding] --> B[Disable affected key/secret version]
B --> C{Manual cleartext operation involved?}
C -->|Yes, e.g. HSM security domain restore| D[Split-knowledge custodians act under quorum]
C -->|No| E[Automated rotation policy issues new version]
D --> F[Re-point consumers via versionless reference or redeploy]
E --> F
F --> G[Confirm old version unusable: disabled and RBAC-blocked]
G --> H[Write audit trail: who, what, when, approval ticket]
Worked example
A retail platform's tokenization key, used to convert primary account numbers into tokens before storage, lives in Managed HSM. An internal audit flags that a departing contractor held Crypto Officer access for three weeks longer than their engagement required. The emergency path: disabling the key itself is not needed, since there is no evidence of use, but the contractor's role assignment is revoked immediately (a role-assignment removal, seconds, no key operation needed). Because the incident review cannot fully rule out access, the key is rotated as a precaution: a new key version is generated, the application, which references the key by its versionless URI, picks up the new version on its next token-generation call with no redeploy, and the previous version is disabled, not deleted, so tokens already generated with it and needing detokenization later still resolve correctly, under the now-revoked contractor's removed access. The audit trail records the role-removal timestamp, the rotation timestamp, and the ticket number, exactly the who, when, what record a PCI-DSS assessor asks for.
Trade-offs & pitfalls
Rotating a KEK is cheap under envelope encryption; rotating a DEK directly is not, because it requires re-encrypting the underlying data. Confusing the two in a rotation policy is the single most common design mistake: teams write "rotate keys every 90 days" without specifying which layer, then discover the policy implies re-encrypting a multi-terabyte cardholder data store every quarter.
Pinning applications to a specific key or secret version for change-control reasons, a legitimate practice for high-risk changes, directly trades away emergency-rotation speed. Decide deliberately, per secret, which property matters more.
Managed HSM's single-tenant isolation and local RBAC are real security wins but come at a real cost and operational premium over shared Key Vault tiers. Reserve it for cardholder-data-scope keys specifically, not as a default for every secret in the environment, or the cost and the quorum-management overhead becomes its own operational burden.
A rollback plan that only covers "re-enable the previous version" is incomplete if the previous version was already disabled for a reason, for instance because it was the thing being rotated away from after a suspected compromise. The runbook needs an explicit rule for when rollback is not the safe choice, not just a default revert action.
Describe how to enable encryption at rest and encryption in transit for Azure Storage, Azure SQL Database, and Azure VM disks. Contrast platform-managed keys (Microsoft-managed) versus customer-managed keys (BYOK) stored in Key Vault / Managed HSM, and discuss when a customer might require BYOK or a Managed HSM.
Sample Answer
Direct answer
Azure Storage, Azure SQL Database, and Azure VM disks all encrypt data at rest by default using Microsoft-managed AES-256 keys and enforce encryption in transit through TLS 1.2 or higher, with no application changes required. The real design decision most teams make is whether to layer a customer-managed key (CMK, also called bring-your-own-key or BYOK) on top, held in Azure Key Vault or Azure Key Vault Managed HSM (a dedicated hardware security module service), which becomes a requirement rather than a preference under most regulatory or data-sovereignty mandates.
Structured elaboration
| Service | At rest (default) | In transit | Customer-managed key option |
|---|---|---|---|
| Azure Storage (Blob/Files/Queue/Table) | Storage Service Encryption, AES-256, cannot be disabled, applied transparently below the storage stack | HTTPS enforced by "secure transfer required"; SMB 3.0+ encryption for Azure Files | Wrap the account's data encryption key with a key in Key Vault or Managed HSM; disabling that key makes the account unreadable |
| Azure SQL Database | Transparent Data Encryption (TDE), on by default, encrypts the database, logs, and backups at the page level | TLS enforced by default for client connections | TDE with a customer-managed key: the database encryption key is protected by your asymmetric key in Key Vault or Managed HSM instead of the Microsoft-managed service key |
| Azure VM disks | Server-side encryption (SSE) on managed disks, on by default, applied at the storage-cluster level | Encryption applied to data moving between compute and disk storage | SSE with customer-managed keys, same Key Vault/Managed HSM wrapping pattern; optionally layer Azure Disk Encryption (BitLocker on Windows, dm-crypt on Linux) inside the guest OS for an additional layer |
Azure SQL Database also supports Always Encrypted, a separate, complementary control that encrypts specific sensitive columns client-side, so the data stays encrypted even from a database administrator with full access to the server, which is a different property than TDE's page-level, at-rest-only protection.
Platform-managed versus customer-managed keys. With platform-managed (Microsoft-managed) keys, Microsoft creates, rotates, and safeguards the key; there is zero operational burden, and it satisfies baseline compliance for most SOC 2 (an independent audit of how a service provider protects customer data) or ISO 27001 (an international standard for information security management) needs, but you cannot independently revoke access or prove exclusive control of the key to an auditor. With customer-managed keys, you control the key's creation, rotation cadence, and revocation; disabling or deleting the key in Key Vault is a real cryptographic kill switch that also cuts off Microsoft's own ability to read your data. That control comes with real operational responsibility: if the key becomes unavailable, so does your data, so the key needs the same high-availability discipline (soft delete, purge protection, an offline backup of the key material) as the data it protects.
Managed HSM specifically. Azure Key Vault Premium and Azure Key Vault Managed HSM are both FIPS 140-3 (the U.S. government standard for validating cryptographic hardware and software) Level 3 validated today (the Standard tier, by contrast, uses software-protected keys with no dedicated hardware at all). The real distinction between Premium and Managed HSM is not the compliance level, it is tenancy: Key Vault Premium runs your HSM-protected keys on a shared, multi-tenant HSM pool, while Managed HSM gives you a single-tenant HSM pool dedicated entirely to your organization, with its own security domain and its own administrative control plane (the management interface used to configure and administer the resource, distinct from the data plane that handles actual key operations), separate from Key Vault's role-based access control (RBAC) model. Choose Managed HSM when a regulator or contract requires a dedicated, non-shared hardware boundary, not merely "an HSM," or when you need that independent administrative control plane; otherwise Key Vault Premium already gives you HSM-backed keys without Managed HSM's higher cost floor.
Worked example
A healthcare ISV storing patient records in Azure SQL Database and Blob Storage needs to show, for a HIPAA (the U.S. law governing protection of patient health information) business-associate audit, that it can revoke Microsoft's technical ability to decrypt its data within a bounded time. With platform-managed keys this is not possible, since Microsoft holds the only key. With a customer-managed key in Key Vault, disabling or deleting that key makes the data unreadable once every dependent service's cached key-validity check next expires (each service's own caching interval, not an instantaneous cutoff, so the exact bound should be confirmed against current Microsoft Learn documentation for Storage and SQL specifically before it is written into a compliance commitment). If the auditor's requirement is a dedicated, non-shared hardware boundary rather than Key Vault Premium's shared HSM pool, the ISV moves that same key into Managed HSM instead, which changes the tenancy model without changing anything about how Storage or SQL consume the key.
Trade-offs and pitfalls
Treating a customer-managed key as a free security upgrade rather than an operational commitment is the most common mistake: an inaccessible or accidentally deleted key takes production data down with it. Assuming Managed HSM is "just a bigger Key Vault" misses that it is a separate resource type with its own RBAC model and a materially higher cost floor, so most teams should default to Key Vault Premium and only move to Managed HSM when a specific requirement forces the single-tenant boundary. Finally, conflating "encrypted in transit" with "authenticated" is a real gap: TLS alone does not stop a valid but compromised credential from reading the data once it's decrypted server-side.
Compare Terraform, ARM templates, and Bicep for managing enterprise Azure infrastructure across multiple teams and environments. Discuss module reuse, state management, drift detection, policy enforcement (Azure Policy), testing strategies, secret handling, and CI/CD integration. Which would you choose for large, cross-team deployments and why?
Sample Answer
Direct answer
For large, cross-team Azure deployments, Terraform is the default recommendation: it is provider-agnostic (useful the moment the enterprise touches even one non-Azure system), has the most mature module registry and testing ecosystem, and its state model gives explicit, reviewable drift detection that Azure Resource Manager (ARM) templates and Bicep do not have an equivalent for. Choose Bicep instead when the organization is genuinely Azure-only and values Bicep's tighter native integration and simpler onboarding over Terraform's ecosystem breadth. Avoid authoring new infrastructure directly in raw ARM JSON; treat it as Bicep's compiled output, not a hand-written format.
Structured elaboration
| Concern | Terraform | Bicep | ARM templates |
|---|---|---|---|
| Module reuse | Versioned module registry, explicit input/output contracts, easy cross-team sharing | Modules and parameter files work well for Azure-specific constructs, but the ecosystem for sharing them across teams is less mature than Terraform's registry | Nested and linked templates exist but are verbose and error-prone to reuse at scale |
| State management | Explicit state file in a remote backend (Azure Storage with blob-lease locking, or Terraform Cloud), giving a reviewable plan before every apply | No separate state file: Azure itself is the source of truth, and deployments are idempotent, but there is no local plan artifact with the same preview fidelity as a Terraform plan | Same as Bicep: no separate state model |
| Drift detection | terraform plan shows exactly what changed against the last known state, independent of what Azure Policy separately reports | what-if deployments compare the template against live resources; steadily improving but historically less precise than a Terraform plan | what-if exists but the experience is clunkier than Bicep's |
| Policy enforcement | Azure Policy evaluation happens server-side regardless of which tool deployed the resource; Terraform can also deploy the policy definitions themselves as code | Native, first-class integration since both are Microsoft products | Native, same as Bicep |
| Testing | Static analysis (tflint, a Terraform-specific linter that catches syntax and best-practice issues before a plan even runs), policy-as-code scanning (checkov and terrascan, tools that check your Terraform against security and compliance rules before anything deploys), and real integration tests (Terratest, a Go testing framework that actually deploys the code to a real, throwaway environment and asserts the result works) against ephemeral environments | ARM Template Test Toolkit and the Bicep linter, plus what-if as a dry run; true integration testing usually means standing up real resources and validating with scripts | Same testing story as Bicep, with a clunkier authoring experience |
| Secret handling | Never store secrets in state or variables in plain text; reference Azure Key Vault through the Key Vault data source, and encrypt the remote state backend itself | Reference Key Vault secrets through secure parameters at deploy time, backed by a managed identity, never a plaintext parameter file | Same pattern as Bicep, but more verbose to express |
| CI/CD integration | Well-established plan-then-apply pattern with a human approval gate between them; Terraform Cloud or Enterprise adds policy checks and run history for larger orgs | Integrates cleanly into Azure DevOps or GitHub Actions using az deployment what-if followed by the actual deployment command | Same pipeline shape as Bicep |
Azure DevOps versus GitHub Actions as the pipeline engine. Both support the same plan-review-apply pattern for either tool; the practical decision usually comes down to where the rest of the organization's source control and existing pipelines already live, not a technical limitation of either tool. Azure DevOps has slightly deeper native integration with Azure service connections and environments-with-approvals out of the box; GitHub Actions has a larger third-party action ecosystem and is the natural choice if the code already lives in GitHub. For a regulated enterprise, the more important factor than which engine is whether the pipeline authenticates using workload identity federation (short-lived, no stored secret) rather than a long-lived service principal credential, since that is the control an auditor will actually ask about.
Regulated-enterprise secret handling. In a regulated environment, go further than "use Key Vault": require that CI/CD pipelines authenticate to Azure using workload identity federation instead of a stored client secret, require that Terraform's remote state backend itself is encrypted with a customer-managed key and accessed only through role-based access control (not a shared access signature token that anyone with the string can use), and require an audit log of every plan and apply, not just every Key Vault secret access, since the state file and the pipeline execution are both part of the sensitive surface, not only the secrets referenced inside them.
Worked example
Consider an enterprise with 15 application teams and one platform team, currently authoring ARM JSON templates that have grown to thousands of lines per environment and are rarely touched without introducing a typo. Migrating to Terraform: the platform team publishes versioned modules (networking, a standard app-service pattern, a standard database pattern) to an internal module registry; each application team's pipeline runs terraform plan on every pull request, posts the plan as a reviewable comment, and requires a second approver before terraform apply runs against production. Drift is caught the next time anyone runs a plan, since the state file records exactly what Terraform believes exists, and a manual portal change shows up as an unexpected diff rather than going unnoticed until it causes an incident.
Trade-offs & pitfalls
The pitfall this question is really testing is picking a tool based on which one the platform team already knows, rather than the organization's actual shape: a genuinely single-cloud, Azure-only organization with a small platform team may get more value from Bicep's simpler onboarding and native tooling than from standing up and operating Terraform's remote state backend and module registry. The second is treating Azure Policy as optional once an IaC tool is chosen; policy evaluation happens server-side no matter which tool deploys the resource, so skipping policy-as-code review in the pipeline just means violations are caught after deployment instead of before it.
Unlock Full Question Bank
Get access to all Microsoft Azure Services and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.