Microsoft Azure Services and Architecture Questions
Microsoft Azure's core service catalog and architectural patterns: Virtual Machines, managed Kubernetes (AKS), App Service and Azure Functions, Storage accounts and managed disks, Azure SQL and Cosmos DB, VNets with hybrid connectivity and global load balancing, Microsoft Entra ID and RBAC, and Key Vault secrets and encryption. Covers Azure service selection, infrastructure as code (ARM, Bicep, Terraform), observability with Azure Monitor and Kusto queries, cost governance and Azure Policy, the Azure Well-Architected design principles, and hybrid management via Azure Arc, common in enterprise Azure estates. For provider-agnostic trade-offs, see the cross-cloud entries.
Design an enterprise governance model for 100 Azure subscriptions across multiple regions. Define the management group hierarchy, policies (Azure Policy) to restrict resource creation, use of Blueprints or Terraform modules for standardization, a tagging strategy, and how to migrate existing ungoverned subscriptions into this model with minimal disruption.
Sample Answer
Direct answer
Adopt an Azure landing-zone management-group hierarchy: Platform (Identity, Connectivity, and Management subscriptions run by the central platform team), Landing Zones (pre-configured, governed subscription environments where actual workloads run) split by workload type (Corp for internal-facing, Online for internet-facing), and Sandbox for anything not yet governed. Enforce standardization with versioned Terraform or Bicep modules and Azure Policy assignments at each management-group level rather than Azure Blueprints, which Microsoft has been retiring since mid-2026 in favor of Template Specs and Deployment Stacks. Enforce a mandatory tagging taxonomy through policy, and migrate ungoverned subscriptions by moving them into Sandbox first, where guardrails run in audit mode, and progressively tightening policy as each one is remediated, never by moving all 100 into strict enforcement at once.
Structured elaboration
Management group hierarchy.
flowchart TD
Root[Tenant root group] --> Platform[Platform management group]
Root --> LZ[Landing zones management group]
Root --> Sandbox[Sandbox management group]
Platform --> IdentitySub[Identity subscription]
Platform --> ConnSub[Connectivity subscription]
Platform --> MgmtSub[Management subscription]
LZ --> Corp[Corp landing zones]
LZ --> Online[Online landing zones]
Corp --> DevSub[Dev subscriptions]
Corp --> ProdSub[Prod subscriptions]
Sandbox --> UngovSub[Migrating ungoverned subscriptions]
At 100 subscriptions, this split is what keeps policy assignment tractable: a policy is assigned once per management group and applies to every current and future subscription placed under it, rather than to 100 subscriptions individually.
Azure Policy as the standardization mechanism. Assign policies at the tenant root for anything universal, such as denying resources outside approved regions or requiring diagnostic settings to the central workspace. Assign platform-specific rules at the Platform group, such as restricting VPN or ExpressRoute (a private, dedicated network circuit connecting on-premises to Azure) gateway creation to the Connectivity subscription. Assign workload-appropriate rules at each Landing Zone sub-group, such as requiring a web application firewall in front of any public endpoint in Online but not in Corp, which has no public endpoint. Use Policy initiatives, grouped sets of related policies such as the built-in Azure Security Benchmark initiative, rather than assigning dozens of individual policies separately, since that is what stays maintainable at 100-subscription scale.
Standardization via versioned modules, not Blueprints. Azure Blueprints is in Microsoft's own phased retirement: new blueprint definitions could not be created as of mid-2026 per Microsoft Learn's blueprint-retirement guidance, with full retirement scheduled for January 2027. A governance model designed now should not adopt it even though it is a named option in the question. The direct replacement Microsoft recommends is Template Specs, a versioned, role-based access control (RBAC)-controlled ARM or Bicep template published as its own Azure resource, together with Deployment Stacks, which additionally track and can clean up everything a deployment created as a unit, or an equivalent versioned Terraform module in a private registry if the estate already standardized on Terraform. Either way, the underlying idea Blueprints offered, a versioned, reusable definition of what a compliant landing-zone subscription looks like, is delivered through currently-supported tooling instead.
Tagging strategy. A small mandatory set enforced by an Azure Policy deny effect at subscription creation, for example cost-center, environment, owner, and data-classification, with everything else left team-optional. Keeping the mandatory set small is what makes teams actually comply with it rather than routing around an over-specified schema, and it is what makes cost allocation and access review possible across 100 subscriptions without chasing every team individually.
Migrating ungoverned subscriptions with minimal disruption. Move a subscription into Sandbox first, which changes only Policy and RBAC inheritance, not the resources inside it, so the move itself causes zero downtime. Run the relevant policies in audit mode initially to see what full enforcement would flag without actually breaking anything. Remediate flagged non-compliant resources, or grant a time-boxed, documented exemption for anything that genuinely cannot be fixed immediately, then flip those policies to deny or modify and move the subscription into its correct permanent Landing Zone group. Doing this one subscription, or a small batch, at a time on a published schedule, rather than moving all 100 at once, keeps any single migration's blast radius (how much would be affected if something goes wrong) small enough to roll back.
Worked example
A newly discovered ungoverned subscription running a customer-facing web app moves into Sandbox. The tenant-wide "require diagnostic settings" policy, running in audit mode there, flags that none of its 12 resources send logs anywhere; over two weeks the team adds diagnostic settings to all 12. The "require approved region" policy, also in audit mode, flags 1 resource in a non-standard region; since it is a legacy dependency that cannot move without a separate project, it gets a documented, time-boxed policy exemption with an expiry date and a named owner rather than blocking the whole migration. Once the other 11 resources are compliant and the 1 exception is documented, the subscription moves from Sandbox into the Online Landing Zone group, where the same policies now run in deny mode, and the team's ongoing work is governed the same way as every other subscription in that group, without the web app ever going offline.
Trade-offs and pitfalls
Recommending Azure Blueprints in a plan being designed now, without checking its current status, ships stale advice on the exact tool the question names, since it is already mid-retirement. Running every policy straight to deny mode on day one of a migration, instead of audit-first, turns a governance rollout into an unplanned outage the moment a policy catches something nobody knew was non-compliant. And an over-broad mandatory tagging schema, a dozen required tags instead of a handful, predictably gets satisfied with placeholder values that make the data worse than having no tags at all, since nobody can tell a real value from a placeholder without auditing each one by hand.
Design a secure key rotation and secret lifecycle for a high-compliance environment (e.g., PCI-DSS) using Azure Key Vault, Managed HSM, Azure AD, and automation. Include rotation frequency, emergency rotation plans, secret versioning, access revocation, auditing, role separation, and rollback mechanisms to meet audit requirements.
Sample Answer
Direct answer
Split the design into two independent problems that only merge at policy level: which vault tier holds which secret (Key Vault for general secrets and shared-tenant HSM-backed keys, Managed HSM for keys that must never leave a single-tenant, customer-controlled cryptographic boundary), and what governs when a key or secret changes (a routine rotation schedule tied to a defined cryptoperiod, versus an emergency path triggered by suspected compromise). PCI-DSS (Payment Card Industry Data Security Standard) does not mandate a universal rotation calendar; it mandates that you define and document a cryptoperiod per key based on its strength, algorithm, and exposure, and that you can prove role separation and an audit trail for every key-management action. Key Vault's built-in role-based access control (RBAC), version history, and Managed HSM's customer-held security domain give you exactly that natively, if you design around them instead of bolting audit on afterward.
Structured elaboration
Vault tiering
| Asset | Where it lives | Why |
|---|---|---|
| Cardholder-data-scope encryption keys (transparent data encryption root keys, tokenization keys) | Managed HSM | Single-tenant, FIPS (Federal Information Processing Standard) 140-3 Level 3 validated hardware, with a customer-controlled security domain that even Microsoft cannot access; its local RBAC model lets designated HSM administrators lock out even subscription- or management-group-level administrators, the strongest form of role separation Azure offers |
| Application secrets, connection strings, API keys | Key Vault (standard or premium tier) | Not cryptographic key material in the HSM sense; needs versioning and rotation-policy support, not a dedicated single-tenant HSM |
| TLS certificates | Key Vault Certificates | Built-in integration with managed certificate authorities for automated renewal ahead of expiry |
Rotation frequency and cryptoperiod
PCI-DSS Requirement 3.7 is the one that requires documented key-management processes and procedures covering the full key lifecycle, generation, distribution, storage, cryptoperiod-based changes, retirement, and custodian acknowledgment, with sub-requirement 3.7.6 specifically covering manual cleartext key operations and requiring split knowledge and dual control wherever a human ever touches key material in the clear. Requirement 3.6 is the companion requirement covering how the keys themselves must be secured, for example stored encrypted, inside a secure cryptographic device, or as full-length key components, with cleartext key access restricted to the fewest custodians necessary. Neither requirement hands you a fixed number of days; you set the cryptoperiod based on key strength, algorithm, and exposure, following NIST (National Institute of Standards and Technology) Special Publication 800-57 guidance, and you write that reasoning down.
In practice, for a PCI-DSS environment: rotate data-encryption keys (DEKs) on a defined cadence, commonly quarterly to annually, using envelope encryption, meaning the DEK is wrapped by a key-encryption key (KEK) held in the vault. Rotating the KEK does not require re-encrypting the underlying cardholder data, only re-wrapping the DEK, which is what makes frequent KEK rotation operationally cheap. Rotate application secrets and API keys more aggressively (30 to 90 days is a common baseline) since they are lower-value, higher-exposure targets and cheaper to rotate. Key Vault supports an automated rotation policy natively for keys, setting a rotation time and emitting an Event Grid notification, rather than requiring you to build a scheduler.
Certificates should auto-renew at a fixed percentage of their remaining lifetime (Key Vault's integrated-CA renewal typically triggers well before expiry) so rotation is never a manual fire drill.
Emergency rotation
Trigger conditions: confirmed or suspected key compromise, an offboarding event for someone who held cleartext access, or an audit finding. The first action is always to disable the affected key or secret version, not delete it: disabling stops new use immediately while preserving the version, so anything that legitimately still needs to decrypt old data with the old key version can still do so under access control, whereas deletion, even with soft-delete and purge protection, is a recovery operation, not a revocation one.
Whether the emergency rotation is seamless depends on a design decision made long before the emergency: applications that reference a secret by its versionless URI pick up a new version automatically on the next fetch, while applications pinned to a specific version ID require a redeploy or config change to move. Anything in the PCI-DSS cardholder data environment (CDE, the network segment that stores, processes, or transmits cardholder data) that might need emergency rotation should default to versionless references specifically so an emergency does not turn into an outage.
For any manual, cleartext key-management step, which in practice comes up specifically during a Managed HSM security-domain download or restore (the encrypted backup artifact that lets you recover the HSM in a disaster), PCI-DSS 3.7.6's split-knowledge and dual-control requirement maps directly onto Managed HSM's own quorum design: the security domain is protected using Shamir's Secret Sharing among a minimum of 3, and up to 10, customer-held RSA key pairs, so no single custodian can reconstruct it alone. Use that native feature as your compliance control rather than inventing a separate manual process, since it is the same requirement Managed HSM was built to satisfy.
Versioning, revocation, and rollback
Both Key Vault and Managed HSM keep an automatic, immutable version history for every key and secret; this is what "rollback" actually means in this context, since you cannot literally restore a value, only re-point consumers at, or re-enable, a prior version. Design the emergency-rotation runbook to include the rollback path explicitly: if a new key version breaks a downstream integration, the fastest safe recovery is re-enabling the previous version, not disabling the vault or the application.
Access revocation is a role-assignment change, not a secret change, whenever possible: prefer Managed Identities (an Entra ID, Microsoft's cloud identity and access management service formerly called Azure AD, identity tied to an Azure resource itself, with no long-lived credential for an application to hold or leak) over service principals with client secrets for anything talking to the vault, because there is then no long-lived secret to rotate or revoke on the application side at all, only a role assignment to remove.
Role separation in practice: a small, break-glass-only Key Vault or HSM Administrator role that manages the vault and its access policy, a Crypto Officer role that can create and rotate keys, and a Crypto User role, assigned to applications via managed identity, that can only use keys (encrypt, decrypt, wrap, unwrap) and can never export or manage them. Auditing every operation under these roles through Azure Monitor and a Log Analytics workspace, retained per your PCI-DSS retention requirement, is what turns "we have RBAC" into "we can prove who did what, when" for an assessor.
flowchart TD
A[Compromise suspected: alert, audit finding, or offboarding] --> B[Disable affected key/secret version]
B --> C{Manual cleartext operation involved?}
C -->|Yes, e.g. HSM security domain restore| D[Split-knowledge custodians act under quorum]
C -->|No| E[Automated rotation policy issues new version]
D --> F[Re-point consumers via versionless reference or redeploy]
E --> F
F --> G[Confirm old version unusable: disabled and RBAC-blocked]
G --> H[Write audit trail: who, what, when, approval ticket]
Worked example
A retail platform's tokenization key, used to convert primary account numbers into tokens before storage, lives in Managed HSM. An internal audit flags that a departing contractor held Crypto Officer access for three weeks longer than their engagement required. The emergency path: disabling the key itself is not needed, since there is no evidence of use, but the contractor's role assignment is revoked immediately (a role-assignment removal, seconds, no key operation needed). Because the incident review cannot fully rule out access, the key is rotated as a precaution: a new key version is generated, the application, which references the key by its versionless URI, picks up the new version on its next token-generation call with no redeploy, and the previous version is disabled, not deleted, so tokens already generated with it and needing detokenization later still resolve correctly, under the now-revoked contractor's removed access. The audit trail records the role-removal timestamp, the rotation timestamp, and the ticket number, exactly the who, when, what record a PCI-DSS assessor asks for.
Trade-offs & pitfalls
Rotating a KEK is cheap under envelope encryption; rotating a DEK directly is not, because it requires re-encrypting the underlying data. Confusing the two in a rotation policy is the single most common design mistake: teams write "rotate keys every 90 days" without specifying which layer, then discover the policy implies re-encrypting a multi-terabyte cardholder data store every quarter.
Pinning applications to a specific key or secret version for change-control reasons, a legitimate practice for high-risk changes, directly trades away emergency-rotation speed. Decide deliberately, per secret, which property matters more.
Managed HSM's single-tenant isolation and local RBAC are real security wins but come at a real cost and operational premium over shared Key Vault tiers. Reserve it for cardholder-data-scope keys specifically, not as a default for every secret in the environment, or the cost and the quorum-management overhead becomes its own operational burden.
A rollback plan that only covers "re-enable the previous version" is incomplete if the previous version was already disabled for a reason, for instance because it was the thing being rotated away from after a suspected compromise. The runbook needs an explicit rule for when rollback is not the safe choice, not just a default revert action.
You're handed an enterprise Azure invoice that shows unexpectedly high compute and storage costs. Describe a step-by-step approach to analyze the bill, identify hotspots (e.g., idle VMs, oversized disks, retention policies), estimate savings potential from reserved instances and spot VMs, and propose a prioritized implementation plan with ROI for each recommendation.
Sample Answer
Direct answer
Work top-down from the bill, not bottom-up from guesses: use Cost Management's cost analysis to find which resource groups and services actually drove the increase, drill into the specific resources inside those, then apply resource-specific checks (idle compute, oversized disks, stale retention) only to the hotspots the data pointed at, and size the savings and the implementation order by expected dollar impact divided by effort, not by which fix is most familiar.
Structured elaboration: a step-by-step approach
1. Establish the baseline and the delta. In Cost Management, group cost by service and by resource group over the last 3 months, and identify which specific service or resource group's cost curve actually inflected upward, and when. This turns "the bill is unexpectedly high" into "compute in resource group X roughly doubled starting the week of the 14th," which is a findable, explainable fact instead of a vague impression.
2. Attribute the delta to specific resources. Within the resource group or service that inflected, group by resource (not just service) to find the specific VM, disk, or database instance(s) responsible. A jump that is spread evenly across 40 VMs suggests a platform-wide change (a SKU-wide price change, a new region's pricing, a policy that stopped a discount from applying); a jump concentrated in 2 VMs suggests those two specifically changed behavior (someone resized them, a workload moved onto them, they stopped getting shut down on schedule).
3. Check the classic hotspots, informed by what step 2 found, not blindly.
- Idle or oversized compute: pull average CPU/memory utilization for the flagged VMs over the same window (Azure Monitor metrics, or Azure Advisor's own "low-utilization VM" recommendations, which already do this analysis for you). A VM sized for peak load that never approaches that peak is the single most common and highest-value finding in this kind of review.
- Oversized or orphaned disks: managed disks billed at their provisioned size regardless of how full they are, and a disk detached from a deleted VM keeps billing until it is explicitly deleted; both are easy to miss because neither shows up as "a resource that's obviously broken," only as a line item that quietly never goes away.
- Retention and duplication: Log Analytics/Application Insights ingestion at a default or forgotten-about long retention, unused snapshots, or backup vaults retaining far more restore points than the actual recovery policy requires, are all silent, compounding monthly costs with no functional owner actively watching them.
- Missing reservations or savings plans on stable, always-on workloads: the highest-value structural fix is often not "turn something off" but "the always-on baseline is paying full pay-as-you-go price when it has been running steadily for months."
4. Estimate savings potential, grounded in real numbers, not a memorized percentage. Microsoft's own published pricing states Reserved Instances (RIs) save "up to 72%" compared to pay-as-you-go, with typical realized savings in the 36-72% range depending on term length, region, and instance family; Spot VMs are marketed as "up to 90%" off pay-as-you-go, though the realized discount is workload- and region-dependent and can be smaller. As one concrete, currently-measured data point (Azure Retail Prices API, eastus, checked live for this analysis): a Standard_D4s_v5 Linux VM lists at $0.192/hour pay-as-you-go and $0.040531/hour as Spot, a 78.9% discount for that specific SKU and region at the time of the check, which sits inside Microsoft's advertised "up to 90%" range but is a real, not a rounded-headline, number. Apply the RI/Spot decision correctly: RIs (or savings plans, which are more flexible across SKU/family changes) suit steady, predictable, always-on workloads; Spot suits interruptible, fault-tolerant, or batch/dev workloads that can absorb an eviction, never a stateful production database.
5. Prioritize and propose, with ROI shown per recommendation, not just a total. Rank recommendations by (estimated monthly savings) against (implementation effort and risk), and separate "safe to do immediately" (deleting an orphaned disk with zero attached VM, right-sizing a VM that Advisor already flagged at 3% average CPU) from "needs validation before committing" (buying a 3-year RI locks in a capacity commitment; right-sizing a VM that occasionally does need its current size for an unmodeled peak needs a load-pattern check first, not just an average-CPU glance).
Worked example: sizing one recommendation
An idle VM found in step 3, a Standard_D4s_v5 running 24/7 at under 5% average CPU for the last 60 days, currently on pay-as-you-go: its current cost is 0.192/hour x 730 hours/month = $140.16/month. Two candidate fixes: right-size it down to a Standard_D2s_v5 (roughly half the vCPU/memory of a D4s_v5, since the workload clearly does not need 4 vCPUs at 5% utilization) or move it onto a 1-year Reserved Instance at its current size if the workload genuinely needs to stay always-on at this size for a reason the utilization graph does not show (a fixed baseline requirement, not just historical accident). If right-sizing is correct, the fix saves roughly half of $140.16 (~$70/month, ~$840/year) for one VM; if this pattern repeats across 15 similarly oversized VMs found in the same resource group, that is roughly $12,600/year from a single, low-risk, low-effort change category, which is exactly the kind of number that should be presented and prioritized ahead of a more complex, higher-risk architectural change with a similar dollar impact.
Trade-offs and pitfalls
The most common mistake in this kind of review is starting from a checklist of "things that are usually wasteful" and applying it uniformly across the whole subscription, instead of letting the actual cost delta point at where to look; that wastes the reviewer's time auditing resources that were never the problem and can miss the real driver entirely if it is not on the generic checklist (a forgotten large Log Analytics retention setting, for instance, rarely makes a generic "reduce cloud waste" checklist but can be a bigger line item than every idle VM combined). A second common mistake is recommending a 3-year Reserved Instance purchase for a workload whose long-term shape is not actually known yet; a savings plan (more flexible, applies across a broader set of matching compute regardless of exact SKU/family/region within its scope) or a shorter 1-year term is the more defensible recommendation until the workload's stability is proven over a longer observation window.
Compare Azure Functions hosting plans: Consumption, Premium, and Dedicated (App Service Plan). For a bursty event-driven API with sensitivity to cold starts and a need for VNet access, which hosting plan would you recommend and what mitigations would you apply to reduce cold-start impact?
Sample Answer
Direct answer
For a bursty, event-driven API that is sensitive to cold starts and needs Virtual Network (VNet) access, recommend Flex Consumption, the current recommended plan for new serverless Azure Functions apps, sized with a small number of always-ready instances to cover baseline traffic. If the workload's language or region isn't yet supported on Flex Consumption, fall back to Premium with a minimum instance count of at least 1. The older Consumption plan is disqualified outright: it doesn't support outbound virtual network integration at all, so it can't meet the VNet requirement no matter how cold starts are handled.
Hosting plan comparison
| Plan | Scaling model | Cold starts | VNet | Cost model |
|---|---|---|---|---|
| Consumption (legacy) | Fully event-driven autoscale from zero | Expected on every scale-from-zero instance, with no always-ready option to avoid it | Not supported: the plan has no outbound VNet integration at all | Pay per execution only |
| Flex Consumption (current default) | Event-driven autoscale, plus configurable "always ready" instances | Eliminated for traffic within the always-ready capacity; burst instances beyond it still cold-start | Native, no extra hop | Pay per execution, plus cost for always-ready instances |
| Premium (Elastic Premium) | Pre-warmed instances (a configurable minimum count) plus autoscale for bursts | Eliminated on the pre-warmed capacity | Full support | Billed for pre-warmed capacity continuously, plus scale-out usage |
| Dedicated (App Service Plan) | Runs on VM capacity you already pay for; not autoscaled for event volume unless you add App Service autoscale rules | None, as long as "Always On" is set | Full support | Continuous VM cost regardless of function invocations; idle capacity is wasted spend |
Cold-start mitigations, regardless of plan
- Keep the deployment package small and the dependency tree lean; fewer assemblies or packages to load on a cold instance means a shorter cold start.
- Keep the function app's own startup code (static constructors, dependency-injection container setup) minimal, since it reruns on every cold start.
- The legacy Consumption plan cannot be paired with VNet integration at all, since the plan doesn't support it; move VNet-plus-cold-start-sensitive workloads to Flex Consumption or Premium instead.
- If stuck on Consumption for now, a scheduled timer-triggered "warm-up" ping is a workaround, not a fix; treat it as temporary.
Worked example
A payment-webhook receiver expects a steady 5 requests/second baseline and bursts to 200 requests/second for a few minutes at a time, with a hard requirement that outbound calls to a Key Vault (Azure's managed service for storing secrets, keys, and certificates) and a database go through a private VNet. Configure Flex Consumption with an instance memory size of 2048 MB and an always-ready instance count sized to comfortably absorb the 5 requests/second baseline, so steady traffic never cold-starts. Let the 200 requests/second burst scale out on top of that; the newly spun-up burst instances, not the always-ready ones, absorb the cold-start cost. If even burst traffic can't tolerate a cold start, size the always-ready count closer to the burst peak instead, which trades cost for latency consistency.
Trade-offs and pitfalls
- Premium and Flex Consumption both cost money at zero traffic once you configure pre-warmed or always-ready capacity, so "serverless" here is a spectrum, not a binary; sizing that capacity to the true peak defeats the pay-per-use benefit that made a consumption model attractive in the first place.
- Sizing always-ready or minimum instances too low reproduces the exact cold-start problem the plan change was meant to solve; validate against real traffic percentiles, not the average.
- Legacy Consumption has no outbound VNet integration at all, so it is disqualified immediately for any VNet-plus-cold-start-sensitive scenario, not merely a slower option; don't reach for it out of familiarity.
Design an Azure architecture for a healthcare customer subject to HIPAA that processes PHI. Cover encryption at rest/in-transit, Key Vault with HSM, identity enforcement (PIM, conditional access), network segmentation, logging and auditing retention, and Azure Policy/Azure Blueprints for continuous compliance. Identify residual risks and mitigations.
Sample Answer
Direct answer
HIPAA doesn't certify Azure services directly; it obligates the covered entity (or business associate) to implement specific administrative, physical, and technical safeguards, and Azure's role is to sign a Business Associate Agreement (BAA) and provide services capable of supporting those safeguards. The architecture has to satisfy the Security Rule's requirements end to end (encryption, access control, audit logging, retention) and be continuously enforced, not just correct on day one, which is why Azure Policy and Blueprints matter as much as the initial design.
Controls by requirement
- Encryption at rest: Azure Storage and Azure SQL both encrypt at rest by default (platform-managed keys), but for PHI (protected health information) specifically, use customer-managed keys (CMK) stored in Key Vault with HSM-backed keys (HSM: Hardware Security Module, a dedicated, tamper-resistant hardware device that generates and stores the key so it is never exposed as plain data in software memory) (Managed HSM or Premium-tier Key Vault), so the organization controls key lifecycle and can prove it independently of Microsoft. This matters for HIPAA's technical safeguard requirement to control access to ePHI, since key control is a genuine additional access boundary, not just a compliance checkbox.
- Encryption in transit: enforce TLS 1.2+ on every endpoint (Storage, SQL, App Service) and disable any legacy protocol versions at the resource level, not just at the application layer, so a misconfigured client can't silently negotiate a weaker protocol.
- Identity enforcement: PIM (Privileged Identity Management) for any standing access to systems touching PHI, so admin access is time-boxed and justified rather than persistent, and Conditional Access requiring MFA and a compliant device for any access to PHI-handling resources. Role assignments should follow least privilege at the resource level (a specific database, not the whole subscription).
- Network segmentation: PHI-handling resources live in their own subnet(s) with NSGs (Network Security Groups: rule sets, similar to a firewall, that allow or block network traffic in and out of a subnet or network interface) default-denying anything not explicitly required, PaaS services (Storage, SQL) reachable only via Private Endpoint, and no PHI-handling resource exposed with a public endpoint at all. A hub-spoke topology with Azure Firewall inspecting anything crossing the boundary gives one place to enforce and log this.
- Logging and audit retention: HIPAA's audit control requirement means every access to ePHI needs to be logged with enough detail to reconstruct who accessed what, when. Route diagnostic logs (Key Vault
AuditEvent, SQL Auditing, Storage access logs, Entra ID sign-in/audit logs, Entra ID being Microsoft's cloud identity service, the directory of an organization's users, groups, and their sign-ins) into a Log Analytics workspace with a retention period that meets the organization's actual policy (commonly 6 years for HIPAA-relevant records, which is a business/legal retention decision to confirm with compliance, not an engineering default to assume). Immutable/append-only storage for the audit trail itself (a separate, more locked-down storage account with a retention/legal-hold policy) protects against the audit log being tampered with by the same access it's supposed to catch. - Continuous compliance: Azure Policy definitions enforcing the above as code (deny creation of a storage account without CMK encryption in the PHI resource group, deny a public IP on anything tagged as PHI-handling, require diagnostic settings on every new resource in scope) so drift is caught and blocked at deployment time, not discovered in a quarterly audit. Azure Blueprints is in Microsoft's retirement window (phased retirement began July 31, 2026, full retirement January 31, 2027), so it is not the right choice for a new design today. Package this policy set plus the RBAC (role-based access control: granting permissions by assigning a role, such as Contributor or a custom role, to a user or identity at a specific scope, instead of handing out one-off individual permissions)/network baseline using Deployment Stacks (Microsoft's recommended replacement, which groups a set of resources for coordinated lifecycle management and can deny out-of-stack changes to protect against drift) plus Template Specs (versioned, shareable ARM/Bicep templates for the underlying resources), giving the same "repeatable, compliant-from-creation artifact" outcome Blueprints used to provide, without building on a service that is being phased out.
Worked example
A patient-portal application stores PHI in Azure SQL Database. The database has TDE (Transparent Data Encryption) using a customer-managed key from a Managed HSM, is reachable only via Private Endpoint from the app tier's subnet, and has SQL Auditing enabled, writing to a Log Analytics workspace with a 7-year retention policy set at the workspace level (chosen conservatively above the commonly-cited 6-year figure, per the organization's own legal guidance, which is the kind of number that has to come from compliance, not be assumed by engineering). An Azure Policy assignment at the PHI resource group's scope denies deployment of any SQL Database without Transparent Data Encryption enabled and without a diagnostic setting configured, so a developer cannot accidentally stand up a non-compliant database even if they skip the documented process. Access to the database's admin role requires PIM activation with a business justification and a 4-hour maximum window, and Conditional Access blocks any access attempt from a non-compliant device regardless of MFA status.
Residual risks and mitigations
- A compliant architecture does not guarantee compliant application code. SQL injection, over-broad application-level data access (the app's own service account can read more PHI than any single user session needs), or PHI logged into application logs by mistake are all real risks Azure Policy and network controls don't catch; application-level security review and data-loss-prevention scanning of logs are necessary complements.
- Key control is only as strong as who can manage the Key Vault/Managed HSM itself. If the same broadly-privileged admins who could access the data can also manage the encryption keys, CMK provides less real separation than it appears to; consider a separate, more restricted admin group for key management specifically.
- BAA scope gaps: confirm the BAA and the organization's own configuration actually cover every Azure service in the architecture that touches PHI, including ones added later (a new service added to the pipeline six months in that isn't covered by the existing compliance review is a common, quiet source of exposure).
- Backup and disaster recovery copies of PHI need the same encryption, access control, and audit requirements as the primary; a backup that's encrypted with a different key management story, or restorable by a broader set of admins than the primary, is a real and commonly-missed gap.
Trade-offs and pitfalls
The most consequential mistake is treating HIPAA compliance as a one-time architecture review rather than an enforced, continuously-checked policy set; an architecture that was compliant at launch drifts the first time someone adds a resource outside the documented pipeline, which is exactly what Azure Policy's deny-by-default assignments are for. The second is assuming "Azure has a BAA available" means every service is automatically in scope and covered; the BAA covers services the organization has actually configured and attested to using appropriately, and using an out-of-scope service for PHI without updating that assessment is a compliance gap regardless of the underlying technical security.
Unlock Full Question Bank
Get access to all Microsoft Azure Services and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.