Microsoft Azure Services and Architecture Questions
Microsoft Azure's core service catalog and architectural patterns: Virtual Machines, managed Kubernetes (AKS), App Service and Azure Functions, Storage accounts and managed disks, Azure SQL and Cosmos DB, VNets with hybrid connectivity and global load balancing, Microsoft Entra ID and RBAC, and Key Vault secrets and encryption. Covers Azure service selection, infrastructure as code (ARM, Bicep, Terraform), observability with Azure Monitor and Kusto queries, cost governance and Azure Policy, the Azure Well-Architected design principles, and hybrid management via Azure Arc, common in enterprise Azure estates. For provider-agnostic trade-offs, see the cross-cloud entries.
Describe how to implement end-to-end encryption for an application across compute, storage, and database in Azure using Azure Key Vault, Managed HSM, and Bring-Your-Own-Key (BYOK). Discuss key lifecycle management, rotation policies, access controls, performance impact, and recovery options if a key is accidentally deleted or rotated incorrectly.
Sample Answer
Direct answer
Use envelope encryption end to end: a per-resource data encryption key (DEK, the key that actually encrypts your bytes) wrapped by a key encryption key (KEK) that never leaves Azure Key Vault or Managed HSM (a dedicated, single-tenant hardware security module; Microsoft updated the Managed HSM fleet to FIPS 140-3 Level 3 validated firmware in 2026, superseding the earlier FIPS 140-2 Level 3 validation, a US government cryptographic hardware standard). Bring-your-own-key (BYOK) means the KEK itself originates from your own key management infrastructure and is imported, not generated, in Azure, which is what gives you the ability to revoke Azure's access to your data by revoking the key rather than trusting a delete request alone.
Structured elaboration
Compute, storage, and database integration.
- Compute (virtual machines, Azure Kubernetes Service (AKS), Azure's managed Kubernetes container-orchestration service): fetch or unwrap keys through a managed identity at startup; never write a plaintext key to disk, only hold it in memory for the life of the process.
- Storage: enable Storage Service Encryption with a customer-managed key (CMK) stored in Key Vault or Managed HSM, so Microsoft's platform-level encryption is now keyed by a key your team controls.
- Database: use Transparent Data Encryption (TDE) with a customer-managed key for Azure SQL, or application-level envelope encryption where the database stores ciphertext and the application unwraps it using a KEK from Key Vault or Managed HSM.
Key lifecycle and rotation. Rotate the KEK on a defined schedule (commonly 90 days, or shorter for the most sensitive data classes); rotating the KEK means re-wrapping the existing DEKs with the new KEK, not re-encrypting the underlying data, which is what keeps rotation fast even for large datasets. Automate this with a scheduled function or runbook: generate the new KEK version, re-wrap every affected DEK, update the reference, and only then retire the old KEK version (never delete it immediately, since backups and archived data may still reference it).
Access controls and governance. Apply Azure RBAC plus Key Vault's own access model at the vault level, use managed identities for every service that needs a key operation, and require Privileged Identity Management (just-in-time elevation) for any human account that needs administrative access to keys. Enable diagnostic logging on every key operation and alert on unexpected deletion or rotation events.
Performance impact. Hardware security module operations are inherently slower than software cryptography; envelope encryption keeps this cost bounded because the KEK is only invoked to wrap or unwrap a DEK at key lifecycle events (creation, rotation, cache refresh), not on every read or write of application data. Cache the unwrapped DEK in memory with a time-to-live and measure the latency this adds to the critical path before assuming it is negligible.
Recovery from accidental deletion or bad rotation. Enable soft-delete and purge protection on both Key Vault and Managed HSM so an accidental deletion is recoverable inside the retention window rather than permanent. Back up key objects on a schedule and store the backup in a separate, encrypted, immutable location. If a rotation was performed incorrectly (for example, the old KEK was deleted before every DEK was re-wrapped), recover the previous key version from soft-delete or backup, re-run the re-wrap step against the recovered key, and only then retire it again. For the worst case, an air-gapped, offline copy of the root key material (an escrow) with a documented, rehearsed restore runbook is the last line of defense.
Worked example
Trace a rotation event concretely: the KEK is due for its 90-day rotation. The automation creates a new KEK version in Managed HSM, then iterates over every DEK currently wrapped by the old version, unwraps each one using the old KEK version (which is still valid during the transition), and re-wraps it with the new KEK version, updating only the small wrapped-DEK metadata, not the (potentially terabytes of) underlying encrypted data. If this process is interrupted halfway through, some DEKs are wrapped with the new KEK and some with the old, which is why both key versions must remain valid and available until the automation confirms every DEK has been re-wrapped, at which point (and not before) the old KEK version is scheduled for retirement, never immediate deletion.
Trade-offs & pitfalls
Managed HSM with a fully independent BYOK root key gives the strongest control and compliance posture but adds real operational complexity and cost; Key Vault's own key management is simpler and is the right default when the compliance requirement does not specifically mandate hardware-isolated, tenant-dedicated key storage. The pitfall that causes real incidents is deleting an old key version immediately after rotation "to clean up," before confirming every dependent DEK and every backup that might reference it has actually been re-wrapped; that turns a routine rotation into an unrecoverable data-loss event.
Describe Azure Key Vault core capabilities: secrets, keys, certificates, access policies versus Azure RBAC, soft-delete and purge protection, and best practices for secret storage and rotation. When would you choose Key Vault over application configuration or environment variables for secret management?
Sample Answer
Direct answer
Azure Key Vault centralizes three kinds of protected objects, secrets (arbitrary small values like connection strings or API keys), keys (cryptographic keys that stay inside the vault or an attached hardware security module for sign and wrap operations, never leaving in plaintext), and certificates (X.509 certificates, which Key Vault can also auto-renew through an integrated certificate authority). Choose Key Vault over environment variables or plain application configuration the moment more than one person or system needs controlled, audited access to a secret, or the secret needs to be rotated without redeploying the application.
Access model: access policies versus Azure role-based access control (RBAC)
Access policies are the older, vault-scoped model: a per-vault list mapping a principal directly to a set of permissions (get, list, set, and so on), coarse-grained and specific to that one vault. Azure RBAC is the current recommended model: standard Azure role assignments (for example, "Key Vault Secrets User", "Key Vault Crypto Officer") scoped at the vault, resource-group, or subscription level like any other Azure resource, auditable through the same Activity Log as everything else, and consistent with tools like Privileged Identity Management (which grants elevated roles only for a limited, just-in-time window instead of standing access) and Conditional Access (which enforces sign-in requirements, such as multi-factor authentication, based on user, location, or device) that only understand RBAC. New vaults should default to the RBAC permission model rather than access policies.
Soft-delete and purge protection
Soft-delete is no longer optional: as of a platform-wide change in 2025, Microsoft enabled soft-delete on all key vaults and removed the ability to opt out. A deleted vault or object is retained and recoverable for a configurable window, from 7 to 90 days, set once at vault creation and fixed thereafter, with 90 days as the default. Purge protection is a separate, still-optional setting on top of soft-delete: without it, someone with sufficient permission can permanently purge a soft-deleted vault or object immediately; turning purge protection on blocks any purge, by anyone, until the retention window elapses, protecting against a malicious or mistaken permanent deletion, at the cost that you cannot force-delete the vault yourself during that window either, worth knowing before enabling it in a fast-moving development environment where vaults get created and torn down often.
Secret rotation, worked example
For a database credential:
- Generate a new password value.
- Update it on the actual database server.
- Write the new value as a new version of the existing secret in Key Vault, not a new secret name, so applications referencing "the latest version" pick it up automatically without a code or configuration change.
- Confirm dependent applications refresh: either they restart, or their in-memory secret cache expires on its own Time-To-Live (TTL) and re-fetches.
- The old version stays in the secret's version history (recoverable, since it's just superseded, not deleted) unless a compliance requirement demands it be removed outright, and even then it remains recoverable during the soft-delete window.
For cryptographic keys specifically, Key Vault's built-in rotation policy can rotate them on a schedule automatically; for a secret backing an external system like a database, Key Vault doesn't know how to change that system's actual password by itself, only how to store the new value, so pair an Event Grid notification on near-expiry with an Azure Function or Automation runbook that performs steps 1 and 2 above and writes the result back as a new version.
When to choose Key Vault over application configuration or environment variables
Environment variables and plain application settings are visible to anyone with read access to the application's configuration, often a broader audience than "should see the database password", and carry no audit trail, no rotation mechanism, and no per-secret access control. Key Vault centralizes secrets with per-secret RBAC, a full access audit log, version history (so a rotated secret's prior value isn't gone, just superseded), and integrates with managed identities so an application fetches a secret at runtime with no static, long-lived credential ever materialized in configuration at all. The trade-off is an added network dependency and a small latency and availability consideration, an outage or network block reaching Key Vault becomes an outage for anything reading secrets at startup, mitigated by caching retrieved secrets in memory for the application's lifetime rather than calling Key Vault on every request.
Trade-offs and pitfalls
- Treating access policies as equivalent to RBAC because "it still restricts access" misses that access policies don't produce the same centralized, auditable trail RBAC does, and current guidance has moved past them for new vaults.
- Enabling purge protection without understanding it blocks even your own force-deletion during the retention window can turn a routine environment teardown into an unexpected multi-day wait.
- Rotating a secret by creating a brand-new secret name instead of a new version of the existing one breaks every application still referencing the old name, and defeats the point of Key Vault's version-based rotation model.
Describe how to use Azure Policy to enforce resource-level constraints such as requiring an enforced 'cost-center' tag, permitted VM SKUs, and encryption at rest. Provide an example policy effect (Audit, Deny, DeployIfNotExists) for tagging and explain how to remediate existing non-compliant resources without breaking production.
Sample Answer
Direct answer
Azure Policy enforces resource-level rules declaratively: a policy definition states a condition and an effect, and an assignment applies it at a scope (management group, subscription, or resource group). For the three examples in the question: a required cost-center tag uses Deny (or Modify/DeployIfNotExists if you want to auto-remediate rather than just block) at creation time; permitted VM SKUs use Deny against any SKU outside an allowed list; encryption at rest is usually already enforced by the platform for many resource types, but where it is optional (some storage or database configurations), Audit or Deny verifies it is turned on. Remediating existing non-compliant resources without breaking production means using Audit first to find the population, then a scoped, reviewed remediation task (for tags) or a planned, tested change (for anything that risks an outage) rather than a blanket auto-fix.
Structured elaboration
The three effects asked about, and what each actually does.
- Audit: allows the operation to proceed but flags the resource as non-compliant in Policy's compliance dashboard. Zero risk of breaking anything, since nothing is blocked or changed; the trade-off is that non-compliant resources keep getting created until someone acts on the audit findings.
- Deny: blocks the create/update operation outright if it violates the condition. Effective for stopping new non-compliance immediately, but has zero effect on resources that already exist and already violate the rule; a
Denypolicy assigned today does not retroactively fix yesterday's resources. - DeployIfNotExists (DINE): after a resource is created or updated, checks whether a related resource or configuration exists (a diagnostic setting, an extension, a specific configuration) and if not, deploys it using a specified template, via a managed identity attached to the policy assignment. This is how you get "every VM automatically gets a monitoring agent" or "every storage account automatically gets a diagnostic setting pointed at the central workspace" without anyone manually configuring it per resource. (A close relative,
Modify, directly changes properties of the resource itself, such as adding or updating a tag, rather than deploying a related resource.)
Example: enforcing the cost-center tag. A Deny policy with a condition like field 'tags[cost-center]' exists 'false' blocks any new resource created without that tag, which is simple and effective for new resources but has the sharp edge that automation, IaC pipelines, or a team's existing deployment templates that do not set this tag will start failing outright the moment the policy is assigned, potentially in the middle of an unrelated urgent deployment. A more commonly recommended pattern for tags specifically is a Modify effect with DeployIfNotExists-style remediation: instead of blocking the create, automatically add a default cost-center value (e.g. unassigned) to any resource missing the tag, paired with a separate Audit policy or a Workbook that surfaces every resource still tagged unassigned for a human to correct later. This keeps deployments from breaking while still guaranteeing the tag is always present in some form.
Remediating existing non-compliant resources without breaking production. For a tag fix, this is low-risk: a Policy remediation task (built on the Modify or DeployIfNotExists effect) applies to the existing non-compliant population directly and adding or correcting a tag has no runtime impact on the resource. For something with real blast radius (how much could break, and how far the damage could spread, if the change goes wrong), such as retroactively enforcing a VM SKU restriction on VMs that are already running a now-disallowed SKU, do not remediate automatically: Deny going forward stops the bleeding, and existing non-compliant VMs go into a planned, scheduled migration (resize during a maintenance window, tested against the specific workload) rather than an automated remediation task that could resize or replace a running production VM without warning. The general rule: remediate automatically when the fix is purely metadata or purely additive (a missing diagnostic setting, a missing tag); remediate manually and deliberately when the fix changes the resource's actual runtime configuration or availability.
Permitted VM SKUs and encryption at rest. A SKU allowlist is typically a built-in policy (Allowed virtual machine SKUs) parameterized with your list, assigned with Deny. For encryption at rest, note that many Azure storage/database services encrypt at rest by default with platform-managed keys and cannot be turned off, so a policy here is more often about enforcing customer-managed keys where that's a compliance requirement, or catching a specific resource type where encryption is genuinely optional; check what the specific resource type's actual default is before writing a policy to enforce something that may already be non-optional.
Worked example
A subscription-scoped Deny policy definition targeting Microsoft.Compute/virtualMachines and Microsoft.Compute/virtualMachineScaleSets, condition not (field 'tags[cost-center]' exists 'true'), effect Deny, assigned to the subscription with an exemption for a specifically named "sandbox" resource group where short-lived experimentation is allowed to skip tagging. Rollout sequence to avoid breaking production: assign the same policy as Audit first for two weeks, review the compliance report to find every pipeline/team that would have been blocked, fix those pipelines' templates to include the tag, then flip the assignment to Deny. This order (Audit before Deny) is the single most reliable way to introduce a blocking policy without an unplanned outage caused by the policy itself.
Trade-offs and pitfalls
Deny policies assigned without an Audit-first rollout are the most common source of "why did my deployment suddenly start failing" incidents in a governed environment, especially when the policy is assigned at a management-group scope covering many subscriptions a central platform team does not operate day-to-day. DeployIfNotExists remediation runs under a system-assigned managed identity with specific role assignments granted at policy-assignment time; forgetting to grant that identity the roles its remediation template actually needs is a common, silent failure mode where the policy shows resources as "non-compliant, remediation in progress" indefinitely because the identity cannot actually perform the deployment. Finally, policies stack: a resource can be subject to policies assigned at the management group, the subscription, and the resource group simultaneously, and the most restrictive Deny anywhere in that chain wins, which means testing a new policy's effect at the exact scope and resource types it will actually apply to, not just in isolation, matters before a wide rollout.
You must ensure customer data is stored only in EU regions to meet GDPR requirements. Design Azure architecture and controls for storage, databases, backups, Key Vault keys, and logs to enforce data residency. Include automated enforcement using Azure Policy, tagging conventions, access controls, and audit trails to demonstrate compliance.
Sample Answer
Direct answer
Enforce European Union (EU) data residency preventively with Azure Policy denying any resource creation, replication target, or diagnostic destination outside an approved EU region list, detect drift with continuous compliance scanning, and produce an auditable trail (policy compliance reports plus immutable log exports) that proves it rather than merely asserting it. In plain terms for a non-technical stakeholder: nothing that stores, backs up, encrypts, or logs your data is allowed to exist outside the EU, and the platform checks this automatically every time something is deployed, not just when someone remembers to look.
Structured elaboration
Governance boundary. Group every subscription that may hold customer data under a single Azure management group ("EU-Data"), and assign the enforcement policies at that management group so no subscription inside it can opt out individually.
Core Azure Policy controls, as one initiative.
- Deny resource creation outside an explicit list of EU regions (a built-in "allowed locations" policy).
- Deny storage or database configurations whose replication target falls outside that same EU list. In practice this control matters most for services where an administrator explicitly chooses the secondary region, such as Cosmos DB's global distribution or SQL Database active geo-replication, since those accept any region, paired or not: an admin adding UK South as a secondary read region (the United Kingdom left the EU and is not covered by an EU-only residency policy) is a real, common way personal data leaves the boundary. Azure Storage's own geo-redundant replication is different: Microsoft fixes the GRS secondary to a deliberately same-geography paired region (for example, West Europe pairs with North Europe, and both are EU member states), so GRS on an EU primary does not by itself put data outside the EU; the policy still needs to check it because a future or reconfigured pairing could change, but the risk is concentrated in the services that let an operator pick an arbitrary region.
- Require a small tagging convention on every relevant resource, not a single flag: data-residency (EU), data-classification (personal-data or non-personal), and owner, all enforced by policy with automatic remediation to append the standard values where it is safe to do so, and denied creation where a human would otherwise have to guess the right value.
- Require diagnostic settings to export logs only to EU-based Log Analytics workspaces or storage accounts.
- Require Key Vault instances to sit in an EU region with soft-delete and purge protection enabled.
Storage, database, backup, and key placement. Restrict creation location for storage accounts, Cosmos DB, and relational databases to the approved EU list. For geo-redundancy, choose only EU-to-EU region pairs; if the only available paired region for a given service crosses outside the EU, that is a deliberate architecture exception requiring documented approval, not a default configuration. Recovery Services vaults for backup must themselves be created in an EU region, since a backup vault outside the residency boundary would defeat the whole control even if the primary data never left. Encryption keys used for customer-managed encryption live in an EU Key Vault or Managed HSM, so decryption capability itself never exists outside the boundary.
Logs and audit trails. Route Activity Logs and every resource's diagnostic logs to EU-based Log Analytics and, where a security information and event management (SIEM) tool is used, an EU-based Microsoft Sentinel workspace. Apply immutability policies to the log storage so the audit trail itself cannot be altered after the fact, which is what actually lets you prove compliance to a regulator rather than merely claim it.
Data deletion pathways (the right-to-erasure requirement). A residency control only proves where data lives, not that it can be removed on request, and General Data Protection Regulation (GDPR) includes a right to erasure. Build an explicit deletion workflow that removes a data subject's records from the primary store and propagates that deletion through every EU replica and every backup's retention policy, with a completion record logged to the same immutable audit trail used for residency evidence, so a single workflow answers both "where is this data" and "can you actually delete it."
Pseudonymization for cross-border analytics. Raw, identifiable customer data stays inside the EU boundary under the controls above. Where the business genuinely needs cross-border aggregate analytics (for example, a global reporting dashboard), pseudonymize or aggregate the data before it leaves the EU boundary, so what crosses the border is no longer personal data under GDPR's definition; this is a materially different (and defensible) design from simply replicating raw customer records to a non-EU analytics region, which the policy set above blocks entirely and correctly so.
flowchart TD
MG[Management group: EU-Data] --> POL[Azure Policy initiative: allowed EU regions, tag, diagnostics, KV]
POL --> ST[Storage and databases, EU only]
POL --> KV[Key Vault or Managed HSM, EU only]
POL --> LOG[Log Analytics, EU only]
ST --> BACKUP[Backup vaults, EU only]
LOG --> AUDIT[Immutable audit trail]
ANON[Pseudonymization step] --> GLOBAL[Cross-border aggregate analytics]
ST --> ANON
Worked example
Trace a specific policy catch: an engineer configures a Cosmos DB account's global distribution and adds UK South as an additional read region, intending only to widen availability for a nearby customer base. UK South is not an EU region (the United Kingdom left the EU in 2020), even though it sits in the same general part of the region list as the EU regions the account already uses. The custom policy inspecting each configured region against the approved EU list denies the deployment outright, with a message naming the specific non-compliant region, rather than allowing it to succeed and only catching the violation in a later audit. The engineer removes UK South and adds an EU region instead, and the deployment succeeds; this is the preventative control working as designed, not the detective control catching a mistake after data has already left the boundary. Note that Azure Storage's own GRS would not have produced this exact failure mode on a West Europe account, since its automatic secondary is North Europe, itself in the EU; the preventative check earns its keep specifically on services like Cosmos DB and SQL Database where the secondary region is a free choice, not a fixed pairing.
Trade-offs & pitfalls
The recurring pitfall seen in real incidents is checking only the primary resource's region and missing its replication target, backup vault, or log destination, each of which is a separate place personal data can end up outside the EU even when the primary looks compliant. The second is conflating pseudonymized aggregate data with raw customer data when deciding what can cross the border; the moment an aggregate can be re-identified (a small enough cohort, or a join back to an identifiable source), it is personal data again and the residency controls apply to it just the same.
Describe Azure Confidential Compute and how it protects data in use. For a customer processing encrypted PII in-memory who cannot expose plaintext to host administrators, design a solution using Confidential VMs or TEEs, discuss integration with Key Vault, attestation, and performance trade-offs.
Sample Answer
Direct answer
Azure confidential computing protects data in use, meaning while it's being processed in memory, by running the workload inside a hardware-isolated trusted execution environment (TEE) that even Azure's own hypervisor and host operators cannot read or tamper with, which is a different guarantee than encryption at rest (disk) or in transit (TLS), both of which leave data in plaintext in memory during processing. For a customer processing encrypted PII who cannot expose plaintext to host administrators, the right building block is an Azure confidential VM, using AMD SEV-SNP or Intel TDX hardware (CPU-level security extensions built by AMD and Intel respectively that encrypt a virtual machine's memory and check its integrity in hardware, so not even someone with full administrative control of the physical host can read or silently alter that memory), not a standard VM with disk encryption, because disk encryption alone does nothing once the process is running and the data is in RAM.
Design
- Compute: deploy on a confidential VM size (the DCasv5/DCesv6 general-purpose or ECasv5/ECesv6 memory-optimized families, using AMD SEV-SNP or Intel TDX depending on the series). These VMs get a hardware-enforced memory-isolation boundary between the guest and the hypervisor/host, plus Confidential OS disk encryption with keys bound to the VM's own dedicated virtual TPM (vTPM: a virtual Trusted Platform Module, a small hardware-backed component, virtualized here, that securely generates and stores cryptographic keys and records measurements of the boot process), so even the disk-encryption keys are inaccessible to the host.
- Attestation: on boot, Azure Attestation verifies the platform's security state (SEV-SNP/TDX actually enabled, correct firmware measurements) before the VM is allowed to start at all; if the host is missing the required isolation settings, attestation fails and the VM simply won't boot, rather than silently running unprotected. Inside the running VM, the application can also request its own attestation report to prove to a relying party (or to itself, before releasing a decryption key) that it's genuinely running in the attested confidential environment, not a spoofed one.
- Key Vault integration ("secure key release"): the pattern for encrypted-PII-in-memory is to keep the data encrypted at rest with a key that Key Vault will only release after checking the requesting VM's attestation report. Concretely: Key Vault (Premium tier, with an HSM-backed key, HSM meaning Hardware Security Module: a dedicated, tamper-resistant hardware device that generates and stores the key so it never exists as plain data in software memory, not even Key Vault's own) holds a release policy tied to specific attestation claims (SEV-SNP/TDX enabled, a specific measured boot state); the confidential VM fetches an attestation token from Azure Attestation, presents it to Key Vault, and only then receives the decryption key, which is used to decrypt the PII into memory that's hardware-isolated from the host the whole time. Without a valid, current attestation, Key Vault never releases the key, so a compromised or non-confidential host has no path to the plaintext even if it somehow obtained the ciphertext.
- What this does not cover: confidential VMs protect against a curious or compromised host administrator or hypervisor; they do not protect against a vulnerability in the guest OS or application itself, or against a legitimate, authenticated user of the application misusing data they're authorized to see. Scope the security claim accurately: "the cloud provider's own operators cannot read this data," not "no one can misuse this data."
Worked example: attestation and performance trade-offs
Consider a service processing encrypted patient records: records are stored in Blob Storage encrypted with a customer-managed key held in Key Vault Premium (HSM-backed). The processing VM is a DCasv5 confidential VM. On startup, the app requests an attestation token from Azure Attestation, which checks that SEV-SNP is active and the VM's measured boot state matches an expected value; that token is presented to Key Vault as part of the key-release request. Key Vault's release policy checks the token's claims and, only if they match, releases the decryption key to the VM's memory. The VM decrypts each record just before processing it and re-encrypts (or discards) it before writing back out, so plaintext never exists outside the confidential VM's isolated memory and never touches disk unencrypted.
Performance trade-offs to be upfront about, since this question explicitly asks for them: confidential VMs carry real overhead versus a standard VM of the same size, from memory encryption/integrity checking done by the CPU (AMD SEV-SNP and Intel TDX both add per-memory-access overhead) and from Confidential OS disk encryption adding to VM boot time. They also come with real feature limitations worth knowing before committing to them: no Accelerated Networking, no Live Migration, no Azure Site Recovery support, and confidential disk encryption itself is only supported below 128 GB per disk (larger disks need a different storage approach). None of this rules confidential VMs out for a PII-processing workload, since the security requirement (host operators cannot see plaintext) genuinely cannot be met by any of the cheaper alternatives, but it does mean the cost/performance conversation with the customer needs those specific numbers, not a vague "there's some overhead."
Trade-offs and pitfalls
The most common mistake is treating "confidential VM" as a checkbox that alone satisfies "protect PII in use," without also wiring up the attestation-gated key release; a confidential VM that decrypts data using a key it always has standing access to, regardless of attestation state, has the hardware isolation but not the actual security property the customer is asking for (proof that only an attested, genuine confidential environment ever saw the key). The second pitfall is assuming confidential VMs are a drop-in replacement for any workload; the feature-support gaps above (no live migration, no ASR, no accelerated networking) can be real operational blockers for workloads that depend on them, and should be checked against the actual requirements before committing to the design.
Unlock Full Question Bank
Get access to all 20 Microsoft Azure Services and Architecture interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.