Cloud Service and Deployment Models Questions
The foundational service models (IaaS, PaaS, SaaS, FaaS) and deployment models (public, private, and hybrid cloud) and when each is appropriate. Covers the shared-responsibility boundary, the core value proposition of cloud versus on-premises, how service-model choice shifts operational ownership, vendor lock-in risks and mitigation, and when a single-cloud, multi-cloud, or hybrid-cloud strategy is the better choice. The conceptual entry point before any provider-specific or architectural depth.
Define vendor lock-in in the context of cloud platforms. List five common lock-in vectors (APIs, managed services, data formats, tooling, identity) and propose practical mitigation techniques for each vector that a cloud architect might implement during evaluation and design.
Sample Answer
Direct answer
Vendor lock-in is the cost, in effort, time and money, of switching away from a provider once you have adopted its services, and it grows precisely where you have taken on the most provider-specific convenience. Five common vectors are proprietary application programming interfaces (APIs), managed services with no equivalent elsewhere, provider-specific data formats, provider-specific tooling, and identity. A cloud architect can mitigate each of these during evaluation and design, before the dependency is load-bearing, at a fraction of the cost of untangling it later.
The five vectors and their mitigations
APIs. Calling a provider's proprietary API surface directly throughout the application ties every call site to that vendor. Mitigation: introduce a thin abstraction layer, an internal client wrapping the provider's software development kit, so a future swap touches one module instead of every call site. Where a genuinely portable standard already exists, prefer it over a proprietary equivalent even when the proprietary option is marginally more convenient today.
Managed services with no equivalent elsewhere. Adopting a highly specific proprietary service means there is no "just switch providers" path, only "rebuild it." Mitigation: at design time, evaluate whether an open-source or multi-cloud-available alternative meets the need at an acceptable extra operational cost, and treat the proprietary option as a deliberate, documented trade rather than a default.
Data formats. Exporting data only in a provider-specific format, or having no bulk-export path at all, means migration starts with a data-transformation project before anything else can move. Mitigation: build and periodically test a bulk-export path to an open format (CSV, Parquet, or a standard SQL dump) as part of normal operations, not as a one-time migration afterthought, so it stays trustworthy.
Tooling. Building deployment automation entirely around a provider's proprietary infrastructure-as-code or continuous integration and continuous delivery (CI/CD) product makes the pipeline itself a migration cost. Mitigation: prefer a portable infrastructure-as-code tool such as Terraform over a provider-only equivalent for anything you might need to reproduce elsewhere, and keep build and deploy logic in provider-agnostic tooling where practical.
Identity. Wiring every application's authentication and authorization directly to a provider's identity and access management (IAM) system makes identity itself hard to unwind, since every permission and integration would need recreating. Mitigation: front identity with a standard protocol such as OpenID Connect (OIDC) so the provider's identity system sits behind a portable interface, and keep an inventory of every place a provider-specific role or permission is referenced.
Trade-offs and pitfalls
Every mitigation above has a real, upfront cost, and sometimes worse day-one ergonomics, in exchange for optionality you may never use. The practical stance is to spend mitigation effort proportional to how core and how hard to replace a dependency is, not apply every mitigation to every service uniformly. A low-stakes utility service can reasonably be adopted with zero abstraction; a system-of-record data store deserves the full treatment. The pitfall is doing the opposite: abstracting the trivial dependency for the sake of tidiness while leaving the genuinely core, hard-to-replace one wired in directly because "it was easier to ship that way."
Design a simple production architecture for a customer-facing web application expected to serve 100k daily active users. The team prefers minimal OS management but needs some control over scaling rules and custom middleware. Choose an appropriate cloud service model (or combination) and justify how your choice balances control, operational overhead, and scalability, citing vendor examples. Then describe a past project where you made a similar service-model trade-off and what the measurable outcome was.
Sample Answer
Direct answer
For a customer-facing application at 100,000 daily active users, with the team wanting minimal OS management but real control over scaling rules and custom middleware, I would choose a managed application platform (PaaS), specifically a container-based one such as AWS App Runner, Azure App Service, or Google Cloud Run, sitting in front of a managed database. Container-based PaaS gives up direct OS control, matching the team's stated preference, while still letting them ship whatever middleware (authentication, rate limiting, logging) they need baked into the container image, and it exposes autoscaling as configuration rather than code.
Why this, and not the alternatives
Raw IaaS would over-deliver control the team explicitly said it does not want, at the cost of OS patching and instance-fleet management they would have to staff for no real benefit at this scale. Pure serverless functions could technically work, but the team's requirement for custom middleware, which today typically runs as a layer wrapping the whole request path, would need real re-architecture into a function-per-endpoint shape, adding engineering cost for a workload that is not described as bursty. A container-based PaaS product avoids both problems: it takes the container image unchanged, including the middleware, and the platform still owns the OS and scaling.
A quick sanity check on scale: 100,000 daily active users translating to a handful of requests per user per session is on the order of a few hundred thousand requests a day, which is not an exotic scale. It argues for favoring operability over squeezing out a raw performance ceiling.
flowchart LR
U["User traffic"] --> LB["Load balancer / content delivery network"]
LB --> APP["PaaS app tier: autoscaling containers with custom middleware"]
APP --> CACHE["Managed cache for session and rate-limit state"]
APP --> DB["Managed database"]
A content delivery network (CDN) in front of the load balancer caches static assets close to users and takes load off the app tier; the app tier autoscales on request volume or queue depth, using the platform's built-in policy rather than custom scaling scripts; and the managed database removes OS and patch ownership from that tier too, consistent with the team's stated preference.
Balancing control, overhead and scalability
Scaling rules: container PaaS platforms expose autoscaling policy (CPU or request-based) as configuration you tune, not infrastructure you build. Custom middleware: bringing your own container image means you keep full control over the middleware stack, unlike a "just push source code" PaaS product that would constrain you to its supported frameworks. Minimal OS management: the team never patches a kernel or manages an OS image, since the platform owns that layer entirely.
A past project with a measurable outcome
A team migrating an internal tool off self-managed virtual machines moved its app tier to a container-based PaaS platform specifically to get out of OS patch management, which had been consuming a real, recurring slice of on-call time. Framed as a short story: the team was losing meaningful after-hours time to OS-level patch cycles across a small fleet; the goal was to cut that load without giving up the custom authentication middleware already built into the app; the action was repackaging that middleware into the container image unchanged and switching to the platform's built-in autoscaling instead of custom scripts; the illustrative result was eliminating the OS-patching workstream entirely, with the team reporting roughly a third fewer after-hours pages in the months that followed, since OS-version drift across the fleet stopped being something that needed attention.
Trade-offs and pitfalls
Container PaaS is not unlimited freedom: you are still bound by whatever networking model and container runtime the platform allows, so a workload needing raw kernel modules or unusual hardware access would need to fall back to IaaS. A common pitfall is applying one model uniformly across the whole application when part of it, such as a nightly batch job, fits a scheduled compute pattern better run alongside the main system rather than inside it.
Explain the differences between IaaS, PaaS, and SaaS from a systems administrator's perspective. For each model, name two example services from AWS, Azure, or GCP, describe one operational responsibility that shifts as you move from IaaS toward PaaS, and note one monitoring or backup implication of that shift.
Sample Answer
Direct answer
From a systems administrator's chair, the IaaS-to-SaaS ladder isn't a technical abstraction exercise, it's a description of which daily tasks disappear or move to a different team at each step, starting with patching and ending with almost the entire job shifting from "keep the infrastructure running" to "manage user access and vendor relationships."
IaaS: the starting point
Example services: AWS Elastic Compute Cloud (EC2), Azure Virtual Machines.
The sysadmin still owns OS patch management: choosing a patch cadence, testing patches, and applying them across the fleet, exactly as with on-premises servers, just on rented hardware. The monitoring and backup implication is direct: you must build and maintain your own OS-level and application-level monitoring and backup jobs, since the provider guarantees the underlying hardware and network, not that your specific virtual machine's disk gets backed up or that a runaway process gets alerted on. A sysadmin moving from on-premises to IaaS who assumes "the cloud backs things up for me" is making the single most common early mistake in this transition.
PaaS: the first real shift
Example services: Azure App Service, Google App Engine.
OS-level patch management moves entirely to the provider. The sysadmin's job shifts from "patch the box" to "verify the platform's automatic runtime and OS updates haven't broken application compatibility," and to managing deployment configuration and scaling policy instead of server configuration. The monitoring and backup implication: infrastructure-level monitoring, is the OS healthy, is disk full, is now the provider's concern, so monitoring effort moves up the stack to application-level health checks and request and error-rate metrics. For backup, "backing up a server" stops being a meaningful task, since there's no persistent server to back up, and the real backup concern moves entirely to whatever managed database or storage the application actually uses.
SaaS: the full shift
Example services: Salesforce, Google Workspace.
There's no infrastructure left for the sysadmin to touch. Operational responsibility shifts to identity and access management, who has an account, what permissions they hold, how quickly access is revoked when someone leaves, and to vendor management, tracking the vendor's service-level agreement (SLA) commitments and uptime history. The monitoring and backup implication: "monitoring" becomes watching the vendor's status page and your own usage and license metrics rather than any system you operate, and "backup" becomes verifying, often by actually testing it rather than trusting a marketing claim, that the vendor's own data-export or retention policy actually meets your organization's recovery needs, since you generally have no independent backup mechanism of your own unless you build one on top of the vendor's export capabilities.
Worked example: a sysadmin's week, across the shift
Under IaaS, a real chunk of a sysadmin's week might go to reviewing patch reports and confirming backup jobs completed successfully across a virtual machine fleet. After a move to PaaS for the same application, that time gets reallocated to reviewing application-level error-rate dashboards and adjusting an autoscaling policy, since there's no OS layer left to patch. After a further move of an adjacent capability, internal email for example, to a SaaS product, the equivalent time goes to a quarterly access review, confirming former employees' accounts were actually deactivated, and confirming that the vendor's exported backup of mailbox data can actually be restored, since that's now the only backup lever left.
Trade-offs and pitfalls
The pitfall specific to this transition is treating it as "less work" rather than "different work." A SaaS-era sysadmin still carries real, auditable responsibility: identity governance, vendor SLA tracking, and verified export or backup capability. An organization that lets go of headcount assuming "SaaS runs itself" typically discovers the gap first during an access-related security incident or a data-recovery request that the vendor's default retention policy doesn't actually cover.
You discover that a proposed architecture relies heavily on a single cloud-provider managed service with proprietary APIs. As solutions architect you need to document vendor lock-in risk and propose mitigations. List technical and contractual mitigations you would recommend, and how you'd estimate the migration effort later if needed.
Sample Answer
Direct answer
Documenting lock-in risk on a real architecture means going past "we depend on Service X" to a risk-register entry that names the specific capability you actually use, counts how many integration points touch it, and states plainly what breaks and how fast if the vendor changes terms or deprecates the feature. Mitigations split into technical (retrofit an abstraction boundary, keep a tested export path) and contractual (notice periods, price caps, exit rights), and a later migration-effort estimate should be built bottom-up from countable units of work, not guessed as a lump sum.
Documenting the risk
For the proprietary managed service in question, write down: which specific capability you use (not just the service's name); what a competing or open-source alternative would look like; how many integration points touch it, meaning call sites, data formats it produces, and downstream consumers; and a plain-language severity statement such as "if this vendor doubles price or retires this feature, these specific customer-facing behaviors break, and it would take roughly N weeks to recover."
Mitigations
Technical. Retrofit an abstraction boundary so internal callers use your own interface rather than the vendor's software development kit directly. Identify a concrete reference alternative, an actual competing service or open-source equivalent, and note which of its own proprietary features you are deliberately not using today, which tells you how close you already are to portable. Add automated, regularly tested data export to a standard format, so a future migration doesn't discover mid-cutover that the export path was never actually exercised.
Contractual. Negotiate an explicit deprecation-notice clause (a committed minimum notice period before a relied-upon feature is retired), price-increase caps or predictable renewal terms, and, where there is credible negotiating leverage, a right to extract data in a specified format on exit.
Estimating migration effort later: a worked example
A workable approach is a surface-area estimate: enumerate every integration point, classify each as trivial, moderate, or complex based on how provider-specific its logic is, price each class using a per-item basis from a comparable past migration, sum it, and add contingency weighted toward the high side, since migrations of this shape run over more often than under.
Suppose the application uses a proprietary managed queue-and-workflow product across 12 background jobs. On inspection, 8 use only generic enqueue and dequeue semantics (trivial, about 1 person-day each to re-point at a standards-based alternative), 3 use the provider's specific retry and dead-letter queue (DLQ, where failed messages are routed for inspection) configuration (moderate, about 4 person-days each to re-implement the same semantics elsewhere), and 1 relies on a proprietary workflow-orchestration feature with no direct equivalent (complex, scoped separately at 15 person-days based on a comparable past redesign).
8×1+3×4+15=8+12+15=35 person-days (base estimate)
Applying a 40 percent contingency for the migration-overrun pattern:
35×1.4=49 person-days
That gives a planning range of roughly 35 to 49 person-days, presented to stakeholders as "about 7 to 10 weeks of one engineer's time" (35 divided by 5 working days per week is 7 weeks, 49 divided by 5 is 9.8 weeks), rather than a single, falsely precise number.
Trade-offs and pitfalls
The biggest pitfall in documenting lock-in risk is treating it as binary, locked in or not, rather than as a spectrum of how many integration points exist, how proprietary each one is, and how reversible the data is; a risk-register entry that just says "high risk, uses a proprietary service" without that breakdown doesn't actually help anyone prioritize or estimate later. The second pitfall is doing the estimate once, at design time, and never revisiting it: treat the trivial and moderate and complex breakdown as living documentation that gets rechecked whenever new integration points are added, or the estimate silently rots and becomes useless exactly when it is needed most.
Serverless functions are often described as a form of PaaS. Explain whether serverless (e.g., AWS Lambda, Azure Functions, GCP Cloud Functions) is best classified as PaaS, something between PaaS and SaaS, or a distinct model. Discuss responsibilities for runtime, scaling, and application code management.
Sample Answer
Direct answer
Serverless functions, also called Function as a Service (FaaS: AWS Lambda, Azure Functions, Google Cloud Run functions, the platform formerly known as Cloud Functions), are best understood as a further point along the PaaS spectrum rather than a wholesale new category or a step toward SaaS. The provider takes over everything PaaS already took over (OS, runtime, middleware) plus two things classic PaaS still leaves you: capacity planning and idle cost. What makes serverless feel distinct is a difference of granularity and billing, not a different responsibility layer.
Why not SaaS, and why "further along PaaS"
SaaS gives you a finished application. Serverless gives you a place to run your application code, so it fails the core SaaS test immediately: does the provider own the application logic? No, you still write the function.
PaaS gives you a managed runtime you deploy a service into and typically leave running continuously; you still think in terms of "how many instances" and pay for a provisioned server whether or not it's handling a request. Serverless removes the instance concept from your mental model entirely: the platform decides how many copies of your function to run, scales all the way to zero when idle, and bills per invocation and execution duration rather than per hour of a running server. That is a difference in degree along the same "who manages what" ladder PaaS is already on, not a new category.
The three responsibilities the question names
- Runtime: entirely the provider's. It picks the execution environment, patches the language runtime, and you only choose a supported version and keep your code compatible with it.
- Scaling: entirely the provider's, and automatic. It scales per invocation, including to zero, and the levers you get are reserved concurrency (a cap that both guarantees and limits how many copies of your function can run at once, useful for protecting a downstream dependency like a database connection pool, at no extra charge) and, on some platforms, provisioned concurrency (an additional-cost option that keeps a guaranteed number of copies pre-initialized and ready, so a request doesn't have to wait for a fresh one to start) to reduce cold starts.
- Application code: yours, same as PaaS. You still own business logic, dependencies, and any bugs in your function.
Worked example
Contrast two workloads. An internal admin dashboard that gets hit sporadically all day is a natural PaaS fit: a provisioned instance sits warm between requests, and you pay for that idle time regardless of traffic. An event-driven job that resizes an uploaded image whenever one lands in storage is a natural serverless fit: the function runs for, say, a few hundred milliseconds per image and you pay only for that time, at the cost of a cold-start penalty the first time an idle function wakes up to handle a request.
Trade-offs and pitfalls
Serverless trades away two things PaaS still gives you. First, statelessness: nothing reliably persists in memory between invocations, so any state the workload needs has to live in an external database or cache. Second, a hard maximum execution duration, which makes serverless a poor fit for long-running or streaming work. The common pitfall is treating "serverless" as free of operational burden entirely; in practice you trade server-patching burden for observability burden, since tracing a request across many short-lived, independently scaling function instances is genuinely harder to reason about than reading one long-lived process's logs.
Unlock Full Question Bank
Get access to all 15 Cloud Service and Deployment Models interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.