☁️

Cloud & Infrastructure Topics

Cloud platform services, infrastructure architecture, Infrastructure as Code, environment provisioning, and infrastructure operations. Covers cloud service selection, infrastructure provisioning patterns, container orchestration (Kubernetes), multi-cloud and hybrid architectures, infrastructure cost optimization, and cloud platform operations. For CI/CD pipeline and deployment automation, see DevOps & Release Engineering. For cloud security implementation, see Security Engineering & Operations. For data infrastructure design, see Data Engineering & Analytics Infrastructure.

Observability and Monitoring Architecture

Building visibility into infrastructure and services: metrics, logs, and traces, dashboards and alerting, SLIs/SLOs, and the design of an observability stack. Covers instrumenting systems for actionable signal, reducing alert noise, and diagnosing production issues from telemetry. Infrastructure-wide observability, distinct from network-specific monitoring.

1 questions

Infrastructure as Code and Automation

Defining, provisioning, and automating infrastructure programmatically. Covers declarative IaC with Terraform and comparable tools like CloudFormation (resource and provider model, state management and remote backends, module design and reuse, workspaces, drift detection, and safe plan/apply workflows), plus the broader automation discipline: provisioning pipelines, golden-image and machine-image building, scripting glue, self-service platforms, and end-to-end environment stand-up. The authoring, lifecycle, and automation of infrastructure code that reduces manual toil across provisioning workflows.

13 questions

Infrastructure Scaling, Capacity Planning, and High Availability

How to make running infrastructure scale and stay available: the operational and architectural mechanics of doing it, not the growth-modeling math behind it. Covers autoscaling policy design and Kubernetes cluster scaling (target-tracking, step, scheduled, and predictive triggers, cooldowns, warm pools, HPA, VPA, Cluster Autoscaler, GPU/node scheduling, bin-packing), load balancing algorithms and architecture (round-robin, least-connections, consistent hashing, L4 vs L7, health checks, connection draining, session affinity), horizontal versus vertical scaling choices, and high-availability and redundancy design (multi-AZ and multi-region failover, active-active vs active-passive, split-brain and leader election). Also covers turning a given demand or growth figure into concrete provisioning: headroom and safety-margin sizing, back-of-envelope instance, IOPS, and replica-count math, right-sizing, and multi-year procurement planning; validating a sizing or scaling change through load, stress, soak, and chaos testing; scaling stateful tiers such as databases, caches, and message queues; cost-aware trade-offs including reserved versus spot capacity and managed versus self-hosted infrastructure; and the observability needed to catch capacity saturation before it breaches an SLO.

33 questions

AWS Core Services and Architecture

Amazon Web Services' core service catalog and how the pieces compose into a working system: EC2, Lambda, S3, VPC, IAM, RDS, and the managed-service ecosystem. Covers service selection within AWS, common reference architectures, the AWS Well-Architected Framework pillars, and operational patterns specific to the platform. For provider-agnostic compute or storage trade-offs, see the cross-cloud entries.

0 questions

Cloud Compute Options and Trade-offs

Choosing among compute abstractions independent of provider: virtual machines, containers, managed container services, serverless functions, and bare metal. Covers the cost, control, cold-start, scaling, and operational trade-offs of each model; how to pick an instance family or hardware accelerator (general-purpose, compute-optimized, memory-optimized, GPU, TPU) and a purchasing model (on-demand, reserved, spot); and how workload characteristics (latency, statefulness, burstiness) drive the decision. Managed-versus-self-managed reasoning lives here.

3 questions

Infrastructure Strategy and Technology Selection

Setting technical direction for infrastructure and deciding what to build on. Covers infrastructure vision and long-term roadmap, modernization and technical-debt strategy, platform decisions, and organizational and governance considerations, alongside the decision framework for adopting or retiring technology: build-versus-buy-versus-cloud-versus-on-premises trade-offs, vendor and platform evaluation, technology-portfolio rationalization, requirements-driven service selection, and weighing total cost of ownership and risk. It also covers managing an existing vendor relationship: quantifying and mitigating lock-in, negotiating contract terms, and planning an exit from a service the organization depends on. The leadership-altitude discipline of directing an infrastructure estate over time, building consensus among stakeholders with conflicting priorities, and persuading leadership or a team to accept a major technology change.

50 questions

Cloud Architecture Design Principles and Trade-offs

The cross-pillar reasoning skill for architecting cloud systems: weighing reliability, scalability, security, performance, and cost against each other to justify ONE architectural choice over another under real constraints (budget, team size, timeline, existing systems). Covers well-architected-style design reviews, resilience and failure-mode reasoning (blast radius, graceful degradation, idempotency), consistency-versus-availability trade-offs (CAP/PACELC), and scenario-based decisions such as choosing a managed versus self-hosted component or an architectural style (monolithic, microservices, or serverless) for one system. Provider-agnostic: no specific cloud vendor's service catalog. This topic is the JUSTIFICATION layer, not a subsystem deep dive: a full design of observability, disaster recovery, identity and access management, networking, caching, or Kubernetes orchestration belongs to that subsystem's own topic. Comparing compute abstractions (VM versus container versus serverless versus GPU/TPU) belongs to compute options and trade-offs. Choosing an architectural style is covered here, but the internal implementation patterns of that style (service mesh, sagas, two-phase commit, event sourcing) belong to microservices architecture and service design. Multi-year roadmaps, vendor evaluation, and governance belong to infrastructure strategy and technology selection. Spanning multiple cloud providers or bridging on-premises and cloud belongs to multi-cloud and hybrid cloud architecture. The IaaS/PaaS/SaaS delivery-model taxonomy belongs to cloud service and deployment models. Region-crossing replication and failover design belongs to multi-region and geo-distributed systems.

5 questions

Cloud Cost Optimization and FinOps

Controlling and optimizing cloud spend: cost modeling and forecasting, rightsizing, reserved capacity and savings plans, autoscaling for cost, tagging and chargeback, and the FinOps operating model. Covers building the business justification for infrastructure spend and continuously driving efficiency at scale without sacrificing reliability. Cost as a first-class architectural concern.

12 questions

Cloud Migration Strategy and Execution

Planning and executing a move to the cloud: the migration strategies (rehost, replatform, refactor, repurchase, retire, retain), legacy assessment, dependency mapping, cutover planning, and rollback. Covers phased migration roadmaps, workload modernization, risk management during cutover, and validating success post-migration. The end-to-end migration lifecycle, not steady-state operations.

9 questions