Cloud & Infrastructure Topics
Cloud platform services, infrastructure architecture, Infrastructure as Code, environment provisioning, and infrastructure operations. Covers cloud service selection, infrastructure provisioning patterns, container orchestration (Kubernetes), multi-cloud and hybrid architectures, infrastructure cost optimization, and cloud platform operations. For CI/CD pipeline and deployment automation, see DevOps & Release Engineering. For cloud security implementation, see Security Engineering & Operations. For data infrastructure design, see Data Engineering & Analytics Infrastructure.
Cloud Networking and VPC Design
Designing networks inside a cloud provider: VPC/VNet topology, subnets, route tables, gateways, NAT, and peering, plus private connectivity through VPC endpoints and cloud load balancers. Covers segmentation, security groups and network ACLs, hybrid connectivity to on-premises data centers over VPN or dedicated links like Direct Connect and ExpressRoute, IP address planning across many VPCs and accounts, and how cloud network design differs from traditional data-center networking.
Observability and Monitoring Architecture
Building visibility into infrastructure and services: metrics, logs, and traces, dashboards and alerting, SLIs/SLOs, and the design of an observability stack. Covers instrumenting systems for actionable signal, reducing alert noise, and diagnosing production issues from telemetry. Infrastructure-wide observability, distinct from network-specific monitoring.
Networking Fundamentals and Protocols
The core model of how networks move data: the OSI and TCP/IP layers, the Internet Protocol suite, transport protocols (TCP versus UDP), encapsulation, and TCP behavior including congestion control. Covers the protocol foundations every networking and infrastructure discussion builds on, from link layer through transport. The conceptual bedrock beneath addressing, routing, and switching.
Infrastructure as Code and Automation
Defining, provisioning, and automating infrastructure programmatically. Covers declarative IaC with Terraform and comparable tools like CloudFormation (resource and provider model, state management and remote backends, module design and reuse, workspaces, drift detection, and safe plan/apply workflows), plus the broader automation discipline: provisioning pipelines, golden-image and machine-image building, scripting glue, self-service platforms, and end-to-end environment stand-up. The authoring, lifecycle, and automation of infrastructure code that reduces manual toil across provisioning workflows.
Google Cloud Platform Services and Architecture
Google Cloud Platform's core services and architecture: Compute Engine, Cloud Run, GKE, Cloud Storage, VPC, managed databases (Cloud SQL, Spanner, Firestore, Bigtable), and BigQuery-adjacent data services. Covers GCP service selection, networking, IAM and security specifics, cost and quota management, and reference patterns for building on the platform. For provider-agnostic compute, storage, or networking concepts, see the cross-cloud entries.
DNS, DHCP, and Name Resolution
How names and addresses are served and kept working on a network. DNS: resolution flow (stub, recursive, authoritative, root/TLD referrals), host-side resolver configuration and why one machine or process resolves differently from another, record types and their pitfalls (including apex aliasing, MX, SRV and CAA use), zones, delegation and glue, caching, TTL and negative caching, split-horizon and private zones, reverse DNS, zone transfers (AXFR/IXFR, TSIG), DNSSEC and key rollover, resolver and authoritative fleet design, forwarding versus running recursion, encrypted DNS (DoH/DoT), truncation, TCP fallback and EDNS, DNS-layer attacks (cache poisoning, amplification, spoofed-source floods), DNS health from the user's point of view, and DNS changes, mail and registrar migrations and outages. DHCP: address assignment and leases, scopes and sizing, reservations, relay across VLANs including relay agent information (option 82), redundancy and failover, and rogue or exhausted-scope failures. Includes diagnosing resolution failures that masquerade as wider outages. Excludes the layered network fault-isolation method and general packet capture, Active Directory-integrated DNS on domain controllers, DNS-based service registries for microservices, load-balancing algorithm design and global traffic steering, incident-command process and SLO or error-budget design, and generic scripting or infrastructure-as-code tooling.
Kubernetes Architecture, Operations, and Troubleshooting
How Kubernetes works, how to run it, and how to debug it. Covers control-plane and node components, the scheduler and API server, cluster design, high availability and multi-cluster topologies, and platform-level operations; the workload primitives (pods, deployments, services, controllers), cluster upgrades, and designing Kubernetes as an internal platform; and the operational depth inside a cluster including pod and service networking, ingress and the CNI model, service mesh, persistent volumes and storage classes, resource requests and limits, and systematically diagnosing scheduling, networking, and storage failures. The full architecture-through-day-two-operations span of Kubernetes.
Infrastructure Scaling, Capacity Planning, and High Availability
How to make running infrastructure scale and stay available: the operational and architectural mechanics of doing it, not the growth-modeling math behind it. Covers autoscaling policy design and Kubernetes cluster scaling (target-tracking, step, scheduled, and predictive triggers, cooldowns, warm pools, HPA, VPA, Cluster Autoscaler, GPU/node scheduling, bin-packing), load balancing algorithms and architecture (round-robin, least-connections, consistent hashing, L4 vs L7, health checks, connection draining, session affinity), horizontal versus vertical scaling choices, and high-availability and redundancy design (multi-AZ and multi-region failover, active-active vs active-passive, split-brain and leader election). Also covers turning a given demand or growth figure into concrete provisioning: headroom and safety-margin sizing, back-of-envelope instance, IOPS, and replica-count math, right-sizing, and multi-year procurement planning; validating a sizing or scaling change through load, stress, soak, and chaos testing; scaling stateful tiers such as databases, caches, and message queues; cost-aware trade-offs including reserved versus spot capacity and managed versus self-hosted infrastructure; and the observability needed to catch capacity saturation before it breaches an SLO.
Cloud Service and Deployment Models
The foundational service models (IaaS, PaaS, SaaS, FaaS) and deployment models (public, private, and hybrid cloud) and when each is appropriate. Covers the shared-responsibility boundary, the core value proposition of cloud versus on-premises, how service-model choice shifts operational ownership, vendor lock-in risks and mitigation, and when a single-cloud, multi-cloud, or hybrid-cloud strategy is the better choice. The conceptual entry point before any provider-specific or architectural depth.