Situation: Our on-prem monitoring (Nagios + custom scripts) struggled with Kubernetes visibility and alert noise as we migrated services to EKS. Incidents were increasing and SLOs slipped.
Task: I led evaluation of a new monitoring solution to improve Kubernetes metrics, reduce alert fatigue, and support SLO tracking.
Action:
- Defined evaluation criteria: native k8s integration, metrics + logs + traces, alerting & deduplication, SLO support, scalability, TCO (licensing + infra), vendor lock-in, security/compliance, and team learning curve.
- Shortlisted three options (Prometheus + Loki + Tempo self-managed, Datadog, and New Relic).
- Ran a 4-week PoC: deployed each in a sandbox cluster, ingested representative workloads, implemented 5 key alerts, built an SLO dashboard for a critical service, measured query latency and storage costs, and timed onboarding for two engineers.
- Trade-offs: self-managed stack minimized vendor cost but increased ops burden and slower feature delivery; SaaS gave faster time-to-value and better k8s auto-instrumentation but higher recurring cost and some vendor lock-in; Datadog had best UX and fewer false alerts but was most expensive.
- I compiled quantitative results (query latency, monthly cost estimates, MTTR on simulated incidents) and qualitative feedback (dev onboarding time, support responsiveness).
Result: Recommended Datadog for initial rollout to reach reliability goals quickly while negotiating a pilot discount and contract clauses to limit lock-in (data export guarantees). Presented a migration plan: 3-month phased rollout, training sessions, and a long-term plan to evaluate hybrid/self-managed components if costs rose. Stakeholders approved based on clear cost-vs-benefit numbers and reduced incident risk.