A production-focused guide for engineers running EKS in anger — not for beginners. Covers two core jobs: building resilient workloads and diagnosing live incidents. Includes a triage sequence to identify failure domains in under 2 minutes, deep dives into EKS Tier-0 components (CoreDNS, VPC CNI, kube-proxy, EBS CSI, Load Balancer Controller), and common failure modes like conntrack exhaustion, ephemeral port exhaustion, NLB/NAT idle timeouts, DNS query limits, and security group connection tracking. Provides concrete kubectl commands, evidence collection scripts, and a step-by-step incident playbook for Tier-0 failures.

2h 23m read timeFrom samof76.space
Post cover image
Table of contents
0. Quick start (emergency edition)Introduction1. How to dive into an EKS cluster2. EKS Tier-0 components (what must stay healthy)3. Networking4. Workload identity and security5. Storage and persistent volumes6. Observability and monitoring7. Scaling and performance8. Upgrades and maintenance9. Disaster recovery10. Cost optimization11. Troubleshooting cookbook12. EKS at scaleAppendices
306 Impressions