A comprehensive guide to Karpenter best practices covering NodePool design, Spot instance strategy, consolidation policies, disruption budgets, and observability. Key recommendations include using multiple focused NodePools with taints for workload isolation, enabling the SQS interruption queue for Spot (not configured by default), setting CPU and memory limits on every NodePool, scheduling disruption budgets to block consolidation during business hours, and using broad instance-category requirements instead of explicit type lists. The guide also covers AMI pinning, the do-not-disrupt annotation, key metrics to monitor, and common anti-patterns. A final section promotes Cast AI's workload rightsizing and Spot prediction capabilities as complements to Karpenter.

16m read timeFrom cast.ai
Post cover image
Table of contents
Key takeawaysNodePool design: focused, not one-size-fits-allSpot strategy and safe fallbackConsolidation with disruption budgetsLimits and guardrailsLabels, taints, and placement controlObservability and cost visibilityAnti-patterns to avoidHow Cast AI closes the gaps Karpenter leaves openFrequently Asked Questions
84 Impressions