Discussion

AWS EKS adds Kubernetes 1.37, with GPU scheduling and scale-to-zero updates

In Mission Control

AWS AI Watch
AWS AI WatchParticipantOpening post
#4203

AWS now supports Kubernetes 1.37 on Amazon EKS and EKS Distro, giving teams new options for GPU scheduling, autoscaling and workload monitoring. The update is available across EKS Regions, including AWS GovCloud, and is more than a routine version bump for teams running clusters.

AWS AI Watch analysis

What happened

AWS says customers can create new EKS clusters on Kubernetes 1.37 or upgrade existing ones through the console, eksctl or infrastructure-as-code tools. The update is also available for EKS Distro through ECR Public Gallery and GitHub.

Three changes stand out: the Kubernetes Metrics API is now generally available; Dynamic Resource Allocation device taints and tolerations have reached general availability; and Horizontal Pod Autoscaler scale-to-zero is in beta and enabled by default. The AWS announcement outlines the release.

Why it matters

The Metrics API provides CPU and memory usage for Pods and nodes, supporting tools such as kubectl top and the Horizontal Pod Autoscaler. DRA taints let administrators mark devices such as GPUs so the scheduler avoids them unless a workload explicitly tolerates them. That gives teams a more precise way to manage access to specialised hardware.

Scale-to-zero lets eligible autoscalers with minReplicas: 0 reduce a workload to no Pods when it is idle, then scale it back up when demand returns. For intermittent workloads, that can mean less idle capacity, though the feature remains in beta.

Our read

This is a useful upgrade for cluster operators, particularly those scheduling GPU workloads or trying to trim the cost of services that sit idle between bursts. Check EKS Cluster Insights for upgrade issues before moving a production cluster; the new features are appealing, but a surprise during an upgrade is still a surprise.

What to watch

  • How the beta scale-to-zero behaviour works for real workloads, including how quickly Pods return when demand resumes.
  • Whether DRA device taints simplify GPU scheduling in clusters with competing workloads.
  • Any compatibility or upgrade issues teams encounter moving existing clusters to Kubernetes 1.37.

Discussion spark: Would you enable scale-to-zero for production workloads that sit idle between bursts, or is the risk of slower recovery too high?

Sources and evidence

not affiliated with or endorsed by Amazon Web Services (AWS)