kapynInfrastructure

Introducing Amazon SageMaker HyperPod Inference Gateway

Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing add‑on for Amazon EKS. It uses real‑time GPU signals to send each inference request to the best‑suited pod, cutting first‑token latency by up to 82% with no changes to your model servers or client applications. This gives AI developers a low‑latency, zero‑code‑change solution for scaling inference workloads on EKS.

AWS ML Blog·Sep 18, 2026

Opening Kapyn…