Evidence-linked product lifecycle intelligenceSUPPORT · SECURITY · RETIREMENT
← Lifecycle catalogue
SERVICE UPDATEAMAZON WEB SERVICESVERIFIED

PUBLISHER UPDATE · AMAZON-WEB-SERVICES-25EFD1851AE14B0C

AMAZON-WEB-SERVICES-25EFD1851AE14B0C release notes and known issues

Amazon SageMaker HyperPod Inference Gateway for scalable LLM inference

Scope: AWS services. This update record adds version, fix and known-issue context. Its publication date is not a lifecycle boundary.

Summary

Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing system that deploys as a single EKS managed add-on on existing SageMaker HyperPod infrastructure with zero application changes. By replacing unintelligent round-robin load balancing with real-time inference-signal-driven routing, it reduces first-token latency by up to 82% and p99 TTFT reductions of 97–98% in mixed-hardware and burst traffic scenarios. The Gateway is built around 3 core components. The Envoy Endpoint terminates HTTPS traffic and exposes a single private endpoint per cluster. The Body-Based Router reads the model name directly from each incoming request and routes it to the correct GPU pool - enabling one gateway to serve many models from a single endpoint URL with no client-side changes required. The Endpoint Picker continuously scores every model server pod in real time across 6 inference-level signals - KV cache utilization, queue depth, LoRA adapter residency, prefix cache hit rate, predicted latency and running requests - selecting the optimal pod for each individual request. The gateway works with any OpenAI-compatible model server, including vLLM and SGLang, requiring no application code changes. Per-cluster routing is available today in all AWS Regions where the SageMaker HyperPod inference add-on is supported. Coming soon - cross-cluster and cross-region routing with a centralized fleet gateway, global rate limiting, and cost-tier-aware traffic shaping. To learn more, read the launch blog and explore the documentation

Improvements and security content

  • Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing system that deploys as a single EKS managed add-on on existing SageMaker HyperPod infrastructure with zero application changes. By replacing unintelligent round-robin load balancing with real-time inference-signal-driven routing, it reduces first-token latency by up to 82% and p99 TTFT reductions of 97–98% in mixed-hardware and burst traffic scenarios. The Gateway is built around 3 core components. The Envoy Endpoint terminates HTTPS traffic and exposes a single private endpoint per cluster. The Body-Based Router reads the model name directly from each incoming request and routes it to the correct GPU pool - enabling one gateway to serve many models from a single endpoint URL with no client-side changes required. The Endpoint Picker continuously scores every model server pod in real time across 6 inference-level signals - KV cache utilization, queue depth, LoRA adapter residency, prefix cache hit rate, predicted latency and running requests - selecting the optimal pod for each individual request. The gateway works with any OpenAI-compatible model server, including vLLM and SGLang, requiring no application code changes. Per-cluster routing is available today in all AWS Regions where the SageMaker HyperPod inference add-on is supported. Coming soon - cross-cluster and cross-region routing with a centralized fleet gateway, global rate limiting, and cost-tier-aware traffic shaping. To learn more, read the launch blog and explore the documentation

Known issues

Publisher statement

Not stated. The verified publisher record does not contain a known-issues statement.

Affected products and versions

Products

  • AWS services

Affected versions

  • See the official publisher source for applicability.

Fixed versions or updates

  • No fixed version is stated in this record.

Recommended action

Review the official publisher document before deployment.

Official publisher evidence

AMAZON WEB SERVICESVERIFIED

Amazon SageMaker HyperPod Inference Gateway for scalable LLM inference

Checked 30 Sep 2026. BlackTree preserves the last verified facts if a later source check is temporarily unavailable.

Open the official publisher source