kapynInfrastructure

Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM

Qwen3.8‑2.4T‑A95B, a 2.4‑trillion‑parameter open‑weight model, is deployed on Amazon SageMaker HyperPod using vLLM. The guide covers cluster provisioning, NVFP4 quantization, and an OpenAI‑compatible endpoint that supports built‑in reasoning, tool calling, and native MTP speculative decoding. This setup enables developers to run massive models efficiently on SageMaker’s high‑performance GPU infrastructure.

AWS ML Blog·Sep 9, 2026

Opening Kapyn…