Skip to main content

Amazon SageMaker HyperPod

Ship models, we'll keep the training and inference GPUs working

Maximize every GPU hour building domain and specialized intelligence, from BioFMs and Physical AI to World Models and Sovereign AI. From notebook to production, through training, fine-tuning, serving, and observability, no re-platforming or pipeline assembly required.

Trusted by leading Startups, ISVs, and AI Enterprises building domain intelligence

Missing alt text value
Missing alt text value
Missing alt text value
Missing alt text value
Allen-Institute Logo

Cost-Optimized GPU Utilization

Up to 40% cost savings with Task Governance

Up to 65% savings with Flexible Training Plans

Trainium delivers up to 40% better price performance than comparable GPU-based instances on AWS

Every GPU hour working. Every capacity shift absorbed.

  • Task Governance automates prioritization, allocation, and lending/borrowing of compute across tasks and teams, maximizing utilization and reducing costs by up to 40%.
  • Self-service capacity reservations help achieve up to 65% savings vs. on-demand.
  • Training jobs automatically scale based on resource availability and priority, freeing up compute for higher-priority workloads and saving hours of engineering time.
  • Run on Trainium or NVIDIA GPUs on the same service with one-line recipe switching.
Customer Proof
"Articul8 achieved 5x total cost of ownership and increased 35% productivity. With automated task prioritization, they have seen a dramatic improvement in GPU utilization, thereby reducing idle time and accelerating their model development process by optimizing tasks ranging from training and fine-tuning to inference."
Missing alt text value

Maximize Goodput at Scale

Training goodput

Less than 2 minute recovery time

Perplexity achived 40% faster foundation model training and doubled training throughput.

Train for weeks. Maximize every hour. Never lose a run.

  • Achieve up to 95% training goodput on clusters with thousands of accelerators.
  • Never stop a training run because your cluster changed shape.
  • When hardware fails, recover in minutes, not hours, without losing training progress.
  • Checkpoint-restart cycles are eliminated entirely, maintaining forward momentum despite failures and reducing recovery time to minutes.
  • Health monitoring continuously detects, diagnoses, and recovers from faults with zero manual intervention.
Missing alt text value

Train-to-Serve on One Cluster

From data prep to production serving. Nothing to assemble. Governance built in.
  • Recipes and agent skills generate fine-tuning code and auto-diagnose NCCL failures, GPU hardware faults, and cluster issues, reducing time to deploy and support.
  • The only managed service that offers both Slurm and Amazon EKS orchestration.
  • JupyterLab and VSCode run directly on HyperPod clusters, enabling development on existing compute without switching infrastructure.
  • Training jobs run on shared infrastructure with fractional GPU allocations (NVIDIA MIG).
  • FSx for Lustre is built in, providing the fastest storage performance for GPU instances in the cloud.
  • Observability is built in across model development tasks and compute resources, from open source to managed (MLflow, Grafana, Prometheus, CloudWatch).
Missing alt text value

High-Performance Inference

Up to 25% Cost Savings and Throughput

Up to 40% Latency Reduction

Same resilient cluster for training and inference. No re-platforming to serve the models you just trained.

  • Deploy and serve models at production scale with the same HyperPod resilience you rely on for training. 
  • Scale-to-zero, multi-instance GPUs (MIG) for GPU sharing, Spot Instances, and intelligent routing pack requests onto fewer instances leading to up to 25% lower costs. 
  • Tiered KV Cache cuts latency up to 40%, intelligent routing reuses caches for 25% more throughput, model weight caching and image cache together nearly eliminate cold start on new nodes.
  • Multi-instance type deployment with prioritized fallback, auto-recovery on node failures, and dual-layer autoscaling (KEDA + Karpenter) scales on demand with near instantaneously.
Missing alt text value

Proven at Scale. Chosen by Leaders.

Luma AI

Luma AI built frontier visual models on Amazon SageMaker HyperPod

TwelveLabs

From horse videos to HyperPod: How Twelve Labs is unlocking video AI

Noetik

AWS Founder Spotlight: Noetik
Thomson-Reuters Logo
The path from training to production with full governance was decisive for our selection.

Thomson Reuters

HyperPod's out-of-the-box observability features provide us with a comprehensive set of metrics across multiple dimensions (cluster, node, task, etc.)

Sony Honda Mobility

3x accelerated model iteration cycles, 90% reduction in training pipeline failures, and zero manual intervention in workload distribution. The infrastructure has proved exceptionally resilient.

Writer

RESTRICTIONS: Stability.ai logo
With SageMaker HyperPod's managed infrastructure and optimization libraries, we can reduce training time and costs by over 50%.

Stability AI

Ray on SageMaker HyperPod

Run fully compatible open-source Ray at production scale with managed development environments, built-in observability, resilient training, and accelerated serving.  

  • Out-of-the-box Observability — Auto-provisioned Grafana dashboards and web-accessible Ray Dashboard. No manual setup, no port-forwarding. 
  • Managed Developer Environments — Run VSCode, Kiro, Cursor, or JupyterLab powered by a Ray Cluster to iterate interactively without submitting a job for every change. 
  • Resilient Training — Take advantage of the self-healing features of HyperPod so your training continues 
  • Reinforcement Learning — Run RL workloads with Ray-based frameworks like RLlib, verl, OpenRLHF, and NeMo-RL. 
  • Performant Serving — Managed Tiered KV Cache with L1 (local CPU) and L2 (cluster-wide storage) for accelerated TTFT.  
  • Purpose-built web interface for Data Scientists — Create Ray clusters, submit and track jobs, and manage model deployments directly from SageMaker Studio.
  • Efficient Capacity Sharing and job queueing — Task Governance works on Ray to allocates cluster capacity across teams so GPUs stay busy.
Missing alt text value

Common questions about SageMaker HyperPod

FAQs

Open all

    Teams building domain and specialized intelligence models including BioFMs, World Models, and Sovereign AI, lose days per month to hardware failures and checkpoint recovery, weeks assembling separate tools from data prep through training to serving, and GPU hours wasted when capacity sits siloed across teams. SageMaker HyperPod eliminates all of it. Automatic healing in minutes with no checkpoint reload, elastic training that adjusts to capacity changes mid-job, shared GPU governance across teams and workloads, built-in observability, and self-service developer environments. And, with Flexible Training Plans, customers can achieve up to 65% cost savings compared to on-demand instances. The result: recovery from failures in minutes, models hitting production on schedule, and engineers building models instead of managing infrastructure.

    SageMaker HyperPod integrates with Amazon EKS, giving platform teams a consistent Kubernetes-based experience to manage and operate clusters for training, fine-tuning, and inference. You can use native kubectl commands, Helm charts, and Karpenter-based autoscaling while HyperPod adds Task Governance, automatic healing, and built-in observability on top. Teams already running Kubernetes workflows can adopt HyperPod easily. 

    SageMaker HyperPod provides native Slurm orchestration with SBATCH workflows and multi-node coordination, so HPC and ML teams can use their existing Slurm scripts and scheduling policies. HyperPod is the only managed training service that offers both Slurm and Kubernetes on the same cluster, letting you share compute capacity across both orchestrators. 

    HyperPod continuously monitors GPU and network health and automatically detects, isolates, and replaces faulty nodes without human intervention. With checkpointless training, recovery happens in minutes through peer-to-peer state transfer from healthy accelerators, eliminating the need for checkpoint-restart cycles. This enables over 95% training goodput on clusters with thousands of accelerators.

    Task Governance automates prioritization, allocation, and lending/borrowing of compute across teams, ensuring the most critical tasks complete first while idle capacity is automatically redistributed. Administrators set per-team quotas and priorities, and HyperPod handles queuing, preemption, and resource lending without manual coordination. This maximizes GPU utilization and reduces model development costs by up to 40%.

    Yes. HyperPod supports training and inference on the same cluster with shared Task Governance and observability, eliminating the need to re-platform models onto separate infrastructure for serving. 

Join the frontier AI teams that chose HyperPod.

Every training run lost to a hardware failure is money and time lost. Your team was hired to build and deploy models, not to rebuild from checkpoints and explain slipped roadmaps to the board. Start using SageMaker HyperPod and failures heal themselves in minutes. No checkpoints. No restarts. Your training and inference keep running. Your models hit production on schedule, and your business moves forward instead of recovering backward.

Loading
Loading
Loading
Loading
Loading