SageMaker HyperPod adds model caching for inference
🚀 Amazon SageMaker HyperPod now supports model caching to speed up LLM inference by pre-loading model weights and container images onto cluster nodes. This reduces cold-start delays for scale-out events, allowing pods to start in seconds rather than minutes. The feature includes a weights cache on local NVMe and an image cache that pre-pulls container images, with automatic fallback to the original source if a cache is unavailable. Model caching is GA across all SageMaker HyperPod regions and is enabled via the HyperPod Inference Operator.
