SageMaker JumpStart Adds Optimized Deployments for FMs
🚀 SageMaker JumpStart now offers optimized deployments for foundation models, providing pre-configured, task-aware settings tailored to specific use cases and performance goals. Customers can choose cost-, throughput-, latency-optimized, or balanced configurations and preview P50 latency, time-to-first-token, and throughput metrics before deployment. Supported models include variants from Meta, Microsoft, Mistral AI, Qwen, Google, and TII, and deployments target SageMaker AI Managed Inference endpoints or HyperPod clusters. The feature leverages VPC deployment capabilities and is available in all regions where JumpStart is supported.
