Build SageMaker endpoint configurations with production variants, auto-scaling, data capture, and canary deployments.
Build SageMaker endpoint configurations with production variants, auto-scaling, data capture, and canary deployments.
Required Fields
EndpointConfigNameProductionVariantsOutput will appear here...Build a SageMaker real-time inference endpoint configuration with multiple production variants, each an independently-scaled model deployment sharing the same endpoint, weighted by InitialVariantWeight for canary or A/B traffic splitting. ManagedInstanceScaling with its own MinInstanceCount/MaxInstanceCount lets SageMaker autoscale a variant's instance count directly (newer than the older Application Auto Scaling-based approach), while DataCaptureConfig's InitialSamplingPercentage logs a percentage of live inference requests/responses to S3 for model quality monitoring, at the storage and slight latency cost of writing every sampled request.
InitialVariantWeight sets the starting traffic distribution at endpoint creation, SageMaker's routing then honors that weight ratio for incoming inference requests, but 'Initial' in the name is a hint: you can update variant weights later via UpdateEndpointWeightsAndCapacities for a gradual shift (a canary rollout progressing from 90/10 to 50/50 to 100/0) without recreating the endpoint.
The default routing is essentially round-robin across a variant's instances. LEAST_OUTSTANDING_REQUESTS instead tracks how many requests are currently in-flight per instance and routes new requests to the least-busy one, which matters a lot for inference workloads with highly variable per-request latency (an LLM generating a long versus short response), where round-robin can pile up requests on an instance that's still working through a slow one.
For sampled requests (governed by InitialSamplingPercentage), there's a small overhead to write the captured input/output to S3 asynchronously, but it's designed to not block the response back to the caller. The bigger practical cost is usually S3 storage and the downstream cost of running Model Monitor jobs against the captured data, not per-request latency.
The builder validates that EndpointConfigName and ProductionVariants resolve before accepting the JSON as a valid CreateEndpointConfig request, the minimum SageMaker needs to know the config's name and at least one model variant to serve; it can't verify the referenced ModelName values actually exist as registered SageMaker models, that's checked only against the live account.
ManagedInstanceScaling is the newer, SageMaker-native autoscaling path, some existing endpoints still use Application Auto Scaling policies layered on top, mixing both approaches on the same variant produces confusing, competing scaling decisions, pick one mechanism per variant.
A canary variant with a low InitialInstanceCount (1, as in the example) means that variant has zero redundancy, if its single instance fails a health check, 10% of traffic briefly has nowhere healthy to go, size canary capacity with at least minimal redundancy for anything beyond a very short-lived test.
DataCaptureConfig's DestinationS3Uri needs its own lifecycle policy, captured inference data at even a modest sampling rate accumulates continuously and is easy to forget about until an S3 storage cost review surfaces it months later.
Was this tool helpful?
Disclaimer: This tool runs entirely in your browser. No data is sent to our servers. Always verify outputs before using them in production. AWS, Azure, and GCP are trademarks of their respective owners.