Build EMR cluster configurations with instance fleets, Spot allocation, Spark/Hive settings, and auto-termination.
Build EMR cluster configurations with instance fleets, Spot allocation, Spark/Hive settings, and auto-termination.
Required Fields
NameReleaseLabelApplicationsInstancesOutput will appear here...The builder validates that Name, ReleaseLabel, Applications, and Instances all resolve before accepting the JSON as a valid RunJobFlow request, the fields EMR needs to know the cluster's name, software release, installed applications, and instance configuration; it can't verify subnet/security group IDs or IAM role names actually exist, those are checked only against the live account.
Build an EMR cluster using instance fleets (rather than the older instance groups model), which let a single fleet definition list multiple instance types with individual BidPriceAsPercentageOfOnDemandPrice values, so EMR can fall back across types when Spot capacity for the preferred type is unavailable. TargetSpotCapacity alongside TargetOnDemandCapacity on the same CORE fleet (as in the example's 2 on-demand + 4 spot) is a genuine mixed-capacity pattern, not an either/or choice, letting a cluster keep a guaranteed on-demand floor while bursting additional capacity on cheaper Spot.
capacity-optimized AllocationStrategy for Spot within a fleet picks instance pools with the most available capacity (lowest interruption risk) rather than the cheapest, which is usually the right tradeoff for a data-processing cluster where a mid-job interruption is more costly than a slightly higher Spot price.
AutoTerminationPolicy's IdleTimeout is a genuinely useful cost control that's easy to forget to set, a cluster left running 'just in case' after the actual job finishes is a common source of surprise EMR bills, especially on clusters people spin up for ad-hoc exploration.
Mixing too many dissimilar instance types (different generations, drastically different WeightedCapacity) in one fleet can create uneven task distribution across the cluster, keep listed types reasonably comparable in actual compute capacity even if they differ in family.
The SpotSpecification's TimeoutDurationMinutes and TimeoutAction control this: after the timeout window elapses without fulfilling the requested Spot capacity, TimeoutAction either switches the unfulfilled portion to on-demand (SWITCH_TO_ON_DEMAND) or terminates the cluster launch, depending on which action is configured, EMR doesn't wait indefinitely for Spot capacity by default.
It's a genuine mixed-capacity pattern, a single instance fleet (like the CORE fleet in the example) can specify both a target on-demand capacity and a target Spot capacity simultaneously, EMR fulfills both targets from the fleet's listed instance types, giving you a guaranteed on-demand floor plus supplemental cheaper Spot capacity, not a choice between the two.
They serve different scopes and can coexist: KeepJobFlowAliveWhenNoSteps controls whether the cluster stays alive between discrete steps (jobs) submitted to it, while AutoTerminationPolicy's IdleTimeout is a newer mechanism that terminates the cluster after a period of genuine idleness regardless of the KeepJobFlowAlive setting. Using both means the cluster survives between steps but still eventually self-terminates if truly idle for the configured duration, a sensible combination for cost control on a shared cluster.
Was this tool helpful?
Disclaimer: This tool runs entirely in your browser. No data is sent to our servers. Always verify outputs before using them in production. AWS, Azure, and GCP are trademarks of their respective owners.