Build Synapse Spark pool configs with autoscale, dynamic executor allocation, Spark properties, and library requirements.
Build Synapse Spark pool configs with autoscale, dynamic executor allocation, Spark properties, library requirements, and session-level packages.
Required Fields
workspaceNamesparkPoolNamenodeSizesparkVersionautoScaleOutput will appear here...Build an Azure Synapse Spark pool with autoScale bounds, autoPause (deallocating the pool entirely after delayInMinutes of inactivity, the primary cost control for an intermittently-used pool), and dynamicExecutorAllocation as a separate, Spark-level scaling mechanism operating within whatever nodes autoScale has currently provisioned. autoScale and dynamicExecutorAllocation solve genuinely different scaling problems and are easy to conflate: autoScale changes the actual number of underlying compute nodes in the pool (a slower, infrastructure-level scaling operation), while dynamicExecutorAllocation adjusts how many Spark executors run within the currently-provisioned nodes (a faster, Spark-application-level scaling operation), a pool can have dynamicExecutorAllocation scaling executors up and down rapidly within a completely stable node count that autoScale hasn't changed at all.
Don't conflate autoScale and dynamicExecutorAllocation as the same scaling knob, they operate at genuinely different layers (infrastructure nodes versus Spark executors within those nodes) and tuning one without understanding the other can leave a workload bottlenecked at whichever layer wasn't actually adjusted.
Set autoPause's delayInMinutes based on actual observed usage gaps, too short a delay causes frequent cold-start penalties for a pool used somewhat regularly throughout the day, too long wastes cost on idle compute between genuinely infrequent usage sessions.
Match nodeSizeFamily to the workload's actual bottleneck (memory versus compute), a memory-bottlenecked Spark job on a compute-optimized node family hits performance ceilings that adding more nodes doesn't fix, since the constraint is per-node memory, not aggregate compute.
The builder validates that workspaceName, sparkPoolName, nodeSize, sparkVersion, and autoScale all resolve before accepting the JSON as a valid Spark pool definition, the fields Synapse needs to provision the pool with its sizing and scaling behavior; it can't verify the specified node size is available in the target region or that library requirements (pip packages, custom whl files) are mutually compatible, those are only confirmed when a session actually starts against the pool.
Was this tool helpful?
Disclaimer: This tool runs entirely in your browser. No data is sent to our servers. Always verify outputs before using them in production. AWS, Azure, and GCP are trademarks of their respective owners.