Build Databricks cluster configs with autoscaling, Photon runtime, Spark tuning, instance pools, and cluster policies.
Build Databricks cluster configs with autoscaling, Photon runtime, Spark tuning, instance pools, init scripts, libraries, and cluster policies.
Required Fields
workspaceNameresourceGroupclusterConfig.clusterNameclusterConfig.sparkVersionclusterConfig.nodeTypeIdOutput will appear here...The builder validates that workspaceName, resourceGroup, clusterConfig.clusterName, clusterConfig.sparkVersion, and clusterConfig.nodeTypeId all resolve before accepting the JSON as a valid combined cluster, instance pool, and policy definition; it can't verify the specified nodeTypeId is actually available in the target region or that library package versions (PyPI, Maven) are compatible with the specified Spark version, those are only confirmed when the cluster actually attempts to start.
Build an Azure Databricks cluster configuration combining autoscale bounds, the Photon runtime engine (a native, vectorized query engine that substantially accelerates SQL and DataFrame operations over the standard JVM-based Spark engine for compatible workloads), an instancePool for faster cluster startup via pre-warmed idle instances, and a clusterPolicies block constraining what configuration values are actually allowed. dataSecurityMode set to USER_ISOLATION enables per-user data isolation on a shared cluster (multiple users can attach to the same cluster while each only accessing data they're individually authorized for), a materially different security posture from a cluster with no isolation mode, where any user attached to the cluster can potentially access any data any other attached process can reach.
Verify Photon's actual benefit for your specific workload pattern before assuming it's worth enabling broadly, SQL/DataFrame-heavy workloads benefit meaningfully, but custom UDF-heavy or RDD-level code may see little improvement despite Photon's compute cost premium.
USER_ISOLATION (or an equivalent isolation mode) is worth deliberately enabling for any genuinely shared multi-user cluster, don't assume workspace-level access controls alone provide meaningful data isolation between users actively attached to the same running cluster.
Size instancePool's minIdleInstances against actual observed concurrent cluster-launch demand, an under-sized pool still leaves many launches paying the full cold-start cost, while an oversized pool pays for idle capacity that's rarely actually used.
No, Photon's performance benefit is most pronounced for SQL and DataFrame-style operations that map well to its vectorized execution engine, workloads dominated by custom Python/Scala UDFs or RDD-level operations that don't route through Photon's execution path see less or no benefit. Evaluate Photon's actual impact on your specific workload's query patterns rather than assuming it's a universal performance win worth enabling by default for every cluster.
It enforces that each user attached to the shared cluster can only access data and resources they're individually authorized for, rather than a shared cluster where any attached user's code could potentially access any data reachable by any other process on that same cluster. Without an isolation mode configured, a genuinely shared multi-user cluster has a much weaker data-access boundary between users than most organizations assume when just seeing 'the cluster has access controls' at the workspace level.
It significantly reduces startup time by drawing from already-running, idle instances (up to minIdleInstances) rather than provisioning new VMs from scratch, but if the pool's idle capacity is already exhausted (more concurrent cluster requests than minIdleInstances covers), additional instances still need to be provisioned from cold, so startup speed benefits depend on actual pool utilization relative to minIdleInstances at any given moment, not an absolute guarantee.
Was this tool helpful?
Disclaimer: This tool runs entirely in your browser. No data is sent to our servers. Always verify outputs before using them in production. AWS, Azure, and GCP are trademarks of their respective owners.