Build Dataproc Serverless Spark batch configurations with runtime settings and dynamic allocation.
Build Dataproc Serverless Spark batch configurations with runtime settings, dynamic allocation, networking, and metastore integration.
Required Fields
nameruntimeConfig.versionenvironmentConfig.executionConfig.serviceAccountOutput will appear here...Build a Dataproc Serverless Spark batch that runs without you ever provisioning or tearing down a cluster: sparkBatch (or pysparkBatch) points at your job artifact, runtimeConfig.version pins the Spark runtime, and environmentConfig.executionConfig.ttl bounds how long the ephemeral infrastructure lives before auto-teardown, defaulting to 4 hours in this example. Autoscaling here means spark.dynamicAllocation, not a node-pool setting; minExecutors/maxExecutors in runtimeConfig.properties control it directly since there's no persistent cluster to resize.
No, that's the entire point of Serverless: you set Spark-level resource properties (executor/driver memory and cores) plus min/max executor bounds for dynamic allocation, and Dataproc provisions and tears down the underlying infrastructure automatically per job. There's no persistent cluster to size or keep patched.
The ttl in executionConfig is a hard cap; Dataproc terminates the batch's infrastructure once it's reached regardless of job completion state, so a job needs to genuinely finish within that window. Size ttl generously enough for the job's expected runtime plus a safety margin, not exactly to the expected duration.
Yes, via peripheralsConfig.metastoreService pointing at a Dataproc Metastore service, which both ephemeral Serverless batches and persistent clusters can attach to. That's the standard way to keep a consistent table catalog across ephemeral and long-lived Spark workloads without duplicating metadata.
The builder requires name, runtimeConfig.version, and environmentConfig.executionConfig.serviceAccount to resolve before accepting the JSON as a valid Batch resource, the fields the Dataproc Serverless API needs to know what to run, which Spark runtime to use, and what identity to run the ephemeral infrastructure as.
spark.dynamicAllocation.minExecutors above 0 keeps some executor capacity warm from the start, useful for latency-sensitive batches, but every extra idle minExecutor is billed compute even during a slow startup phase where they're doing nothing yet.
A kmsKey set on executionConfig for CMEK encryption must be in the same region as the batch; cross-region key references fail at submission rather than silently falling back to Google-managed encryption.
The stagingBucket in executionConfig needs to already exist with appropriate permissions for the executionConfig.serviceAccount, a missing bucket fails the batch at the very start with an error that reads like a permissions issue even when it's really a missing-resource issue.
Was this tool helpful?
Disclaimer: This tool runs entirely in your browser. No data is sent to our servers. Always verify outputs before using them in production. AWS, Azure, and GCP are trademarks of their respective owners.