Build Glue ETL job configurations with worker sizing, Spark UI, auto-scaling, job bookmarks, and custom Python modules.
Build Glue ETL job configurations with worker sizing, Spark UI, auto-scaling, job bookmarks, and custom Python modules.
Required Fields
NameRoleCommand.NameCommand.ScriptLocationOutput will appear here...Build a Glue Spark ETL job with worker sizing (G.2X workers, each roughly 8 vCPU/32GB, versus the smaller G.1X), job bookmarks for incremental processing across runs, and auto-scaling enabled alongside a fixed NumberOfWorkers, where auto-scaling actually overrides that fixed count as a maximum ceiling rather than the exact worker count once enabled. --job-bookmark-option job-bookmark-enable is what makes a job incremental (processing only new data since the last successful run) rather than reprocessing the entire source dataset every single run, but bookmarks are tied to the specific job and script logic, a change to the script's transformation logic without also resetting bookmarks can produce inconsistent results mixing old bookmark state with new logic.
Yes, but its meaning changes, once --enable-auto-scaling is set, NumberOfWorkers acts as the maximum number of workers Glue can scale up to, not a fixed worker count the job runs with the whole time. The job may run with fewer workers when the workload doesn't need the full capacity, and scale up toward NumberOfWorkers as needed, this is a meaningfully different behavior than the same field's meaning when auto-scaling is disabled (a genuinely fixed count).
The bookmark itself just tracks which source data has already been processed (by file/partition, generally), it has no awareness of script logic changes. If you change the transformation logic significantly, the next run using the existing bookmark only processes new data with the new logic, existing already-processed data isn't reprocessed with the updated logic unless you explicitly reset the job bookmark, which can leave historical data processed under old logic sitting alongside newly-processed data under new logic.
It enables writing Spark event logs to the configured S3 path (--spark-event-logs-path) on every run, which has a small ongoing storage cost, and a minor overhead for writing the logs, but it's generally low compared to the debugging value when you actually need to investigate a slow or failing job. It's reasonable to leave on by default for jobs where debugging via Spark UI has real value, but the accumulating S3 logs need their own lifecycle policy so they don't grow unbounded.
The builder validates that Name, Role, Command.Name, and Command.ScriptLocation all resolve before accepting the JSON as a valid CreateJob request, the fields Glue needs to know the job's identity, IAM execution role, and where its Spark/Python script lives; it can't verify the referenced script actually exists at that S3 location or that DefaultArguments match parameters the script genuinely expects, those are only surfaced when the job actually runs.
Understand that NumberOfWorkers changes meaning once auto-scaling is enabled, it becomes a ceiling, not a fixed count, sizing it based on peak expected load rather than typical load is the right mental model once auto-scaling is on.
A job bookmark reset needs to happen deliberately whenever a transformation logic change would produce meaningfully different output for already-processed data, don't assume bookmarks 'just work' across every code change without considering whether historical data needs reprocessing under the new logic.
Pin --additional-python-modules to specific versions, not open-ended constraints, an unpinned dependency in a Glue job can silently pick up a new, potentially breaking version whenever the underlying Glue runtime environment is refreshed.
Was this tool helpful?
Disclaimer: This tool runs entirely in your browser. No data is sent to our servers. Always verify outputs before using them in production. AWS, Azure, and GCP are trademarks of their respective owners.