Build EMR Serverless Spark/Hive job configurations with driver/executor sizing, monitoring, and Glue catalog integration.
Build EMR Serverless Spark/Hive job configurations with driver/executor sizing, monitoring, and Glue catalog integration.
Required Fields
ApplicationIdExecutionRoleArnJobDriverOutput will appear here...Build an EMR Serverless Spark job run referencing a pre-created Application, with SparkSubmitParameters tuning executor/driver sizing and dynamic allocation bounds directly in the spark-submit argument string rather than through separate structured fields, plus ConfigurationOverrides for logging and Glue Data Catalog integration via the AWSGlueDataCatalogHiveClientFactory Hive metastore client. ExecutionTimeoutMinutes is a hard kill switch, a job still legitimately running past this timeout gets forcibly terminated regardless of progress, so it needs to be sized with real margin above the job's expected worst-case runtime, not just its typical runtime, or an occasional larger-than-usual data volume causes a false job failure rather than a slower-but-successful completion.
The builder validates that ApplicationId, ExecutionRoleArn, and JobDriver all resolve before accepting the JSON as a valid StartJobRun request, the fields EMR Serverless needs to know which pre-created application runs the job, under what IAM role, and what to actually execute; it can't verify the referenced ApplicationId exists, has sufficient configured capacity for the requested SparkSubmitParameters, or that the entry point script is valid, those are only confirmed once the job run is actually submitted.
Size ExecutionTimeoutMinutes generously above typical runtime, not tightly against it, a hard timeout kill on an otherwise-successful-but-slower-than-usual run is a self-inflicted false failure that's easy to avoid with reasonable margin.
Check the target Application's own configured maximum capacity before tuning a job's SparkSubmitParameters resource requests upward, a job can be correctly configured on its own terms and still fail to get the requested executors if it exceeds the application-level capacity ceiling.
The AWSGlueDataCatalogHiveClientFactory Hive metastore integration is what lets Spark SQL/Hive queries in the job resolve table names against the Glue Data Catalog, omitting it when the job actually depends on catalog-registered tables produces table-not-found errors that look like a data problem but are actually a metastore configuration gap.
Was this tool helpful?
Disclaimer: This tool runs entirely in your browser. No data is sent to our servers. Always verify outputs before using them in production. AWS, Azure, and GCP are trademarks of their respective owners.