Build Glue crawler configurations with S3/JDBC targets, schema change policies, recrawl behavior, and Lake Formation integration.
Build Glue crawler configurations with S3/JDBC targets, schema change policies, recrawl behavior, and Lake Formation integration.
Required Fields
NameRoleDatabaseNameTargetsOutput will appear here...Build a Glue crawler that scans S3 (or JDBC/DynamoDB) targets and populates the Glue Data Catalog, with SchemaChangePolicy controlling what happens when a re-crawl finds a schema that's drifted from the catalog's existing definition, UPDATE_IN_DATABASE for updates versus LOG (not DELETE) for removed columns/tables by default, meaning a source that deletes data doesn't silently delete the corresponding catalog entries, it just logs the discrepancy, requiring an explicit DeleteBehavior change if you actually want removed sources reflected as removed in the catalog. RecrawlBehavior set to CRAWL_NEW_FOLDERS_ONLY is a meaningful cost and time optimization for an append-only, partition-growing dataset (like daily-partitioned event data), skipping a full re-scan of already-cataloged historical partitions on every scheduled run.
The builder validates that Name, Role, DatabaseName, and Targets all resolve before accepting the JSON as a valid CreateCrawler request, the fields Glue needs to know the crawler's identity, its IAM execution role, the target catalog database, and what to scan; it can't verify the referenced S3 paths are actually accessible by the given role or that JDBC connection details are correct, those are only confirmed on the crawler's first actual run.
The default LOG delete behavior is a safety feature, not an oversight, don't assume a crawler will keep your catalog perfectly in sync with actual data deletions unless you've deliberately opted into DELETE_FROM_DATABASE and understand the risk of a transient access issue being mistaken for real deletion.
CRAWL_NEW_FOLDERS_ONLY is the right optimization for append-only partitioned data but is actively wrong for a dataset where historical partitions can be updated in place, verify your data's actual mutation pattern before assuming this setting is safe to enable.
Event-based crawling via EventQueueArn adds real infrastructure (an SQS queue, S3 event notification configuration) that needs its own monitoring, a silently-broken event notification pipeline means new data stops appearing in the catalog promptly with no obvious crawler-level error.
Was this tool helpful?
Disclaimer: This tool runs entirely in your browser. No data is sent to our servers. Always verify outputs before using them in production. AWS, Azure, and GCP are trademarks of their respective owners.