Build Macie sensitive data discovery job configurations with S3 bucket scoping and custom data identifiers.
Build Macie sensitive data discovery job configurations with S3 bucket scoping and custom data identifiers.
Required Fields
NameJobTypeS3JobDefinition.BucketDefinitionsOutput will appear here...The builder validates that Name, JobType, and S3JobDefinition.BucketDefinitions resolve before treating the JSON as a valid CreateClassificationJob request, the minimum Macie needs to know what to name the job, whether it's a one-time or recurring scan, and which buckets are in scope; scoping filters, custom identifiers, and sampling are passed through as-is since their correctness depends on your actual data layout, not something the tool can validate.
Build a Macie sensitive-data discovery job that scans S3 buckets against both managed data identifiers (built-in patterns for SSNs, credit cards, credentials) and custom regex-based identifiers, scoped with SimpleScopeTerm filters on object key prefix, extension, or size. SamplingPercentage under 100 trades completeness for cost, since Macie bills per GB scanned, a 25% sample on a multi-terabyte bucket finds proportionally fewer sensitive-data hits but at a quarter of the scan cost, useful for a first-pass exploratory job before committing to a full scan.
ManagedDataIdentifierSelector: "ALL" runs every built-in identifier Macie has, including many irrelevant to your data (foreign national ID formats, for instance), which inflates both scan time and the false-positive review burden, most teams get better signal-to-noise scoping to a curated identifier list once they know what's actually in their data.
SamplingPercentage is a cost lever, not a security control, don't rely on a sampled job as your compliance evidence for 'we scanned this bucket for PII', a full-percentage run is what most audit frameworks actually expect as proof.
Custom data identifiers using overly broad regex (like a loose SSN pattern matching any 9-digit number) generate a flood of false positives that bury the genuine findings, test custom identifiers against a small sample bucket before wiring them into a production scheduled job.
It samples at the object level, a percentage of the eligible objects in scope get scanned in full, not a percentage of bytes within each object. This means a low sampling percentage on a bucket with a small number of very large files can still miss entire files' worth of sensitive data, sampling is most reliable for buckets with many small, relatively uniform objects.
By default a scheduled job re-evaluates the bucket each run according to InitialRun and its scope, which can mean re-scanning previously-scanned objects and re-incurring cost for them, not an incremental-only scan. For a large, slowly-growing bucket, tightening the Scoping filters (by key prefix or a date-based convention) is the practical way to control recurring cost rather than relying on the job to skip unchanged objects on its own.
Its extraction capability covers a defined set of supported file types (common text formats, several document and archive formats), it doesn't have generic support for arbitrary binary or proprietary formats. Scoping to specific OBJECT_EXTENSION values, as the example config does, is partly about controlling which files Macie can actually extract meaningful content from, not just about controlling cost.
Was this tool helpful?
Disclaimer: This tool runs entirely in your browser. No data is sent to our servers. Always verify outputs before using them in production. AWS, Azure, and GCP are trademarks of their respective owners.