Build Purview data source scan configs with schedule, scope filters, classification rules, and managed identity authentication.
Build Purview data source scan configs with schedule, scope filters, classification rules, collections, and managed identity authentication.
Required Fields
accountNamedataSourcesdataSources[0].namedataSources[0].kindscansscans[0].namescans[0].dataSourceNameOutput will appear here...Build a Microsoft Purview scan configuration linking dataSources (Azure SQL, ADLS Gen2, Synapse) to scans with ManagedIdentity authentication, schedule-based scanLevel Incremental runs, scopeFilters narrowing what's actually scanned within a large data source, and custom classificationRules using regex patterns to detect organization-specific sensitive data types the built-in classifiers don't cover. scanLevel Incremental only re-scans data that's changed since the last successful scan, dramatically reducing scan time and cost for a large, mostly-static data source, but it depends on the source correctly exposing change-detection metadata Purview can use, a source or configuration where that metadata isn't reliably available effectively forces every 'incremental' scan to behave like a full scan anyway, silently losing the expected time/cost benefit.
In most cases yes, since it only processes data changed since the last successful scan, but this depends on the data source reliably exposing change-detection metadata Purview can use to identify what's actually new or modified. If that metadata isn't consistently available for a given source or configuration, an 'incremental' scan can end up behaving like a full scan every time, silently losing the expected speed and cost benefit without an obvious error indicating why.
Yes, a regex pattern like EMP-[0-9]{6} matches any string fitting that shape, including a coincidental match that isn't actually an employee ID, minimumMatchCount and confidenceLevel settings help tune sensitivity, but any regex-based custom classifier carries some false-positive risk proportional to how generic or coincidentally-matchable the pattern is, test against a representative data sample before relying on the classification results for compliance-sensitive decisions.
No, scopeFilters affects what a scan actually processes going forward, assets from those excluded schemas already cataloged from a prior scan (before the exclusion was added) generally remain in the catalog until an explicit cleanup, adding an exclusion filter narrows future scan scope, it doesn't automatically retroactively remove previously-cataloged data matching the newly-excluded criteria.
The builder validates that accountName, dataSources, the first data source's name/kind, scans, and the first scan's name/dataSourceName all resolve before accepting the JSON as a valid combined data source and scan configuration; it can't verify the managed identity actually has the required permissions on each target data source, or that custom classification regex patterns are syntactically valid, those are only confirmed when a real scan actually runs.
Verify Incremental scans are actually behaving incrementally (check scan duration/volume trends over time) rather than assuming the setting guarantees the expected performance benefit, a source with unreliable change-detection metadata can silently degrade to full-scan behavior every time.
Test custom regex-based classification rules against a representative sample of real data before trusting the results for a compliance-sensitive use case, an overly generic pattern produces false positives that dilute trust in the classification results.
Design the collections hierarchy deliberately around how governance policies actually need to be applied (by sensitivity, by team, by data domain), retrofitting a meaningful collection structure after assets are already scattered across a flat or poorly-organized hierarchy is more work than designing it upfront.
Was this tool helpful?
Disclaimer: This tool runs entirely in your browser. No data is sent to our servers. Always verify outputs before using them in production. AWS, Azure, and GCP are trademarks of their respective owners.