Compare ETL services (Glue, Data Factory, Dataproc, OCI Data Integration) across clouds.
Output will appear here...The comparison table is a static, hand-maintained dataset of feature rows grouped by category (overview, processing, governance, pricing) with free-text search across all fields; it's a reference snapshot, not a live specs feed, verify current format support and pricing directly with the provider before an ETL platform migration decision.
A comparison of managed ETL services across AWS Glue, Azure Data Factory/Synapse Pipelines, GCP (Dataproc/Dataform/Dataflow), and OCI (Data Integration/Data Flow), and the governance/cataloging layer differs meaningfully in maturity: AWS Glue Data Catalog is a Hive-compatible centralized metastore tightly coupled to the ETL engine itself, Azure routes cataloging through the separate, broader Microsoft Purview product, GCP splits catalog responsibility across Dataplex/Data Catalog, and OCI Data Catalog handles metadata harvesting and glossary management as its own distinct service, meaning the catalog and ETL engine aren't always the same product, which affects how tightly schema discovery integrates with your actual transform jobs.
A catalog-and-ETL-engine coupling that's convenient on AWS (Glue Data Catalog feeding Glue jobs directly) doesn't map one-to-one onto Azure's separate Purview product, budget integration work explicitly when porting a Glue-catalog-dependent pipeline to Azure.
GCP's three-way split between Dataproc/Dataform/Dataflow means 'GCP ETL' isn't one tool, picking the wrong one for your transform style (e.g., using Dataflow for what's really a simple SQL warehouse transform better suited to Dataform) adds unnecessary complexity.
Format support gaps (like Delta Lake's absence from OCI's listed formats) are easy to miss when comparing high-level feature checklists, verify the specific format/connector list for your actual pipeline's dependencies, not just the general ETL capability.
It's Hive-metastore-compatible and primarily built to serve Glue ETL jobs and other AWS analytics services (Athena, Redshift Spectrum) that need to read the same catalog, more tightly scoped than Azure's Purview, which is positioned as a broader, standalone data governance and cataloging product spanning more than just ETL job metadata. If your need is comprehensive data governance across many systems, not just ETL job metadata, Purview's broader scope may be a better fit even outside a primarily-Azure environment.
AWS Glue, Azure Data Factory, and GCP all list Delta Lake support in their supported formats; OCI's format list in this comparison doesn't explicitly include Delta Lake, focusing instead on JSON, CSV, Parquet, ORC, Avro, and Oracle-native formats. If Delta Lake is a hard requirement for a lakehouse architecture and OCI is a target platform, verify current Delta Lake support directly rather than assuming parity with the other three.
They serve different transform styles: Dataform is a SQL-based transformation tool (closer to dbt) for in-warehouse BigQuery transforms, Dataproc is managed Spark/Hadoop for traditional big-data ETL jobs, and Dataflow is Apache Beam-based for unified batch and streaming pipelines. Choosing among them depends on whether your transforms are SQL-native warehouse transforms, traditional Spark jobs, or need unified batch/stream processing, they're not three interchangeable options for the same job.
Was this tool helpful?
Disclaimer: This tool runs entirely in your browser. No data is sent to our servers. Always verify outputs before using them in production. AWS, Azure, and GCP are trademarks of their respective owners.