Build Kinesis Firehose delivery stream configurations with S3/Redshift destinations, Parquet conversion, and dynamic partitioning.
Build Kinesis Firehose delivery stream configurations with S3/Redshift destinations, Parquet conversion, and dynamic partitioning.
Required Fields
DeliveryStreamNameDeliveryStreamTypeExtendedS3DestinationConfiguration.BucketARNOutput will appear here...Build a Kinesis Firehose (Amazon Data Firehose) delivery stream with a Kinesis Data Stream source, DataFormatConversionConfiguration transforming incoming JSON into Parquet via a Glue Data Catalog schema reference, DynamicPartitioningConfiguration for content-based S3 prefixing, and an optional Lambda ProcessingConfiguration for custom record transformation before delivery. BufferingHints controls delivery latency versus file-size efficiency as a real tradeoff, Firehose delivers whichever threshold (SizeInMBs or IntervalInSeconds) is hit first, so a 128MB/300s config on a lower-throughput stream might deliver every 300 seconds on the time threshold long before accumulating 128MB, producing many small files, which hurts downstream query efficiency in exactly the way DataFormatConversionConfiguration's Parquet output was meant to avoid.
The builder validates that DeliveryStreamName, DeliveryStreamType, and ExtendedS3DestinationConfiguration.BucketARN all resolve before accepting the JSON as a valid CreateDeliveryStream request, the fields Firehose needs to know the stream's identity, source type, and destination; it can't verify the referenced Glue schema, KMS key, or Lambda processor ARNs are correctly configured and mutually compatible, those are only confirmed once real records start flowing through the stream.
Small-file proliferation from a time-threshold-dominated buffering pattern is a common, quiet cost and performance problem for downstream Athena/Redshift Spectrum queries, monitor actual delivered file sizes, not just the configured BufferingHints values, to catch this.
Always monitor the ErrorOutputPrefix location for a stream using DataFormatConversionConfiguration, schema-mismatched records land there without necessarily triggering an obvious alert, and can silently accumulate if the source starts producing malformed data.
When chaining a Lambda processor with a format-converting destination, tune both the processor's own buffering parameters and the destination's BufferingHints together, treating them as one combined latency/file-size budget rather than optimizing each in isolation.
Was this tool helpful?
Disclaimer: This tool runs entirely in your browser. No data is sent to our servers. Always verify outputs before using them in production. AWS, Azure, and GCP are trademarks of their respective owners.