Compare HPC compute offerings across AWS ParallelCluster, Azure CycleCloud, GCP HPC Toolkit, and OCI HPC.
Output will appear here...A comparison of HPC-oriented compute, interconnect, and job scheduling across AWS (ParallelCluster/Batch/EFA), Azure (CycleCloud/InfiniBand HB-HC-ND series), GCP (HPC Toolkit/Cloud Batch), and OCI (HPC bare-metal shapes/RDMA cluster networking), and the interconnect technology is the detail that actually determines whether a tightly-coupled MPI workload will scale: AWS's EFA delivers up to 3,200 Gbps on P5 instances via a non-blocking fabric, Azure uses InfiniBand HDR/NDR at 200-400 Gbps for its HB/ND series, GCP uses its Jupiter fabric plus GPUDirect-TCPXO specifically for A3 GPU instances, and OCI uses 100 Gbps RoCE v2 RDMA cluster networking, genuinely different numbers that matter for a workload bottlenecked on inter-node communication rather than raw compute.
The comparison table is a static, hand-maintained dataset of feature rows grouped by category (overview, compute, networking, storage, pricing) with free-text search across all fields; it's a reference snapshot, not a live specs or pricing feed, since HPC instance families and interconnect generations change relatively often, verify current generation and bandwidth figures directly with the provider before a procurement decision.
Interconnect bandwidth numbers in vendor marketing are peak, theoretical figures, always validate actual achievable bandwidth and latency for your specific instance count and placement group configuration with a real benchmark (like OSU MPI benchmarks) before sizing a production HPC cluster around the headline number.
HPC-optimized parallel filesystems (Lustre-based options especially) have real minimum provisioned throughput/capacity tiers that get expensive fast, size storage to the workload's actual I/O pattern rather than defaulting to the largest tier out of caution.
A checkpoint-restart strategy is close to mandatory for using spot/preemptible capacity on any long-running tightly-coupled HPC job, without it, a single interruption late in a multi-day MPI run can waste the entire run's compute rather than just the interrupted node's share.
Was this tool helpful?
Disclaimer: This tool runs entirely in your browser. No data is sent to our servers. Always verify outputs before using them in production. AWS, Azure, and GCP are trademarks of their respective owners.