Build Azure ML compute instance and AmlCompute cluster configs with GPU VMs, autoscaling, idle shutdown, and VNet integration.
Build Azure ML compute instance and AmlCompute cluster configs with GPU VMs, autoscaling, idle shutdown, VNet integration, and schedules.
Required Fields
workspaceNameresourceGroupcomputeTargetscomputeTargets[0].namecomputeTargets[0].computeTypecomputeTargets[0].properties.vmSizeOutput will appear here...The builder validates that workspaceName, resourceGroup, computeTargets, and the first target's name/computeType/properties.vmSize all resolve before accepting the JSON as a valid combined compute target definition, the fields Azure ML needs to provision either compute type into the workspace; it can't verify the referenced subnet has sufficient IP capacity for the requested node count or that the specified GPU VM size is actually available in the target region, those are only confirmed against the live Azure ML workspace API.
Build Azure ML compute targets combining ComputeInstance (a per-user, always-named development VM with idleTimeBeforeShutdown and cron-based start/stop schedules) and AmlCompute (an autoscaling cluster with minNodeCount/maxNodeCount for shared training workloads, supporting vmPriority LowPriority for genuinely cheaper but preemptible GPU capacity). minNodeCount: 0 on an AmlCompute cluster means it scales fully to zero nodes (and zero compute cost) when idle, which is the entire economic case for using AmlCompute over a dedicated always-on VM for training, but it also means the first job submitted after idle time pays a node-provisioning cold-start delay before training actually begins.
Always pair a ComputeInstance with both idleTimeBeforeShutdown and explicit cron schedules, relying on just one leaves a gap, idle timeout alone doesn't stop a VM someone leaves running with an active-but-unused connection, and schedules alone don't help if someone's actively using it outside the expected hours.
LowPriority is the right default for fault-tolerant, checkpointed training workloads given the meaningful cost savings, but verify your training code actually checkpoints before assuming the cost savings are risk-free.
minNodeCount: 0's cold-start delay is a real UX cost for anyone waiting on a training job to start, for a team running frequent, time-sensitive training jobs, weigh a small non-zero minNodeCount against the idle cost of keeping that baseline capacity always warm.
No, when scaled to zero nodes, submitting a new job requires Azure to provision fresh compute nodes before the job can actually start running, this cold-start provisioning delay (often several minutes for GPU SKUs) is the tradeoff for not paying for idle capacity. For a workload needing consistently fast job start times, keeping minNodeCount above zero avoids this delay at the cost of paying for that baseline capacity continuously.
Azure can reclaim LowPriority (spot-equivalent) compute capacity at any time when it's needed for higher-priority workloads, meaning a training job can be interrupted mid-run. This is a reasonable tradeoff for jobs that checkpoint their progress and can resume, but a long training job without checkpointing can lose significant progress on preemption, size the risk against your job's actual checkpointing behavior, not just the cost savings alone.
No, a ComputeInstance is fundamentally a single-user development VM, tied to one user's identity for Jupyter/VS Code-based development work. For shared, multi-user or scaled training workloads, AmlCompute clusters are the appropriate compute type instead, they're not interchangeable despite both being 'Azure ML compute'.
Was this tool helpful?
Disclaimer: This tool runs entirely in your browser. No data is sent to our servers. Always verify outputs before using them in production. AWS, Azure, and GCP are trademarks of their respective owners.