The cheapest A40 on the table right now is $0.320 per GPU-hour at Vast.ai (spot), out of 6 offers from 4 providers. The A40 carries 48 GB of VRAM on NVIDIA's Ampere generation. That is peer/marketplace capacity — rented from other users' machines, and not directly comparable to a dedicated instance.
On-demand pricing for the same card spans $0.336 to $0.650 per GPU-hour — a 1.9× spread between the cheapest and the most expensive provider for identical silicon. That gap is the entire reason this table exists.
Cheapest by pricing model: on-demand $0.336 at Vast.ai · spot $0.320 at Vast.ai · reserved $0.380 at LeaderGPU.
| Provider | Tier | $/GPU-hr | Pricing | Node | Commit | In stock | |
|---|---|---|---|---|---|---|---|
| Vast.ai | marketplace | $0.320 | spot | 1× GPU | — | yes | rent |
| Vast.ai | marketplace | $0.336 | on-demand | 1× GPU | — | yes | rent |
| LeaderGPU | dedicated | $0.380 | reserved | 8× GPU | 1 mo | — | rent |
| Runpod | dedicated | $0.440 | on-demand | 1× GPU | — | yes | rent |
| Runpod | dedicated | $0.440 | spot | 1× GPU | — | yes | rent |
| Denvr Dataworks | dedicated | $0.650 | on-demand | 4× GPU | — | — | rent |
How much does an A40 cost per hour?
As of the latest scrape, $0.320 per GPU-hour at Vast.ai is the cheapest A40 offer across 4 tracked cloud providers. The cheapest on-demand rate is $0.336 per GPU-hour at Vast.ai, and the most expensive on-demand rate is $0.650 at Denvr Dataworks.
Which cloud is cheapest for the A40?
Vast.ai at $0.320 per GPU-hour on spot capacity. Prices move constantly, so this page is regenerated from the live scrape every 15 minutes. Note that marketplace tiers are peer hardware and are not directly comparable to dedicated capacity.
Is spot A40 capacity cheaper?
Yes — the cheapest spot A40 is $0.320 per GPU-hour at Vast.ai, -5% versus the cheapest on-demand rate of $0.336. Spot instances can be reclaimed by the provider at any time, so they suit checkpointed training and batch inference rather than long-lived services.
How many GB of VRAM does the A40 have?
48 GB, on the Ampere architecture. Multi-GPU instances multiply that: an 8× A40 node exposes 384 GB of GPU memory in total.