The NVIDIA A40 is available on Clore.ai as bare metal from $0.53 per GPU-hour, in blocks from 8 GPUs on terms from 30 days, hosted in the USA, France or Japan and payable in BTC, USDT or USDC. Each card is the same Ampere GA102 silicon as the RTX A6000, with 48 GB of GDDR6 and ECC, 696 GB/s of bandwidth and NVLink in pairs, built without a fan of its own to sit in a chassis that supplies the airflow.
Everything that distinguishes the A40 from its workstation twin is physical rather than computational: how it is cooled, how it is connected, and where that lets it live.
The A40 has no fan. It is designed for a chassis that forces air across it, which is why it appears in rack servers and not in towers. That single property decides who can host one, and it is the reason A40 supply behaves differently from GeForce supply on this marketplace.
Two bridged A40s talk to each other at roughly 112.5 GB/s, which beats PCIe for tensor and pipeline parallelism. They do not become one 96 GB card. Splitting a model larger than 48 GB across the pair is your framework's job; the interconnect only makes it cheaper.
Forty-eight gigabytes is enough for a 70B model at INT4 or a 32B model at INT8 with cache to spare. What is not on the menu is FP8, which arrived with Ada, and MIG, which this card does not have despite being a datacenter part.
The A40 against its workstation twin, the smaller A5000, and the HBM part people reach for when bandwidth is the constraint. Tensor figures are dense throughout.
none of these NVLink links pools memory · only the A100 in this table supports MIG
A40 capacity is sold on fixed terms from eight GPUs up, and the only variable that moves the price is how long you commit for. Below are the two ends of that ladder. There is no per-minute column here because nothing with an A40 in it was listed on 21 September 2026, and inventing a starting price for capacity that was not there is exactly the habit this page was rewritten to remove.
The A40 and the A6000 do the same arithmetic. Which one you want comes down to four questions, in this order.
A 70B model at INT4 is about 40 GB and fits with cache to spare. At FP16 it is roughly 140 GB and does not fit on one card, or on a bridged pair.
If generation speed is your constraint rather than capacity, the extra memory bandwidth of an A100 will do more for you than the extra gigabytes here.
With no per-minute listings for this card, short jobs point at the A6000 or the L40S. A month or more is what the A40 rate card is built for.
Eight GPUs is the floor and thirty days the shortest term. Bitcoin, USDT and USDC settle bare-metal orders.
Cooling and bandwidth, and that is close to the whole list. Both are Ampere GA102 with 10,752 CUDA cores and 48 GB of GDDR6 with ECC at 300 W. The A6000 runs 768 GB/s, the A40 696, so the A40 gives up about nine percent of the memory bandwidth. In exchange it is built as a passively cooled datacenter card that expects chassis airflow rather than carrying its own. On CLORE the practical difference is price: bare-metal A40 runs $0.53 to $0.75 per GPU-hour against $0.67 to $0.95 for the A6000.
It means they are rack servers, not desktops. A card with no fan of its own needs a chassis that pushes air through it, which rules out the tower under someone's desk and limits A40 supply to operators with real server hardware. For a renter that is mostly good news, since the population skews toward datacenter deployments. It also explains the availability picture on this page: fewer people can host this card than can host a GeForce part.
It does not, and the surprise is the point worth dwelling on: this is a datacenter card, and datacenter cards are exactly where people expect to find the feature. Ampere shipped it on two products, the A100 and the A30. The A40 is neither. It runs GA102, the die it shares with the RTX A6000, and GA102 was never given partitioning. NVIDIA's supported-GPU list, checked on 21 September 2026, is the reference. So an A40 rental is always the whole card, and a workload that genuinely needs isolated instances belongs on an A100 node.
On a single-tenant rental, nothing you can point at, and it is worth saying so rather than dressing it up. The features that make this card attractive to a virtualization platform are about carving one GPU between many desktops, and a rented server hands you the whole card. What does carry over is the design intent behind it: ECC memory, passive cooling and a duty cycle meant for machines that stay on, which is a reasonable thing to want from hardware you are renting for a month.
All of the ones that fit, which is the honest answer. At roughly 0.5 GB per billion parameters a 70B model at INT4 is about 40 GB, leaving 8 GB for the KV cache. Dividing 696 GB/s by 40 GB gives about 17 full weight passes per second as a ceiling before the cache and kernel overheads take their cut. FP8 is not an option here at all, since Ampere has no FP8 tensor cores, and FP16 at roughly 140 GB is far outside 48 GB. If bandwidth is your binding constraint rather than capacity, an A100 at 1,555 GB/s is the card that fixes it.
No, and this is the most common misreading of the spec. NVLink on the A40 is a peer-to-peer link between two cards at roughly 112.5 GB/s, useful for tensor and pipeline parallelism because it moves data between cards faster than PCIe. It does not merge two 48 GB cards into one 96 GB address space, and no NVLink implementation on any of these Ampere parts does. A model that needs more than 48 GB has to be split across the two cards by your framework, not by the interconnect.
Capacity from parameter counts, ceilings from 696 GB/s, and one card about what the interconnect does not do. None of these are benchmark results.
Half a gigabyte per billion parameters gets a 70B model onto one A40. The same model at FP16 is roughly 140 GB and is not a single-card proposition anywhere in this tier.
Read the guide →Every generated token reads the weights once, so this is the arithmetic ceiling on a 70B INT4 deployment before cache traffic and kernel overheads take their share.
Read the guide →The link is faster than PCIe for parallel work across cards. It is not a memory pool, so a model over 48 GB still has to be partitioned by the framework.
Read the guide →Same capacity, four different answers on bandwidth, interconnect and low-precision support. The last column is what a block of eight costs per GPU-hour on a contract.
NVLink here is a card-to-card link, never a memory pool · none of the four supports MIG · rates are the 30-day to 360-day band for a minimum of 8 GPUs
Walkthroughs on docs.clore.ai. Everything here assumes capacity rather than bandwidth is your advantage, and none of it assumes FP8.
Put them on Clore.ai and earn up to ~$500/mo per card, whether you list per minute or supply bare-metal contracts. You set the price and are paid for every rented minute.
The A40 is the cheapest 48 GB card CLORE offers on contract, at $0.53 to $0.75 per GPU-hour depending on the term. Every number on this page is dated and sourced.