A100 80GB bare metal on Clore.ai starts at $1.54 per GPU-hour, in blocks from 8 GPUs on terms from 30 days, in the USA, Japan and Slovenia, paid in BTC, USDT or USDC. The card is contracted as bare metal rather than listed per minute, and what the block buys is the memory line that separates it from the 40 GB part: 80 GB of HBM2e at 1,935 GB/s, enough to hold a 70B at INT8 on one board, or at FP16 across a sharded NVLink pair. Ampere, so BF16 and TF32 but no FP8 and no Transformer Engine.
The 40 GB and 80 GB A100 share a die, a clock and a TFLOPS number. Every real difference between them is memory: how much, and how fast it can be read. That is the whole case for paying more.
Seventy billion parameters at 8 bits is roughly 70 GB of weights. That number fits inside 80 GB and does not fit inside 40 GB, and no amount of clever offloading changes which side of the line a model falls on. Quantise the same model to INT4 and it drops to about 40 GB, which is where the smaller board becomes viable again, with almost no room left over.
Generating a token reads the full weight set once, so decode throughput tracks memory bandwidth almost linearly. The 80 GB part reads 25 percent faster than the 40 GB part for the same compute, which is the second reason to prefer it and the one people forget when they compare only capacity.
The 80 GB board partitions the same seven ways as the 40 GB one, but each instance is about 10 GB instead of about 5. That is the difference between a slice that holds a 1B model and a slice that holds an 8B at FP16, which changes what you can sell a tenant.
Read this table down the VRAM and bandwidth rows. The compute row barely moves between the two A100s, which is exactly the point. Listing counts are the CLORE marketplace on 21 September 2026.
specs from the NVIDIA A100, H200 and V100 datasheets · memory footprints at 1 byte per parameter for INT8, before KV cache · listing counts from the CLORE marketplace, 21 Sep 2026
A per-minute rate is whatever a host charges, and on 21 September 2026 no host had an A100 80GB listed. The published price is the bare-metal contract, from $1.54 per GPU-hour, and it moves with term length rather than with demand.
With nothing listed per minute, the sequence is a procurement flow rather than a click. It is still measured in days, not quarters.
Weights, then KV cache, then activations. If the total clears 40 GB you need this board; if it clears 80 GB you need more than one, or an H200.
The configurator takes a GPU count from 8 upward and a term from 30 days upward. The rate falls as the term lengthens and settles at $1.54 per GPU-hour from a year out.
One SKU covers both memory sizes in the configurator. If your model needs 80 GB, say so before the contract is signed rather than after the machines arrive.
Per-minute supply is hosts deciding to list. It was zero at the last snapshot, which is a fact about that instant, not a permanent state.
Not much, and that is the honest answer. Seventy gigabytes of weights against an 80 GB board leaves roughly 10 GB for the KV cache, the activations and the CUDA context, which at 70B scale buys you a few thousand tokens across a handful of concurrent requests rather than a long-context service. If you want real context headroom on one board, quantise further: the same model at INT4 lands near 40 GB and leaves half the card free. If you want both the precision and the context, you are looking at two boards.
Yes, with tensor or pipeline parallelism, and no, not as one flat 160 GB pool. NVLink is a fast interconnect, not a memory merger. The framework still shards the model across two devices and moves activations over the link, and 600 GB/s is fast enough that the sharding overhead stays modest on a two-way split. What you get is a 70B at FP16 that runs; what you do not get is the illusion of a single 160 GB GPU.
On memory-bound decode, close to the ratio: the H100 reads weights about 1.7 times faster, so tokens per second scale roughly with that. Compute-bound phases widen the gap further, because the H100 does 989 dense BF16 TFLOPS to the A100 312, and FP8 widens it again on models that support it. The counterweight is price. Compare on dollars per token on your own traffic rather than on either headline.
FP8 execution and the automatic per-layer scaling that goes with it. The Transformer Engine picks FP8 or BF16 per layer and keeps the loss scaling correct, which is where a large part of the H100 training advantage comes from. On an A100 you stay in BF16 or TF32 for training and reach for INT8 or INT4 weight quantisation when serving. Checkpoints published in FP8 will either need converting or will fall back, and FlashAttention-3, which targets Hopper, is not an option either.
The 600 GB/s figure is the SXM4 board in an HGX baseboard, where every GPU talks to every other GPU through NVSwitch. A PCIe A100 80GB has the same die and the same 80 GB but reaches its neighbour through an NVLink bridge across a card pair, or through PCIe if no bridge is fitted, which is a different number entirely. Check the listing before you plan a multi-GPU job around peer bandwidth, and ask the host if the machine description does not say.
At the 21 September 2026 snapshot zero servers carrying an A100 80GB were listed on the per-minute marketplace, so there was nothing to rent by the hour. Two routes remain. CLORE sells A100 bare metal in blocks of at least 8 GPUs on a 30-day minimum term at $1.54 to $2.28 per GPU-hour depending on length, in the USA, Japan and Slovenia, and the configurator lists a single A100 SKU with no 40 GB or 80 GB split, so confirm the memory with the team before you commit. Or take the A100 40GB, which did have 6 servers listed, if your model fits. The marketplace is the live source for both.
Memory footprints below are weights only, at 2 bytes per parameter for FP16 and 1 byte for INT8. Add somewhere between 10 and 30 percent for KV cache and activations before you decide anything fits.
One board, one model, a thin margin for context. Drop to INT4 if you need concurrency more than you need precision.
Read the guide →Two SXM4 boards at 600 GB/s peer bandwidth. The framework splits the model; the link carries activations between the halves.
Read the guide →Weights, gradients and optimiser state together are where 40 GB runs out and 80 GB does not, even on a model that serves comfortably on far less.
Read the guide →Weights only, before the KV cache: 2 bytes per parameter at FP16, 1 at INT8, half a byte at INT4. A row that says no is a model that will not load, not one that runs slowly.
Most of these assume more memory than a consumer card has and no FP8 path. That combination is what an A100 80GB is for.
List it on Clore.ai at your own price, or supply eight-card nodes to the bare-metal contract route that sets the top of that range. The hosting page covers both, from the one-command install onward.
Nothing was listed per minute at the last snapshot, so the route to this card is a bare-metal block: 8 GPUs or more, 30 days or more, $1.54 to $2.28 per GPU-hour depending on term.