Log in A100 bare metal
A100 80GB · bare metal from $1.54/GPU-hour, blocks from 8 GPUs

The A100 80GB
holds a 70B.
Rented by the minute.

A100 80GB bare metal on Clore.ai starts at $1.54 per GPU-hour, in blocks from 8 GPUs on terms from 30 days, in the USA, Japan and Slovenia, paid in BTC, USDT or USDC. The card is contracted as bare metal rather than listed per minute, and what the block buys is the memory line that separates it from the 40 GB part: 80 GB of HBM2e at 1,935 GB/s, enough to hold a 70B at INT8 on one board, or at FP16 across a sharded NVLink pair. Ampere, so BF16 and TF32 but no FP8 and no Transformer Engine.

●Blocks from 8 GPUs ●Terms from 30 days ●USA, Japan, Slovenia ●BTC, USDT, USDC
80GB
HBM2e at 1,935 GB/s per board
3
Regions: USA, Japan and Slovenia
8
GPUs in the smallest bare-metal block
$1.54/GPU-h
Bare metal at the longest contract length

What 80 GB buys
that 40 GB does not.

The 40 GB and 80 GB A100 share a die, a clock and a TFLOPS number. Every real difference between them is memory: how much, and how fast it can be read. That is the whole case for paying more.

A 70B at INT8 loads on one board

Seventy billion parameters at 8 bits is roughly 70 GB of weights. That number fits inside 80 GB and does not fit inside 40 GB, and no amount of clever offloading changes which side of the line a model falls on. Quantise the same model to INT4 and it drops to about 40 GB, which is where the smaller board becomes viable again, with almost no room left over.

Llama 3 70B at INT8 ~70 GB of weights

1,935 GB/s, and why decode cares

Generating a token reads the full weight set once, so decode throughput tracks memory bandwidth almost linearly. The 80 GB part reads 25 percent faster than the 40 GB part for the same compute, which is the second reason to prefer it and the one people forget when they compare only capacity.

Against the 40 GB board +25% bandwidth

MIG, at ten gigabytes a slice

The 80 GB board partitions the same seven ways as the 40 GB one, but each instance is about 10 GB instead of about 5. That is the difference between a slice that holds a 1B model and a slice that holds an 8B at FP16, which changes what you can sell a tenant.

MIG instances per board up to 7 × ~10 GB

Memory first,
FLOPS second.

Read this table down the VRAM and bandwidth rows. The compute row barely moves between the two A100s, which is exactly the point. Listing counts are the CLORE marketplace on 21 September 2026.

A100 80GB A100 40GB H200 141GB Tesla V100
Architecture Ampere GA100 Ampere GA100 Hopper Volta GV100
VRAM 80 GB HBM2e 40 GB HBM2e 141 GB HBM3e 32 GB HBM2
Memory bandwidth 1,935 GB/s 1,555 GB/s 4,800 GB/s 900 GB/s
FP16 / BF16 (dense) 312 TFLOPS 312 TFLOPS 989 TFLOPS 125 TFLOPS
Llama 3 70B at INT8 (~70 GB) fits on one board does not fit fits, with room does not fit
MIG instances up to 7 up to 7 up to 7 none
Servers listed 21 Sep 2026 0 6 0 186

specs from the NVIDIA A100, H200 and V100 datasheets · memory footprints at 1 byte per parameter for INT8, before KV cache · listing counts from the CLORE marketplace, 21 Sep 2026

The hourly price is the contract rate,
not a listing.

A per-minute rate is whatever a host charges, and on 21 September 2026 no host had an A100 80GB listed. The published price is the bare-metal contract, from $1.54 per GPU-hour, and it moves with term length rather than with demand.

Per-minute marketplace

— no listings
0 servers carrying this card, as of 21 Sep 2026
  • Supply here is whatever hosts choose to list
  • It can reappear without notice, so check the marketplace
  • The A100 40GB did have 6 servers listed at the same snapshot
  • Billed per minute when it is available
Check the marketplace
AVAILABLE ROUTE

Bare metal

$1.54 – $2.28 / GPU-hour
$2.28 at 30 days, $1.98 at 90, $1.81 at 180, $1.54 from 360
  • Minimum 8 GPUs, minimum 30-day term
  • USA, Japan and Slovenia
  • The configurator sells one A100 SKU with no memory split, so confirm 80 GB before signing
  • Paid in BTC, USDT or USDC
Configure a block
Pay with
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

How you actually get one.

With nothing listed per minute, the sequence is a procurement flow rather than a click. It is still measured in days, not quarters.

01 / SIZE

Count the gigabytes first

Weights, then KV cache, then activations. If the total clears 40 GB you need this board; if it clears 80 GB you need more than one, or an H200.

02 / CONFIGURE

Pick a block and a term

The configurator takes a GPU count from 8 upward and a term from 30 days upward. The rate falls as the term lengthens and settles at $1.54 per GPU-hour from a year out.

8 GPUs · 30 days → $2.28 / GPU-hour
03 / CONFIRM

Ask which A100

One SKU covers both memory sizes in the configurator. If your model needs 80 GB, say so before the contract is signed rather than after the machines arrive.

04 / OR WAIT

Watch the marketplace

Per-minute supply is hosts deciding to list. It was zero at the last snapshot, which is a fact about that instant, not a permanent state.

Questions about 80 GB.

Llama 3 70B at INT8 is about 70 GB. What context length does 80 GB leave after the weights?

Not much, and that is the honest answer. Seventy gigabytes of weights against an 80 GB board leaves roughly 10 GB for the KV cache, the activations and the CUDA context, which at 70B scale buys you a few thousand tokens across a handful of concurrent requests rather than a long-context service. If you want real context headroom on one board, quantise further: the same model at INT4 lands near 40 GB and leaves half the card free. If you want both the precision and the context, you are looking at two boards.

Two A100 80GB over NVLink at 600 GB/s. Does 70B at FP16 (about 140 GB) actually work?

Yes, with tensor or pipeline parallelism, and no, not as one flat 160 GB pool. NVLink is a fast interconnect, not a memory merger. The framework still shards the model across two devices and moves activations over the link, and 600 GB/s is fast enough that the sharding overhead stays modest on a two-way split. What you get is a 70B at FP16 that runs; what you do not get is the illusion of a single 160 GB GPU.

This card runs at 1,935 GB/s against the H100 3,350. How much serving throughput does that cost me?

On memory-bound decode, close to the ratio: the H100 reads weights about 1.7 times faster, so tokens per second scale roughly with that. Compute-bound phases widen the gap further, because the H100 does 989 dense BF16 TFLOPS to the A100 312, and FP8 widens it again on models that support it. The counterweight is price. Compare on dollars per token on your own traffic rather than on either headline.

Ampere has no Transformer Engine. What do I actually lose against an H100 on the same model?

FP8 execution and the automatic per-layer scaling that goes with it. The Transformer Engine picks FP8 or BF16 per layer and keeps the loss scaling correct, which is where a large part of the H100 training advantage comes from. On an A100 you stay in BF16 or TF32 for training and reach for INT8 or INT4 weight quantisation when serving. Checkpoints published in FP8 will either need converting or will fall back, and FlashAttention-3, which targets Hopper, is not an option either.

SXM4 or PCIe: which A100 80GB am I renting, and does the 600 GB/s NVLink figure apply?

The 600 GB/s figure is the SXM4 board in an HGX baseboard, where every GPU talks to every other GPU through NVSwitch. A PCIe A100 80GB has the same die and the same 80 GB but reaches its neighbour through an NVLink bridge across a card pair, or through PCIe if no bridge is fitted, which is a different number entirely. Check the listing before you plan a multi-GPU job around peer bandwidth, and ask the host if the machine description does not say.

Nothing is listed per minute today. What are my actual options?

At the 21 September 2026 snapshot zero servers carrying an A100 80GB were listed on the per-minute marketplace, so there was nothing to rent by the hour. Two routes remain. CLORE sells A100 bare metal in blocks of at least 8 GPUs on a 30-day minimum term at $1.54 to $2.28 per GPU-hour depending on length, in the USA, Japan and Slovenia, and the configurator lists a single A100 SKU with no 40 GB or 80 GB split, so confirm the memory with the team before you commit. Or take the A100 40GB, which did have 6 servers listed, if your model fits. The marketplace is the live source for both.

Three jobs that need the bigger board.

Memory footprints below are weights only, at 2 bytes per parameter for FP16 and 1 byte for INT8. Add somewhere between 10 and 30 percent for KV cache and activations before you decide anything fits.

Serving a quantised 70B
vLLM or TensorRT-LLM, INT8 weights
~70 GB of 80 GB used

One board, one model, a thin margin for context. Drop to INT4 if you need concurrency more than you need precision.

Read the guide →
70B at FP16 across a pair
Tensor parallel over NVLink 3
~140 GB, sharded not pooled

Two SXM4 boards at 600 GB/s peer bandwidth. The framework splits the model; the link carries activations between the halves.

Read the guide →
Full fine-tune of an 8B
FSDP or ZeRO-3, BF16, sharded optimiser
312 dense BF16 TFLOPS

Weights, gradients and optimiser state together are where 40 GB runs out and 80 GB does not, even on a model that serves comfortably on far less.

Read the guide →

What fits on one board.

Weights only, before the KV cache: 2 bytes per parameter at FP16, 1 at INT8, half a byte at INT4. A row that says no is a model that will not load, not one that runs slowly.

GPU
VRAM
Mem BW (GB/s)
8B FP16 (~16 GB)
70B INT4 (~40 GB)
70B INT8 (~70 GB)
Listed 21 Sep 2026
Bare metal $/GPU-h
A100 80GB / this page
80 GB HBM2e
1,935
yes
yes
yes, tight
0 servers
$1.54–$2.28
A100 40GB
40 GB HBM2e
1,555
yes
no headroom
no
6 servers
$1.54–$2.28
H200 141GB
141 GB HBM3e
4,800
yes
yes
yes
0 servers
$2.09–$4.18
RTX 5090
32 GB GDDR7
1,792
yes
no
no
260 servers
$0.64–$1.06

Guides for a board this size.

Most of these assume more memory than a consumer card has and no FP8 path. That combination is what an A100 80GB is for.

Training
DeepSpeed multi-GPU training
ZeRO-2/3 training across multiple cards.
Training
LLM fine-tuning
LoRA / QLoRA fine-tuning workflow.
Language Models
vLLM serving
High-throughput LLM serving with PagedAttention.
Language Models
Llama 3.3 on CLORE.AI
Run Meta's flagship Llama-3.3 on your rented card.
Training
HF Transformers training
Train and fine-tune with the Trainer API.
Language Models
Mistral / Mixtral
Run Mistral 7B and Mixtral 8x7B / 8x22B.
Advanced
Multi-GPU setup
Configure NVLink, NCCL, and distributed training.
See all guides →

If this one is not available.

A100 40GB
Half the memory, 6 servers listed
Compare →
H100 80GB
Same capacity, FP8 and 3,350 GB/s
Compare →
H200 141GB
When 80 GB is still not enough
Compare →

Eighty gigabytes,
on contract.

Nothing was listed per minute at the last snapshot, so the route to this card is a bare-metal block: 8 GPUs or more, 30 days or more, $1.54 to $2.28 per GPU-hour depending on term.