Log in Rent H100
H100 80GB · bare metal · 8 cards, 30-day floor

Rent an H100 80GB.
70B fits at FP8,
and at FP16.

Clore.ai supplies the H100 80GB as a bare-metal datacenter contract: 8 cards minimum, 30 days minimum, quoted between $1.81 and $2.62 per GPU-hour depending on region and term as of 21 Sep 2026. Alongside it, per-minute listings from $0.99 per GPU-hour appear on the marketplace, priced by the hosts who run them. Size the node in the configurator and the quote comes back for the exact shape you asked for.

●8 GPUs minimum ●30-day minimum term ●Five countries ●Per-minute listings from $0.99/GPU-hour
80GB
HBM3 on the SXM5 module
3,350GB/s
Memory bandwidth, SXM5
8
Cards in the smallest contract
$1.81/GPU-hr
Lowest term tier, 21 Sep 2026

What 80 GB of HBM3
is actually bought for.

Three jobs justify the price of this card. All three are capacity arguments before they are speed arguments, which is why the memory figure matters more than the FLOPS figure.

FP8 and the Transformer Engine

Hopper introduced FP8 together with the Transformer Engine, which carries per-tensor scaling so a layer can run in eight bits without the numerics collapsing. The consequence is capacity. Llama 3 70B quantised to FP8 is roughly 70 GB of weights and sits inside an 80 GB card with the remainder left for the KV cache. The same model at FP16 is about 140 GB and needs a second card.

Llama 3 70B, FP8 weights ~70 GB

SXM5 and PCIe are different products

Both are sold as "H100 80GB". The SXM5 module carries 16,896 CUDA cores and 3,350 GB/s of HBM3 at 700 W. The PCIe card carries 14,592 cores and roughly 2,000 GB/s of HBM2e. On a memory-bound decode that gap lands directly in tokens per second, so read the module type before you read the price.

Bandwidth, SXM5 over PCIe about 1.7×

Seven tenants on one card

NVIDIA's MIG supported-GPUs table lists H100-SXM5 and H100-PCIE at both 80 GB and 94 GB with up to seven instances each. Each slice gets its own memory and its own SMs, which is hardware isolation rather than time slicing. It is a property of the silicon, not something a marketplace layers on top.

MIG instances per card up to 7

Two cards,
one product name.

Almost every H100 listing anywhere says "H100 80GB" and stops. Figures below come from NVIDIA's Hopper architecture resources and NVIDIA's MIG supported-GPUs table. Tensor throughput is quoted dense, never with sparsity.

H100 SXM5 H100 PCIe A100 80GB H200 SXM5
Architecture Hopper GH100 Hopper GH100 Ampere GA100 Hopper GH100
CUDA cores 16,896 14,592 6,912 16,896
VRAM 80 GB HBM3 80 GB HBM2e 80 GB HBM2e 141 GB HBM3e
Memory bandwidth 3,350 GB/s ~2,000 GB/s 1,935 GB/s 4,800 GB/s
Board power 700 W — 400 W 700 W
FP8 + Transformer Engine yes yes no yes
MIG instances up to 7 up to 7 up to 7 up to 7

SXM5 FP16/BF16 dense tensor throughput is 989 TFLOPS (1,979 with structured sparsity). This page quotes dense figures only. NVIDIA's current H100 page publishes board power for the SXM part (up to 700 W, configurable) and for the 94 GB NVL (350 to 400 W, configurable), and not for the 80 GB PCIe card, so that cell is left empty.

A contract, or a
per-minute listing.

The contract is the route Clore.ai sells H100 capacity through. The marketplace is where independent hosts list their own machines at their own prices, billed by the minute.

Per-minute marketplace

$0.99 / GPU-hr, from
Per-minute on-demand, priced by each host
  • Every host sets its own price; live figures live on the marketplace
  • Billed per minute, no commitment
  • Spot listings can be pre-empted by an on-demand renter
  • Availability follows what hosts list, so check the marketplace
Check what is listed
MAIN ROUTE

Bare-metal contract

$1.81-$2.62 / GPU-hr
8 GPUs minimum · 30-day minimum term · 21 Sep 2026
  • USA, Japan, Slovenia, Thailand and China
  • 30 to 360-day tiers; flat beyond 360 days
  • China opens at 90 days, Thailand at 360
  • Settled in BTC, CLORE, USDT or USDC
Configure a node
Settle in
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Four steps to an H100 node.

The configurator at the top of this page is the front door. It prices the exact shape you ask for, and a person comes back with the quote.

01 / SIZE

Pick the card count

Anything from 8 to 1,000 H100s. Eight is the floor because the unit Clore contracts is a full node, not a single card.

02 / TERM

Pick the term

30 days is the floor in the USA, Japan and Slovenia. China opens at 90 days, Thailand at 360. Each longer tier drops the per-GPU-hour rate.

8 × H100 · 30 days · USA
03 / QUOTE

Send the configuration

Submit it and Clore returns a contract quote for that region, that card count and that term.

04 / SETTLE

Pay from the same balance

Bare-metal orders draw on the same BTC, CLORE, USDT and USDC balances the per-minute marketplace uses.

H100 questions worth an answer.

SXM5 or PCIe: what changes besides the 700 W versus 350 W?

Core count and bandwidth, mostly. The SXM5 module has 16,896 CUDA cores and 3,350 GB/s of HBM3; the PCIe card has 14,592 cores and roughly 2,000 GB/s of HBM2e. Both carry 80 GB and both do FP8 through the Transformer Engine, so what PCIe costs you is throughput, not capability. SXM5 also carries NVLink 4 at 900 GB/s for card-to-card traffic, which only starts to matter once a single job spans the whole node. One caveat on the wattages: NVIDIA's current H100 page publishes board power for the SXM part (up to 700 W, configurable) and for the 94 GB NVL (350 to 400 W, configurable), but not for the 80 GB PCIe card, so treat the 350 W figure in wide circulation as something to confirm on the datasheet for the exact board you are offered.

Llama 3 70B at FP8 is about 70 GB. What does the Transformer Engine do that plain FP8 casting does not?

Casting a model to eight bits and hoping is how accuracy goes missing. The Transformer Engine keeps per-tensor scaling factors and moves layers between FP8 and higher precision as the run proceeds, so dynamic range survives where it has to. Budget the rest of the card deliberately: at FP8 the weights are around 70 GB, leaving roughly 10 GB for the KV cache, and that remainder sets your concurrency ceiling far more than any FLOPS number does.

When does an 8-GPU NVLink-Switch pod beat 8 independent H100s?

Whenever the eight cards have to agree on something. Gradient all-reduce under FSDP or DeepSpeed, tensor parallelism across a model too large for one card, pipeline stages handing activations forward: all of it is card-to-card traffic, and NVLink 4 moves it at 900 GB/s inside the node. Eight separate H100s in eight separate machines negotiate over the network instead, and a compute-bound run becomes a communication-bound one. If your workload is eight independent inference workers, the pod buys you nothing.

H100 MIG gives 7 slices of about 10 GB. Is that enough for an 8B model per tenant?

At FP16 an 8B model is roughly 16 GB of weights, so no. At FP8 or INT8 it is around 8 GB, which fits a slice with almost nothing left for the KV cache, so per-tenant concurrency will be low. At INT4 it is about 4 GB and the arithmetic gets comfortable. NVIDIA's MIG table lists up to seven instances on H100-SXM5 and H100-PCIE at both 80 GB and 94 GB, so the slice count is real; the only question is whether your model plus its cache fits in one slice.

What does a 94 GB H100 NVL change versus the 80 GB part?

Fourteen gigabytes, and they are a KV-cache story rather than a weights story. A model that needed 80 GB still needs 80 GB. What 94 GB buys is longer contexts, or more concurrent requests before eviction starts. NVIDIA's MIG table lists the 94 GB variants with the same seven instances, so slicing behaviour is unchanged. If weights rather than cache are your bottleneck, the jump that matters is not 80 to 94; it is 80 to the H200's 141 GB.

Can I rent a single H100 by the minute on Clore.ai today?

Yes, whenever a host has one listed. Per-minute H100 listings start from $0.99 per GPU-hour, and the prices there are set by each host rather than by Clore, so the marketplace itself is the live source for them. If you need H100 capacity you can plan a quarter around, the bare-metal contract is the route built for it: 8 cards minimum, 30 days minimum.

Three shapes of H100 job.

Capacity arithmetic rather than benchmark numbers nobody can reproduce. Each of these is a reason the card is worth its contract price.

Llama 3 70B served at FP8
vLLM or TensorRT-LLM, FP8 weights
~70 GB of the 80

The whole argument for the card. Weights fit, and whatever is left over is your KV cache, which is what decides how many requests you can hold at once.

Read the guide →
One node, eight cards, one run
PyTorch FSDP or DeepSpeed ZeRO
900 GB/s on NVLink 4

Sharded training keeps the cards in constant conversation. Inside one node that traffic stays on NVLink; the moment it leaves the node, it does not.

Read the guide →
Several tenants, one card
MIG, up to 7 instances
~10 GB per slice

Hardware partitioning, not time slicing. Each slice holds its own memory and SMs, so the per-slice budget is the design constraint you work backwards from.

Read the guide →

Where the H100 contract is priced.

Customer-facing rates per GPU-hour returned by the bare-metal configurator on 21 Sep 2026. The shortest term available differs by region, which is why some cells are empty.

Region
Shortest term
30 days
90 days
180 days
360 days
720 days
Min GPUs
USA
30 days
$2.38-$2.62
$2.07-$2.28
$1.90
$1.81-$2.07
$1.81-$2.07
8
Slovenia
30 days
$2.38-$2.62
$2.07-$2.28
$1.90
$1.81-$2.07
$1.81-$2.07
8
Japan
30 days
$2.50
$2.17
$1.81
$1.81
$1.81
8
China
90 days
—
$2.31
—
$2.31
$2.31
8
Thailand
360 days
—
—
—
$2.60
$2.35
8

The USA and Slovenia each carry two H100 SKUs at different rates, which is where the ranges come from; the configurator shows both. Contracts on the other Blackwell and Hopper parts are priced separately: see H200, B200 and B300.

Guides written for this much memory.

Each one assumes a datacenter card rather than a desktop one, and says where the 80 GB starts to bind.

Language Models
vLLM serving
Paged KV cache and continuous batching, where most FP8 deployments begin.
Training
DeepSpeed multi-GPU training
ZeRO stages 2 and 3 across the eight cards inside one node.
Language Models
Llama 3.3 on CLORE.AI
Meta's 70B release, the model this card is sized around.
Training
LLM fine-tuning
LoRA and QLoRA runs, where the adapter is what you actually train.
Language Models
Qwen 2.5
Alibaba's 0.5B to 72B family, including the sizes one card holds.
Language Models
DeepSeek Coder
Completion and refactoring models served off a rented node.
Advanced
Multi-GPU setup
NCCL, NVLink topology, and what torchrun needs to see.
See all guides →

The rest of the datacenter tier.

A100 80GB
80 GB HBM2e · contract from $1.54/GPU-hr
Compare →
H200
141 GB HBM3e · contract from $2.09/GPU-hr
Compare →
B200
180 GB HBM3e · contract from $3.93/GPU-hr
Compare →

Eight cards,
thirty days, one quote.

Choose the region and the term in the configurator at the top of this page. Clore comes back with the contract price for that exact shape.