Log in Rent H200
H200 141GB · bare metal · from $2.09/GPU-hour, five regions

Rent an H200 141GB:
141 GB holds 70B at FP16.
So does 80 GB.

Clore.ai supplies the H200 as bare metal from $2.09 per GPU-hour: 8-GPU blocks, terms from 30 days, five regions (Iceland, the USA, India, Japan and Slovenia), settled in Bitcoin, USDT or USDC. It is contracted rather than listed per minute. The card itself is a Hopper GH100 with a different memory system attached: 141 GB of HBM3e at 4,800 GB/s, where the H100 has 80 GB at 3,350, on the same compute units.

●From $2.09/GPU-hour ●Blocks from 8 GPUs ●Five regions ●H200 and H200 SXM
141GB
HBM3e per card
+43%
Bandwidth over an H100 SXM5
5
Regions, from Iceland to Japan
$2.09/GPU-hr
Bare metal from, per GPU-hour

Every argument here
reduces to memory.

Nothing on this page will tell you the H200 computes faster than an H100, because it does not. The case for the card is what it can hold and how quickly it can read it.

The 70B model stops being a two-card problem

Llama 3 70B at FP16 is roughly 140 GB of weights. On an 80 GB card that is two GPUs and a tensor-parallel split before you have served a single token. On 141 GB it loads onto one card. Be honest about what is left afterwards: about a gigabyte, so the practical serving configuration is still FP8, at which point the 70 GB of freed capacity becomes KV cache instead of a second card.

Llama 3 70B, FP16 weights ~140 GB

Decode is a bandwidth problem, not a FLOPS problem

Generating one token reads the entire weight set once. That makes decode throughput track memory bandwidth far more closely than it tracks tensor throughput, and it is why two cards with identical compute can serve at visibly different rates. HBM3e at 4,800 GB/s moves about 43 percent more bytes per second than the H100's HBM3.

Memory bandwidth 4,800 GB/s

Frontier models land on fewer cards

Llama 3.1 405B quantised to INT4 is around 200 GB of weights: two H200s, where the same quantisation needs three H100s. At FP8 it is three cards against six. Fewer cards in a tensor-parallel group means less time spent inside each collective and more of the node's clock spent on arithmetic.

Llama 3.1 405B at INT4 2 cards, not 3

Card counts,
not benchmark scores.

The only question this card answers differently from its neighbours is how many of them a given model needs. Weight sizes below are 2 bytes per parameter at FP16, 1 at FP8 or INT8, 0.5 at INT4.

H200 141GB H100 80GB A100 80GB B200
VRAM 141 GB HBM3e 80 GB HBM3 80 GB HBM2e 180 GB HBM3e
Memory bandwidth 4,800 GB/s 3,350 GB/s 1,935 GB/s 8,000 GB/s
Llama 3 70B at FP16 1 card 2 cards 2 cards 1 card
Llama 3 70B at FP8 1 card 1 card no FP8 1 card
Llama 3.1 405B at INT4 2 cards 3 cards 3 cards 2 cards
Low-precision formats FP8 FP8 neither FP8 and FP4
MIG instances up to 7 up to 7 up to 7 up to 7

Card counts cover weights only. Add 10 to 30 percent for KV cache and activations before you size a node. MIG counts from NVIDIA's supported-GPUs table.

The configurator prices
H200 and H200 SXM apart.

They are not the same line item, they are not in the same regions, and they do not cost the same. Rates below are the customer-facing figures the bare-metal route returned on 21 Sep 2026.

H200 SXM

$2.81-$4.18 / GPU-hr
Iceland and Japan · 8 GPUs minimum · 21 Sep 2026
  • 30-day floor, then 180 and 360-day tiers
  • $4.18 at 30 days, $2.81 from 360 days on
  • Iceland pricing is identical to Japan's
  • No 90-day tier on this SKU
Configure H200 SXM
WIDEST COVERAGE

H200

$2.09-$3.86 / GPU-hr
USA, Japan, Slovenia, India · 8 GPUs minimum · 21 Sep 2026
  • 30, 90, 180 and 360-day tiers
  • Two price bands in the USA and Slovenia; the cheaper one floors at $2.09
  • India is a 360-day commitment at $3.01, flat
  • Settled in BTC, USDT or USDC
Configure H200
Invoiced in
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Four decisions, then a quote.

There is no self-serve checkout for a datacenter contract. The configurator on this page captures the four things that set the price, and a person answers with terms.

01 / SKU

H200 or H200 SXM

They appear as separate entries because they are priced separately and sit in different regions. Pick before you compare rates, or you will compare the wrong two numbers.

02 / REGION

Where the cards sit

Iceland and Japan carry the SXM entry. The USA, Japan and Slovenia open at 30 days on the standard entry; India is available on a 360-day commitment only.

8 × H200 SXM · 180 days · Iceland
03 / TERM

How long you are holding it

Rates fall through the 90, 180 and 360-day tiers and then stop. Anything past 360 days is priced the same as 360 on every H200 entry.

04 / QUOTE

Send it and wait for terms

Submit the configuration and Clore returns a contract quote. Invoicing runs against the same BTC, USDT and USDC balances the marketplace uses.

What people get wrong about the H200.

The H200 has the same compute as an H100. So what does 141 GB actually unlock?

Nothing on the compute side. It is the same GH100 die, the same 16,896 CUDA cores, the same 989 TFLOPS of dense FP16/BF16 tensor throughput as an H100 SXM5. What changes is the memory system: 141 GB of HBM3e at 4,800 GB/s instead of 80 GB of HBM3 at 3,350. So the card does not make a job that already fitted run faster in any arithmetic sense. It makes jobs that did not fit, fit, and it feeds memory-bound work more quickly. If your model already sits comfortably in 80 GB and your bottleneck is arithmetic, this is the wrong upgrade.

Llama 3 70B at FP16 is ~140 GB. Does it really fit on one H200, and what is left for the KV cache?

It fits, and then there is nothing left. Roughly 140 GB of weights against 141 GB of capacity leaves the KV cache about a gigabyte, which means a very short context and a batch size of one. Treat single-card FP16 70B as proof that it loads rather than as a serving configuration. The version people actually serve is FP8, around 70 GB, which leaves roughly 70 GB for cache, and that is the point at which this card becomes a comfortable 70B server rather than a party trick.

4,800 GB/s versus 3,350: how much of that translates into tokens per second on a memory-bound decode?

Decoding reads the whole weight set once per token, so a genuinely memory-bound decode scales close to linearly with bandwidth. 4,800 against 3,350 is about 43 percent more bytes per second, and that is the neighbourhood to expect, not a doubling. Prefill will not move at all, because prefill is compute-bound and the compute is identical. Which regime you are in depends on prompt length and batch size, so measure your own workload rather than taking either number on faith.

Llama 3.1 405B at INT4 across 4 H200s instead of 8 H100s: does the interconnect math work out?

The card count works, and it is smaller than four. At INT4 the 405B weights are roughly 200 GB, which lands on two H200s against three H100s; at FP8 it is about 405 GB, so three H200s against six. Inside an 8-GPU node every one of those layouts keeps its traffic on NVLink 4 at 900 GB/s, so the interconnect is not what decides it. What decides it is that a smaller tensor-parallel group spends less time inside each all-reduce, leaving more of the node's clock for arithmetic.

What is DeepSeek-V3's actual context window, and what does 141 GB let me do with it?

128K tokens. The one-million figure that circulates on GPU rental pages is not DeepSeek-V3's published window. Against a 128K window the extra memory buys cache depth rather than window length: once the 671B mixture-of-experts weights are quantised, the KV cache becomes the dominant consumer, and the 61 GB an H200 holds over an H100 is what lets you keep several long sessions resident at once instead of one.

Why is there no H200 on the per-minute marketplace?

Because Clore.ai supplies this card as contracted capacity rather than as individually listed hardware. On 21 September 2026 the per-minute marketplace held zero servers carrying an H200. The route that exists is the bare-metal contract: 8 cards minimum, 30 days minimum in the USA, Japan, Slovenia and Iceland, and a 360-day minimum in India. If per-minute H200 hardware is listed later it will appear under the marketplace's GPU filter like any other card.

Jobs that need the extra 61 GB.

Each of these is a memory argument. None of them gets faster because of a FLOPS figure, because the FLOPS figure is the H100's.

Llama 3 70B loaded at FP16
One card, no tensor-parallel split
~140 GB of 141

Useful as a reference run and for evaluating quantisation loss against the unquantised original. It leaves no cache, which is exactly why FP8 is what gets deployed.

Read the guide →
DeepSeek-V3 at its full window
671B mixture of experts, quantised
128K tokens

The published context window is 128K, not a million. At that depth the KV cache dominates the memory budget and the card's capacity sets how many sessions stay resident.

Read the guide →
Llama 3.1 405B at INT4
Tensor parallel across two cards
~200 GB of weights

Two H200s hold what three H100s hold. A narrower parallel group means shorter collectives, which is where large-model serving usually loses its efficiency.

Read the guide →

Identical everywhere except two rows.

Put the two Hopper parts side by side and the shape of the upgrade is unmistakable. Specifications from NVIDIA; contract rates and listing counts from Clore.ai on 21 Sep 2026.

H200 SXM5
H100 SXM5
Architecture
Hopper GH100
Hopper GH100
CUDA cores
16,896
16,896
FP16 / BF16 tensor, dense
989 TFLOPS
989 TFLOPS
GPU memory
141 GB HBM3e
80 GB HBM3
Memory bandwidth
4,800 GB/s
3,350 GB/s
NVLink
900 GB/s
900 GB/s
Board power
700 W
700 W
Contract floor, per GPU-hour
$2.09
$1.81
Per-minute servers, 21 Sep 2026
0
4

The full H100 contract matrix, region by region, is on the H100 page. If the memory ceiling is still the problem, the next steps up are B200 and B300.

Guides for the memory-bound half of the job.

Serving stacks, sharded training and the vision models where activation memory, not weights, is what runs out first.

Language Models
vLLM serving
How the paged cache uses whatever memory the weights leave behind.
Training
DeepSpeed multi-GPU training
Sharding optimizer state when a bigger card lets you shard less.
Language Models
DeepSeek-V3
A 671B mixture of experts with a 128K window, and what it costs in memory.
Language Models
Llama 3.3 on CLORE.AI
The 70B release, at the precision that leaves room for a cache.
Training
LLM fine-tuning
Where optimizer state, not weights, decides how big a card you need.
Vision Models
Llama Vision
Multimodal inference, where image tokens inflate the cache quickly.
Advanced
Multi-GPU setup
Choosing a tensor-parallel degree once you no longer have to maximise it.
See all guides →

One step down, two steps up.

H100 80GB
Same compute · contract from $1.81/GPU-hr
Compare →
B200
FP4 and 8 TB/s · contract from $3.93/GPU-hr
Compare →
B300
288 GB · 720-day minimum term
Compare →

141 GB a card,
eight cards a contract.

Choose the SKU, the region and the term in the configurator above. What comes back is a quote for that shape, not a checkout page.