Log in Rent B200
B200 · 8-GPU pod · from $3.93/GPU-hour, three regions

Rent a B200.
The unit is a pod,
or a single card.

B200 pods are available on Clore.ai from $3.93 per GPU-hour, contracted whole: 8 GPUs and 30 days at minimum, in the USA, Japan and Slovenia on one shared rate card, with payment in BTC, USDT or USDC. They are sold as bare metal rather than listed per minute, and what a pod buys is Blackwell's FP4 path, 8,000 GB/s of HBM3e per GPU, and NVLink 5 at 1.8 TB/s between the eight of them.

●From $3.93/GPU-hour ●Pods from 8 GPUs ●Terms from 30 days ●USA, Japan, Slovenia
8
GPUs in the smallest pod
8,000GB/s
HBM3e bandwidth per GPU
1.8TB/s
NVLink 5, GPU to GPU
$3.93/GPU-hr
Bare metal from, per GPU-hour

The card is not the product.
The pod is.

Three properties justify contracting eight of these at once, and none of them survives being sold one card at a time.

FP4 exists on Blackwell and nowhere earlier

Hopper stops at FP8. Blackwell adds a four-bit tensor path, which halves the bytes per parameter one more time. NVIDIA's HGX B200 page quotes 144 PFLOPS of FP4 with sparsity and 72 PFLOPS dense across the eight-GPU system. In capacity terms that is the difference between a 671B mixture of experts needing four cards and needing two.

HGX B200, FP4 dense 72 PFLOPS per 8 GPUs

Eight terabytes a second, per card

HBM3e at 8,000 GB/s is roughly 1.7 times an H200 and about 2.4 times an H100 SXM5. Since token generation reads the model once per token, that is the figure inference throughput tracks. It is also the figure that makes mixture-of-experts routing bearable, because the experts a token wakes up are scattered across memory.

Bandwidth over an H100 SXM5 about 2.4×

A fabric, not a memory pool

NVLink 5 moves 1.8 TB/s between any two GPUs, and NVIDIA quotes 14.4 TB/s of aggregate NVLink bandwidth for the eight-GPU HGX B200 system. That is what stops per-layer collectives from dominating a tensor-parallel step. It does not turn eight cards into one address space: each GPU still owns its own memory, and the fabric only makes crossing between them cheap.

Aggregate NVLink, 8 GPUs 14.4 TB/s

Per card,
generation by generation.

Capacity, bandwidth, power and the minimum contract, for the four parts Clore.ai offers at this end of the range. Memory and MIG figures come from NVIDIA's MIG supported-GPUs table; contract minimums from Clore's bare-metal configurator on 21 Sep 2026.

B200 H200 H100 SXM5 B300
Architecture Blackwell GB100 Hopper GH100 Hopper GH100 Blackwell Ultra
VRAM 180 GB HBM3e 141 GB HBM3e 80 GB HBM3 288 GB HBM3e
Memory bandwidth 8,000 GB/s 4,800 GB/s 3,350 GB/s 8,000 GB/s
Board power 1,000 W 700 W 700 W 1,400 W
NVLink, GPU to GPU 1.8 TB/s 900 GB/s 900 GB/s 1.8 TB/s
Low-precision formats FP8 and FP4 FP8 FP8 FP8 and FP4
Smallest contract 8 GPUs, 30 days 8 GPUs, 30 days 8 GPUs, 30 days 8 GPUs, 720 days

B200 memory is the 180 GB figure NVIDIA's MIG supported-GPUs table lists. The B300 is not on that table, so no MIG claim is made for it here.

One rate card,
three regions, four steps.

On 21 Sep 2026 the USA, Japan and Slovenia all returned the same B200 pricing, so choosing a region is a latency and jurisdiction decision rather than a cost one. Term is where the money is.

Short term

$5.59 / GPU-hr
30 days · drops to $4.86 at 90 days · 21 Sep 2026
  • 30 days is the shortest B200 commitment offered
  • Suits a single pretraining run or an evaluation cycle
  • Same rate in all three regions
  • The most expensive way to hold the hardware
Price 30 days
FLOOR RATE

Long term

$3.93 / GPU-hr
360 days and beyond · $4.05 at 180 days · 21 Sep 2026
  • The rate stops falling at 360 days and stays flat to 1,800
  • 30 percent below the 30-day tier
  • Pod sizes from 8 to 1,000 GPUs
  • Settled in BTC, USDT or USDC
Price 360 days
Contracts settle in
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Four inputs, then a quote.

The configurator above takes the four variables that move the price. Nothing about a B200 contract is self-serve past that point, and any page telling you otherwise is guessing.

01 / POD SIZE

Eight, or a multiple of it

The configurator accepts 8 to 1,000 B200s. Eight is not an arbitrary floor: it is one NVLink 5 chassis, and splitting it would sell away the fabric.

02 / REGION

USA, Japan or Slovenia

All three quoted identically on 21 Sep 2026, so pick for latency to your data and for the jurisdiction you need to sit in.

03 / TERM

Where the price actually moves

$5.59 per GPU-hour at 30 days, $4.86 at 90, $4.05 at 180, $3.93 from 360 days onward. That is the entire spread.

8 × B200 · 360 days · Japan
04 / QUOTE

A contract, then an invoice

Clore returns terms for the pod you described. Invoicing runs against your BTC, USDT or USDC balance.

Pod questions, answered plainly.

Why is B200 sold as an 8-GPU pod rather than a single card?

Because the product is a node. One B200 draws 1,000 W and lives on an SXM board inside an eight-GPU chassis wired with NVLink 5, and you cannot sell one of those eight to one customer and seven to another without giving away the fabric that made the machine worth building. Clore's bare-metal configurator reflects that directly: minimum quantity 8, maximum 1,000, minimum term 30 days. There is no per-minute listing to fall back on either, since the marketplace held zero B200 servers on 21 September 2026.

What does FP4 change for serving a 671B MoE model that FP8 could not do?

It halves the bytes per parameter one more time. DeepSeek-V3 is a 671B mixture of experts: roughly 671 GB of weights at FP8 and about 336 GB at FP4, before any KV cache. Against the 180 GB per card that NVIDIA's MIG table lists, that is four cards at FP8 and two at FP4. Blackwell is the generation that introduced the four-bit tensor path, and NVIDIA's HGX B200 page quotes 144 PFLOPS of FP4 with sparsity and 72 PFLOPS dense across the eight-GPU system. Hopper has no FP4 at all, which is why this is a generation argument rather than a faster-card argument.

8 TB/s of HBM3e per GPU: which workloads are actually bandwidth-bound at that rate?

Decoding, almost always. Generating a token reads every weight that token touches, so autoregressive decode is bandwidth-bound by construction and takes close to linear returns from 8,000 GB/s. Mixture-of-experts routing is worse still, because the experts a token activates are scattered through memory rather than contiguous. Prefill, dense training steps and anything with high arithmetic intensity per byte are compute-bound and will not notice. The honest summary is that B200 bandwidth pays for inference far more reliably than it pays for training.

NVLink 5 at 1.8 TB/s between 8 GPUs: what does that make possible that PCIe cannot?

It changes what tensor parallelism costs. Splitting one model across eight cards means a collective at every layer, and over PCIe those collectives can consume more of the step than the arithmetic does. NVLink 5 moves 1.8 TB/s between GPUs, and NVIDIA quotes 14.4 TB/s of aggregate NVLink bandwidth for the eight-GPU HGX B200 system. What it does not do is merge the eight cards into a single 1.4 TB memory space: every GPU still owns its own memory, and the fabric only makes crossing between them cheap.

What is the real VRAM per B200, and why do published figures differ?

NVIDIA's MIG supported-GPUs table lists the B200 at 180 GB with up to seven instances, and NVIDIA's HGX B200 page lists 1.4 TB of total memory across eight GPUs, which agrees with 180 GB a card at the precision given. The 192 GB figure that circulates widely matches neither of those NVIDIA sources, so this page uses 180 GB and cites the MIG table. When you contract a pod, the per-GPU capacity for that particular SKU is what the contract states, and that is the number to build your memory budget on.

How does a B200 pod compare with an H200 pod on price per gigabyte of resident memory?

Badly, and that is the point. Taking each card's lowest contract tier on 21 September 2026: an H200 at $2.09 for 141 GB is about $0.0148 per GB-hour; a B300 at $5.02 for 288 GB is about $0.0174; a B200 at $3.93 for 180 GB is about $0.0218; an H100 at $1.81 for 80 GB is about $0.0226. If resident memory is the only thing you are buying, the H200 is the efficient choice. What the B200 adds is FP4, 8,000 GB/s instead of 4,800, and NVLink 5. Buy it for those, not for the gigabytes.

Three jobs that need eight cards.

Weight budgets at the precision each job is actually served in, against the 180 GB a card that NVIDIA's MIG table lists.

DeepSeek-V3 671B at FP4
Mixture of experts on the Blackwell four-bit path
~336 GB of weights

Two cards at FP4 where FP8 needed four. The routing pattern is what makes the 8,000 GB/s worth paying for.

Read the guide →
Llama 3.1 405B at FP8
Tensor parallel across part of the pod
~405 GB of weights

Three cards hold it, leaving five for cache, for a second replica, or for the retrieval stack sitting in front of it.

Read the guide →
Training inside one NVLink domain
FSDP or DeepSpeed across all eight
14.4 TB/s aggregate

Sharded training is a communication workload wearing a compute workload's clothes. Keeping it inside one chassis is the whole trick.

Read the guide →

The contract, not the card.

Every datacenter part Clore.ai offers, ranked by what you have to commit to before you can touch it. Figures from the bare-metal configurator and the marketplace payload on 21 Sep 2026.

GPU
VRAM
Min GPUs
Min term
Contract $/GPU-hr
Regions
Per-minute servers
A100 80GB
80 GB
8
30 days
$1.54-$2.28
USA, Japan, Slovenia
0
H100 80GB
80 GB
8
30 days
$1.81-$2.62
5 countries
4
H200 141GB
141 GB
8
30 days
$2.09-$4.18
5 regions
0
B200 / this page
180 GB
8
30 days
$3.93-$5.59
USA, Japan, Slovenia
0
B300
288 GB
8
720 days
$5.02-$5.16
USA, EU, Japan
0

Four of these five carried no per-minute servers in the 21 Sep 2026 snapshot. Hardware of this class is sold on contract, which is why Clore.ai quotes it through the configurator instead of listing it per minute.

What to run once the pod is yours.

Eight cards is a small cluster, and it behaves like one. These are the stacks that assume that.

Training
DeepSpeed multi-GPU training
Partitioning a run across a single NVLink domain rather than a network.
Language Models
DeepSeek-V3
The 671B mixture of experts this pod is most often bought to serve.
Language Models
vLLM serving
Setting tensor-parallel degree when you have eight cards to spend.
Language Models
Llama 3.3 on CLORE.AI
A 70B model that leaves most of the pod free for something else.
Language Models
Qwen 2.5
Running several sizes of the family concurrently across the cards.
Training
LLM fine-tuning
Adapter training on a frontier base model without reshaping the pod.
Advanced
Multi-GPU setup
Checking that NCCL sees the NVLink topology it is supposed to see.
See all guides →

Cheaper, or larger.

H200
141 GB · cheapest per GB-hour of the four
Compare →
B300
288 GB · 720-day minimum term
Compare →
H100 80GB
80 GB · the only one also listed per minute
Compare →

Eight cards,
one fabric, one contract.

Pick the pod size, the region and the term above. What comes back is a quote for exactly that, with the rate the term earns you.