B200 pods are available on Clore.ai from $3.93 per GPU-hour, contracted whole: 8 GPUs and 30 days at minimum, in the USA, Japan and Slovenia on one shared rate card, with payment in BTC, USDT or USDC. They are sold as bare metal rather than listed per minute, and what a pod buys is Blackwell's FP4 path, 8,000 GB/s of HBM3e per GPU, and NVLink 5 at 1.8 TB/s between the eight of them.
Three properties justify contracting eight of these at once, and none of them survives being sold one card at a time.
Hopper stops at FP8. Blackwell adds a four-bit tensor path, which halves the bytes per parameter one more time. NVIDIA's HGX B200 page quotes 144 PFLOPS of FP4 with sparsity and 72 PFLOPS dense across the eight-GPU system. In capacity terms that is the difference between a 671B mixture of experts needing four cards and needing two.
HBM3e at 8,000 GB/s is roughly 1.7 times an H200 and about 2.4 times an H100 SXM5. Since token generation reads the model once per token, that is the figure inference throughput tracks. It is also the figure that makes mixture-of-experts routing bearable, because the experts a token wakes up are scattered across memory.
NVLink 5 moves 1.8 TB/s between any two GPUs, and NVIDIA quotes 14.4 TB/s of aggregate NVLink bandwidth for the eight-GPU HGX B200 system. That is what stops per-layer collectives from dominating a tensor-parallel step. It does not turn eight cards into one address space: each GPU still owns its own memory, and the fabric only makes crossing between them cheap.
Capacity, bandwidth, power and the minimum contract, for the four parts Clore.ai offers at this end of the range. Memory and MIG figures come from NVIDIA's MIG supported-GPUs table; contract minimums from Clore's bare-metal configurator on 21 Sep 2026.
B200 memory is the 180 GB figure NVIDIA's MIG supported-GPUs table lists. The B300 is not on that table, so no MIG claim is made for it here.
On 21 Sep 2026 the USA, Japan and Slovenia all returned the same B200 pricing, so choosing a region is a latency and jurisdiction decision rather than a cost one. Term is where the money is.
The configurator above takes the four variables that move the price. Nothing about a B200 contract is self-serve past that point, and any page telling you otherwise is guessing.
The configurator accepts 8 to 1,000 B200s. Eight is not an arbitrary floor: it is one NVLink 5 chassis, and splitting it would sell away the fabric.
All three quoted identically on 21 Sep 2026, so pick for latency to your data and for the jurisdiction you need to sit in.
$5.59 per GPU-hour at 30 days, $4.86 at 90, $4.05 at 180, $3.93 from 360 days onward. That is the entire spread.
Clore returns terms for the pod you described. Invoicing runs against your BTC, USDT or USDC balance.
Because the product is a node. One B200 draws 1,000 W and lives on an SXM board inside an eight-GPU chassis wired with NVLink 5, and you cannot sell one of those eight to one customer and seven to another without giving away the fabric that made the machine worth building. Clore's bare-metal configurator reflects that directly: minimum quantity 8, maximum 1,000, minimum term 30 days. There is no per-minute listing to fall back on either, since the marketplace held zero B200 servers on 21 September 2026.
It halves the bytes per parameter one more time. DeepSeek-V3 is a 671B mixture of experts: roughly 671 GB of weights at FP8 and about 336 GB at FP4, before any KV cache. Against the 180 GB per card that NVIDIA's MIG table lists, that is four cards at FP8 and two at FP4. Blackwell is the generation that introduced the four-bit tensor path, and NVIDIA's HGX B200 page quotes 144 PFLOPS of FP4 with sparsity and 72 PFLOPS dense across the eight-GPU system. Hopper has no FP4 at all, which is why this is a generation argument rather than a faster-card argument.
Decoding, almost always. Generating a token reads every weight that token touches, so autoregressive decode is bandwidth-bound by construction and takes close to linear returns from 8,000 GB/s. Mixture-of-experts routing is worse still, because the experts a token activates are scattered through memory rather than contiguous. Prefill, dense training steps and anything with high arithmetic intensity per byte are compute-bound and will not notice. The honest summary is that B200 bandwidth pays for inference far more reliably than it pays for training.
It changes what tensor parallelism costs. Splitting one model across eight cards means a collective at every layer, and over PCIe those collectives can consume more of the step than the arithmetic does. NVLink 5 moves 1.8 TB/s between GPUs, and NVIDIA quotes 14.4 TB/s of aggregate NVLink bandwidth for the eight-GPU HGX B200 system. What it does not do is merge the eight cards into a single 1.4 TB memory space: every GPU still owns its own memory, and the fabric only makes crossing between them cheap.
NVIDIA's MIG supported-GPUs table lists the B200 at 180 GB with up to seven instances, and NVIDIA's HGX B200 page lists 1.4 TB of total memory across eight GPUs, which agrees with 180 GB a card at the precision given. The 192 GB figure that circulates widely matches neither of those NVIDIA sources, so this page uses 180 GB and cites the MIG table. When you contract a pod, the per-GPU capacity for that particular SKU is what the contract states, and that is the number to build your memory budget on.
Badly, and that is the point. Taking each card's lowest contract tier on 21 September 2026: an H200 at $2.09 for 141 GB is about $0.0148 per GB-hour; a B300 at $5.02 for 288 GB is about $0.0174; a B200 at $3.93 for 180 GB is about $0.0218; an H100 at $1.81 for 80 GB is about $0.0226. If resident memory is the only thing you are buying, the H200 is the efficient choice. What the B200 adds is FP4, 8,000 GB/s instead of 4,800, and NVLink 5. Buy it for those, not for the gigabytes.
Weight budgets at the precision each job is actually served in, against the 180 GB a card that NVIDIA's MIG table lists.
Two cards at FP4 where FP8 needed four. The routing pattern is what makes the 8,000 GB/s worth paying for.
Read the guide →Three cards hold it, leaving five for cache, for a second replica, or for the retrieval stack sitting in front of it.
Read the guide →Sharded training is a communication workload wearing a compute workload's clothes. Keeping it inside one chassis is the whole trick.
Read the guide →Every datacenter part Clore.ai offers, ranked by what you have to commit to before you can touch it. Figures from the bare-metal configurator and the marketplace payload on 21 Sep 2026.
Four of these five carried no per-minute servers in the 21 Sep 2026 snapshot. Hardware of this class is sold on contract, which is why Clore.ai quotes it through the configurator instead of listing it per minute.
Eight cards is a small cluster, and it behaves like one. These are the stacks that assume that.
That is the highest per-card figure on Clore.ai, earned by supplying whole 8-GPU pods to the contract route. The host page covers terms, regions and listing per minute.
Pick the pod size, the region and the term above. What comes back is a quote for exactly that, with the rate the term earns you.