This is the longest commitment on Clore.ai and the largest card on it. Blackwell Ultra carries 288 GB of HBM3e at 8,000 GB/s and draws 1,400 W. Clore contracts it at $5.02 to $5.16 per GPU-hour, minimum 8 GPUs, minimum 720 days, in the USA, the EU and Japan as of 21 Sep 2026. Every longer tier is priced the same, so the term buys certainty rather than a discount. Zero B300 servers were listed on the per-minute marketplace that day.
Every other card on Clore.ai can be had for a month. This one cannot, and that changes which question you should be asking first.
Eight cards at the EU rate for the minimum term is roughly $694,000 of committed spend, and at the USA and Japan rate roughly $713,000. Clore prices 720, 1,080, 1,440 and 1,800 days identically, so extending the commitment buys certainty of supply and nothing off the rate. Work out whether you want this silicon for two years before you read another line of this page.
288 GB moves several models out of the tensor-parallel category entirely. Llama 3.1 405B at INT4 is roughly 200 GB of weights and lands on one card. DeepSeek-V3's 671B mixture of experts is about 336 GB at FP4, so two. What that buys is not speed, it is the absence of a collective at every layer.
NVIDIA's HGX page gives the eight-GPU B300 system 144 PFLOPS of FP4 with sparsity and 108 PFLOPS dense, where HGX B200 is 144 and 72. FP8, FP6, FP16, BF16 and TF32 are identical on both platforms. The generational step is dense four-bit throughput, half again as much, and nothing else.
Rates and minimums as the bare-metal configurator returned them on 21 Sep 2026. The row worth staring at is the last one.
// Customer-facing rates per GPU-hour from Clore's bare-metal route, 21 Sep 2026. Savings are the drop from each card's shortest tier to its longest, measured on that card's highest-priced entry.
B300 has no term ladder, so region is the only variable that moves the price. Over eight cards and the minimum term the gap works out at roughly $19,000.
A 720-day contract is a procurement exercise, not a checkout. The configurator above collects the inputs; the first decision is not one it can make for you.
There is no 30-day tier, no 90 and no 360. If your horizon is shorter than 720 days, a B200 or H200 contract is the honest answer and both are on this site.
Eight is the minimum and a thousand the maximum. Eight cards is one Blackwell Ultra chassis on an NVLink 5 fabric.
$5.02 against $5.16 per GPU-hour. Across eight cards over the minimum term that difference is roughly $19,000, which may or may not outweigh where you need the data to sit.
Clore returns a contract quote. Invoicing runs against your BTC, CLORE, USDT or USDC balance, the same as every other order on the platform.
Someone with a workload they already know the shape of. The commitment is roughly $694,000 in the EU or $713,000 in the USA and Japan for the minimum eight cards, before power, staff or anything else, and it does not get cheaper if you extend it. That suits an inference service with committed customers, a lab with a funded two-year research programme, or a company replacing a cloud bill it can already measure. It does not suit exploratory work, a single training run, or anyone who cannot say what they will be running in eighteen months.
The ones where splitting the model is the cost you are trying to avoid. Llama 3.1 405B at INT4 is roughly 200 GB of weights: one B300 card, against two H200s. DeepSeek-V3's 671B mixture of experts is about 336 GB at FP4: two cards against three. Trillion-parameter-class mixtures of experts are the class this capacity exists for. If your largest model fits in 141 GB with cache to spare, the extra capacity is idle and you are paying for it every hour of both years.
In dense FP4, and only there. NVIDIA's HGX platform page gives the eight-GPU B300 system 144 PFLOPS of FP4 with sparsity and 108 PFLOPS dense, against HGX B200's 144 and 72. FP8 and FP6 are 72 PFLOPS on both, FP16 and BF16 are 36 PFLOPS on both, TF32 is 18 PFLOPS on both. Per GPU, memory goes from 180 GB to 288 GB and board power from 1,000 W to 1,400 W. So the upgrade is capacity plus half again as much dense four-bit throughput, bought with 40 percent more power.
NVIDIA lists 2.1 TB of total memory for the eight-GPU HGX B300 system against 1.1 TB for an eight-card H200 node at 141 GB each. In practice that is the difference between holding one frontier model plus a deep KV cache and holding two frontier models, or one model and the retrieval and reranking stack in front of it. Note that NVIDIA's 2.1 TB is below eight times the 288 GB part figure, so confirm the per-GPU capacity of the SKU you are contracting rather than multiplying.
No. The bare-metal configurator offers 720, 1,080, 1,440 and 1,800 days for B300 and nothing below, and the per-minute marketplace held zero B300 servers on 21 September 2026. If you need Blackwell for less than two years, the B200 contract starts at 30 days at $5.59 per GPU-hour, and if you need capacity rather than FP4 throughput, the H200 contract also starts at 30 days and is the cheapest of these per gigabyte held.
Because Blackwell Ultra rebalanced the die toward low-precision AI work. NVIDIA's HGX platform page lists INT8 tensor throughput at 3 POPS for the eight-GPU B300 system against 72 POPS for B200, and FP64 at 10 TFLOPS against 296. FP16, BF16, TF32 and FP8/FP6 are identical between them. So if your workload is INT8-quantised inference from before the FP4 era, or double-precision scientific computing, the B300 is a step backwards and a B200 contract is the right call. If your workload is four-bit inference, it is the other way round.
Each of these is a job that stops needing a tensor-parallel split. Weight budgets are 1 byte per parameter at FP8 and 0.5 at FP4 or INT4, before cache.
No all-reduce per layer, no parallel degree to tune, and seven cards still free for replicas or for whatever sits in front of the model.
Read the guide →Dense FP4 is the one axis where Blackwell Ultra pulls ahead of Blackwell, and a large mixture of experts is where that axis is loaded hardest.
Read the guide →The contract length is the feature here. Nothing gets reclaimed between runs, so the cluster you tuned in month three is the cluster you still have in month twenty.
Read the guide →Transcribed from NVIDIA's HGX platform specifications, both columns as printed. Two rows go the wrong way, and they are the reason this card is not automatically the better one.
The 2.1 TB system figure is below eight times the 288 GB part specification, so treat per-GPU capacity as something the contract states rather than something to multiply. If the INT8 or FP64 rows rule this card out for you, the B200 contract starts at 30 days; if capacity per dollar is what you are optimising, the H200 is cheaper per gigabyte than either.
Guides for a cluster you keep rather than one you rent for an afternoon: pipelines, batch jobs and the training stacks that assume persistent storage.
The host page treats the same term from the operator's side: depreciation, an 11 kW power line, and what a two-year commitment exposes you to.
Set the pod size and the region above and Clore returns contract terms. If two years is longer than your horizon, the B200 and H200 contracts start at thirty days.