Clore.ai supplies the H200 as bare metal from $2.09 per GPU-hour: 8-GPU blocks, terms from 30 days, five regions (Iceland, the USA, India, Japan and Slovenia), settled in Bitcoin, USDT or USDC. It is contracted rather than listed per minute. The card itself is a Hopper GH100 with a different memory system attached: 141 GB of HBM3e at 4,800 GB/s, where the H100 has 80 GB at 3,350, on the same compute units.
Nothing on this page will tell you the H200 computes faster than an H100, because it does not. The case for the card is what it can hold and how quickly it can read it.
Llama 3 70B at FP16 is roughly 140 GB of weights. On an 80 GB card that is two GPUs and a tensor-parallel split before you have served a single token. On 141 GB it loads onto one card. Be honest about what is left afterwards: about a gigabyte, so the practical serving configuration is still FP8, at which point the 70 GB of freed capacity becomes KV cache instead of a second card.
Generating one token reads the entire weight set once. That makes decode throughput track memory bandwidth far more closely than it tracks tensor throughput, and it is why two cards with identical compute can serve at visibly different rates. HBM3e at 4,800 GB/s moves about 43 percent more bytes per second than the H100's HBM3.
Llama 3.1 405B quantised to INT4 is around 200 GB of weights: two H200s, where the same quantisation needs three H100s. At FP8 it is three cards against six. Fewer cards in a tensor-parallel group means less time spent inside each collective and more of the node's clock spent on arithmetic.
The only question this card answers differently from its neighbours is how many of them a given model needs. Weight sizes below are 2 bytes per parameter at FP16, 1 at FP8 or INT8, 0.5 at INT4.
Card counts cover weights only. Add 10 to 30 percent for KV cache and activations before you size a node. MIG counts from NVIDIA's supported-GPUs table.
They are not the same line item, they are not in the same regions, and they do not cost the same. Rates below are the customer-facing figures the bare-metal route returned on 21 Sep 2026.
There is no self-serve checkout for a datacenter contract. The configurator on this page captures the four things that set the price, and a person answers with terms.
They appear as separate entries because they are priced separately and sit in different regions. Pick before you compare rates, or you will compare the wrong two numbers.
Iceland and Japan carry the SXM entry. The USA, Japan and Slovenia open at 30 days on the standard entry; India is available on a 360-day commitment only.
Rates fall through the 90, 180 and 360-day tiers and then stop. Anything past 360 days is priced the same as 360 on every H200 entry.
Submit the configuration and Clore returns a contract quote. Invoicing runs against the same BTC, USDT and USDC balances the marketplace uses.
Nothing on the compute side. It is the same GH100 die, the same 16,896 CUDA cores, the same 989 TFLOPS of dense FP16/BF16 tensor throughput as an H100 SXM5. What changes is the memory system: 141 GB of HBM3e at 4,800 GB/s instead of 80 GB of HBM3 at 3,350. So the card does not make a job that already fitted run faster in any arithmetic sense. It makes jobs that did not fit, fit, and it feeds memory-bound work more quickly. If your model already sits comfortably in 80 GB and your bottleneck is arithmetic, this is the wrong upgrade.
It fits, and then there is nothing left. Roughly 140 GB of weights against 141 GB of capacity leaves the KV cache about a gigabyte, which means a very short context and a batch size of one. Treat single-card FP16 70B as proof that it loads rather than as a serving configuration. The version people actually serve is FP8, around 70 GB, which leaves roughly 70 GB for cache, and that is the point at which this card becomes a comfortable 70B server rather than a party trick.
Decoding reads the whole weight set once per token, so a genuinely memory-bound decode scales close to linearly with bandwidth. 4,800 against 3,350 is about 43 percent more bytes per second, and that is the neighbourhood to expect, not a doubling. Prefill will not move at all, because prefill is compute-bound and the compute is identical. Which regime you are in depends on prompt length and batch size, so measure your own workload rather than taking either number on faith.
The card count works, and it is smaller than four. At INT4 the 405B weights are roughly 200 GB, which lands on two H200s against three H100s; at FP8 it is about 405 GB, so three H200s against six. Inside an 8-GPU node every one of those layouts keeps its traffic on NVLink 4 at 900 GB/s, so the interconnect is not what decides it. What decides it is that a smaller tensor-parallel group spends less time inside each all-reduce, leaving more of the node's clock for arithmetic.
128K tokens. The one-million figure that circulates on GPU rental pages is not DeepSeek-V3's published window. Against a 128K window the extra memory buys cache depth rather than window length: once the 671B mixture-of-experts weights are quantised, the KV cache becomes the dominant consumer, and the 61 GB an H200 holds over an H100 is what lets you keep several long sessions resident at once instead of one.
Because Clore.ai supplies this card as contracted capacity rather than as individually listed hardware. On 21 September 2026 the per-minute marketplace held zero servers carrying an H200. The route that exists is the bare-metal contract: 8 cards minimum, 30 days minimum in the USA, Japan, Slovenia and Iceland, and a 360-day minimum in India. If per-minute H200 hardware is listed later it will appear under the marketplace's GPU filter like any other card.
Each of these is a memory argument. None of them gets faster because of a FLOPS figure, because the FLOPS figure is the H100's.
Useful as a reference run and for evaluating quantisation loss against the unquantised original. It leaves no cache, which is exactly why FP8 is what gets deployed.
Read the guide →The published context window is 128K, not a million. At that depth the KV cache dominates the memory budget and the card's capacity sets how many sessions stay resident.
Read the guide →Two H200s hold what three H100s hold. A narrower parallel group means shorter collectives, which is where large-model serving usually loses its efficiency.
Read the guide →Put the two Hopper parts side by side and the shape of the upgrade is unmistakable. Specifications from NVIDIA; contract rates and listing counts from Clore.ai on 21 Sep 2026.
The full H100 contract matrix, region by region, is on the H100 page. If the memory ceiling is still the problem, the next steps up are B200 and B300.
Serving stacks, sharded training and the vision models where activation memory, not weights, is what runs out first.
Clore.ai contracts 141 GB cards in 8-GPU nodes on terms from 30 days, the route that pays that figure, and per-minute listing is open as well. The host page covers terms, regions and setup.
Choose the SKU, the region and the term in the configurator above. What comes back is a quote for that shape, not a checkout page.