The V100 is the largest datacenter fleet CLORE carries and the hardest to get onto. On 21 September 2026, 186 servers held 1,389 cards and exactly one server was unrented. What keeps them booked is the price of the bandwidth: 900 GB/s of HBM2 at a $0.065 median per GPU-hour, which nothing else here comes close to. The catch is the generation. Volta is compute capability 7.0, with no BF16, no FP8, no MIG and no FlashAttention-3.
A fleet this size sitting at one free server is not an accident of pricing. Something keeps renting these boards, and it is not the feature list.
At the 21 September 2026 snapshot a V100 was renting at a median $0.065 per GPU-hour against $0.417 for an A100 40GB. The A100 reads memory 1.7 times faster and costs more than six times as much. For work that is bound by how fast weights move rather than by what the tensor cores support, that ratio is the entire argument for this board.
Code that was written against compute capability 7.0 and never migrated keeps running here without a rewrite. That is a real and stubborn category: training scripts, CUDA extensions built for sm_70, and FP32 simulation kernels that never needed a tensor core in the first place.
An 8B at FP16 is about 16 GB and fits with room. A 14B at FP16 is about 28 GB and fits with none. A 70B does not fit at any precision this card can execute, and neither does anything that needs FP8 or BF16. Read that as a hard boundary, not a performance note.
The V100 loses every column here except one, and that one is the reason its fleet is booked solid. Specs from the NVIDIA Volta, Ampere and Ada datasheets; all TFLOPS figures are dense, never with sparsity.
specs from the NVIDIA V100, A100 and GeForce datasheets · rates are the median per-GPU hourly price across CLORE listings on 21 Sep 2026
On a fleet that is 185 of 186 servers rented, the discount for accepting preemption has mostly evaporated. Both medians landed on $0.065 per GPU-hour on 21 September 2026. Hosts set their own rates and live prices are on the marketplace.
If waiting for a free server is not an option, V100 bare metal is sold from 8 GPUs on a 30-day minimum term: $0.48 per GPU-hour at 30 days, $0.43 at 90, $0.40 at 180 and $0.36 from 360, in the USA, France and Japan. Configure bare metal →
With one server free out of 186, the difference between getting a V100 and not getting one is mostly about how you watch and what you will accept.
Filter by Tesla V100 and sort by availability, not by price. Confirm the memory: the fleet is overwhelmingly the 32 GB SXM2 part, but 16 GB boards exist.
Both medians sat at $0.065 per GPU-hour, so spot is not buying you much of a discount here. Take on-demand if an interruption would cost you a checkpoint.
Run one short job first. If a wheel ships only sm_80 kernels, you will find out in seconds rather than after an hour of billing.
Bare metal removes the availability problem: 8 GPUs and 30 days minimum, $0.36 to $0.48 per GPU-hour by term, in three countries.
Anything whose kernels are gated on Ampere or newer. In practice that means FlashAttention-2 and FlashAttention-3, every FP8 code path, and every BF16 code path, because Volta has neither data type in hardware. Frameworks themselves still build for Volta, but individual wheels increasingly ship only sm_80 and above kernels, so the failure shows up as an unsupported-architecture error at import or at the first kernel launch rather than at install time. Check that the wheel you plan to use still ships sm_70 before you book a long run.
It means you go back to managing dynamic range by hand. BF16 keeps the exponent range of FP32 and trades mantissa bits for it, which is why Ampere-and-later recipes can mostly ignore overflow. FP16 has a much narrower range, so you need loss scaling, and with a badly tuned scale you get either silent overflow to infinity or gradients that underflow to zero. Automatic mixed precision handles the common cases, but a recipe written and tuned in BF16 will not always transfer unchanged.
Work that reads a lot of memory and does comparatively little arithmetic on it. Decoding tokens from a model small enough to fit in 32 GB, embedding or reranking passes over a large corpus, FP32 simulation kernels, and anything that spent its life being throttled by a consumer card memory bus. What it is not good for is anything that needs modern data types or a model too big for 32 GB, where a cheaper hourly rate cannot buy back the capability.
Both exist and the CLORE fleet is overwhelmingly the 32 GB SXM2 part, but overwhelmingly is not always. The listing carries the GPU name reported by the host agent, and the two variants appear under distinct names, so read it before you rent. On the machine itself, nvidia-smi prints the memory total in the first block. If your job needs 32 GB, confirm from the listing, not from the model name alone, and remember there has never been an 80 GB V100 whatever a spec sheet elsewhere tells you.
Let PyTorch choose. Its scaled_dot_product_attention dispatches across several backends, and on a Volta card it lands on the memory-efficient one rather than the flash one, which needs Ampere or newer. That gets you most of the memory saving without hand-written kernels. If a library hard-codes a FlashAttention import, look for a configuration switch back to the reference implementation before assuming the model will not run at all.
Not reliably, at that moment. The 21 September 2026 snapshot found 186 servers carrying 1,389 V100 cards, the largest datacenter fleet on the platform, and 185 of those servers were rented, leaving exactly one free with 4 free cards on it. That is a fleet that clears. Practical options are to watch the marketplace for a release, to bid on spot and accept that an on-demand renter can take the machine back, or to take a bare-metal V100 block: 8 GPUs minimum, 30 days minimum, $0.36 to $0.48 per GPU-hour by term, in the USA, France and Japan.
All three are bound by memory traffic rather than by tensor-core features, which is exactly where an eight-year-old HBM board still competes. Throughput depends on your model and batch, so measure before you budget.
Half the board for weights, half for context. Decode speed follows the 900 GB/s bus, not the 125 TFLOPS number.
Read the guide →An offline queue where latency does not matter and hourly cost does. Nothing in the pipeline needs BF16 or FP8.
Read the guide →Code that predates BF16 runs unchanged. Code written for BF16 usually needs its loss scaling put back before it will converge.
Read the guide →One snapshot of the CLORE marketplace, 21 September 2026. Free means not rented at that instant, which is a single reading rather than an availability rate. Rates are median per-GPU hourly prices derived from whole-server daily prices.
Nothing here depends on BF16, FP8 or FlashAttention. If a guide elsewhere assumes an Ampere kernel, that is the part you will have to replace.
Renters book almost every V100 listed. Bring yours to Clore.ai and earn up to ~$320/mo per card on a bare-metal contract, or put it on the per-minute marketplace at your own rate.
That was the reading on 21 September 2026. Supply moves, so check what is open now, or take a bare-metal block if you need the cards to still be yours next month.