Log in V100 listings
Tesla V100 SXM2 · 32 GB HBM2 1,389 cards listed, 1 server free, as of 21 Sep 2026

The cheapest HBM
on the platform.
And easy to get hold of.

The V100 is the largest datacenter fleet CLORE carries and the hardest to get onto. On 21 September 2026, 186 servers held 1,389 cards and exactly one server was unrented. What keeps them booked is the price of the bandwidth: 900 GB/s of HBM2 at a $0.065 median per GPU-hour, which nothing else here comes close to. The catch is the generation. Volta is compute capability 7.0, with no BF16, no FP8, no MIG and no FlashAttention-3.

●Billed per minute ●SSH, Docker and Jupyter ●Volta, compute capability 7.0 ●Bare metal from 8 GPUs
1,389
Cards listed on 186 servers, 21 Sep 2026
1
Of those servers unrented at that instant
900GB/s
HBM2 bandwidth per card
$0.065/GPU-h
Median rate, spot and on-demand alike

Eight-year-old silicon,
still fully booked.

A fleet this size sitting at one free server is not an accident of pricing. Something keeps renting these boards, and it is not the feature list.

Bandwidth per dollar, and nothing else

At the 21 September 2026 snapshot a V100 was renting at a median $0.065 per GPU-hour against $0.417 for an A100 40GB. The A100 reads memory 1.7 times faster and costs more than six times as much. For work that is bound by how fast weights move rather than by what the tensor cores support, that ratio is the entire argument for this board.

Median rate against an A100 40GB $0.065 vs $0.417

Pipelines pinned to Volta

Code that was written against compute capability 7.0 and never migrated keeps running here without a rewrite. That is a real and stubborn category: training scripts, CUDA extensions built for sm_70, and FP32 simulation kernels that never needed a tensor core in the first place.

Compute capability 7.0

What 32 GB will not hold

An 8B at FP16 is about 16 GB and fits with room. A 14B at FP16 is about 28 GB and fits with none. A 70B does not fit at any precision this card can execute, and neither does anything that needs FP8 or BF16. Read that as a hard boundary, not a performance note.

BF16, FP8, MIG none of them

Volta against everything
that replaced it.

The V100 loses every column here except one, and that one is the reason its fleet is booked solid. Specs from the NVIDIA Volta, Ampere and Ada datasheets; all TFLOPS figures are dense, never with sparsity.

Tesla V100 SXM2 A100 40GB RTX 3090 RTX 4090
Architecture Volta GV100 Ampere GA100 Ampere GA102 Ada Lovelace
CUDA cores 5,120 6,912 10,496 16,384
VRAM 32 GB HBM2 40 GB HBM2e 24 GB GDDR6X 24 GB GDDR6X
Memory bandwidth 900 GB/s 1,555 GB/s 936 GB/s 1,008 GB/s
FP16 / BF16 (dense) 125 TFLOPS 312 TFLOPS 142 TFLOPS 165 TFLOPS
BF16 tensor cores no yes yes yes
Board power 300 W 400 W 350 W 450 W
Median on-demand 21 Sep 2026 $0.065 /GPU-h $0.417 $0.155 $0.375

specs from the NVIDIA V100, A100 and GeForce datasheets · rates are the median per-GPU hourly price across CLORE listings on 21 Sep 2026

Spot and on-demand
have converged.

On a fleet that is 185 of 186 servers rented, the discount for accepting preemption has mostly evaporated. Both medians landed on $0.065 per GPU-hour on 21 September 2026. Hosts set their own rates and live prices are on the marketplace.

Spot

$0.065 / GPU-hour
median of 186 listings · cheapest seen $0.038 · cheapest unrented $0.063
  • Same median as on-demand, at this snapshot
  • Billed per minute
  • An on-demand renter can take the machine back
  • Checkpoint often on a fleet this tight
Browse spot listings
NOT PREEMPTIBLE

On-demand

$0.065 / GPU-hour
median of 186 listings · cheapest seen $0.042 · cheapest unrented $0.083
  • Yours until you release it
  • No preemption
  • The unrented listings were priced above the median, which is what a tight fleet looks like
  • Billed per minute
Rent on demand

If waiting for a free server is not an option, V100 bare metal is sold from 8 GPUs on a 30-day minimum term: $0.48 per GPU-hour at 30 days, $0.43 at 90, $0.40 at 180 and $0.36 from 360, in the USA, France and Japan. Configure bare metal →

Pay with
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Renting into a tight fleet.

With one server free out of 186, the difference between getting a V100 and not getting one is mostly about how you watch and what you will accept.

01 / CHECK

See what is actually free

Filter by Tesla V100 and sort by availability, not by price. Confirm the memory: the fleet is overwhelmingly the 32 GB SXM2 part, but 16 GB boards exist.

02 / DECIDE

Spot or on-demand

Both medians sat at $0.065 per GPU-hour, so spot is not buying you much of a discount here. Take on-demand if an interruption would cost you a checkpoint.

marketplace filter → Tesla V100
03 / VERIFY

Check the toolchain before the run

Run one short job first. If a wheel ships only sm_80 kernels, you will find out in seconds rather than after an hour of billing.

04 / OR CONTRACT

Take a block instead

Bare metal removes the availability problem: 8 GPUs and 30 days minimum, $0.36 to $0.48 per GPU-hour by term, in three countries.

Questions about Volta in 2026.

Volta is compute capability 7.0. Which modern libraries refuse to build for it?

Anything whose kernels are gated on Ampere or newer. In practice that means FlashAttention-2 and FlashAttention-3, every FP8 code path, and every BF16 code path, because Volta has neither data type in hardware. Frameworks themselves still build for Volta, but individual wheels increasingly ship only sm_80 and above kernels, so the failure shows up as an unsupported-architecture error at import or at the first kernel launch rather than at install time. Check that the wheel you plan to use still ships sm_70 before you book a long run.

No BF16 on a V100. What does that mean for training stability against FP16 plus loss scaling?

It means you go back to managing dynamic range by hand. BF16 keeps the exponent range of FP32 and trades mantissa bits for it, which is why Ampere-and-later recipes can mostly ignore overflow. FP16 has a much narrower range, so you need loss scaling, and with a badly tuned scale you get either silent overflow to infinity or gradients that underflow to zero. Automatic mixed precision handles the common cases, but a recipe written and tuned in BF16 will not always transfer unchanged.

900 GB/s of HBM2 at a $0.065 median per GPU-hour. What is that cheap bandwidth genuinely good for?

Work that reads a lot of memory and does comparatively little arithmetic on it. Decoding tokens from a model small enough to fit in 32 GB, embedding or reranking passes over a large corpus, FP32 simulation kernels, and anything that spent its life being throttled by a consumer card memory bus. What it is not good for is anything that needs modern data types or a model too big for 32 GB, where a cheaper hourly rate cannot buy back the capability.

16 GB or 32 GB: which V100 am I getting, and how do I tell?

Both exist and the CLORE fleet is overwhelmingly the 32 GB SXM2 part, but overwhelmingly is not always. The listing carries the GPU name reported by the host agent, and the two variants appear under distinct names, so read it before you rent. On the machine itself, nvidia-smi prints the memory total in the first block. If your job needs 32 GB, confirm from the listing, not from the model name alone, and remember there has never been an 80 GB V100 whatever a spec sheet elsewhere tells you.

FlashAttention-3 is Hopper-only. What attention kernel should I use on Volta instead?

Let PyTorch choose. Its scaled_dot_product_attention dispatches across several backends, and on a Volta card it lands on the memory-efficient one rather than the flash one, which needs Ampere or newer. That gets you most of the memory saving without hand-written kernels. If a library hard-codes a FlashAttention import, look for a configuration switch back to the reference implementation before assuming the model will not run at all.

186 servers listed and one free. Can I actually get on one?

Not reliably, at that moment. The 21 September 2026 snapshot found 186 servers carrying 1,389 V100 cards, the largest datacenter fleet on the platform, and 185 of those servers were rented, leaving exactly one free with 4 free cards on it. That is a fleet that clears. Practical options are to watch the marketplace for a release, to bid on spot and accept that an on-demand renter can take the machine back, or to take a bare-metal V100 block: 8 GPUs minimum, 30 days minimum, $0.36 to $0.48 per GPU-hour by term, in the USA, France and Japan.

Three things Volta is still good at.

All three are bound by memory traffic rather than by tensor-core features, which is exactly where an eight-year-old HBM board still competes. Throughput depends on your model and batch, so measure before you budget.

Serving a model that fits
llama.cpp or a GGUF server, FP16 weights
8B at FP16 ≈ 16 GB of 32

Half the board for weights, half for context. Decode speed follows the 900 GB/s bus, not the 125 TFLOPS number.

Read the guide →
Batch transcription
Whisper large, FP16, large batches
32 GB for model plus batch

An offline queue where latency does not matter and hourly cost does. Nothing in the pipeline needs BF16 or FP8.

Read the guide →
An old training script
FP16 with loss scaling, sm_70 extensions
125 dense FP16 TFLOPS

Code that predates BF16 runs unchanged. Code written for BF16 usually needs its loss scaling put back before it will converge.

Read the guide →

What was listed, and what was free.

One snapshot of the CLORE marketplace, 21 September 2026. Free means not rented at that instant, which is a single reading rather than an availability rate. Rates are median per-GPU hourly prices derived from whole-server daily prices.

GPU
Memory
Mem BW (GB/s)
BF16
Servers listed
Free at the snapshot
Median on-demand
Median spot
Tesla V100 / this page
32 GB HBM2
900
—
186
1
$0.065
$0.065
RTX 3090
24 GB GDDR6X
936
yes
219
67
$0.155
$0.125
RTX 4090
24 GB GDDR6X
1,008
yes
243
130
$0.375
$0.365
A100 40GB
40 GB HBM2e
1,555
yes
6
4
$0.417
$0.333
Tesla T4
16 GB GDDR6
320
—
13
2
$0.018
$0.014

Guides that work on compute 7.0.

Nothing here depends on BF16, FP8 or FlashAttention. If a guide elsewhere assumes an Ampere kernel, that is the part you will have to replace.

Language Models
llama.cpp server
GGUF quantized inference with HTTP/OpenAI-compatible API.
Training
HF Transformers training
Train and fine-tune with the Trainer API.
Training
Jupyter for ML training
Notebook-driven training and experimentation.
Audio Voice
Whisper transcription
OpenAI Whisper-large for speech-to-text.
Image Generation
A1111 WebUI on CLORE.AI
The classic SD WebUI with extensions and LoRA.
Language Models
Mistral / Mixtral
Run Mistral 7B and Mixtral 8x7B / 8x22B.
Advanced
CLORE API integration
Programmatic order creation via the public API.
See all guides →

When the V100 fleet is full.

RTX 3090
936 GB/s, BF16, 67 servers free
Compare →
A100 40GB
1,555 GB/s, MIG, 4 servers free
Compare →
Tesla T4
Cheaper still, at 320 GB/s and 16 GB
Compare →

One free server
out of 186.

That was the reading on 21 September 2026. Supply moves, so check what is open now, or take a bare-metal block if you need the cards to still be yours next month.