Log in Rent RTX 5090
RTX 5090 · marketplace snapshot · 260 servers, 21 Sep 2026

Rent an RTX 5090
for its bandwidth.
For 70B.

The RTX 5090 pairs 32 GB of GDDR7 with 1,792 GB/s of memory bandwidth, roughly 1.78 times what a 4090 moves. On a decoder-bound job that bandwidth, rather than the extra 8 GB, is what you are renting. On 21 Sep 2026 the marketplace held 260 servers carrying 539 of these cards, 137 servers not rented, at a median of $0.542 per GPU-hour on demand. Billing is per minute; settlement is in BTC, CLORE, USDT or USDC.

●Billed per minute ●Root SSH inside your own image ●Spot & on-demand ●539 cards listed on 21 Sep 2026
$0.542/hr
Median on-demand price per GPU-hour, 21 Sep 2026
32GB
GDDR7 on every card
1,792GB/s
Memory bandwidth per card
137
Servers not rented at that snapshot, out of 260

What 32 GB actually holds,
and what it does not.

The memory budget decides which jobs belong on this card. Weights at FP16 cost about 2 GB per billion parameters, FP8 and INT8 about 1 GB, INT4 about 0.5 GB, and the KV cache and activations add 10 to 30 per cent on top.

A 32B model at INT4 is the single-card ceiling

Qwen2.5 32B quantised to 4-bit lands near 18 GB of weights, which leaves genuine room for context on a 32 GB card. Llama 3 70B at the same quantisation needs about 40 GB before any cache, so it does not fit here no matter how the listing is worded. Two cards, split with tensor parallelism, is the honest route to 70B.

Qwen2.5 32B at INT4 ~18 GB of weights

Diffusion that never touches system RAM

Flux.1 dev at fp16 is about 24 GB, so it stays resident with headroom instead of streaming weights over PCIe. Wan2.1 and HunyuanVideo at 720p are the two video models this card is most often rented for, and 32 GB is the reason they run without an offload path.

Flux.1 dev at fp16 ~24 GB resident

A small model with an unusually large cache

Llama 3 8B at FP16 is roughly 16 GB, so on this card half the memory is left over for KV cache, which is what raises concurrency rather than single-stream speed. Drop the same model to FP4 on the fifth-generation tensor cores and the weight budget falls again.

Llama 3 8B at FP16 ~16 GB of weights

1.78× the bandwidth,
1.45× the rent.

Hardware figures are NVIDIA's published specifications. The price row is the median per-GPU hourly rate observed across the marketplace on 21 September 2026; the live figure is always on the marketplace itself.

RTX 5090 RTX 4090 RTX 5080 vs 4090
Architecture Blackwell GB202 Ada AD102 Blackwell GB203 new generation
VRAM 32 GB GDDR7 24 GB GDDR6X 16 GB GDDR7 +8 GB
Memory bandwidth 1,792 GB/s 1,008 GB/s 960 GB/s 1.78×
CUDA cores 21,760 16,384 10,752 1.33×
Board power 575 W 450 W 360 W +125 W
FP4 tensor path yes no yes added by Blackwell
Median on-demand / GPU-hr $0.542 $0.375 $0.292 1.45×

no TFLOPS row: NVIDIA's published tensor figures for the 50-series disagree across sources, so this page quotes none

Hosts price their own cards.
Here is where they landed.

Both figures below are medians per GPU-hour across the 247 RTX 5090 servers that carried a quote on 21 September 2026. Individual listings sat well either side of them, and the current numbers are on the marketplace.

Spot

$0.521 / GPU-hr
median of 247 quoted servers · lowest quote $0.175
  • You bid; a higher bid can take the card back
  • Billed per minute of runtime
  • Marketplace fee 2.5%, split evenly with the host
  • Suited to work you can checkpoint and resume
See spot listings
FIXED PRICE

On-demand

$0.542 / GPU-hr
median of 247 quoted servers · $0.792 among those actually free
  • A fixed rate the host sets, no bidding
  • Cannot be preempted while your balance covers it
  • Marketplace fee 10%, split evenly with the host
  • The cheapest 5090s were already rented at the snapshot
See on-demand listings

Bare metal is the other route: RTX 5090 blocks start at 8 GPUs, from a 14-day term in the UK and 30 days elsewhere, quoted $0.64 to $1.06 per GPU-hour across the USA, the UK, Japan and Slovenia on 21 Sep 2026. Configure a bare-metal block →

Pay with
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Four steps, no procurement call.

Marketplace prices are quoted for the whole server, so the first thing worth doing is dividing by the card count.

01 / FILTER

Sort by price per card

Filter on RTX 5090, then compare per-GPU rates rather than per-server ones. Multi-card rigs look expensive until you divide.

02 / RENT

Bid, or take the fixed rate

Spot is an auction and can be taken back; on-demand is fixed and cannot. The same server often carries both.

$ clore rent --gpu "RTX 5090"
03 / CONNECT

Bring your own image

Any registry the host can reach works. You get root inside the container with the card passed through.

04 / STOP

Close the order

Billing accrues by the minute and stops when the order does, so a short experiment costs what it took.

The 5090 questions worth answering.

32 GB is not enough for 70B at INT4. What is the largest model that genuinely fits?

A 70-billion-parameter model at 4-bit needs roughly 40 GB for weights alone, at about 0.5 GB per billion parameters, before any KV cache. That does not fit 32 GB at any usable quantization. The largest genuinely single-card class here is a 32B model at INT4, Qwen2.5 32B for example, which lands near 18 GB of weights and leaves real room for context. For 70B you rent two cards and split the model across them.

How much of the 5090's advantage over a 4090 is bandwidth rather than compute?

Most of it. The 5090 moves 1,792 GB/s against the 4090's 1,008 GB/s, a 1.78x gap, while the CUDA core count rises only from 16,384 to 21,760, a 1.33x gap. Token-by-token generation is memory-bound, so it tracks the bandwidth number far more closely than the core count. Prefill and image diffusion are compute-bound and see much less of that 1.78x.

Is one 5090 cheaper per token than two cheaper cards?

It depends on where the job is bound, because what matters is throughput per dollar of rent. At the 21 September 2026 medians a 5090 was $0.542 per GPU-hour and a 4090 $0.375, so you pay a 1.45x premium for a 1.78x bandwidth advantage. When the job is bandwidth-bound and fits in 32 GB, one 5090 is the cheaper unit. When it fits in 24 GB and is compute-bound, two smaller cards can win.

Qwen2.5 32B at INT4 on one card, what context length does that leave?

Weights take about 18 GB, which leaves roughly 12 to 13 GB once the runtime and activations are accounted for. How much context that buys depends on the model's layer and head geometry and on the precision you keep the KV cache in, so the honest answer is that you get either a long context or high concurrency, not both, and you should measure it on the exact checkpoint you intend to serve rather than trust a table.

Video diffusion at 720p on 32 GB: what actually runs without offload?

Wan2.1 and HunyuanVideo at 720p are the two this card is usually rented for, and 32 GB is what keeps them resident instead of streaming weights through system RAM. Flux.1 dev at fp16 is about 24 GB and also fits with room to spare. Longer clips and higher resolutions still push past 32 GB, and there the answer is a larger card rather than a quantization trick.

Does the RTX 5090 support MIG partitioning?

No. MIG exists only on A100, A30, H100, H200, the B200 and GB200 class and the RTX PRO Blackwell parts. No GeForce card has it, checked against NVIDIA's MIG supported-GPU list on 21 September 2026. Isolation on CLORE.AI comes from renting the whole card inside your own container, not from partitioning it.

Three memory budgets on 32 GB.

Weight footprints below are arithmetic, not benchmarks: parameter count multiplied by bytes per parameter. Throughput depends on your runtime, batch size and context, so measure it on the card rather than trusting a headline.

Qwen2.5 32B at INT4
vLLM or TensorRT-LLM
~18 GB of weights, ~13 GB left

The largest model class that is genuinely single-card here. What remains goes to KV cache, so the trade is context length against concurrency.

Read the guide →
Flux.1 dev at fp16
ComfyUI
~24 GB resident, no offload path

Full-precision Flux sits inside the card with headroom instead of paging weights over PCIe, which is where the offload penalty usually comes from.

Read the guide →
Wan2.1 and HunyuanVideo at 720p
ComfyUI
the reason to take 32 GB over 24

Video diffusion is where the extra capacity earns its price. Longer clips and higher resolutions still overflow the card, and that is a bigger-GPU problem.

Read the guide →

Three cards this page can cite.

Hardware figures come from NVIDIA's published specifications. Price columns are medians per GPU-hour across quoted listings on 21 September 2026, not a promise about today.

GPU
VRAM
Mem BW (GB/s)
CUDA cores
Board power
FP4 path
Median spot $/GPU-hr
Median on-demand $/GPU-hr
RTX 5090 / this page
32 GB GDDR7
1,792
21,760
575 W
yes
$0.521
$0.542
RTX 4090
24 GB GDDR6X
1,008
16,384
450 W
no
$0.365
$0.375
RTX 5080
16 GB GDDR7
960
10,752
360 W
yes
$0.289
$0.292

Recipes for a 32 GB card.

Each guide names the image to pull and the settings to change. The video and image-generation ones are the pair that most needs the capacity this card has.

Image Generation
Flux.1 on CLORE.AI
Run Black Forest Labs' Flux for state-of-the-art image gen.
Language Models
vLLM serving
High-throughput LLM serving with PagedAttention.
Training
LLM fine-tuning
LoRA / QLoRA fine-tuning workflow.
Video Generation
Wan Video
Alibaba's Wan-2.1 text/image-to-video.
Video Generation
Hunyuan Video
Tencent's open video generation model.
Training
Kohya SS LoRA training
The standard SDXL LoRA training pipeline.
Comparisons
LLM serving stacks compared
Which serving framework to pick, and what each one costs you in memory.
See all guides →

If 32 GB is the wrong size.

RTX 4090
24 GB GDDR6X · cheaper per hour
Rent →
RTX 6000 Ada
Workstation Ada · more memory per card
Rent →

Rent the bandwidth,
not the headline.

137 of the 260 listed RTX 5090 servers were unrented at the 21 Sep 2026 snapshot. Prices move, so the marketplace is where you check what a card costs today.