The RTX 5090 pairs 32 GB of GDDR7 with 1,792 GB/s of memory bandwidth, roughly 1.78 times what a 4090 moves. On a decoder-bound job that bandwidth, rather than the extra 8 GB, is what you are renting. On 21 Sep 2026 the marketplace held 260 servers carrying 539 of these cards, 137 servers not rented, at a median of $0.542 per GPU-hour on demand. Billing is per minute; settlement is in BTC, CLORE, USDT or USDC.
The memory budget decides which jobs belong on this card. Weights at FP16 cost about 2 GB per billion parameters, FP8 and INT8 about 1 GB, INT4 about 0.5 GB, and the KV cache and activations add 10 to 30 per cent on top.
Qwen2.5 32B quantised to 4-bit lands near 18 GB of weights, which leaves genuine room for context on a 32 GB card. Llama 3 70B at the same quantisation needs about 40 GB before any cache, so it does not fit here no matter how the listing is worded. Two cards, split with tensor parallelism, is the honest route to 70B.
Flux.1 dev at fp16 is about 24 GB, so it stays resident with headroom instead of streaming weights over PCIe. Wan2.1 and HunyuanVideo at 720p are the two video models this card is most often rented for, and 32 GB is the reason they run without an offload path.
Llama 3 8B at FP16 is roughly 16 GB, so on this card half the memory is left over for KV cache, which is what raises concurrency rather than single-stream speed. Drop the same model to FP4 on the fifth-generation tensor cores and the weight budget falls again.
Hardware figures are NVIDIA's published specifications. The price row is the median per-GPU hourly rate observed across the marketplace on 21 September 2026; the live figure is always on the marketplace itself.
no TFLOPS row: NVIDIA's published tensor figures for the 50-series disagree across sources, so this page quotes none
Both figures below are medians per GPU-hour across the 247 RTX 5090 servers that carried a quote on 21 September 2026. Individual listings sat well either side of them, and the current numbers are on the marketplace.
Bare metal is the other route: RTX 5090 blocks start at 8 GPUs, from a 14-day term in the UK and 30 days elsewhere, quoted $0.64 to $1.06 per GPU-hour across the USA, the UK, Japan and Slovenia on 21 Sep 2026. Configure a bare-metal block →
Marketplace prices are quoted for the whole server, so the first thing worth doing is dividing by the card count.
Filter on RTX 5090, then compare per-GPU rates rather than per-server ones. Multi-card rigs look expensive until you divide.
Spot is an auction and can be taken back; on-demand is fixed and cannot. The same server often carries both.
Any registry the host can reach works. You get root inside the container with the card passed through.
Billing accrues by the minute and stops when the order does, so a short experiment costs what it took.
A 70-billion-parameter model at 4-bit needs roughly 40 GB for weights alone, at about 0.5 GB per billion parameters, before any KV cache. That does not fit 32 GB at any usable quantization. The largest genuinely single-card class here is a 32B model at INT4, Qwen2.5 32B for example, which lands near 18 GB of weights and leaves real room for context. For 70B you rent two cards and split the model across them.
Most of it. The 5090 moves 1,792 GB/s against the 4090's 1,008 GB/s, a 1.78x gap, while the CUDA core count rises only from 16,384 to 21,760, a 1.33x gap. Token-by-token generation is memory-bound, so it tracks the bandwidth number far more closely than the core count. Prefill and image diffusion are compute-bound and see much less of that 1.78x.
It depends on where the job is bound, because what matters is throughput per dollar of rent. At the 21 September 2026 medians a 5090 was $0.542 per GPU-hour and a 4090 $0.375, so you pay a 1.45x premium for a 1.78x bandwidth advantage. When the job is bandwidth-bound and fits in 32 GB, one 5090 is the cheaper unit. When it fits in 24 GB and is compute-bound, two smaller cards can win.
Weights take about 18 GB, which leaves roughly 12 to 13 GB once the runtime and activations are accounted for. How much context that buys depends on the model's layer and head geometry and on the precision you keep the KV cache in, so the honest answer is that you get either a long context or high concurrency, not both, and you should measure it on the exact checkpoint you intend to serve rather than trust a table.
Wan2.1 and HunyuanVideo at 720p are the two this card is usually rented for, and 32 GB is what keeps them resident instead of streaming weights through system RAM. Flux.1 dev at fp16 is about 24 GB and also fits with room to spare. Longer clips and higher resolutions still push past 32 GB, and there the answer is a larger card rather than a quantization trick.
No. MIG exists only on A100, A30, H100, H200, the B200 and GB200 class and the RTX PRO Blackwell parts. No GeForce card has it, checked against NVIDIA's MIG supported-GPU list on 21 September 2026. Isolation on CLORE.AI comes from renting the whole card inside your own container, not from partitioning it.
Weight footprints below are arithmetic, not benchmarks: parameter count multiplied by bytes per parameter. Throughput depends on your runtime, batch size and context, so measure it on the card rather than trusting a headline.
The largest model class that is genuinely single-card here. What remains goes to KV cache, so the trade is context length against concurrency.
Read the guide →Full-precision Flux sits inside the card with headroom instead of paging weights over PCIe, which is where the offload penalty usually comes from.
Read the guide →Video diffusion is where the extra capacity earns its price. Longer clips and higher resolutions still overflow the card, and that is a bigger-GPU problem.
Read the guide →Hardware figures come from NVIDIA's published specifications. Price columns are medians per GPU-hour across quoted listings on 21 September 2026, not a promise about today.
Each guide names the image to pull and the settings to change. The video and image-generation ones are the pair that most needs the capacity this card has.
Set your own price and get paid for every rented minute in BTC, CLORE, USDT or USDC, or supply 5090s to Clore on bare-metal contracts.
137 of the 260 listed RTX 5090 servers were unrented at the 21 Sep 2026 snapshot. Prices move, so the marketplace is where you check what a card costs today.