Log in Rent RTX 5080
RTX 5080 · 21 Sep 2026 snapshot · 141 servers, 267 cards

RTX 5080 rentals:
GDDR7 on 16 GB.
Extra capacity.

The RTX 5080 is a bandwidth and format upgrade, not a memory one. The card holds the same 16 GB as a 4080 but feeds it at 960 GB/s instead of 716, and its fifth-generation tensor cores execute FP4, the same Blackwell feature the 5090 has. On 21 Sep 2026 the marketplace listed 141 servers carrying 267 of these cards, 74 servers unrented, at a median of $0.292 per GPU-hour on demand. Billing is per minute; settlement in BTC, CLORE, USDT or USDC.

●Charged by the minute ●Your container, root inside it ●Bid on spot or take the fixed rate ●Bare metal from two cards
MARKETPLACE SNAPSHOT 21 Sep 2026 · 18:05 UTC
# RTX 5080 inventory, counted from the per-server payload servers listed 141 cards listed 267 servers unrented 74 (184 cards) servers rented 67 # per GPU-hour, across the 129 servers carrying a quote on-demand median $0.292 · lowest quote $0.094 spot median $0.289 · lowest quote $0.094 # prices move; the marketplace carries today's
GPU
RTX 5080
VRAM
16 GB
Median on-demand
$0.292/GPU-hr
Unrented
74 of 141
$0.292/hr
Median hourly rate per card, on demand (21 Sep 2026)
16GB
GDDR7, and that is the whole budget
960GB/s
Bandwidth, about 34% over a 4080
74
Free servers out of 141 at the snapshot

Sixteen gigabytes,
spent carefully.

Quantization is the lever on this card. FP8 halves a model against FP16 and FP4 halves it again, so what changes between formats is not speed alone but whether the job fits at all.

FP4 is what makes 16 GB workable

Llama 3 8B costs roughly 16 GB at FP16, about 8 GB at FP8 and about 5 GB at FP4. Only the last two leave room for a KV cache worth serving from, which is why the format matters more than the clock on this particular card. FP4 came with Blackwell as a generation, so the 5090 has it too; the 5080 is the cheaper way in.

Llama 3 8B at FP4 ~5 GB of weights

Image work that fits comfortably

SDXL at 1024 by 1024 with a batch of four is the workload this capacity suits, and Flux.1 at FP4 is the one the generation adds. Flux at fp16 is about 24 GB and belongs on a larger card.

SDXL at 1024² batch 4 fits

Serving small models to many callers

With weights compressed to FP8 or FP4, most of the 16 GB is left for KV cache, and cache is what concurrency is made of. TensorRT-LLM is the stack that actually uses the FP4 path.

Llama 3 8B at FP8 ~8 GB of weights

Same capacity as Ada,
fed a third faster.

Hardware from NVIDIA's published specifications. The price row is the median per-GPU hourly rate across quoted listings on 21 September 2026; today's number lives on the marketplace.

RTX 5080 RTX 5090 RTX 4090 vs 5090
Architecture Blackwell GB203 Blackwell GB202 Ada AD102 same family
VRAM 16 GB GDDR7 32 GB GDDR7 24 GB GDDR6X half
Memory bandwidth 960 GB/s 1,792 GB/s 1,008 GB/s 0.54×
CUDA cores 10,752 21,760 16,384 0.49×
Board power 360 W 575 W 450 W -215 W
FP4 tensor path yes yes no identical
Median on-demand / GPU-hr $0.292 $0.542 $0.375 0.54×

there is no TFLOPS row here: NVIDIA's published 50-series tensor figures are inconsistent between sources, so this page declines to pick one

Two order types,
and a two-card bare-metal floor.

The figures below are medians per GPU-hour across the 129 RTX 5080 servers carrying a quote on 21 September 2026. Spot and on-demand sat within half a cent of each other that day, which is unusual and will not always hold.

Spot

$0.289 / GPU-hr
median of 129 quotes · lowest $0.094 · $0.125 among free servers
  • An auction: outbid, and you lose the card
  • Metered to the minute
  • 2.5% marketplace fee, half of it on the host
  • Sensible for restartable batch jobs
Look at spot listings
FIXED PRICE

On-demand

$0.292 / GPU-hr
median of 129 quotes · $0.417 among the 74 free servers
  • Priced by the host, no bidding involved
  • Held for you while the balance lasts
  • 10% marketplace fee, half of it on the host
  • Free servers priced above the all-listings median
Look at on-demand listings

The 5080 carries the lowest bare-metal entry of the Blackwell parts: blocks start at two GPUs on a 30-day term, quoted $0.40 to $0.61 per GPU-hour in the USA, the EU and Japan on 21 Sep 2026, with the lower end reserved for longer commitments. Price a two-card block →

Pay with
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Quantize first, then rent.

On a 16 GB card the order of operations matters: decide the format before you pick the server, because it decides whether the job fits.

01 / BUDGET

Count the gigabytes

Parameters times bytes per parameter, plus 10 to 30 per cent for cache and activations. If that exceeds 16, change the format or the card.

02 / RENT

Choose the order type

Spot can be outbid and reclaimed; on-demand cannot. Both were within half a cent of each other at the snapshot.

$ clore rent --gpu "RTX 5080"
03 / CONNECT

Ship your own image

Pull from any registry the host can reach. The card is passed into the container and you are root inside it.

04 / VERIFY

Check the format took

Read the memory the process reserves. If it did not fall, the runtime widened your FP4 weights and you are paying for a format you are not using.

Five things people get wrong about this card.

What does FP4 change in practice on a 16 GB card, quality, throughput, or both?

Mostly footprint, and throughput follows from it. FP4 stores a weight in half a byte, so a model weighs about a quarter of its FP16 size and half its FP8 size, and the fifth-generation tensor cores execute in that format instead of widening it first. On a 16 GB card the memory saving is the whole point, because it is what lets the weights and a useful batch live on the card at once. Quality is not free: 4-bit quantization discards information, and whether that matters is settled by evaluating your own model, not by reading a spec sheet.

GDDR7 at 960 GB/s versus GDDR6X at 716, where does that actually show up?

In anything that has to read the entire model once per token. Generation is memory-bound, so it tracks bandwidth closely, and 960 GB/s against a 4080 at 716 GB/s is about 34 per cent more traffic at identical 16 GB of capacity. Prefill, image diffusion and rendering are compute-bound and see far less of that gap. If you are serving batch-1 text, the bandwidth row is the one that decides the card. If you are grinding through diffusion steps in a large batch, it barely moves.

Is the 5080 a better 8B serving card than a 4090, or only a cheaper one?

Cheaper, and better only under a condition. Llama 3 8B at FP16 is roughly 16 GB of weights, which is the entire 5080, so here you serve that model at FP8 or FP4 and spend what is left on KV cache. A 4090 has 24 GB and 1,008 GB/s, so it holds the same weights at FP16 with cache to spare. At the 21 September 2026 medians the 5080 quoted $0.292 per GPU-hour against $0.375 for the 4090. If a quantized 8B passes your evaluation, the 5080 is the cheaper serving unit. If you need full-precision weights or a long context, it is not the card.

Does 16 GB still bottleneck me on Blackwell the way it did on Ada?

Yes. FP4 changes how much model fits into 16 GB; it does not add memory. Llama 3 8B at FP16 leaves almost nothing over for the KV cache, a 13B model does not fit at FP16 at all, and 70B does not fit at any precision. Blackwell makes 16 GB stretch further. It does not make 16 GB behave like 24.

Which frameworks can use the 5080's FP4 path today, and which silently fall back?

TensorRT-LLM is the route NVIDIA ships FP4 through, and checkpoints generally have to be quantized for it rather than loaded as they are. Support elsewhere varies by release, and the failure mode is quiet: the runtime accepts the request, executes in a wider format and returns correct output at FP8 or FP16 speed with no warning. The reliable check is the memory the process reserves, which will not fall if the format never changed.

Was the 5080 the first consumer card with native FP4 tensor cores?

No. FP4 arrived with the Blackwell generation as a whole, and the RTX 5090 carries the same fifth-generation tensor cores. The 5080 is the cheaper way into that generation, not the only one. Any page claiming it was first is wrong, and this one used to be one of them.

One model, three formats.

The same 8B checkpoint at three precisions, priced in gigabytes. Numbers are the arithmetic of parameters times bytes, not measured throughput, which depends on your runtime and batch.

Llama 3 8B at FP16
the format that does not fit
~16 GB, the entire card

Weights alone consume the capacity, leaving almost nothing for a KV cache. Technically loadable, practically not servable.

Read the guide →
Llama 3 8B at FP8
the safe default
~8 GB, half the card free

Ada could do this too. What Blackwell adds is the option below it, and the bandwidth to feed either one faster.

Read the guide →
Llama 3 8B at FP4
TensorRT-LLM
~5 GB, the rest is cache

The generation's new format, and the one that turns this into a concurrency card. Verify the quality loss on your own evaluation set.

Read the guide →

What the cheaper Blackwell gives up.

The 5080 and the 5090 share an architecture and a tensor format. Everything that separates them is in this table, alongside the Ada card they both replace.

GPU
VRAM
Mem BW (GB/s)
CUDA cores
Board power
FP4 path
Median spot $/GPU-hr
Bare-metal floor
RTX 5080 / this page
16 GB GDDR7
960
10,752
360 W
yes
$0.289
2 GPUs / 30 days
RTX 5090
32 GB GDDR7
1,792
21,760
575 W
yes
$0.521
8 GPUs / 14 days
RTX 4090
24 GB GDDR6X
1,008
16,384
450 W
no
$0.365
5 GPUs / 14 days

Setups that respect 16 GB.

Start with the serving and fine-tuning guides: both spend most of their pages on memory, which is the constraint that decides everything else on this card.

Image Generation
Flux.1 on CLORE.AI
Run Black Forest Labs' Flux for state-of-the-art image gen.
Image Generation
ComfyUI on CLORE.AI
Node-based pipeline for SDXL, Flux, and SD3.
Language Models
vLLM serving
High-throughput LLM serving with PagedAttention.
Training
LLM fine-tuning
LoRA / QLoRA fine-tuning workflow.
Video Generation
Wan Video
Alibaba's Wan-2.1 text/image-to-video.
Training
Kohya SS LoRA training
The standard SDXL LoRA training pipeline.
Comparisons
Serving frameworks side by side
What each stack supports, and how much memory it charges you for it.
See all guides →

When 16 GB runs out.

RTX 4080
The Ada card at the same capacity
Rent →
RTX 4090
24 GB GDDR6X, 1,008 GB/s
Rent →
RTX 5090
32 GB GDDR7, same FP4 path
Rent →

Cheapest way into
Blackwell, both routes.

74 of the 141 listed RTX 5080 servers were free at the 21 Sep 2026 snapshot, and bare metal starts at two cards. Check the marketplace for what either costs today.