The RTX 5080 is a bandwidth and format upgrade, not a memory one. The card holds the same 16 GB as a 4080 but feeds it at 960 GB/s instead of 716, and its fifth-generation tensor cores execute FP4, the same Blackwell feature the 5090 has. On 21 Sep 2026 the marketplace listed 141 servers carrying 267 of these cards, 74 servers unrented, at a median of $0.292 per GPU-hour on demand. Billing is per minute; settlement in BTC, CLORE, USDT or USDC.
Quantization is the lever on this card. FP8 halves a model against FP16 and FP4 halves it again, so what changes between formats is not speed alone but whether the job fits at all.
Llama 3 8B costs roughly 16 GB at FP16, about 8 GB at FP8 and about 5 GB at FP4. Only the last two leave room for a KV cache worth serving from, which is why the format matters more than the clock on this particular card. FP4 came with Blackwell as a generation, so the 5090 has it too; the 5080 is the cheaper way in.
SDXL at 1024 by 1024 with a batch of four is the workload this capacity suits, and Flux.1 at FP4 is the one the generation adds. Flux at fp16 is about 24 GB and belongs on a larger card.
With weights compressed to FP8 or FP4, most of the 16 GB is left for KV cache, and cache is what concurrency is made of. TensorRT-LLM is the stack that actually uses the FP4 path.
Hardware from NVIDIA's published specifications. The price row is the median per-GPU hourly rate across quoted listings on 21 September 2026; today's number lives on the marketplace.
there is no TFLOPS row here: NVIDIA's published 50-series tensor figures are inconsistent between sources, so this page declines to pick one
The figures below are medians per GPU-hour across the 129 RTX 5080 servers carrying a quote on 21 September 2026. Spot and on-demand sat within half a cent of each other that day, which is unusual and will not always hold.
The 5080 carries the lowest bare-metal entry of the Blackwell parts: blocks start at two GPUs on a 30-day term, quoted $0.40 to $0.61 per GPU-hour in the USA, the EU and Japan on 21 Sep 2026, with the lower end reserved for longer commitments. Price a two-card block →
On a 16 GB card the order of operations matters: decide the format before you pick the server, because it decides whether the job fits.
Parameters times bytes per parameter, plus 10 to 30 per cent for cache and activations. If that exceeds 16, change the format or the card.
Spot can be outbid and reclaimed; on-demand cannot. Both were within half a cent of each other at the snapshot.
Pull from any registry the host can reach. The card is passed into the container and you are root inside it.
Read the memory the process reserves. If it did not fall, the runtime widened your FP4 weights and you are paying for a format you are not using.
Mostly footprint, and throughput follows from it. FP4 stores a weight in half a byte, so a model weighs about a quarter of its FP16 size and half its FP8 size, and the fifth-generation tensor cores execute in that format instead of widening it first. On a 16 GB card the memory saving is the whole point, because it is what lets the weights and a useful batch live on the card at once. Quality is not free: 4-bit quantization discards information, and whether that matters is settled by evaluating your own model, not by reading a spec sheet.
In anything that has to read the entire model once per token. Generation is memory-bound, so it tracks bandwidth closely, and 960 GB/s against a 4080 at 716 GB/s is about 34 per cent more traffic at identical 16 GB of capacity. Prefill, image diffusion and rendering are compute-bound and see far less of that gap. If you are serving batch-1 text, the bandwidth row is the one that decides the card. If you are grinding through diffusion steps in a large batch, it barely moves.
Cheaper, and better only under a condition. Llama 3 8B at FP16 is roughly 16 GB of weights, which is the entire 5080, so here you serve that model at FP8 or FP4 and spend what is left on KV cache. A 4090 has 24 GB and 1,008 GB/s, so it holds the same weights at FP16 with cache to spare. At the 21 September 2026 medians the 5080 quoted $0.292 per GPU-hour against $0.375 for the 4090. If a quantized 8B passes your evaluation, the 5080 is the cheaper serving unit. If you need full-precision weights or a long context, it is not the card.
Yes. FP4 changes how much model fits into 16 GB; it does not add memory. Llama 3 8B at FP16 leaves almost nothing over for the KV cache, a 13B model does not fit at FP16 at all, and 70B does not fit at any precision. Blackwell makes 16 GB stretch further. It does not make 16 GB behave like 24.
TensorRT-LLM is the route NVIDIA ships FP4 through, and checkpoints generally have to be quantized for it rather than loaded as they are. Support elsewhere varies by release, and the failure mode is quiet: the runtime accepts the request, executes in a wider format and returns correct output at FP8 or FP16 speed with no warning. The reliable check is the memory the process reserves, which will not fall if the format never changed.
No. FP4 arrived with the Blackwell generation as a whole, and the RTX 5090 carries the same fifth-generation tensor cores. The 5080 is the cheaper way into that generation, not the only one. Any page claiming it was first is wrong, and this one used to be one of them.
The same 8B checkpoint at three precisions, priced in gigabytes. Numbers are the arithmetic of parameters times bytes, not measured throughput, which depends on your runtime and batch.
Weights alone consume the capacity, leaving almost nothing for a KV cache. Technically loadable, practically not servable.
Read the guide →Ada could do this too. What Blackwell adds is the option below it, and the bandwidth to feed either one faster.
Read the guide →The generation's new format, and the one that turns this into a concurrency card. Verify the quality loss on your own evaluation set.
Read the guide →The 5080 and the 5090 share an architecture and a tensor format. Everything that separates them is in this table, alongside the Ada card they both replace.
Start with the serving and fine-tuning guides: both spend most of their pages on memory, which is the constraint that decides everything else on this card.
Renters come for Blackwell FP4 on GDDR7. Set your own price, get paid per rented minute in BTC, CLORE, USDT or USDC, and add bare-metal contracts as your rig grows.
74 of the 141 listed RTX 5080 servers were free at the 21 Sep 2026 snapshot, and bare metal starts at two cards. Check the marketplace for what either costs today.