Sixteen gigabytes is the whole story of this card. A Llama 3 8B checkpoint at FP16 weighs about 16 GB, so the RTX 4080 is the smallest GPU on this marketplace where that model loads at full precision, and the first where you have to budget what is left for the KV cache. Quantise the same weights to FP8 and they drop to roughly 8 GB, which is where the card gets comfortable. Billed per minute, paid in BTC, USDT, USDC or CLORE.
The RTX 4080 pairs 9,728 Ada cores with 16 GB of GDDR6X at 716.8 GB/s. Every workload below is sized against that 16 GB, because on this card capacity is the binding constraint long before compute is.
At two bytes per parameter an 8B model is about 16 GB of weights, which is all of it. In practice you serve at FP8 or INT8 instead, around 8 GB, and spend the other half on KV cache and activations. Ada's fourth-generation tensor cores do FP8 in hardware, so the quantised path is the intended one.
Diffusion is where 16 GB stops being tight. SDXL at 1024 by 1024 runs at batch 4, and Flux.1 schnell stays resident without pushing the text encoder into system RAM. Flux.1 dev at fp16 is the one that still needs offloading.
Four-bit weights for an 8B base come to roughly 4 GB, leaving twelve for gradients, optimizer state and a batch size that keeps the card busy. A 13B model at FP16 needs about 26 GB and belongs somewhere else entirely.
Against a 4090 the 4080 gives up 8 GB of capacity, about 290 GB/s of bandwidth and 9,728 cores against 16,384, while drawing 130 W less. Both are Ada Lovelace and both do FP8 in hardware, so there is no feature gap, only a size one.
silicon figures from NVIDIA product pages · medians from the marketplace on 21 Sep 2026
Each host prices their own server, so there is a spread rather than one rate. The two figures below are the medians across the 15 RTX 4080 servers that were listed on 21 September 2026.
There is no quota request and no minimum term on a marketplace rental. The sequence below is the whole of it.
Deposit BTC, USDT, USDC or CLORE. Orders settle against that balance minute by minute, so nothing is committed in advance.
Set the marketplace GPU filter to the model string. It matches as a substring, so RTX 4080 SUPER servers and mixed rigs come back alongside plain 4080s. Read the model name on each listing before you order.
Point the order at any image you can pull, set the ports and environment you need, and connect over SSH once the container is up.
Billing is per minute and ends with the order. A twenty-minute experiment costs twenty minutes.
The weights fit and almost nothing else does. At two bytes per parameter an 8B checkpoint is roughly 16 GB, which is the whole card, so an FP16 load leaves no usable KV cache and falls over as soon as the context grows. The working configuration is FP8 or INT8 weights at roughly 8 GB, which leaves about half the card for cache, activations and CUDA context. Ada's fourth-generation tensor cores support FP8 in hardware, so this is a supported path rather than a memory-saving hack.
Cores and bandwidth. The 4080 carries 9,728 CUDA cores against the 4090's 16,384, and 716.8 GB/s of memory bandwidth against 1,008 GB/s. Both are Ada Lovelace with the same FP8 tensor support, so no feature is missing. The 8 GB capacity gap decides which models load at all; the rest only decides how fast the ones that fit will run.
On 16 GB it decides whether a model loads, which matters more than speed. Halving 8B weights from about 16 GB to about 8 GB turns an impossible load into a comfortable one. On a 24 GB card the same model already fits at FP16, so there FP8 buys throughput and a larger batch instead. The narrower the card, the more of FP8's value is capacity rather than performance.
During token generation. Producing each token reads the entire weight set once, so single-stream generation speed tracks memory bandwidth far more closely than core count. Prompt processing, diffusion sampling and training steps behave the other way and are compute-bound. Serving one user at a time, bandwidth is your ceiling; batching many requests pushes the limit back towards the cores.
On 21 September 2026 the marketplace listed 15 servers carrying 20 RTX 4080 cards, 9 of those servers not rented. That is the thinnest consumer fleet on the platform, so the median price moves with the behaviour of a handful of hosts rather than a broad market. Every count on this page is plain RTX 4080 only. The filtered marketplace view returns more than that, because the substring match also catches the RTX 4080 SUPER, a separate and faster 16 GB product. Treat the figures here as a dated snapshot and read the model name on a listing before you order.
No. MIG is a datacenter feature and NVIDIA lists it only for A100, A30, H100, H200 and B200 class parts. No GeForce card supports it, the 4080 included. If you need hard isolation between two workloads, rent two servers rather than partitioning one.
Each card states what the workload costs in VRAM, because that is the number which decides whether it runs on a 4080 at all. Throughput depends on your settings, so it is not quoted here.
Quantising halves the checkpoint and hands the other half of the card to the KV cache, which is what sets your usable context and concurrency.
Read the guide →UNet, VAE and both text encoders stay resident, so a batch-4 pass never round-trips through system RAM.
Read the guide →Twelve spare gigabytes go to optimizer state, gradients and batch size, which is why 16 GB is a comfortable fine-tuning card and 12 GB is not.
Read the guide →Counts and medians come from a single read of the marketplace on 21 September 2026, and they count plain RTX 4080 cards only. The RTX 4080 SUPER is a separate, faster 16 GB product, and because the marketplace GPU filter is a substring match a search for RTX 4080 returns both. Free counts are an instant rather than an average, and any host can change a price at any time.
Each of these lives on docs.clore.ai. The VRAM figures earlier on this page tell you which ones a 16 GB card will carry.
Hosts set their own price and are paid for every minute the card is rented, in BTC, USDT, USDC or CLORE.
That was the RTX 4080 picture on 21 September 2026. The marketplace holds the current one, along with whatever each host is charging today.