Log in Rent RTX 4080
RTX 4080 · marketplace snapshot 21 Sep 2026 · 15 servers, 9 free

Sixteen gigabytes.
One RTX 4080.
One 8B model at FP16.

Sixteen gigabytes is the whole story of this card. A Llama 3 8B checkpoint at FP16 weighs about 16 GB, so the RTX 4080 is the smallest GPU on this marketplace where that model loads at full precision, and the first where you have to budget what is left for the KV cache. Quantise the same weights to FP8 and they drop to roughly 8 GB, which is where the card gets comfortable. Billed per minute, paid in BTC, USDT, USDC or CLORE.

●Per-minute billing ●16 GB GDDR6X, 716.8 GB/s ●FP8 tensor cores ●No MIG on GeForce
MARKETPLACE SNAPSHOT 21 Sep 2026
# plain RTX 4080 listings only, 21 September 2026 servers listed 15 cards on those servers 20 servers not rented 9 on-demand median $0.292 per GPU-hour spot median $0.250 per GPU-hour lowest quoted spot $0.106 per GPU-hour # the RTX 4080 SUPER shares this filter string and is a different card # every price is set by its host and changes when the host edits the listing
GPU
RTX 4080 ×1
VRAM
16 GB
Median on-demand
$0.292/hr
As of
21 Sep 2026
15
Plain RTX 4080 servers listed, 21 Sep 2026
20
RTX 4080 cards across those servers
$0.292/hr
Median on-demand price per GPU-hour
16GB
GDDR6X per card at 716.8 GB/s

What 16 GB fits,
and where it runs out.

The RTX 4080 pairs 9,728 Ada cores with 16 GB of GDDR6X at 716.8 GB/s. Every workload below is sized against that 16 GB, because on this card capacity is the binding constraint long before compute is.

One 8B checkpoint at FP16, and that is the card

At two bytes per parameter an 8B model is about 16 GB of weights, which is all of it. In practice you serve at FP8 or INT8 instead, around 8 GB, and spend the other half on KV cache and activations. Ada's fourth-generation tensor cores do FP8 in hardware, so the quantised path is the intended one.

Llama 3 8B weights, FP16 ~16 GB

SDXL at 1024² and Flux.1 schnell

Diffusion is where 16 GB stops being tight. SDXL at 1024 by 1024 runs at batch 4, and Flux.1 schnell stays resident without pushing the text encoder into system RAM. Flux.1 dev at fp16 is the one that still needs offloading.

SDXL 1024², batch 4 fits in 16 GB

QLoRA on 7B and 8B, with room for batch

Four-bit weights for an 8B base come to roughly 4 GB, leaving twelve for gradients, optimizer state and a batch size that keeps the card busy. A 13B model at FP16 needs about 26 GB and belongs somewhere else entirely.

8B QLoRA, 4-bit weights ~4 GB

The 4090 gives you
eight more gigabytes.

Against a 4090 the 4080 gives up 8 GB of capacity, about 290 GB/s of bandwidth and 9,728 cores against 16,384, while drawing 130 W less. Both are Ada Lovelace and both do FP8 in hardware, so there is no feature gap, only a size one.

RTX 4080 RTX 4090 RTX 5080 RTX 3090
Architecture Ada Lovelace AD103 Ada Lovelace AD102 Blackwell GB203 Ampere GA102
CUDA cores 9,728 16,384 10,752 10,496
VRAM 16 GB GDDR6X 24 GB GDDR6X 16 GB GDDR7 24 GB GDDR6X
Memory bandwidth 716.8 GB/s 1,008 GB/s 960 GB/s 936 GB/s
Board power 320 W 450 W 360 W 350 W
Median on-demand $/GPU-hr $0.292 $0.375 $0.292 $0.155

silicon figures from NVIDIA product pages · medians from the marketplace on 21 Sep 2026

Spot or on-demand.
Billed by the minute either way.

Each host prices their own server, so there is a spread rather than one rate. The two figures below are the medians across the 15 RTX 4080 servers that were listed on 21 September 2026.

Spot

$0.250 / GPU-hr
median of 15 listed servers · lowest quoted $0.106
  • You bid; the highest bid holds the server
  • An on-demand order can take the machine back
  • Marketplace fee 2.5%, split with the host
  • Suited to checkpointed training and rendering
See spot listings
RESERVED

On-demand

$0.292 / GPU-hr
median of 15 listed servers · lowest quoted $0.175
  • Fixed price set by the host, no preemption
  • 9 of the 15 servers were free at the snapshot
  • Marketplace fee 10%, split with the host
  • Suited to serving, interactive work and demos
See on-demand listings
Pay with
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Fund, filter, connect, stop.

There is no quota request and no minimum term on a marketplace rental. The sequence below is the whole of it.

01 / FUND

Put a balance in place

Deposit BTC, USDT, USDC or CLORE. Orders settle against that balance minute by minute, so nothing is committed in advance.

02 / FILTER

Find the 4080s

Set the marketplace GPU filter to the model string. It matches as a substring, so RTX 4080 SUPER servers and mixed rigs come back alongside plain 4080s. Read the model name on each listing before you order.

RTX 4080
03 / CONNECT

Bring your own image

Point the order at any image you can pull, set the ports and environment you need, and connect over SSH once the container is up.

04 / STOP

Cancel when you are done

Billing is per minute and ends with the order. A twenty-minute experiment costs twenty minutes.

Sixteen gigabytes, answered.

Llama 3 8B at FP16 is about 16 GB. Does it actually fit on a 4080, and what is left for the KV cache?

The weights fit and almost nothing else does. At two bytes per parameter an 8B checkpoint is roughly 16 GB, which is the whole card, so an FP16 load leaves no usable KV cache and falls over as soon as the context grows. The working configuration is FP8 or INT8 weights at roughly 8 GB, which leaves about half the card for cache, activations and CUDA context. Ada's fourth-generation tensor cores support FP8 in hardware, so this is a supported path rather than a memory-saving hack.

What does the 4080 lose to the 4090 besides VRAM?

Cores and bandwidth. The 4080 carries 9,728 CUDA cores against the 4090's 16,384, and 716.8 GB/s of memory bandwidth against 1,008 GB/s. Both are Ada Lovelace with the same FP8 tensor support, so no feature is missing. The 8 GB capacity gap decides which models load at all; the rest only decides how fast the ones that fit will run.

Is FP8 on Ada worth more on a 16 GB card than on a 24 GB one?

On 16 GB it decides whether a model loads, which matters more than speed. Halving 8B weights from about 16 GB to about 8 GB turns an impossible load into a comfortable one. On a 24 GB card the same model already fits at FP16, so there FP8 buys throughput and a larger batch instead. The narrower the card, the more of FP8's value is capacity rather than performance.

When is 716 GB/s the bottleneck rather than the compute?

During token generation. Producing each token reads the entire weight set once, so single-stream generation speed tracks memory bandwidth far more closely than core count. Prompt processing, diffusion sampling and training steps behave the other way and are compute-bound. Serving one user at a time, bandwidth is your ceiling; batching many requests pushes the limit back towards the cores.

Why is the 4080 fleet on CLORE so small, and what does that mean for availability?

On 21 September 2026 the marketplace listed 15 servers carrying 20 RTX 4080 cards, 9 of those servers not rented. That is the thinnest consumer fleet on the platform, so the median price moves with the behaviour of a handful of hosts rather than a broad market. Every count on this page is plain RTX 4080 only. The filtered marketplace view returns more than that, because the substring match also catches the RTX 4080 SUPER, a separate and faster 16 GB product. Treat the figures here as a dated snapshot and read the model name on a listing before you order.

Can I split one RTX 4080 between two tenants with MIG?

No. MIG is a datacenter feature and NVIDIA lists it only for A100, A30, H100, H200 and B200 class parts. No GeForce card supports it, the 4080 included. If you need hard isolation between two workloads, rent two servers rather than partitioning one.

Three jobs, sized in gigabytes.

Each card states what the workload costs in VRAM, because that is the number which decides whether it runs on a 4080 at all. Throughput depends on your settings, so it is not quoted here.

Llama 3 8B served at FP8
vLLM or TensorRT-LLM
~8 GB weights, ~8 GB spare

Quantising halves the checkpoint and hands the other half of the card to the KV cache, which is what sets your usable context and concurrency.

Read the guide →
SDXL at 1024 by 1024
ComfyUI, fp16
batch 4 without offload

UNet, VAE and both text encoders stay resident, so a batch-4 pass never round-trips through system RAM.

Read the guide →
QLoRA on an 8B base
PEFT, 4-bit NF4
~4 GB weights, 12 GB free

Twelve spare gigabytes go to optimizer state, gradients and batch size, which is why 16 GB is a comfortable fine-tuning card and 12 GB is not.

Read the guide →

What was listed on 21 September 2026.

Counts and medians come from a single read of the marketplace on 21 September 2026, and they count plain RTX 4080 cards only. The RTX 4080 SUPER is a separate, faster 16 GB product, and because the marketplace GPU filter is a substring match a search for RTX 4080 returns both. Free counts are an instant rather than an average, and any host can change a price at any time.

GPU
VRAM
Mem BW (GB/s)
Board power
Servers listed
Not rented
Median on-demand
Median spot
RTX 4080 / this page
16 GB GDDR6X
716.8
320 W
15
9
$0.292
$0.250
RTX 5080
16 GB GDDR7
960
360 W
141
74
$0.292
$0.289
RTX 4090
24 GB GDDR6X
1,008
450 W
243
130
$0.375
$0.365
RTX 5090
32 GB GDDR7
1,792
575 W
260
137
$0.542
$0.521

Walkthroughs for the workloads above.

Each of these lives on docs.clore.ai. The VRAM figures earlier on this page tell you which ones a 16 GB card will carry.

Image Generation
Flux.1 on CLORE.AI
Run Black Forest Labs' Flux for state-of-the-art image gen.
Image Generation
ComfyUI on CLORE.AI
Node-based pipeline for SDXL, Flux, and SD3.
Language Models
vLLM serving
High-throughput LLM serving with PagedAttention.
Training
LLM fine-tuning
LoRA / QLoRA fine-tuning workflow.
Video Generation
Stable Video Diffusion
Stability's image-to-video model.
Image Processing
ControlNet advanced
Pose, depth, edge guidance for SDXL.
Comparisons
Image Gen: ComfyUI vs SD WebUI vs Fooocus
Choosing an image-generation interface for a rented card.
See all guides →

If 16 GB is the wrong size.

RTX 5080
same 16 GB on GDDR7 · median $0.292
Rent →
RTX 4090
the 24 GB step up · median $0.375
Rent →
RTX 4070 Ti
12 GB, cheaper Ada · median $0.208
Rent →

Fifteen servers.
Nine of them free.

That was the RTX 4080 picture on 21 September 2026. The marketplace holds the current one, along with whatever each host is charging today.