Log in Rent RTX 3090
RTX 3090 · 219 servers listed · snapshot of 21 Sep 2026

Rent an RTX 3090
for 24 GB of 4-bit headroom.
At a 4090 budget.

The 3090 is the cheapest 24 GB consumer board on this marketplace and the only consumer Ampere board on this marketplace with an NVLink connector. That combination decides what it is good for: 4-bit adapter training where the base model, the KV cache and the activations all have to share one pool, and paired serving where a 40 GB quantised 70B is split across two boards. Median on-demand price on 21 Sep 2026 was $0.155 per GPU-hour, spot $0.125. Billing is per minute; settlement is in BTC, USDT, USDC or CLORE.

●24 GB GDDR6X, 936 GB/s ●NVLink 3 on paired boards ●Root access, your own image ●Also sold as a bare-metal contract
MARKETPLACE SNAPSHOT 21 Sep 2026 18:05 UTC
# counted from the per-server marketplace payload, not a live feed servers listing an RTX 3090 ......... 219 cards across those servers .......... 352 servers rented at that instant ...... 152 servers free at that instant ........ 67 median on-demand / GPU-hour ......... $0.155 median spot / GPU-hour .............. $0.125 # prices are per GPU: whole-server day rate ÷ GPUs ÷ 24
VRAM
24 GB GDDR6X
Bandwidth
936 GB/s
Board power
350 W
NVLink
Pairs only
$0.155/hr
Median on-demand per GPU-hour, 21 Sep 2026
$0.125/hr
Median spot per GPU-hour, same snapshot
352
Cards across 219 listed servers
69%
Of those servers were rented at that instant

Count the optimizer,
then pick the card.

Almost every question about a 24 GB board comes down to arithmetic on one pool of memory. Weights at your chosen precision, adapters and their optimizer state, activations, and a KV cache that grows with every token of context. Here is where the 3090's 24 GB lands.

4-bit adapter training with room to spare

A 7B or 8B base quantised to 4 bits occupies roughly 4 GB. The adapters and their optimizer state are sized by rank, not by the frozen base, so the rest of the card is free for activations and context. That margin is the whole reason people rent a 24 GB board for this rather than a 10 or 12 GB one: you can raise sequence length without immediately falling off a cliff.

Llama 3 8B weights at 4-bit ≈4 GB

Paired boards for a quantised 70B

Llama 3 70B at 4 bits is about 40 GB of weights, so it never fits one card. Two 3090s do hold it, split by the serving framework. NVLink 3 carries roughly 112.5 GB/s between the pair, which is what makes tensor-parallel decoding tolerable compared with going over PCIe. It is a bridge between two pools, not a merge into one.

NVLink 3, pairs only ≈112.5 GB/s

Image models that want the whole pool

Flux.1 dev at fp16 is roughly 24 GB on its own, which means a 3090 runs it with sequential CPU offload rather than comfortably resident. SDXL at 1024 by 1024 is the workload the card handles without argument, including a batch of four. If you need Flux without offload, this is the wrong size of card.

Flux.1 dev at fp16 ≈24 GB

Ampere silicon
against its Ada successors.

Three boards a 3090 renter is usually choosing between, on the figures that decide a job: how much fits, how fast the weights can be read, and how much dense tensor throughput is behind them.

RTX 3090 RTX 4090 RTX 4070 Ti RTX 3080
Architecture Ampere GA102 Ada AD102 Ada AD104 Ampere GA102
CUDA cores 10,496 16,384 7,680 8,704
VRAM 24 GB GDDR6X 24 GB GDDR6X 12 GB GDDR6X 10 GB GDDR6X
Memory bandwidth 936 GB/s 1,008 GB/s 504 GB/s 760 GB/s
FP16 tensor, dense 142 TFLOPS 165 TFLOPS 160 TFLOPS 119 TFLOPS
NVLink connector Yes pairs No No No
Board power 350 W 450 W 285 W 320 W

dense tensor figures from NVIDIA's Ampere and Ada material · with-sparsity numbers are double these and are not mixed in here

Medians, not floors.
Each host sets its own ask.

Listings are priced per whole server per day in the order currency; the per-GPU hourly figures below are that rate divided by the card count and by 24. Both numbers are the median across the 208 RTX 3090 servers quoting a price in the snapshot. The cheapest single listing was lower and the dearest was higher; the marketplace has the current spread.

Spot, median

$0.125 / GPU-hour
median of 208 quoting servers · 21 Sep 2026
  • You bid; the highest bid holds the server
  • A higher bid or an on-demand order can take it back
  • Marketplace fee 2.5%, half of it paid by you
  • Suits checkpointed training runs
See spot listings
NOT PREEMPTIBLE

On-demand, median

$0.155 / GPU-hour
median of 208 quoting servers · 21 Sep 2026
  • Fixed price the host set; no bidding
  • Holds while your balance covers it
  • Marketplace fee 10%, half of it paid by you
  • Suits a serving endpoint you cannot lose
See on-demand listings

Need two or more 3090s on a fixed term instead? The bare-metal configurator prices RTX 3090 from a minimum of 2 GPUs and a 30-day minimum term at $0.34 to $0.50 per GPU-hour depending on term length, in Japan and Hong Kong, as of 21 Sep 2026. Configure a bare-metal contract →

Pay with
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Two boards or one,
the path is the same.

If you want an NVLink pair, the decision happens at the filtering step, because it is the server that carries two cards on a bridge, not something you assemble afterwards.

01 / FILTER

Decide one card or two

Filter listings to RTX 3090. A single board gives you 24 GB; a two-card server is what you want if the model needs splitting.

02 / CHOOSE THE ORDER TYPE

Spot or on-demand

Spot is a bid that a higher bid can displace. On-demand is the host's fixed price and is not preemptible. Both bill per minute.

03 / BRING AN IMAGE

Your container, your keys

Point the order at any Docker image you can pull, set your ports and your public key, and you get root inside the container with the GPU passed through.

04 / STOP

Stop when the run ends

Billing is per minute, so a run that finishes in forty minutes costs forty minutes. Nothing keeps charging once the order is closed.

Six things people get wrong
about 24 GB Ampere.

What size model can I actually QLoRA on 24 GB once optimizer state and gradients are counted?

Start from the weights. At 4-bit a 7B or 8B base is roughly 4 GB, around 0.5 GB per billion parameters. The LoRA adapters, their optimizer state and their gradients are sized by the adapter rank rather than by the frozen base, so they stay in the hundreds of megabytes. What consumes the rest of the 24 GB is activations and the KV cache, and both grow with sequence length and batch size. Llama 3 8B and Mistral 7B at 4-bit leave a wide margin. A 70B base at 4-bit is about 40 GB before anything else is allocated, so it does not fit on one 3090.

Does NVLink on a 3090 pair give me 48 GB, or two separate 24 GB cards?

Two separate 24 GB cards. NVLink 3 on the 3090 is a bridge between two boards at about 112.5 GB/s and it works in pairs only. It moves tensors between the two cards faster than PCIe does, which is what tensor-parallel and pipeline-parallel serving need, but nothing merges the two pools into one 48 GB address space. A model larger than 24 GB still has to be split across the pair by the framework.

Why is a 3090 cheaper than a 4070 Ti with half the VRAM?

Every price on CLORE.AI is set by the host that owns the machine, so it follows supply rather than a central list. On 21 September 2026 there were 219 servers carrying RTX 3090 cards against 43 carrying RTX 4070 Ti cards, and the medians were $0.155 and $0.208 per GPU-hour on-demand. The 3090 is older silicon that many hosts already own, so more of it reaches the marketplace at a lower ask. The 4070 Ti is newer and far scarcer here.

Is 936 GB/s enough for 8B FP16 serving at usable concurrency?

Token generation is memory-bound: every new token reads the weights once. An 8B model at FP16 is about 16 GB, so 936 GB/s caps weight traffic at roughly 58 forward passes per second before attention, the KV cache and framework overhead are counted. Batching amortises that read across concurrent requests, and the remaining 8 GB of the card is what pays for the batch. We do not publish measured tokens per second for this card, so benchmark your own stack before you size a deployment.

Is a 3090 or a 4090 the cheaper way to finish a memory-bound job?

The 3090, as a rule. Rentals are billed for time, so the comparison is rental price against work done. On 21 September 2026 the 3090 median was $0.155 per GPU-hour and the 4090 median was $0.375, about 2.4 times as much, for 165 dense FP16 TFLOPS against 142 and 1,008 GB/s against 936. A memory-bound job finishes cheaper on the 3090.

How tight is RTX 3090 supply on CLORE.AI?

Tighter than any other consumer card in this snapshot. On 21 September 2026, 219 servers carrying 352 RTX 3090 cards were listed and 152 of those servers were rented at that instant, leaving 67 free. That is one reading rather than an average over time, and both the counts and the prices move. The current figures live on the marketplace itself.

Three jobs, sized in gigabytes.

Weight footprints below follow the usual rule of thumb: about 2 GB per billion parameters at FP16, 1 GB at 8-bit and 0.5 GB at 4-bit, then add ten to thirty per cent for the KV cache and activations. Throughput depends on your framework and settings, so we quote capacity rather than a benchmark we have not run for you.

Adapter training at 4-bit
PEFT with a 4-bit quantised base
≈4 GB base, 20 GB left over

Llama 3 8B or Mistral 7B. The leftover is what buys you sequence length and batch size, which is where a 10 GB card runs out first.

Read the guide →
Llama 3 70B at 4-bit, two boards
ExLlamaV2, weights split across the pair
≈40 GB, so 2 × 24 GB

Rent a server that already carries two bridged cards. The bridge is a 112.5 GB/s link between them, not a way to address 48 GB as one pool.

Read the guide →
Flux.1 dev at fp16
ComfyUI with sequential CPU offload
≈24 GB, the whole card

This is the job that sits right at the ceiling. It runs, with offload, and it is the honest upper bound of what a single 3090 will hold for image generation.

Read the guide →

The large-VRAM shelf, priced.

Three consumer boards with 24 GB or more, as the marketplace held them on 21 Sep 2026. Counts are servers and cards in that snapshot; prices are the median per GPU-hour across the servers quoting one.

GPU
VRAM
Servers listed
Free at snapshot
Median spot $/GPU-hr
Median on-demand $/GPU-hr
RTX 3090 / this page
24 GB GDDR6X
219
67
$0.125
$0.155
RTX 4090
24 GB GDDR6X
243
130
$0.365
$0.375
RTX 5090
32 GB GDDR7
260
137
$0.521
$0.542

the 3090 is the only row here whose supply was mostly taken: 152 of its 219 servers were rented at that instant, against 113 of 243 for the 4090 and 123 of 260 for the 5090

Set-up guides that fit 24 GB.

Each of these walks through a stack that a single 3090 can hold, or in the ExLlamaV2 case one that a pair can. They live in the CLORE.AI documentation, not on this page.

Image Generation
Flux.1 on CLORE.AI
Run Black Forest Labs' Flux for state-of-the-art image gen.
Image Generation
ComfyUI on CLORE.AI
Node-based pipeline for SDXL, Flux, and SD3.
Language Models
text-gen WebUI
The oobabooga WebUI for chat, RAG, and agents.
Training
DreamBooth training
Fine-tune SDXL on your subject with DreamBooth.
Language Models
ExLlamaV2 fast inference
Fastest GPTQ/EXL2 inference for consumer GPUs.
Training
Kohya SS LoRA training
The standard SDXL LoRA training pipeline.
Comparisons
Which LLM server to run
Side-by-side on throughput, quantisation support and memory use.
See all guides →

If 24 GB is not the right size.

RTX 4090
Same 24 GB, 1,008 GB/s, median $0.375/GPU-hr
Compare →
RTX 5090
32 GB when 24 is not enough, median $0.542/GPU-hr
Compare →
RTX 3080
10 GB if the job never needed 24, median $0.167/GPU-hr
Compare →

24 GB, by the minute,
from whoever is cheapest.

Filter the marketplace to RTX 3090, compare what the hosts are asking today, and pay for the minutes the job actually takes.