The 3090 is the cheapest 24 GB consumer board on this marketplace and the only consumer Ampere board on this marketplace with an NVLink connector. That combination decides what it is good for: 4-bit adapter training where the base model, the KV cache and the activations all have to share one pool, and paired serving where a 40 GB quantised 70B is split across two boards. Median on-demand price on 21 Sep 2026 was $0.155 per GPU-hour, spot $0.125. Billing is per minute; settlement is in BTC, USDT, USDC or CLORE.
Almost every question about a 24 GB board comes down to arithmetic on one pool of memory. Weights at your chosen precision, adapters and their optimizer state, activations, and a KV cache that grows with every token of context. Here is where the 3090's 24 GB lands.
A 7B or 8B base quantised to 4 bits occupies roughly 4 GB. The adapters and their optimizer state are sized by rank, not by the frozen base, so the rest of the card is free for activations and context. That margin is the whole reason people rent a 24 GB board for this rather than a 10 or 12 GB one: you can raise sequence length without immediately falling off a cliff.
Llama 3 70B at 4 bits is about 40 GB of weights, so it never fits one card. Two 3090s do hold it, split by the serving framework. NVLink 3 carries roughly 112.5 GB/s between the pair, which is what makes tensor-parallel decoding tolerable compared with going over PCIe. It is a bridge between two pools, not a merge into one.
Flux.1 dev at fp16 is roughly 24 GB on its own, which means a 3090 runs it with sequential CPU offload rather than comfortably resident. SDXL at 1024 by 1024 is the workload the card handles without argument, including a batch of four. If you need Flux without offload, this is the wrong size of card.
Three boards a 3090 renter is usually choosing between, on the figures that decide a job: how much fits, how fast the weights can be read, and how much dense tensor throughput is behind them.
dense tensor figures from NVIDIA's Ampere and Ada material · with-sparsity numbers are double these and are not mixed in here
Listings are priced per whole server per day in the order currency; the per-GPU hourly figures below are that rate divided by the card count and by 24. Both numbers are the median across the 208 RTX 3090 servers quoting a price in the snapshot. The cheapest single listing was lower and the dearest was higher; the marketplace has the current spread.
Need two or more 3090s on a fixed term instead? The bare-metal configurator prices RTX 3090 from a minimum of 2 GPUs and a 30-day minimum term at $0.34 to $0.50 per GPU-hour depending on term length, in Japan and Hong Kong, as of 21 Sep 2026. Configure a bare-metal contract →
If you want an NVLink pair, the decision happens at the filtering step, because it is the server that carries two cards on a bridge, not something you assemble afterwards.
Filter listings to RTX 3090. A single board gives you 24 GB; a two-card server is what you want if the model needs splitting.
Spot is a bid that a higher bid can displace. On-demand is the host's fixed price and is not preemptible. Both bill per minute.
Point the order at any Docker image you can pull, set your ports and your public key, and you get root inside the container with the GPU passed through.
Billing is per minute, so a run that finishes in forty minutes costs forty minutes. Nothing keeps charging once the order is closed.
Start from the weights. At 4-bit a 7B or 8B base is roughly 4 GB, around 0.5 GB per billion parameters. The LoRA adapters, their optimizer state and their gradients are sized by the adapter rank rather than by the frozen base, so they stay in the hundreds of megabytes. What consumes the rest of the 24 GB is activations and the KV cache, and both grow with sequence length and batch size. Llama 3 8B and Mistral 7B at 4-bit leave a wide margin. A 70B base at 4-bit is about 40 GB before anything else is allocated, so it does not fit on one 3090.
Two separate 24 GB cards. NVLink 3 on the 3090 is a bridge between two boards at about 112.5 GB/s and it works in pairs only. It moves tensors between the two cards faster than PCIe does, which is what tensor-parallel and pipeline-parallel serving need, but nothing merges the two pools into one 48 GB address space. A model larger than 24 GB still has to be split across the pair by the framework.
Every price on CLORE.AI is set by the host that owns the machine, so it follows supply rather than a central list. On 21 September 2026 there were 219 servers carrying RTX 3090 cards against 43 carrying RTX 4070 Ti cards, and the medians were $0.155 and $0.208 per GPU-hour on-demand. The 3090 is older silicon that many hosts already own, so more of it reaches the marketplace at a lower ask. The 4070 Ti is newer and far scarcer here.
Token generation is memory-bound: every new token reads the weights once. An 8B model at FP16 is about 16 GB, so 936 GB/s caps weight traffic at roughly 58 forward passes per second before attention, the KV cache and framework overhead are counted. Batching amortises that read across concurrent requests, and the remaining 8 GB of the card is what pays for the batch. We do not publish measured tokens per second for this card, so benchmark your own stack before you size a deployment.
The 3090, as a rule. Rentals are billed for time, so the comparison is rental price against work done. On 21 September 2026 the 3090 median was $0.155 per GPU-hour and the 4090 median was $0.375, about 2.4 times as much, for 165 dense FP16 TFLOPS against 142 and 1,008 GB/s against 936. A memory-bound job finishes cheaper on the 3090.
Tighter than any other consumer card in this snapshot. On 21 September 2026, 219 servers carrying 352 RTX 3090 cards were listed and 152 of those servers were rented at that instant, leaving 67 free. That is one reading rather than an average over time, and both the counts and the prices move. The current figures live on the marketplace itself.
Weight footprints below follow the usual rule of thumb: about 2 GB per billion parameters at FP16, 1 GB at 8-bit and 0.5 GB at 4-bit, then add ten to thirty per cent for the KV cache and activations. Throughput depends on your framework and settings, so we quote capacity rather than a benchmark we have not run for you.
Llama 3 8B or Mistral 7B. The leftover is what buys you sequence length and batch size, which is where a 10 GB card runs out first.
Read the guide →Rent a server that already carries two bridged cards. The bridge is a 112.5 GB/s link between them, not a way to address 48 GB as one pool.
Read the guide →This is the job that sits right at the ceiling. It runs, with offload, and it is the honest upper bound of what a single 3090 will hold for image generation.
Read the guide →Three consumer boards with 24 GB or more, as the marketplace held them on 21 Sep 2026. Counts are servers and cards in that snapshot; prices are the median per GPU-hour across the servers quoting one.
the 3090 is the only row here whose supply was mostly taken: 152 of its 219 servers were rented at that instant, against 113 of 243 for the 4090 and 123 of 260 for the 5090
Each of these walks through a stack that a single 3090 can hold, or in the ExLlamaV2 case one that a pair can. They live in the CLORE.AI documentation, not on this page.
Renters keep 3090s busy with 4-bit fine-tuning and paired 70B serving. Set your price, get paid for every rented minute, or supply paired servers on bare-metal contracts.
Filter the marketplace to RTX 3090, compare what the hosts are asking today, and pay for the minutes the job actually takes.