Log in Rent RTX 4090
RTX 4090 · counted 21 Sep 2026 · 378 cards across 243 servers

The RTX 4090 fits
Flux exactly.
With a gigabyte to spare.

The RTX 4090 holds 24 GB of GDDR6X, moves it at 1,008 GB/s and does 165 TFLOPS of dense FP16 tensor work. That capacity lands exactly on Flux.1 dev at full precision and stops well short of a 70B model, which needs about 40 GB at 4 bits and therefore two of these cards. On 21 Sep 2026 the marketplace held 243 servers with 378 of them, 130 servers unrented, at a median of $0.375 per GPU-hour on demand. Settlement in BTC, CLORE, USDT or USDC.

●Metered by the minute ●Full root in your own container ●Auction or fixed price ●Bare metal from five cards
$0.375/hr
Median on-demand rate per card (21 Sep 2026)
24GB
GDDR6X per card
165TFLOPS
FP16 and BF16 tensor, dense, not sparse
378
Cards counted across 243 listed servers

Twenty-four gigabytes
is a real boundary.

Everything this card is good at sits under that line, and everything it cannot do sits above it. The three cards below are the ones worth renting it for, stated in gigabytes rather than in adjectives.

Flux.1 dev at full precision, with nothing to spare

Flux.1 dev in fp16 weighs about 24 GB, so on this card it stays resident and never has to be paged across PCIe. That is an exact fit rather than a roomy one, which is why batch size and resolution are the two settings that decide whether the run is fast or stalled on the bus.

Flux.1 dev at fp16 ~24 GB, the whole card

Fine-tuning at 4 bits

QLoRA on Llama 3 8B, Mistral 7B or Qwen2.5 14B all sit inside 24 GB, which is the band this card was bought for. Note the family sizes: Llama 3 is 8B and 70B, so the 13B and 34B runs that turn up in a lot of tutorials belong to other model families.

QLoRA, 4-bit base 8B, 7B and 14B fit

Serving 8B with cache to spare

Llama 3 8B at fp16 takes about 16 GB, leaving roughly 8 GB for the KV cache, which is what supports a long context or several concurrent callers. Drop it to FP8 and the weight budget halves again, though this generation has no FP4 path.

Llama 3 8B at fp16 ~16 GB, ~8 GB left

One card, two cards,
or the newer one.

The decision this card forces is almost always a scaling decision. Hardware comes from NVIDIA's published specifications; the price row is the median per-GPU hourly rate observed on 21 September 2026, and today's is on the marketplace.

RTX 4090 RTX 5090 RTX 5080 2× RTX 4090
Architecture Ada AD102 Blackwell GB202 Blackwell GB203 two Ada AD102
VRAM 24 GB GDDR6X 32 GB GDDR7 16 GB GDDR7 48 GB split
Memory bandwidth 1,008 GB/s 1,792 GB/s 960 GB/s 1,008 GB/s each
CUDA cores 16,384 21,760 10,752 32,768
FP16 tensor, dense 165 TFLOPS not quoted not quoted 330 TFLOPS combined
FP4 tensor path no yes yes no
Median on-demand $0.375 / GPU-hr $0.542 / GPU-hr $0.292 / GPU-hr $0.750 / hr total

two dense 4090s total 330 TFLOPS, which coincidentally equals one card's with-sparsity figure. They are not the same thing, and the 50-series columns are blank because NVIDIA's published numbers for them disagree between sources

What 230 quoted servers
were actually asking.

Both numbers are medians per GPU-hour taken on 21 September 2026. The spread was wide: the cheapest quote of the day was $0.027, which tells you more about one host's pricing than about the market.

Spot

$0.365 / GPU-hr
median of 230 quotes · $0.100 was the cheapest free server
  • Bid-based, and reclaimable by a higher bid
  • Accrues minute by minute
  • Marketplace fee of 2.5%, halved between you and the host
  • Right for renders and training you can resume
Check spot listings
FIXED PRICE

On-demand

$0.375 / GPU-hr
median of 230 quotes · $0.583 across the 130 unrented servers
  • A rate the host publishes and holds
  • Immune to preemption while funded
  • Marketplace fee of 10%, halved between you and the host
  • Expect to pay above the median for a card sitting free
Check on-demand listings

Reserved instead of rented: RTX 4090 bare metal starts at five GPUs on a 14-day term in the UK. Fourteen days is the shortest term Clore quotes and only the 5090 shares it, and of those two the 4090 takes the smaller block. Japan and Hong Kong offer eight-GPU blocks on 30-day terms. Quotes on 21 Sep 2026 spanned $0.40 to $0.94 per GPU-hour. Build a five-card block →

Pay with
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Decide the card count first.

Most 4090 mistakes are scaling mistakes made before the order is placed, so the sequence below starts there rather than at the rent button.

01 / SIZE

One card or a pair

Under 24 GB, take one. Above it, take a multi-GPU server so the cards sit on the same machine rather than across the internet.

02 / RENT

Auction or fixed

Spot can be outbid away from you mid-run; on-demand cannot. The gap between them was one cent at the snapshot.

$ clore rent --gpu "RTX 4090"
03 / CONNECT

Your image, your root

Pull whatever the host can reach. The GPU is handed into the container and nothing inside it is locked down.

04 / MEASURE

Watch for the bus

If throughput collapses when the batch grows, you have spilled past 24 GB and the run is waiting on PCIe, not on the GPU.

Where the 4090 copy usually lies.

Can I run Llama 3 70B on one 4090?

No. A 70B model quantized to 4 bits needs around 40 GB just to hold its weights, reckoning half a gigabyte for every billion parameters, and this card has 24. The working configuration is two cards with the model split across them, which vLLM and ExLlamaV2 both support. Any claim that a single 4090 serves 70B is either describing a heavily offloaded setup, where most of the model lives in system RAM and the GPU waits on PCIe, or it is simply false. Worth adding: Llama 3 ships as 8B and 70B only, so the 13B and 34B sizes that appear in a lot of copy do not exist in that family at all.

Flux.1 dev at fp16 is about 24 GB. Does it fit without offload on a 4090?

It does, and that is precisely what this card is good for. Full-precision Flux.1 dev sits at roughly 24 GB, which is the card exactly, so the weights stay put instead of being streamed across the bus. It is a tight fit rather than a comfortable one: raise the batch or the resolution far enough and you spill, at which point the GPU spends its time waiting on PCIe rather than working. If you want slack above Flux, the next step up is a 32 GB card.

Two 4090s without NVLink, how much does PCIe tensor parallelism actually cost me?

It depends on the shape of the job, and more than most people assume. This card has no NVLink, so a pair communicates over PCIe, and tensor parallelism exchanges activations at every layer boundary. Throughput-oriented serving with large batches absorbs that well, because the transfers overlap with compute. Latency-sensitive single-stream decoding absorbs it badly, because each layer adds a synchronous hop. Splitting the model by pipeline stage instead moves far less data and is frequently the better choice on a PCIe-only pair. The thing to internalise is that two 4090s are not one card with 48 GB, they are two cards that cooperate at a cost.

Is the 4090's 165 TFLOPS dense figure the one I should compare against other cards?

Yes, as long as the other side of the comparison is also dense. 165 TFLOPS is the dense FP16 and BF16 tensor number. NVIDIA publishes roughly 330 for the same silicon with 2:4 structured sparsity, which only applies to a model pruned into that pattern. Comparison tables regularly place one card’s sparse figure beside another’s dense figure, which is how a smaller card ends up appearing to outrun this one. Check which of the two each column is quoting before you draw a conclusion from it.

When does a 4090 stop being the cheapest path to a given throughput?

In three situations. When the job fits in 16 GB after quantization, a smaller card does identical work at a lower hourly rate. When the job is bound by memory bandwidth and would move 1.78 times faster on a 1,792 GB/s card, paying 1.45 times more per hour for that card is cheaper overall. When the job needs more than 24 GB, this card is not a candidate in the first place.

Is there a bare-metal option for 4090s, and what is the smallest one?

There is. RTX 4090 blocks begin at five GPUs on a 14-day term in the UK, with eight-GPU configurations in Japan and Hong Kong on 30-day terms. Fourteen days is the shortest term Clore quotes anywhere and the 5090 is the only other card offered on it, so between those two the 4090 is the smaller block at the same term. Prices on 21 September 2026 ran from $0.40 to $0.94 per GPU-hour, the low end attached to the longest terms. If what you actually want is the smallest possible block rather than the shortest term, three other cards start at two GPUs instead of five. Bare metal is the route when you want hardware reserved outright rather than bid for by the minute.

Three jobs, sized in gigabytes.

Each figure below is parameters multiplied by bytes per parameter, which is arithmetic rather than a benchmark. Actual speed depends on your runtime, batch and context length, so measure it on the card.

Flux.1 dev, full precision
ComfyUI
~24 GB against 24 GB available

An exact fit, which is the best and the most fragile case at once. Push the batch or the resolution and the model starts spilling over PCIe.

Read the guide →
Llama 3 8B, fp16 weights
vLLM
~16 GB used, ~8 GB for cache

The leftover eight gigabytes are where context length and concurrency come from. Quantizing to FP8 buys more of both at a quality cost you should measure.

Read the guide →
Llama 3 70B at 4 bits
the job that needs two cards
~40 GB, so 24 is not enough

Split across a pair with tensor or pipeline parallelism. Two cards on one server, not two servers, because the split runs over PCIe.

Read the guide →

Add a card, or change the card?

Once a job outgrows 24 GB there are two answers, and they cost different amounts. The rates are the medians observed on 21 September 2026, so the arithmetic below is a snapshot rather than a quote.

Configuration
Usable VRAM
Mem BW (GB/s)
CUDA cores
Board power
FP16 dense TFLOPS
Interconnect
Cost per hour
One RTX 4090 / this page
24 GB
1,008
16,384
450 W
165
none needed
$0.375
Two RTX 4090s
48 GB, split
1,008 each
32,768
900 W
330 combined
PCIe, no NVLink
$0.750
One RTX 5090
32 GB
1,792
21,760
575 W
not quoted
none needed
$0.542
One RTX 5080
16 GB
960
10,752
360 W
not quoted
none needed
$0.292

The recipes people actually run here.

Image generation and 4-bit fine-tuning are the two categories that fit this card squarely. The multi-GPU material matters once a job crosses the 24 GB line.

Image Generation
Flux.1 on CLORE.AI
Run Black Forest Labs' Flux for state-of-the-art image gen.
Language Models
vLLM serving
High-throughput LLM serving with PagedAttention.
Training
LLM fine-tuning
LoRA / QLoRA fine-tuning workflow.
Video Generation
Hunyuan Video
Tencent's open video generation model.
Image Generation
ComfyUI on CLORE.AI
Node-based pipeline for SDXL, Flux, and SD3.
Training
Kohya SS LoRA training
The standard SDXL LoRA training pipeline.
Comparisons
Picking a serving framework
Trade-offs between the stacks, including how each one handles two cards.
See all guides →

Either side of 24 GB.

RTX 3090
The Ampere card at the same capacity
Rent →
RTX 5090
32 GB GDDR7, 1,792 GB/s
Rent →

Twenty-four gigabytes,
priced by the minute.

130 of the 243 listed RTX 4090 servers were unrented on 21 Sep 2026. That was a snapshot, not a promise, so the marketplace is where you check what is free now.