The RTX 4090 holds 24 GB of GDDR6X, moves it at 1,008 GB/s and does 165 TFLOPS of dense FP16 tensor work. That capacity lands exactly on Flux.1 dev at full precision and stops well short of a 70B model, which needs about 40 GB at 4 bits and therefore two of these cards. On 21 Sep 2026 the marketplace held 243 servers with 378 of them, 130 servers unrented, at a median of $0.375 per GPU-hour on demand. Settlement in BTC, CLORE, USDT or USDC.
Everything this card is good at sits under that line, and everything it cannot do sits above it. The three cards below are the ones worth renting it for, stated in gigabytes rather than in adjectives.
Flux.1 dev in fp16 weighs about 24 GB, so on this card it stays resident and never has to be paged across PCIe. That is an exact fit rather than a roomy one, which is why batch size and resolution are the two settings that decide whether the run is fast or stalled on the bus.
QLoRA on Llama 3 8B, Mistral 7B or Qwen2.5 14B all sit inside 24 GB, which is the band this card was bought for. Note the family sizes: Llama 3 is 8B and 70B, so the 13B and 34B runs that turn up in a lot of tutorials belong to other model families.
Llama 3 8B at fp16 takes about 16 GB, leaving roughly 8 GB for the KV cache, which is what supports a long context or several concurrent callers. Drop it to FP8 and the weight budget halves again, though this generation has no FP4 path.
The decision this card forces is almost always a scaling decision. Hardware comes from NVIDIA's published specifications; the price row is the median per-GPU hourly rate observed on 21 September 2026, and today's is on the marketplace.
two dense 4090s total 330 TFLOPS, which coincidentally equals one card's with-sparsity figure. They are not the same thing, and the 50-series columns are blank because NVIDIA's published numbers for them disagree between sources
Both numbers are medians per GPU-hour taken on 21 September 2026. The spread was wide: the cheapest quote of the day was $0.027, which tells you more about one host's pricing than about the market.
Reserved instead of rented: RTX 4090 bare metal starts at five GPUs on a 14-day term in the UK. Fourteen days is the shortest term Clore quotes and only the 5090 shares it, and of those two the 4090 takes the smaller block. Japan and Hong Kong offer eight-GPU blocks on 30-day terms. Quotes on 21 Sep 2026 spanned $0.40 to $0.94 per GPU-hour. Build a five-card block →
Most 4090 mistakes are scaling mistakes made before the order is placed, so the sequence below starts there rather than at the rent button.
Under 24 GB, take one. Above it, take a multi-GPU server so the cards sit on the same machine rather than across the internet.
Spot can be outbid away from you mid-run; on-demand cannot. The gap between them was one cent at the snapshot.
Pull whatever the host can reach. The GPU is handed into the container and nothing inside it is locked down.
If throughput collapses when the batch grows, you have spilled past 24 GB and the run is waiting on PCIe, not on the GPU.
No. A 70B model quantized to 4 bits needs around 40 GB just to hold its weights, reckoning half a gigabyte for every billion parameters, and this card has 24. The working configuration is two cards with the model split across them, which vLLM and ExLlamaV2 both support. Any claim that a single 4090 serves 70B is either describing a heavily offloaded setup, where most of the model lives in system RAM and the GPU waits on PCIe, or it is simply false. Worth adding: Llama 3 ships as 8B and 70B only, so the 13B and 34B sizes that appear in a lot of copy do not exist in that family at all.
It does, and that is precisely what this card is good for. Full-precision Flux.1 dev sits at roughly 24 GB, which is the card exactly, so the weights stay put instead of being streamed across the bus. It is a tight fit rather than a comfortable one: raise the batch or the resolution far enough and you spill, at which point the GPU spends its time waiting on PCIe rather than working. If you want slack above Flux, the next step up is a 32 GB card.
It depends on the shape of the job, and more than most people assume. This card has no NVLink, so a pair communicates over PCIe, and tensor parallelism exchanges activations at every layer boundary. Throughput-oriented serving with large batches absorbs that well, because the transfers overlap with compute. Latency-sensitive single-stream decoding absorbs it badly, because each layer adds a synchronous hop. Splitting the model by pipeline stage instead moves far less data and is frequently the better choice on a PCIe-only pair. The thing to internalise is that two 4090s are not one card with 48 GB, they are two cards that cooperate at a cost.
Yes, as long as the other side of the comparison is also dense. 165 TFLOPS is the dense FP16 and BF16 tensor number. NVIDIA publishes roughly 330 for the same silicon with 2:4 structured sparsity, which only applies to a model pruned into that pattern. Comparison tables regularly place one card’s sparse figure beside another’s dense figure, which is how a smaller card ends up appearing to outrun this one. Check which of the two each column is quoting before you draw a conclusion from it.
In three situations. When the job fits in 16 GB after quantization, a smaller card does identical work at a lower hourly rate. When the job is bound by memory bandwidth and would move 1.78 times faster on a 1,792 GB/s card, paying 1.45 times more per hour for that card is cheaper overall. When the job needs more than 24 GB, this card is not a candidate in the first place.
There is. RTX 4090 blocks begin at five GPUs on a 14-day term in the UK, with eight-GPU configurations in Japan and Hong Kong on 30-day terms. Fourteen days is the shortest term Clore quotes anywhere and the 5090 is the only other card offered on it, so between those two the 4090 is the smaller block at the same term. Prices on 21 September 2026 ran from $0.40 to $0.94 per GPU-hour, the low end attached to the longest terms. If what you actually want is the smallest possible block rather than the shortest term, three other cards start at two GPUs instead of five. Bare metal is the route when you want hardware reserved outright rather than bid for by the minute.
Each figure below is parameters multiplied by bytes per parameter, which is arithmetic rather than a benchmark. Actual speed depends on your runtime, batch and context length, so measure it on the card.
An exact fit, which is the best and the most fragile case at once. Push the batch or the resolution and the model starts spilling over PCIe.
Read the guide →The leftover eight gigabytes are where context length and concurrency come from. Quantizing to FP8 buys more of both at a quality cost you should measure.
Read the guide →Split across a pair with tensor or pipeline parallelism. Two cards on one server, not two servers, because the split runs over PCIe.
Read the guide →Once a job outgrows 24 GB there are two answers, and they cost different amounts. The rates are the medians observed on 21 September 2026, so the arithmetic below is a snapshot rather than a quote.
Image generation and 4-bit fine-tuning are the two categories that fit this card squarely. The multi-GPU material matters once a job crosses the 24 GB line.
Hosts set their own rate and are paid for every rented minute, and larger 4090 rigs can also supply bare-metal contracts.
130 of the 243 listed RTX 4090 servers were unrented on 21 Sep 2026. That was a snapshot, not a promise, so the marketplace is where you check what is free now.