Two boards ship under this name and the difference decides whether your job loads. This page is about the 10 GB card, the one the snapshot counted. Ten gigabytes draws a sharp line: SDXL at 1024 by 1024 fits, with a tiled VAE and a batch of one, while Llama 3 8B at FP16 does not, because that is about 16 GB. Median on-demand price on 21 Sep 2026 was $0.167 per GPU-hour, spot $0.156, billed by the minute.
Ten gigabytes is not a rounding error away from twelve or sixteen. It is the point where several very common workloads stop being a question of speed and become a question of whether the allocation succeeds at all. Three examples, in both directions.
SDXL in fp16 at 1024 by 1024 comes to roughly 10 GB once you enable a tiled VAE, which is to say it fits with the tiling and is tight without it. This is the headline use for the card and the reason people pick it over an 8 GB board. Going to batch 2 is where the arithmetic stops cooperating.
Llama 3 8B at FP16 is about 16 GB of weights before the KV cache exists. There is no setting that makes that fit in ten. Quantise it and the picture changes completely: 8-bit brings it to roughly 9 GB, 4-bit to about 5 GB with room left for a long context. Any 13B at FP16 is out of reach on this card, and so is Flux.1 dev at fp16.
Mistral 7B quantised to 4 bits is around 6 GB with its optimizer state, which leaves a usable but not generous margin for activations and context on a 10 GB board. It is a real fine-tuning card at small scale. It is not the card to pick if you expect to raise sequence length later.
The cards a 3080 renter is realistically choosing between are the ones either side of it on memory, not the datacenter parts. Note that the 3080 has more bandwidth than the 12 GB Ada board while holding less.
dense tensor figures only · the with-sparsity numbers NVIDIA also publishes are exactly double and are not mixed in above
208 of the listed 3080 servers were quoting a price in the snapshot. The cheapest of them asked $0.023 per GPU-hour and the medians below sat well above that, which tells you the spread is worth shopping. Listings are priced per whole server per day; these figures are that rate divided by the card count and by 24.
On a card with two memory configurations under one name, the ordinary rental flow gains one extra step, and it is the step that decides whether the job runs.
Open the server's own detail page and read the memory per card there rather than trusting the model string. If it is not stated, assume 10 GB and size the job for that.
Decide the precision first. An 8B model needs 8-bit or 4-bit weights here. SDXL needs a tiled VAE at 1024. Work this out before the meter starts, not on the box.
Spot is a bid that can be displaced. On-demand is the host's fixed price and holds. With 150 of 217 servers free in the snapshot, either type had supply to choose from.
Billing counts the minutes the order is open. Close it when the render queue drains and nothing keeps accruing.
NVIDIA shipped the RTX 3080 in a 10 GB version and later in a 12 GB version, and they are different products with different memory buses. Everything on this page describes the 10 GB card, which is what our snapshot counted. The model string on its own will not settle it, so check the memory figure on the server's own detail page before you place the order, and if you cannot confirm it there, plan for 10 GB. A job sized for 12 GB that lands on a 10 GB board fails at allocation time, not gracefully.
Batch size, nearly always. It multiplies activation memory linearly, while resolution raises it with pixel count. SDXL at 1024 by 1024 in fp16 sits at roughly 10 GB with a tiled VAE, so batch 1 works and batch 2 is already a negotiation. The language-model equivalent is the KV cache: it grows with context length multiplied by concurrent sequences, which is why a long context at batch 1 often survives on this card where a short context at batch 8 does not.
It does not beat it, and the specification says so in every column: the 3090 has 10,496 CUDA cores against 8,704, 936 GB/s of bandwidth against 760, 142 dense FP16 tensor TFLOPS against 119, and more than twice the memory. There was not even a price argument in this snapshot. On 21 September 2026 the median on-demand price was $0.167 per GPU-hour for an RTX 3080 and $0.155 for an RTX 3090. What the 3080 has is slack: 150 of its 217 listed servers were free at that instant, against 67 of 219 for the 3090, and the cheapest individual 3080 listing quoted $0.023 per GPU-hour. Rent one for availability and for the bottom of the price spread, not for throughput.
For a model that fits, yes. Decoding reads the weights once per token, so 760 GB/s across an 8-bit 8B model of roughly 9 GB puts an upper bound near 84 forward passes per second before attention and framework overhead. Bandwidth is not the constraint on this board. The constraint is that you have to quantise to get under 10 GB at all: the same 8B model at FP16 is about 16 GB and will not load.
NVIDIA's GeForce 30 series specification page, the source of the memory figures on this page, lists size, type and bandwidth for the RTX 3080. It does not list ECC. Work on the assumption that there is no memory error reporting you can query: checkpoint long runs, and verify outputs you intend to keep rather than expecting silent corruption to announce itself. If ECC is a requirement and not a preference, the card you want is a professional or datacenter board, not a GeForce one.
As of 21 September 2026 the configurator prices RTX 3080 bare metal from a minimum of 2 GPUs on a minimum 30-day term at $0.23 to $0.36 per GPU-hour in Japan and Hong Kong, with the lower end reached on longer commitments. That is above the $0.167 marketplace median, so the trade is not price. It is that you hold named hardware for a fixed window instead of competing for whatever happens to be free. Bursty work is cheaper on the per-minute marketplace; work that needs the same machines every day for a month is what the contract is for.
Ten gigabytes is a threshold rather than a slider, so the useful question about each stack below is whether it allocates at all, not how quickly it runs once it has. Two of the three sit within a gigabyte of the ceiling, which is why the precision you pick matters more here than the card does.
The card's defining workload and the reason to take it over an 8 GB board. Tiling the VAE is what keeps the decode step inside the budget.
Read the guide →Precision is the lever: roughly 1 GB per billion parameters at 8-bit and 0.5 GB at 4-bit, against 2 GB at FP16, plus ten to thirty per cent for cache and activations. So 8-bit fills the card and leaves little for context, 4-bit buys a long KV cache, and FP16 is not an option here at all.
Read the guide →A conditioning network is loaded alongside the base, so each one you stack eats into a budget that was already tight. Add them one at a time and watch the allocation.
Read the guide →The three consumer boards that CLORE.AI offers both ways, as the marketplace and the bare-metal configurator held them on 21 Sep 2026. The contract price is what a renter pays for reserved hardware; the marketplace median is what independent hosts were asking for the same silicon by the minute.
contract figures from the live bare-metal configurator on 21 Sep 2026 · the low end of each range needs a longer commitment than the 30-day minimum · RTX 3080 and RTX 3090 contracts run in Japan and Hong Kong, RTX 5080 in the USA, the EU and Japan
Diffusion front ends and quantised language-model servers, which is most of what this card is asked to do. The guides live in the CLORE.AI documentation.
The host page covers both routes: per-minute listing at your own price, and bare-metal contracts from two cards.
150 of the 217 listed 3080 servers were free in the 21 Sep 2026 snapshot, and the quoted prices ranged from $0.023 per GPU-hour upwards. There is room to be choosy.