Log in Rent RTX 3080
RTX 3080 10 GB · 217 servers listed · snapshot of 21 Sep 2026

Rent an RTX 3080:
the 10 GB one.
The 12 GB one.

Two boards ship under this name and the difference decides whether your job loads. This page is about the 10 GB card, the one the snapshot counted. Ten gigabytes draws a sharp line: SDXL at 1024 by 1024 fits, with a tiled VAE and a batch of one, while Llama 3 8B at FP16 does not, because that is about 16 GB. Median on-demand price on 21 Sep 2026 was $0.167 per GPU-hour, spot $0.156, billed by the minute.

●10 GB GDDR6X, 760 GB/s ●Quantise, or it will not load ●150 of 217 servers were free ●Bare metal from 2 cards
10GB
GDDR6X per card, 760 GB/s
$0.167/hr
Median on-demand per GPU-hour, 21 Sep 2026
770
Cards across 217 listed servers
150
Of those servers were not rented at that instant

What loads, what does not,
and what you have to quantise.

Ten gigabytes is not a rounding error away from twelve or sixteen. It is the point where several very common workloads stop being a question of speed and become a question of whether the allocation succeeds at all. Three examples, in both directions.

Loads: SDXL at full resolution, batch one

SDXL in fp16 at 1024 by 1024 comes to roughly 10 GB once you enable a tiled VAE, which is to say it fits with the tiling and is tight without it. This is the headline use for the card and the reason people pick it over an 8 GB board. Going to batch 2 is where the arithmetic stops cooperating.

SDXL 1024, fp16, tiled VAE ≈10 GB

Does not load: an 8B model at FP16

Llama 3 8B at FP16 is about 16 GB of weights before the KV cache exists. There is no setting that makes that fit in ten. Quantise it and the picture changes completely: 8-bit brings it to roughly 9 GB, 4-bit to about 5 GB with room left for a long context. Any 13B at FP16 is out of reach on this card, and so is Flux.1 dev at fp16.

Llama 3 8B at FP16 ≈16 GB

Loads, with care: 4-bit adapter training

Mistral 7B quantised to 4 bits is around 6 GB with its optimizer state, which leaves a usable but not generous margin for activations and context on a 10 GB board. It is a real fine-tuning card at small scale. It is not the card to pick if you expect to raise sequence length later.

Mistral 7B, 4-bit plus optimizer ≈6 GB

Where 10 GB sits
among its neighbours.

The cards a 3080 renter is realistically choosing between are the ones either side of it on memory, not the datacenter parts. Note that the 3080 has more bandwidth than the 12 GB Ada board while holding less.

RTX 3080 RTX 3070 RTX 4070 RTX 4080
Architecture Ampere GA102 Ampere GA104 Ada AD104 Ada AD103
CUDA cores 8,704 5,888 5,888 9,728
VRAM 10 GB GDDR6X 8 GB GDDR6 12 GB GDDR6X 16 GB GDDR6X
Memory bandwidth 760 GB/s 448 GB/s 504 GB/s 717 GB/s
FP16 tensor, dense 119 TFLOPS 81 TFLOPS 117 TFLOPS 195 TFLOPS
FP8 tensor cores No No Yes Yes
Board power 320 W 220 W 200 W 320 W

dense tensor figures only · the with-sparsity numbers NVIDIA also publishes are exactly double and are not mixed in above

A wide spread,
and a median near the middle.

208 of the listed 3080 servers were quoting a price in the snapshot. The cheapest of them asked $0.023 per GPU-hour and the medians below sat well above that, which tells you the spread is worth shopping. Listings are priced per whole server per day; these figures are that rate divided by the card count and by 24.

Spot, median

$0.156 / GPU-hour
cheapest quoted spot in the snapshot: $0.023
  • An auction: your bid holds the server until a higher one arrives
  • Expect to be displaced; save state accordingly
  • Marketplace fee 2.5% in total, half of it yours
  • Fits an image-generation queue you can restart
See spot listings
NOT PREEMPTIBLE

On-demand, median

$0.167 / GPU-hour
cheapest quoted on-demand in the snapshot: $0.023
  • The host's fixed asking price, no bidding
  • Runs until you stop it or the balance runs out
  • Marketplace fee 10% in total, half of it yours
  • Fits a session you are sitting in front of
See on-demand listings
Pay with
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Check the memory figure first.

On a card with two memory configurations under one name, the ordinary rental flow gains one extra step, and it is the step that decides whether the job runs.

01 / CONFIRM THE VRAM

10 GB, or something else

Open the server's own detail page and read the memory per card there rather than trusting the model string. If it is not stated, assume 10 GB and size the job for that.

02 / SIZE THE JOB

Quantise before you book

Decide the precision first. An 8B model needs 8-bit or 4-bit weights here. SDXL needs a tiled VAE at 1024. Work this out before the meter starts, not on the box.

03 / PICK THE ORDER TYPE

Bid or fixed price

Spot is a bid that can be displaced. On-demand is the host's fixed price and holds. With 150 of 217 servers free in the snapshot, either type had supply to choose from.

04 / RUN AND STOP

Minutes, not hours

Billing counts the minutes the order is open. Close it when the render queue drains and nothing keeps accruing.

Six questions that only matter
because the number is ten.

Is the 3080 I rent the 10 GB or the 12 GB card, and how do I tell before I pay?

NVIDIA shipped the RTX 3080 in a 10 GB version and later in a 12 GB version, and they are different products with different memory buses. Everything on this page describes the 10 GB card, which is what our snapshot counted. The model string on its own will not settle it, so check the memory figure on the server's own detail page before you place the order, and if you cannot confirm it there, plan for 10 GB. A job sized for 12 GB that lands on a 10 GB board fails at allocation time, not gracefully.

What breaks first on 10 GB, batch size or resolution?

Batch size, nearly always. It multiplies activation memory linearly, while resolution raises it with pixel count. SDXL at 1024 by 1024 in fp16 sits at roughly 10 GB with a tiled VAE, so batch 1 works and batch 2 is already a negotiation. The language-model equivalent is the KV cache: it grows with context length multiplied by concurrent sequences, which is why a long context at batch 1 often survives on this card where a short context at batch 8 does not.

Why does a 3080 beat a 3090 on some diffusion jobs and lose badly on LLMs?

It does not beat it, and the specification says so in every column: the 3090 has 10,496 CUDA cores against 8,704, 936 GB/s of bandwidth against 760, 142 dense FP16 tensor TFLOPS against 119, and more than twice the memory. There was not even a price argument in this snapshot. On 21 September 2026 the median on-demand price was $0.167 per GPU-hour for an RTX 3080 and $0.155 for an RTX 3090. What the 3080 has is slack: 150 of its 217 listed servers were free at that instant, against 67 of 219 for the 3090, and the cheapest individual 3080 listing quoted $0.023 per GPU-hour. Rent one for availability and for the bottom of the price spread, not for throughput.

Is 760 GB/s enough bandwidth for interactive token generation?

For a model that fits, yes. Decoding reads the weights once per token, so 760 GB/s across an 8-bit 8B model of roughly 9 GB puts an upper bound near 84 forward passes per second before attention and framework overhead. Bandwidth is not the constraint on this board. The constraint is that you have to quantise to get under 10 GB at all: the same 8B model at FP16 is about 16 GB and will not load.

Do I get ECC or error reporting on a consumer 3080?

NVIDIA's GeForce 30 series specification page, the source of the memory figures on this page, lists size, type and bandwidth for the RTX 3080. It does not list ECC. Work on the assumption that there is no memory error reporting you can query: checkpoint long runs, and verify outputs you intend to keep rather than expecting silent corruption to announce itself. If ECC is a requirement and not a preference, the card you want is a professional or datacenter board, not a GeForce one.

What does the bare-metal RTX 3080 contract cost, and when is it worth it?

As of 21 September 2026 the configurator prices RTX 3080 bare metal from a minimum of 2 GPUs on a minimum 30-day term at $0.23 to $0.36 per GPU-hour in Japan and Hong Kong, with the lower end reached on longer commitments. That is above the $0.167 marketplace median, so the trade is not price. It is that you hold named hardware for a fixed window instead of competing for whatever happens to be free. Bursty work is cheaper on the per-minute marketplace; work that needs the same machines every day for a month is what the contract is for.

Sized for the card, not for a benchmark.

Ten gigabytes is a threshold rather than a slider, so the useful question about each stack below is whether it allocates at all, not how quickly it runs once it has. Two of the three sit within a gigabyte of the ceiling, which is why the precision you pick matters more here than the card does.

SDXL at 1024 by 1024
fp16 weights with a tiled VAE
≈10 GB at batch 1

The card's defining workload and the reason to take it over an 8 GB board. Tiling the VAE is what keeps the decode step inside the budget.

Read the guide →
Llama 3 8B, quantised
llama.cpp server, GGUF weights
≈9 GB at 8-bit, ≈5 GB at 4-bit

Precision is the lever: roughly 1 GB per billion parameters at 8-bit and 0.5 GB at 4-bit, against 2 GB at FP16, plus ten to thirty per cent for cache and activations. So 8-bit fills the card and leaves little for context, 4-bit buys a long KV cache, and FP16 is not an option here at all.

Read the guide →
ControlNet on top of SDXL
depth, pose or edge conditioning
the base model plus the adapter

A conditioning network is loaded alongside the base, so each one you stack eats into a budget that was already tight. Add them one at a time and watch the allocation.

Read the guide →

Per-minute listing against a fixed-term contract.

The three consumer boards that CLORE.AI offers both ways, as the marketplace and the bare-metal configurator held them on 21 Sep 2026. The contract price is what a renter pays for reserved hardware; the marketplace median is what independent hosts were asking for the same silicon by the minute.

GPU
VRAM
Min GPUs on contract
Min term
Contract $/GPU-hr
Marketplace median on-demand
RTX 3080 / this page
10 GB
2
30 days
$0.23 to $0.36
$0.167
RTX 3090
24 GB
2
30 days
$0.34 to $0.50
$0.155
RTX 5080
16 GB
2
30 days
$0.40 to $0.61
$0.292

contract figures from the live bare-metal configurator on 21 Sep 2026 · the low end of each range needs a longer commitment than the 30-day minimum · RTX 3080 and RTX 3090 contracts run in Japan and Hong Kong, RTX 5080 in the USA, the EU and Japan

Set-up guides that stay under 10 GB.

Diffusion front ends and quantised language-model servers, which is most of what this card is asked to do. The guides live in the CLORE.AI documentation.

Image Generation
A1111 WebUI on CLORE.AI
The classic SD WebUI with extensions and LoRA.
Image Generation
ComfyUI on CLORE.AI
Node-based pipeline for SDXL, Flux, and SD3.
Language Models
llama.cpp server
GGUF quantized inference with HTTP/OpenAI-compatible API.
Language Models
Ollama on CLORE.AI
One-command LLM inference for Llama, Mistral, Phi.
Image Processing
ControlNet advanced
Pose, depth, edge guidance for SDXL.
Audio Voice
XTTS voice cloning
Coqui XTTS for TTS and voice cloning.
Comparisons
Choosing an LLM server
Which runtime handles quantised weights best on a small card.
See all guides →

One step down, one step up.

RTX 3070
8 GB, the floor below this one, same $0.167 median
Compare →
RTX 4070
12 GB with FP8 support, median $0.104/GPU-hr
Compare →
RTX 3090
24 GB and cheaper in this snapshot, median $0.155/GPU-hr
Compare →

Check the memory,
then take the cheapest one.

150 of the 217 listed 3080 servers were free in the 21 Sep 2026 snapshot, and the quoted prices ranged from $0.023 per GPU-hour upwards. There is room to be choosy.