Log in Rent RTX 3070
RTX 3070 · 1,350 cards listed · snapshot of 21 Sep 2026

Rent an RTX 3070:
8 GB to learn on.
24 GB to grow into.

Eight gigabytes is the floor at which real machine-learning work still happens, and this is the consumer card CLORE.AI has most of. 1,350 RTX 3070 cards across 291 servers were listed on 21 Sep 2026 and 213 of those servers were idle, which makes availability an actual argument here rather than a slogan. What fits: Stable Diffusion 1.5, a 4-bit 8B chat model, Whisper transcription. What does not: anything at FP16 above about 3B parameters. This page says which is which.

●8 GB GDDR6, 448 GB/s ●73% of listed servers idle ●220 W, the gentlest board here ●Billed by the minute
FLEET COUNT 21 Sep 2026 18:05 UTC
# counted once from the per-server payload; this is not a live feed RTX 3070 cards on the marketplace ... 1,350 spread across servers ............... 291 of those servers, idle .............. 213 of those servers, busy .............. 78 median on-demand, per GPU-hour ...... $0.167 cheapest quoted, per GPU-hour ....... $0.008 # the largest consumer fleet listed on this marketplace
VRAM
8 GB GDDR6
Bandwidth
448 GB/s
Board power
220 W
NVLink
None
1,350
RTX 3070 cards listed on 21 Sep 2026
213
Of the 291 servers were idle at that instant
8GB
GDDR6 per card, 448 GB/s
$0.167/hr
Median on-demand per GPU-hour, same snapshot

Three workloads that fit.
Four that will not.

Most pages about small cards tell you what is possible. The more useful list is the other one. On 8 GB, Llama 3 8B at FP16 needs about 16 GB, any 13B at FP16 about 26 GB, Flux.1 dev at fp16 about 24 GB, and SDXL at 1024 by 1024 with a batch of two will not go without a tiled VAE. Here is what is left, which is more than it sounds.

Stable Diffusion 1.5 at a batch of four

The fp16 UNet is about 4 GB at 512 by 512, which leaves half the card for the rest of the pipeline. This is the comfortable ceiling for the board rather than its outer limit, and comfort is the point: you can iterate on prompts and LoRAs without watching an allocation meter. Go up to SDXL and the same work becomes a memory-management exercise.

SD 1.5 UNet, fp16, 512 px ≈4 GB

An 8B chat model, quantised to four bits

Llama 3 8B in Q4_K_M through llama.cpp or Ollama comes to roughly 5 GB of weights. That leaves around 2 GB for an 8K context window, which is enough to hold a real conversation and not enough to hold several at once. The same model at FP16 is about 16 GB and is simply not an option on this board.

Llama 3 8B at 4-bit ≈5 GB

Speech, where 8 GB is not a compromise

Whisper large-v3 in fp16 through faster-whisper and CTranslate2 occupies about 3 GB. Transcription is the one job on this list where a small card is not a downgrade, because the model was never large. If you are batch-processing audio rather than generating images, the 3070 is a sensible choice rather than a concession.

Whisper large-v3, fp16 ≈3 GB

What the next step up
actually buys you.

Three cards a 3070 renter would realistically move to, including one professional board that shares this card's GA104 silicon and its 448 GB/s of bandwidth but carries twice the memory in half the slot width.

RTX 3070 RTX 3080 RTX 4070 RTX A4000
Architecture Ampere GA104 Ampere GA102 Ada AD104 Ampere GA104
CUDA cores 5,888 8,704 5,888 6,144
VRAM 8 GB GDDR6 10 GB GDDR6X 12 GB GDDR6X 16 GB GDDR6 ECC
Memory bandwidth 448 GB/s 760 GB/s 504 GB/s 448 GB/s
FP16 tensor, dense 81 TFLOPS 119 TFLOPS 117 TFLOPS 77 TFLOPS
Board power 220 W 320 W 200 W 140 W
Servers listed, 21 Sep 2026 291 217 47 6

dense tensor throughput only, no sparsity figures mixed in · listing counts are the 21 Sep 2026 snapshot, and the A4000 row shows how thin professional-board supply is here

A median, and a very long tail.

289 of the 291 listed servers were quoting a price in the snapshot, and the spread was enormous: the cheapest asked $0.008 per GPU-hour, roughly one twentieth of the median. With 213 servers idle, shopping the tail is worth the two minutes it takes. Listings are priced per whole server per day, so the per-GPU hourly figures here are that rate divided by the card count and by 24.

Spot, median

$0.156 / GPU-hour
cheapest quoted in the snapshot: $0.008
  • A bid; a higher bid takes the machine
  • With 73% of servers idle, being outbid is less likely here
  • Marketplace fee 2.5% in total, half of it yours
  • Fits an afternoon of learning the toolchain
See spot listings
NOT PREEMPTIBLE

On-demand, median

$0.167 / GPU-hour
cheapest quoted in the snapshot: $0.008
  • The host's fixed asking price; nobody outbids you
  • Runs until you stop it or the balance runs out
  • Marketplace fee 10% in total, half of it yours
  • Fits an interactive notebook session
See on-demand listings
Pay with
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Start small on purpose.

If this is the first GPU you have rented rather than owned, the useful habit is to make the first order deliberately short and deliberately cheap, and to find out where the memory wall is before you care about the answer.

01 / SHOP THE TAIL

Sort by price, not by name

With 213 of 291 servers idle in the snapshot and quotes starting at $0.008 per GPU-hour, the cheapest listing is usually available. Take it.

02 / PICK A SMALL IMAGE

The pull is the slow part

A container has to be fetched onto the host before anything runs. A lean image gets you to a prompt faster than a kitchen-sink one, and you are paying from the moment the order opens.

03 / FIND THE WALL

Load the model before the data

Allocate the model first and watch what is left. If it will not fit, you want to know in the first two minutes, not after you have uploaded a dataset.

04 / CLOSE THE ORDER

Idle time still bills

Billing follows the order, not the GPU utilisation. A container sitting at an idle prompt costs the same as one that is working, so close it when you stop.

What people ask
before their first rental.

What actually fits in 8 GB, and what silently swaps to system RAM?

Three things fit with room to work in: Stable Diffusion 1.5 at 512 by 512 with a batch of four, where the fp16 UNet is about 4 GB; Llama 3 8B quantised to 4 bits, about 5 GB of weights leaving roughly 2 GB for an 8K KV cache; and Whisper large-v3 in fp16 through faster-whisper, about 3 GB. Nothing silently swaps to system RAM by default. CUDA allocations fail rather than page out, so an oversized model raises an out-of-memory error instead of quietly running slowly. What does move data across PCIe is a framework you have explicitly told to offload, and that is when throughput collapses without an error message to explain it.

Can I run SDXL on a 3070, or do I need to drop to SD 1.5?

SDXL at 1024 by 1024 with a batch of two does not fit in 8 GB without a tiled VAE. With tiling and fp16 you can get single images out of the card, but you are working against the memory budget the whole time and every extra conditioning network makes it worse. SD 1.5 at 512 by 512 with a batch of four is the size this card was comfortable at, and it is the honest recommendation. If SDXL at full resolution is the point of the exercise, rent a 10 GB or larger board instead of fighting an 8 GB one.

Is a 3070 fast enough for a first LoRA, or am I better off on a 3090?

Fast enough is not the obstacle; capacity is. A 3070 has 81 dense FP16 tensor TFLOPS against the 3090's 142 and 448 GB/s of bandwidth against 936, so a run takes longer but it does run. What it does not give you is headroom: with a 4-bit 8B base at about 5 GB, the remaining three gigabytes have to cover activations, gradients and the KV cache, so you train at short sequence lengths and small batches. On 21 September 2026 the 3090 median was $0.155 per GPU-hour against $0.167 for the 3070, which means the larger card was not the more expensive one. Learn the workflow on a 3070 if that is what is free; size a real training run on 24 GB.

Why are there so many 3070s on CLORE.AI compared to other cards?

Because CLORE.AI is a marketplace of hardware other people already own, and a great deal of what people already own is mid-range Ampere bought for gaming or mining. On 21 September 2026 it was the largest fleet of any single consumer model here: 1,350 cards across 291 servers, against 770 RTX 3080 cards and 352 RTX 3090 cards. The consequence shows up in the other direction too. Only 78 of those 291 servers were rented at that instant, so supply sits well ahead of demand and prices reflect it.

What do I lose going from a 3070 to a 3060 12GB, speed or capacity?

We do not publish verified figures for the RTX 3060 on this page, and we are not going to invent them, so take the shape of the answer rather than the numbers. Within the same Ampere generation, the part that carries more memory on a narrower configuration trades throughput for capacity: you gain the ability to load something that did not fit and you lose some of the rate at which it is processed. There is a practical point that matters more here. The RTX 3060 is not one of the models our marketplace snapshot counted, and the 3070 is the largest consumer fleet on it, so the card you can actually rent in volume today is this one.

Will there be a 3070 free when I want one?

In this snapshot, comfortably. On 21 September 2026, 213 of the 291 listed RTX 3070 servers were not rented, which is 73 per cent of them sitting idle. That is one reading rather than an average over time and it can change, but the 3070 is the least contended consumer card in the set by a wide margin. If you are learning and you want the machine to be there when you sit down, that idleness is the argument for this card.

Three budgets that leave change.

Eight gigabytes has to hold the weights, the activations and the whole context window at the same time, which is why the deciding number on this card is gigabytes occupied rather than anything about speed. All three stacks below leave change out of the budget, and the second one shows the arithmetic that gets them there.

Diffusion at 512 pixels
SD 1.5 in fp16, a batch of four
≈4 GB, half the card spare

The spare half is what lets you add a LoRA or an upscaler without rethinking the pipeline. A zero-configuration front end is the fastest way into this.

Read the guide →
One 8B chat model at four bits
Ollama, Q4_K_M weights
≈5 GB, ≈2 GB for context

Four-bit weights cost about 0.5 GB per billion parameters against 2 GB at FP16, which is the entire reason this fits. Enough for one conversation at an 8K window; two at that length is where the cache runs out, not where the compute does.

Read the guide →
Transcription, the easy win
Whisper large-v3 through CTranslate2
≈3 GB, comfortably inside

The model was never large, so a small card is not a compromise here. This is the workload where a 3070 is the right tool rather than the affordable one.

Read the guide →

How much of each card there actually is.

Specification tables are everywhere. What is harder to find is how many of a card are listed and how many of them are sitting idle, which is what decides whether you get one at the price you want. All figures are the 21 Sep 2026 snapshot.

GPU
VRAM
Cards listed
Servers idle
Median on-demand $/GPU-hr
Cheapest quoted $/GPU-hr
RTX 3070 / this page
8 GB
1,350
213 of 291
$0.167
$0.008
RTX 3080
10 GB
770
150 of 217
$0.167
$0.023
RTX 4070
12 GB
92
21 of 47
$0.104
$0.075

the 4070 was cheaper at the median and carries more memory, but there were 92 of its cards against 1,350 of these · a cheaper median on a fleet fifteen times smaller is not the same offer

Starting points that fit in 8 GB.

Every stack below fits in 8 GB, some of them only at the resolutions and batch sizes this page has been explicit about. They live in the CLORE.AI documentation rather than on this page.

Image Generation
SDXL Turbo on CLORE.AI
Real-time image generation, 1-step inference.
Image Generation
Fooocus on CLORE.AI
Zero-config Stable Diffusion UI for fast prototyping.
Language Models
Ollama on CLORE.AI
One-command LLM inference for Llama, Mistral, Phi.
Audio Voice
Whisper transcription
OpenAI Whisper-large for speech-to-text.
Image Processing
Real-ESRGAN upscaling
4× image upscaling with Real-ESRGAN.
Computer Vision
YOLOv8 detection
Real-time object detection with YOLOv8.
Comparisons
Choosing an image-generation UI
A1111, ComfyUI and Fooocus side by side, including what each costs in memory.
See all guides →

When 8 GB stops being enough.

RTX 3080
10 GB, the step that makes SDXL at 1024 work
Compare →
RTX 4070
12 GB and 200 W, median $0.104/GPU-hr but only 47 servers
Compare →
RTX 3090
24 GB for real fine-tuning, median $0.155/GPU-hr
Compare →

The cheapest way
to learn on real hardware.

213 of the 291 listed 3070 servers were idle in the 21 Sep 2026 snapshot, with quotes from $0.008 per GPU-hour. An hour of finding out where 8 GB runs out costs very little.