Log in Rent RTX 4070
RTX 4070 · marketplace snapshot 21 Sep 2026 · 47 servers, 21 free

RTX 4070: FP8 tensor
cores at two hundred watts.

Ada's fourth-generation tensor cores put FP8 in hardware, and the RTX 4070 is the cheapest card here that has them. Twelve gigabytes of GDDR6X at 504 GB/s inside a 200 W board power, the lowest of any consumer card on this marketplace. On 21 September 2026 its median rate was $0.104 per GPU-hour across 47 listed servers. Only one FP8 card sat lower that day, and it was a fleet of one.

●200 W board power ●FP8 in hardware ●12 GB GDDR6X, 504 GB/s ●Billed per minute
MARKETPLACE SNAPSHOT 21 Sep 2026
# one read of the RTX 4070 listings, 21 September 2026 servers listed 47 cards on those servers 92 servers not rented 21 on-demand median $0.104 per GPU-hour spot median $0.097 per GPU-hour unrented servers, median $0.202 the cheap boxes were busy # board power 200 W, the lowest consumer figure on this marketplace
GPU
RTX 4070 ×1
VRAM
12 GB
Median on-demand
$0.104/hr
As of
21 Sep 2026
47
RTX 4070 servers listed on 21 Sep 2026
200W
Board power, lowest consumer card here
$0.104/hr
Median on-demand price per GPU-hour
FP8
Native on Ada tensor cores, absent on Ampere

A 200 watt card
with modern datatypes.

Ada gives this board FP8 tensor cores inside a 200 W envelope. Twelve gigabytes at 504 GB/s sets the ceiling on what will load, and the three workloads below sit under it.

FP8 on the cheapest Ada board here

Ada handles FP8 as a native datatype in two formats, E4M3 and E5M2, so a quantised model keeps an exponent rather than being flattened into an integer scale. Every Ampere card on this marketplace stops at INT8. That is the one capability the 4070 has and the 3070, 3080 and 3090 do not.

Board power 200 W

SDXL at 1024 by 1024, batch 1

SDXL at fp16 is roughly 10 GB across the UNet, VAE and two text encoders, which fits with a little room to spare. A second image in the same pass works on a bare pipeline and stops working once a ControlNet or a refiner joins the graph.

SDXL fp16 weights ~10 GB

8B quantised, never at FP16

Sixteen gigabytes of FP16 weights do not go into twelve. At one byte per parameter the same 8B model is about 8 GB and leaves roughly 4 GB for the KV cache, which is a single-user context rather than a concurrent one. Mistral 7B at FP16 is about 14 GB and also does not fit.

Llama 3 8B at FP8 ~8 GB

Watts on one side,
gigabytes on the other.

The Ampere cards beside it carry more memory and more bandwidth, and burn 20 to 150 W more doing it. FP8 exists only in the Ada column. Board power is NVIDIA's figure; the price row is the marketplace median on 21 September 2026.

RTX 4070 RTX 3070 RTX 3090 RTX 4090
Architecture Ada Lovelace AD104 Ampere GA104 Ampere GA102 Ada Lovelace AD102
CUDA cores 5,888 5,888 10,496 16,384
VRAM 12 GB GDDR6X 8 GB GDDR6 24 GB GDDR6X 24 GB GDDR6X
Memory bandwidth 504 GB/s 448 GB/s 936 GB/s 1,008 GB/s
Board power 200 W 220 W 350 W 450 W
Median on-demand $/GPU-hr $0.104 $0.167 $0.155 $0.375

on 21 Sep 2026 the smaller, older 3070 had a higher median than the 4070

Ten cents an hour,
and a catch worth knowing.

Across the 47 RTX 4070 servers listed on 21 September 2026 the median on-demand rate was $0.104 per GPU-hour. Among the 21 that were free at that instant it was $0.202, because the cheapest machines were already taken.

Spot

$0.097 / GPU-hr
median of 47 listed servers · lowest quoted $0.065
  • You bid, and a higher bid takes the machine
  • An on-demand order ends a spot one
  • Marketplace fee 2.5%, split with the host
  • Among free servers the spot median was $0.188
See spot listings
47 SERVERS LISTED

On-demand

$0.104 / GPU-hr
median of 47 listed servers · lowest quoted $0.075
  • Fixed by the host and not subject to outbidding
  • 21 of 47 servers were free at the snapshot
  • Marketplace fee 10%, split with the host
  • Median across those free servers was $0.202
See on-demand listings
Pay with
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Cheap enough to leave running.

At the 21 September median, ten hours of this card came to about a dollar. Per-minute billing means you pay for the minutes a job uses, not the hours it occupies.

01 / QUANTISE

Choose the datatype first

Settle on FP8, INT8 or four-bit before you look at listings. On a 12 GB card that decision is what makes a model viable at all.

02 / FILTER

Search the model string

The GPU filter matches as a substring, so this term also returns 4070 SUPER, 4070 Ti and 4070 Ti SUPER machines, and a 4070 Laptop GPU. Read the model name on a listing before you order; no filter isolates the plain card, because all of them carry 12 GB.

RTX 4070
03 / CONNECT

Open the box

The order gives you an SSH endpoint and whatever ports you asked for, running the image you chose.

04 / STOP

Pay for the minutes

Cancel the order and billing ends with it. Nothing carries over and there is no minimum term.

FP8, twelve gigabytes and the price.

What does FP8 on Ada buy me on a 12 GB card that INT8 on Ampere does not?

Dynamic range. Ada's fourth-generation tensor cores implement FP8 as a hardware datatype in two formats, E4M3 and E5M2, so a quantised weight keeps an exponent instead of being flattened into a fixed integer scale. At the same one byte per parameter you get fewer per-channel calibration surprises than INT8 on Ampere. The 4070 is the cheapest Ada card with a real fleet on this marketplace: 47 servers at a $0.104 median on 21 September 2026. The 3070, 3080 and 3090 are Ampere and stop at INT8.

Does Llama 3 8B fit on a 4070?

Not at FP16. Eight billion parameters at two bytes each come to about 16 GB and the card has 12, so the full-precision load fails before it starts. At FP8 or INT8 the same weights are roughly 8 GB and leave about 4 GB for the KV cache and CUDA context, which serves one user at a sensible context length. Four-bit takes the weights to roughly 4 GB with comfortable headroom, at the quality cost four-bit carries. Anyone telling you 8B runs at FP16 on 12 GB has not tried it.

Is a 4070 cheaper per token than a 3090?

Per hour it is: on 21 September 2026 the median 4070 was $0.104 per GPU-hour against $0.155 for a 3090. Per token it depends on the job. The 3090 decodes faster, with 936 GB/s of bandwidth against 504, and carries twice the memory, so whether the cheaper card wins on cost per token depends on how much of that speed advantage your workload actually captures. For a model that fits comfortably in 12 GB, time a short run on each card before you commit a long one.

SDXL at 1024 by 1024: batch 1 or batch 2 on 12 GB?

Batch 1 comfortably, batch 2 depending on what else is loaded. SDXL at fp16 is roughly 10 GB across the UNet, VAE and two text encoders, so one 1024 by 1024 image leaves a little headroom and a second in the same pass usually does not once a refiner or a ControlNet joins the graph. On a bare base pipeline batch 2 is worth trying; on a long node graph, plan for batch 1.

Is the 4070's 504 GB/s a problem for token generation, or only for training?

It shows up first in generation. Producing a token reads every weight once, so 504 GB/s caps how fast a single stream can decode however many cores sit idle. Training and diffusion push data through in large blocks and are limited by arithmetic instead, which is why the 4070 Ti, with the same 504 GB/s and 30% more cores, pulls ahead on those and not on decoding. Serving one chat session, bandwidth is your number. Fine-tuning or sampling images, it is not.

Why was the median price on unrented 4070 servers double the overall median?

Because the cheapest machines are the ones that get taken. On 21 September 2026 the median on-demand price across all 47 listed RTX 4070 servers was $0.104 per GPU-hour, while the median among the 21 that were free at that instant was $0.202. Twenty-six servers were busy and they were disproportionately the cheap ones. If you need a machine immediately, budget closer to the free-server figure than to the headline one.

Three jobs inside twelve gigabytes.

Each entry gives the memory cost, which is what decides feasibility on this card. Speed depends on your settings and is deliberately not quoted.

SDXL with one ControlNet
ComfyUI, fp16
~10 GB base, adapter on top

SDXL's own weights leave a couple of gigabytes spare, which one ControlNet adapter fits into and two generally do not.

Read the guide →
Llama 3 8B quantised
llama.cpp server, GGUF
~8 GB at one byte each

FP8 or INT8 weights leave roughly 4 GB for the KV cache, which is a single-user context rather than a concurrent one.

Read the guide →
DreamBooth on SD 1.5
Diffusers, 8-bit Adam
fits without offload

Subject training on SD 1.5 is the fine-tuning job 12 GB carries outright. SDXL DreamBooth wants a 16 GB board or aggressive memory settings.

Read the guide →

The 4070 beside the cards renters weigh it against.

Memory, board power, FP8 support and availability for four cards side by side. Counts and medians are from the marketplace on 21 September 2026 and cover plain cards only. The 47 figure counts plain 4070s; the filtered marketplace view shows more, because the same substring also matches the SUPER and Ti variants.

GPU
VRAM
Board power
FP8 tensor cores
Servers listed
Not rented
Median on-demand
RTX 4070 / this page
12 GB GDDR6X
200 W
yes
47
21
$0.104
RTX 3070
8 GB GDDR6
220 W
no
291
213
$0.167
RTX 3090
24 GB GDDR6X
350 W
no
219
67
$0.155
RTX 4090
24 GB GDDR6X
450 W
yes
243
130
$0.375

Guides for a 12 GB, 200 W card.

All of these sit on docs.clore.ai. The memory notes higher up tell you which of them need quantising first.

Image Generation
ComfyUI on CLORE.AI
Node-based pipeline for SDXL, Flux, and SD3.
Image Generation
A1111 WebUI on CLORE.AI
The classic SD WebUI with extensions and LoRA.
Language Models
llama.cpp server
GGUF quantized inference with HTTP/OpenAI-compatible API.
Training
DreamBooth training
Fine-tune SDXL on your subject with DreamBooth.
Image Processing
ControlNet advanced
Pose, depth, edge guidance for SDXL.
Audio Voice
Whisper transcription
OpenAI Whisper-large for speech-to-text.
Comparisons
Image Gen: ComfyUI vs SD WebUI vs Fooocus
Which of the three interfaces suits a rented 12 GB card.
See all guides →

One step down, two steps up.

RTX 3070
8 GB Ampere, no FP8 · median $0.167
Rent →
RTX 4070 Ti
same 12 GB, 7,680 cores · median $0.208
Rent →
RTX 4080
16 GB, 8B at FP16 · median $0.292
Rent →

Ten hours for
about a dollar.

That is the 21 September median of $0.104 per GPU-hour, across 47 listed servers with 21 of them free. The marketplace carries today's figures.