Ada's fourth-generation tensor cores put FP8 in hardware, and the RTX 4070 is the cheapest card here that has them. Twelve gigabytes of GDDR6X at 504 GB/s inside a 200 W board power, the lowest of any consumer card on this marketplace. On 21 September 2026 its median rate was $0.104 per GPU-hour across 47 listed servers. Only one FP8 card sat lower that day, and it was a fleet of one.
Ada gives this board FP8 tensor cores inside a 200 W envelope. Twelve gigabytes at 504 GB/s sets the ceiling on what will load, and the three workloads below sit under it.
Ada handles FP8 as a native datatype in two formats, E4M3 and E5M2, so a quantised model keeps an exponent rather than being flattened into an integer scale. Every Ampere card on this marketplace stops at INT8. That is the one capability the 4070 has and the 3070, 3080 and 3090 do not.
SDXL at fp16 is roughly 10 GB across the UNet, VAE and two text encoders, which fits with a little room to spare. A second image in the same pass works on a bare pipeline and stops working once a ControlNet or a refiner joins the graph.
Sixteen gigabytes of FP16 weights do not go into twelve. At one byte per parameter the same 8B model is about 8 GB and leaves roughly 4 GB for the KV cache, which is a single-user context rather than a concurrent one. Mistral 7B at FP16 is about 14 GB and also does not fit.
The Ampere cards beside it carry more memory and more bandwidth, and burn 20 to 150 W more doing it. FP8 exists only in the Ada column. Board power is NVIDIA's figure; the price row is the marketplace median on 21 September 2026.
on 21 Sep 2026 the smaller, older 3070 had a higher median than the 4070
Across the 47 RTX 4070 servers listed on 21 September 2026 the median on-demand rate was $0.104 per GPU-hour. Among the 21 that were free at that instant it was $0.202, because the cheapest machines were already taken.
At the 21 September median, ten hours of this card came to about a dollar. Per-minute billing means you pay for the minutes a job uses, not the hours it occupies.
Settle on FP8, INT8 or four-bit before you look at listings. On a 12 GB card that decision is what makes a model viable at all.
The GPU filter matches as a substring, so this term also returns 4070 SUPER, 4070 Ti and 4070 Ti SUPER machines, and a 4070 Laptop GPU. Read the model name on a listing before you order; no filter isolates the plain card, because all of them carry 12 GB.
The order gives you an SSH endpoint and whatever ports you asked for, running the image you chose.
Cancel the order and billing ends with it. Nothing carries over and there is no minimum term.
Dynamic range. Ada's fourth-generation tensor cores implement FP8 as a hardware datatype in two formats, E4M3 and E5M2, so a quantised weight keeps an exponent instead of being flattened into a fixed integer scale. At the same one byte per parameter you get fewer per-channel calibration surprises than INT8 on Ampere. The 4070 is the cheapest Ada card with a real fleet on this marketplace: 47 servers at a $0.104 median on 21 September 2026. The 3070, 3080 and 3090 are Ampere and stop at INT8.
Not at FP16. Eight billion parameters at two bytes each come to about 16 GB and the card has 12, so the full-precision load fails before it starts. At FP8 or INT8 the same weights are roughly 8 GB and leave about 4 GB for the KV cache and CUDA context, which serves one user at a sensible context length. Four-bit takes the weights to roughly 4 GB with comfortable headroom, at the quality cost four-bit carries. Anyone telling you 8B runs at FP16 on 12 GB has not tried it.
Per hour it is: on 21 September 2026 the median 4070 was $0.104 per GPU-hour against $0.155 for a 3090. Per token it depends on the job. The 3090 decodes faster, with 936 GB/s of bandwidth against 504, and carries twice the memory, so whether the cheaper card wins on cost per token depends on how much of that speed advantage your workload actually captures. For a model that fits comfortably in 12 GB, time a short run on each card before you commit a long one.
Batch 1 comfortably, batch 2 depending on what else is loaded. SDXL at fp16 is roughly 10 GB across the UNet, VAE and two text encoders, so one 1024 by 1024 image leaves a little headroom and a second in the same pass usually does not once a refiner or a ControlNet joins the graph. On a bare base pipeline batch 2 is worth trying; on a long node graph, plan for batch 1.
It shows up first in generation. Producing a token reads every weight once, so 504 GB/s caps how fast a single stream can decode however many cores sit idle. Training and diffusion push data through in large blocks and are limited by arithmetic instead, which is why the 4070 Ti, with the same 504 GB/s and 30% more cores, pulls ahead on those and not on decoding. Serving one chat session, bandwidth is your number. Fine-tuning or sampling images, it is not.
Because the cheapest machines are the ones that get taken. On 21 September 2026 the median on-demand price across all 47 listed RTX 4070 servers was $0.104 per GPU-hour, while the median among the 21 that were free at that instant was $0.202. Twenty-six servers were busy and they were disproportionately the cheap ones. If you need a machine immediately, budget closer to the free-server figure than to the headline one.
Each entry gives the memory cost, which is what decides feasibility on this card. Speed depends on your settings and is deliberately not quoted.
SDXL's own weights leave a couple of gigabytes spare, which one ControlNet adapter fits into and two generally do not.
Read the guide →FP8 or INT8 weights leave roughly 4 GB for the KV cache, which is a single-user context rather than a concurrent one.
Read the guide →Subject training on SD 1.5 is the fine-tuning job 12 GB carries outright. SDXL DreamBooth wants a 16 GB board or aggressive memory settings.
Read the guide →Memory, board power, FP8 support and availability for four cards side by side. Counts and medians are from the marketplace on 21 September 2026 and cover plain cards only. The 47 figure counts plain 4070s; the filtered marketplace view shows more, because the same substring also matches the SUPER and Ti variants.
All of these sit on docs.clore.ai. The memory notes higher up tell you which of them need quantising first.
You set the rate and collect for each rented minute in BTC, USDT, USDC or CLORE.
That is the 21 September median of $0.104 per GPU-hour, across 47 listed servers with 21 of them free. The marketplace carries today's figures.