Eight gigabytes is the floor at which real machine-learning work still happens, and this is the consumer card CLORE.AI has most of. 1,350 RTX 3070 cards across 291 servers were listed on 21 Sep 2026 and 213 of those servers were idle, which makes availability an actual argument here rather than a slogan. What fits: Stable Diffusion 1.5, a 4-bit 8B chat model, Whisper transcription. What does not: anything at FP16 above about 3B parameters. This page says which is which.
Most pages about small cards tell you what is possible. The more useful list is the other one. On 8 GB, Llama 3 8B at FP16 needs about 16 GB, any 13B at FP16 about 26 GB, Flux.1 dev at fp16 about 24 GB, and SDXL at 1024 by 1024 with a batch of two will not go without a tiled VAE. Here is what is left, which is more than it sounds.
The fp16 UNet is about 4 GB at 512 by 512, which leaves half the card for the rest of the pipeline. This is the comfortable ceiling for the board rather than its outer limit, and comfort is the point: you can iterate on prompts and LoRAs without watching an allocation meter. Go up to SDXL and the same work becomes a memory-management exercise.
Llama 3 8B in Q4_K_M through llama.cpp or Ollama comes to roughly 5 GB of weights. That leaves around 2 GB for an 8K context window, which is enough to hold a real conversation and not enough to hold several at once. The same model at FP16 is about 16 GB and is simply not an option on this board.
Whisper large-v3 in fp16 through faster-whisper and CTranslate2 occupies about 3 GB. Transcription is the one job on this list where a small card is not a downgrade, because the model was never large. If you are batch-processing audio rather than generating images, the 3070 is a sensible choice rather than a concession.
Three cards a 3070 renter would realistically move to, including one professional board that shares this card's GA104 silicon and its 448 GB/s of bandwidth but carries twice the memory in half the slot width.
dense tensor throughput only, no sparsity figures mixed in · listing counts are the 21 Sep 2026 snapshot, and the A4000 row shows how thin professional-board supply is here
289 of the 291 listed servers were quoting a price in the snapshot, and the spread was enormous: the cheapest asked $0.008 per GPU-hour, roughly one twentieth of the median. With 213 servers idle, shopping the tail is worth the two minutes it takes. Listings are priced per whole server per day, so the per-GPU hourly figures here are that rate divided by the card count and by 24.
If this is the first GPU you have rented rather than owned, the useful habit is to make the first order deliberately short and deliberately cheap, and to find out where the memory wall is before you care about the answer.
With 213 of 291 servers idle in the snapshot and quotes starting at $0.008 per GPU-hour, the cheapest listing is usually available. Take it.
A container has to be fetched onto the host before anything runs. A lean image gets you to a prompt faster than a kitchen-sink one, and you are paying from the moment the order opens.
Allocate the model first and watch what is left. If it will not fit, you want to know in the first two minutes, not after you have uploaded a dataset.
Billing follows the order, not the GPU utilisation. A container sitting at an idle prompt costs the same as one that is working, so close it when you stop.
Three things fit with room to work in: Stable Diffusion 1.5 at 512 by 512 with a batch of four, where the fp16 UNet is about 4 GB; Llama 3 8B quantised to 4 bits, about 5 GB of weights leaving roughly 2 GB for an 8K KV cache; and Whisper large-v3 in fp16 through faster-whisper, about 3 GB. Nothing silently swaps to system RAM by default. CUDA allocations fail rather than page out, so an oversized model raises an out-of-memory error instead of quietly running slowly. What does move data across PCIe is a framework you have explicitly told to offload, and that is when throughput collapses without an error message to explain it.
SDXL at 1024 by 1024 with a batch of two does not fit in 8 GB without a tiled VAE. With tiling and fp16 you can get single images out of the card, but you are working against the memory budget the whole time and every extra conditioning network makes it worse. SD 1.5 at 512 by 512 with a batch of four is the size this card was comfortable at, and it is the honest recommendation. If SDXL at full resolution is the point of the exercise, rent a 10 GB or larger board instead of fighting an 8 GB one.
Fast enough is not the obstacle; capacity is. A 3070 has 81 dense FP16 tensor TFLOPS against the 3090's 142 and 448 GB/s of bandwidth against 936, so a run takes longer but it does run. What it does not give you is headroom: with a 4-bit 8B base at about 5 GB, the remaining three gigabytes have to cover activations, gradients and the KV cache, so you train at short sequence lengths and small batches. On 21 September 2026 the 3090 median was $0.155 per GPU-hour against $0.167 for the 3070, which means the larger card was not the more expensive one. Learn the workflow on a 3070 if that is what is free; size a real training run on 24 GB.
Because CLORE.AI is a marketplace of hardware other people already own, and a great deal of what people already own is mid-range Ampere bought for gaming or mining. On 21 September 2026 it was the largest fleet of any single consumer model here: 1,350 cards across 291 servers, against 770 RTX 3080 cards and 352 RTX 3090 cards. The consequence shows up in the other direction too. Only 78 of those 291 servers were rented at that instant, so supply sits well ahead of demand and prices reflect it.
We do not publish verified figures for the RTX 3060 on this page, and we are not going to invent them, so take the shape of the answer rather than the numbers. Within the same Ampere generation, the part that carries more memory on a narrower configuration trades throughput for capacity: you gain the ability to load something that did not fit and you lose some of the rate at which it is processed. There is a practical point that matters more here. The RTX 3060 is not one of the models our marketplace snapshot counted, and the 3070 is the largest consumer fleet on it, so the card you can actually rent in volume today is this one.
In this snapshot, comfortably. On 21 September 2026, 213 of the 291 listed RTX 3070 servers were not rented, which is 73 per cent of them sitting idle. That is one reading rather than an average over time and it can change, but the 3070 is the least contended consumer card in the set by a wide margin. If you are learning and you want the machine to be there when you sit down, that idleness is the argument for this card.
Eight gigabytes has to hold the weights, the activations and the whole context window at the same time, which is why the deciding number on this card is gigabytes occupied rather than anything about speed. All three stacks below leave change out of the budget, and the second one shows the arithmetic that gets them there.
The spare half is what lets you add a LoRA or an upscaler without rethinking the pipeline. A zero-configuration front end is the fastest way into this.
Read the guide →Four-bit weights cost about 0.5 GB per billion parameters against 2 GB at FP16, which is the entire reason this fits. Enough for one conversation at an 8K window; two at that length is where the cache runs out, not where the compute does.
Read the guide →The model was never large, so a small card is not a compromise here. This is the workload where a 3070 is the right tool rather than the affordable one.
Read the guide →Specification tables are everywhere. What is harder to find is how many of a card are listed and how many of them are sitting idle, which is what decides whether you get one at the price you want. All figures are the 21 Sep 2026 snapshot.
the 4070 was cheaper at the median and carries more memory, but there were 92 of its cards against 1,350 of these · a cheaper median on a fleet fifteen times smaller is not the same offer
Every stack below fits in 8 GB, some of them only at the resolutions and batch sizes this page has been explicit about. They live in the CLORE.AI documentation rather than on this page.
List it and earn up to ~$115/mo per card, with every rented minute paid in BTC, CLORE, USDT or USDC at a price you set.
213 of the 291 listed 3070 servers were idle in the 21 Sep 2026 snapshot, with quotes from $0.008 per GPU-hour. An hour of finding out where 8 GB runs out costs very little.