The A5000 holds the same 24 GB as an RTX 3090 and reads it more slowly: 768 GB/s against 936. Everything else about the board is a reason that trade might be worth taking. The memory is error-corrected, so a single-bit flip on day four of a long run gets caught instead of quietly poisoning a checkpoint. The board draws 230 W rather than 350. And it has a real NVLink connector, which pairs two cards at roughly 112.5 GB/s without merging their memory. Per-minute supply is sporadic: one server was listed on 21 Sep 2026 and it was free.
Nothing on this page argues the A5000 is fast. It argues that for a particular class of job, the thing you want from memory is not throughput but the guarantee that what you wrote is what you read back.
A bit flip on a card without error correction does not crash anything. It produces a weight that is slightly off, a checkpoint that loads fine and behaves strangely, a frame with one wrong pixel in a sequence nobody inspects. The longer the run, the more exposure it has. This is the whole argument for ECC and it is the only one that matters.
A 7B to 14B base model held at four bits leaves enough of the 24 GB for optimiser state and activations to fine-tune without offloading. Jobs of that shape run for hours rather than minutes, which is exactly where the 230 W envelope and the corrected memory both start to matter.
Llama 3 8B in FP16 is about 16 GB of weights, which fits with genuine room left rather than the sliver a quantised 70B leaves on a bigger board. Diffusion work sits in the same bracket: Flux.1 dev at half precision and SDXL in small batches both live comfortably inside 24 GB.
Capacity is the column where all four tie, so ignore it and read the rest. The 3090 and 4090 are faster and thirstier; the A4000 gives up capacity to stay at 140 W. This is the shape of the decision.
specs from each board's NVIDIA datasheet · listing counts from the marketplace at 18:05 UTC on 21 Sep 2026
The $0.278 per GPU-hour this page quotes is what that one host asked, a starting point rather than a market average. Below: what the snapshot actually found, and what the two order types mean when you do find a card.
With one A5000 listed, the useful question is not how to rent one but whether you should be waiting for one. Four steps, and the first is the one that decides the other three.
Runs measured in days, checkpoints nobody re-verifies, outputs nobody inspects frame by frame. If none of that describes your job, a faster 24 GB card is the better buy and there are hundreds of them listed.
This is a hard wall, not a soft one. FP16 costs about 2 GB per billion parameters, eight-bit about 1 GB, four-bit about half. Add a third on top for the KV cache and activations before you compare against 24.
One server on 21 Sep 2026. Supply this thin means the answer changes daily in both directions, so check rather than assume.
If ECC was the requirement, the A4000 is the listed alternative at 16 GB. If 24 GB was, the 3090 market is deep. Deciding this before you search saves a day of waiting for a card that may not appear.
Three things, and one of them goes the wrong way. You gain error-corrected memory, an NVLink connector the 3090 pairs can also use, and a 230 W board against 350 W. You lose bandwidth: the A5000 reads at 768 GB/s and the 3090 at 936, so on token generation the consumer card is the faster one. If your job is a week-long batch where a silent bit flip would waste the week, that trade is obviously worth making. If it is an afternoon of inference, it probably is not.
No, and this is the single most repeated falsehood about these boards. NVLink is a peer-to-peer link, roughly 112.5 GB/s between two cards, not a memory controller that fuses them. Two bridged A5000s are two 24 GB address spaces that can move data between each other quickly. What that enables is tensor and pipeline parallelism, so a framework can split a model across the pair. What it never enables is a single allocation larger than 24 GB.
Single-bit errors in memory, which on a card without ECC do not announce themselves. They surface later as a weight that is subtly wrong, a loss curve with an unexplained step, an image with a wrong pixel, or a checkpoint that loads and produces nonsense. The failure mode that costs you is not the crash, it is the run that finishes and is quietly wrong. Error correction turns most of those into corrected reads, and the uncorrectable ones into an explicit error you can see.
The power limit is the easy part; 230 W is a modest envelope and the board is designed to hold it. What decides sustained behaviour is the air around it. In a dense chassis the relevant question is whether each card gets its own intake or is breathing the card below it, and that is a property of the machine you rented rather than of the silicon. Compared with a 350 W consumer card in the same slot, the A5000 gives the chassis a considerably easier job.
Whenever the thing you are loading is larger than 24 GB, which is a hard wall rather than a slow decline. A 70B model at four-bit precision is about 40 GB and will not run here at any batch size. A 32B at eight-bit is roughly 35 GB and also will not. Note that moving up does not buy you speed: the A6000 reads memory at the same 768 GB/s. You are buying capacity, and only capacity, so make sure capacity is the thing you are short of.
It depends which property you were actually after. If it was error-corrected 24 GB, the nearest listed alternative is the 16 GB A4000, which trades capacity for the same ECC. If it was 24 GB at a price, the RTX 3090 market on CLORE is deep, 219 servers on 21 September 2026 with 67 of them free, at the cost of losing ECC and gaining 120 W. If it was the NVLink bridge, that narrows you to the A6000 and the A40. Check the marketplace before deciding, because one listing today is not one listing tomorrow.
Weight budgets use the standard rules: about 2 GB per billion parameters at FP16, 1 GB at eight bits, 0.5 GB at four, plus 10 to 30 per cent for the KV cache and activations. Speed belongs to your runtime; capacity belongs to the board, so capacity is what we quote.
Running the original weights rather than a quantised copy removes one variable from any result you are going to defend. On a board with error correction, that is two variables gone.
Read the guide →The base model compresses, the optimiser state does not. What is left of 24 GB after a four-bit 14B is the budget that decides your batch size, and it is generous at this model scale.
Read the guide →Image work that runs overnight and gets reviewed in the morning is the definition of a job where a single wrong bit goes unnoticed. Modest batch sizes, corrected memory, 230 W.
Read the guide →If error correction is your requirement, this is the whole shortlist and the choice reduces to how much memory you need and how much bandwidth comes with it. Listing counts are the full marketplace at 18:05 UTC on 21 Sep 2026.
Container images and commands on the docs site. Nothing here needs more memory than this board has, which is the filter used to pick them.
List it on Clore.ai and earn up to ~$195/mo per card. You set the price, renters pay by the minute, and payouts arrive in BTC, CLORE, USDT or USDC.
One A5000 was listed on 21 Sep 2026 and it was free, which is both good news and a thin market. If error correction is what brought you here, the A4000 is the listed fallback; if it was 24 GB, the 3090 shelf is deep.