This card is not interesting for its compute. It is interesting because it is single-slot and draws 140 watts, so four of them occupy the slots and the power budget that one 450 W consumer board fills on its own. The memory is 16 GB with ECC, which sets a hard ceiling on what you can load and a useful floor under how badly a long run can go wrong. Six servers carrying eight cards were listed on 21 Sep 2026, every one of them rented, at a median of $0.049 per GPU-hour.
Nothing about 6,144 CUDA cores or 448 GB/s is going to win an argument. What wins is that this board occupies one slot and 140 watts, which is a claim about how many of them exist in a machine rather than how fast any one of them is.
An RTX 4090 draws 450 W on its own. Four A4000s draw 560 W between them and occupy four slots instead of two or three. That arithmetic is the reason this card exists in rented machines: it turns a standard chassis and a standard supply into four independently rentable GPUs instead of one fast one.
Cycles and V-Ray jobs on scenes that fit inside 16 GB, running unattended for hours. Error-corrected memory earns its keep on exactly this shape of work, where a single wrong bit becomes a frame nobody re-inspects before it ships.
Llama 3 8B at eight bits is about 9 GB, which leaves genuine room inside 16 GB for the KV cache. At full precision the same model is around 16 GB and leaves effectively nothing, so eight-bit is the working configuration here rather than a compromise.
The last row is the whole page in one number: how many boards a 1,000 W GPU budget supports. It is the only column where this card wins, and it is the reason to rent one.
specs from each board's NVIDIA datasheet · the last row is 1,000 W divided by board power, rounded down, and ignores everything else in the machine
This is the only page in the pro tier with a sample worth quoting, and it is still six machines. The prices below are what those six hosts were asking, not a rate CLORE sets or endorses.
On a board whose selling point is density, the interesting question is how many of them are in the machine you are booking. These four steps are ordered around that.
Four separate jobs at 16 GB each is where this card wins. One job needing 20 GB is where it loses completely, and no amount of cards fixes that, because there is no NVLink and no memory pooling.
Eight cards across six servers on 21 Sep 2026 means some machines carry more than one. If you want parallel jobs on one box, the card count is the filter that matters.
An 8B model at full precision leaves nothing behind in 16 GB. At eight bits it uses around 9 GB and the rest is yours for the KV cache. Decide this before you pick a container image.
Single-slot boards packed together share their neighbour's exhaust. You cannot see the chassis you rented, so if the job runs for days, confirm the clocks hold during the first hour.
A bit in memory flipping without anything telling you. On a board without error correction that shows up much later as a weight that is slightly wrong, an epoch with an unexplained loss spike, or a rendered frame nobody catches. The expensive version is not a crash, it is a job that completes and is quietly incorrect after three days of compute you have already paid for. ECC corrects most single-bit errors outright and turns the rest into a visible failure.
You would not, if the thing you are running needs more than 16 GB in one allocation. Four boards give you 4 by 16 GB, not 64, and no framework changes that. What four boards give you instead is four independent jobs, four times the aggregate bandwidth for embarrassingly parallel work, and a failure that costs a quarter of your capacity instead of all of it. Choose by shape of workload, not by the sum on the sticker.
Being single-slot is not itself the problem; the card is specified at 140 W and the cooler is built to shed that. What costs you clocks is four of them packed together, where each board's intake is whatever the board beside it just exhausted. On a rented machine you cannot see the chassis, so if a long job depends on stable clocks, watch them for the first hour rather than assuming the datasheet applies to someone else's build.
It is a real ceiling and worth doing the arithmetic on. Token generation reads the resident weights once per token, so an 8B model at eight bits, around 9 GB, cannot be read more than about fifty times a second at 448 GB/s no matter what else you optimise. That is the physics, and it sits well below what a 936 GB/s consumer board manages. If throughput per card is your metric, this is the wrong card. If it is jobs per chassis, it is not.
Across PCIe, with the framework splitting the model. Tensor and pipeline parallelism in vLLM or DeepSpeed need an interconnect, not specifically NVLink, and PCIe qualifies. The cost lands hardest on gradient-heavy training where data crosses the link every step, and lightest on pipeline-parallel inference where one activation tensor moves per stage. If a peer bridge is genuinely required, the A5000 and A6000 have connectors and this board does not.
It is a single instant, and it cannot tell the two apart. Six servers carrying eight cards, none free at 18:05 UTC on 21 September 2026, is a snapshot rather than a utilisation rate, and treating it as one would be the most common mistake made with this kind of data. What it does tell you practically is that the A4000 shelf is small enough to empty, so check the marketplace before planning a job around finding one.
Memory budgets use the standard precision arithmetic: about 2 GB per billion parameters at FP16, 1 GB at eight bits, 0.5 GB at four, plus 10 to 30 per cent for the KV cache and activations. Speed depends on the runtime, so what is quoted here is what fits.
At full precision the same model is around 16 GB and leaves nothing for the KV cache. Quantising is not a compromise on this board, it is the only way the model has room to serve anything.
Read the guide →Diffusion at this resolution is comfortable inside 16 GB at modest batch sizes. Unattended runs over thousands of images are the case where corrected memory stops being theoretical.
Read the guide →No NVLink connector, no memory pooling. Four boards in one chassis are four separate 16 GB spaces. A model larger than one of them has to be sharded by the framework over PCIe, or run somewhere else.
Read the guide →The last two columns are the point of this page. Everything else is context for why a slower card can still be the right rental: seven of these run on the power two 4090s need, and each one is a separate rentable GPU. Follow a row for that card.
Container images and commands on the docs site. These were picked because none of them needs more memory than this board has, which rules out a great deal of what the internet suggests running on a GPU.
List them on Clore.ai: up to ~$45/mo per card, with four single-slot boards to a chassis, each one earning for every minute it is rented.
Eight A4000 cards existed across six listed servers on 21 Sep 2026, and all eight were rented. A shelf that small empties and refills quickly, so open the marketplace and see what is there rather than planning a week of work around finding one.