Log in RTX A4000 listings
RTX A4000 · single slot · 140 W · density, not speed

Four of these
fit where one 4090 does.
And pool 64 GB.

This card is not interesting for its compute. It is interesting because it is single-slot and draws 140 watts, so four of them occupy the slots and the power budget that one 450 W consumer board fills on its own. The memory is 16 GB with ECC, which sets a hard ceiling on what you can load and a useful floor under how badly a long run can go wrong. Six servers carrying eight cards were listed on 21 Sep 2026, every one of them rented, at a median of $0.049 per GPU-hour.

●Single-slot blower ●140 W board power ●16 GB GDDR6 with ECC ●No NVLink connector
SPECIFICATION SHEET
RTX A4000 · Ampere GA104
16 GB · 140 W NVIDIA datasheet
ArchitectureAmpere · GA104
Memory16 GB GDDR6 · ECC
Memory bandwidth448 GB/s
CUDA cores6,144
FP16 / BF16 tensor · dense76.7 TFLOPS
FP16 / BF16 tensor · 2:4 sparse153.4 TFLOPS
FP8 tensor coresnone
MIG partitioningnot supported
NVLinkno connector
Board power · slot width140 W · single slot
Comfortable fitLlama 3 8B at INT8 · ~9 GB
Will not hold8B at FP16 with real KV headroom
MEDIAN ON-DEMAND $0.049/GPU-h
LOWEST ASKING $0.014/GPU-h
SAMPLE 6servers
FREE 21 SEP 2026 0cards
140W
Board power, against 450 W on an RTX 4090
6
Servers listed per-minute on 21 Sep 2026, carrying 8 cards
0
Of those cards free at 18:05 UTC that day
$0.049/GPU-h
Median on-demand asking price across those six servers

The spec that matters here
is measured in slots.

Nothing about 6,144 CUDA cores or 448 GB/s is going to win an argument. What wins is that this board occupies one slot and 140 watts, which is a claim about how many of them exist in a machine rather than how fast any one of them is.

Four boards, one power budget

An RTX 4090 draws 450 W on its own. Four A4000s draw 560 W between them and occupy four slots instead of two or three. That arithmetic is the reason this card exists in rented machines: it turns a standard chassis and a standard supply into four independently rentable GPUs instead of one fast one.

Four boards draw 560 W total

Renders that run while nobody watches

Cycles and V-Ray jobs on scenes that fit inside 16 GB, running unattended for hours. Error-corrected memory earns its keep on exactly this shape of work, where a single wrong bit becomes a frame nobody re-inspects before it ships.

Scene ceiling 16 GB with ECC

An 8B model, quantised, running all day

Llama 3 8B at eight bits is about 9 GB, which leaves genuine room inside 16 GB for the KV cache. At full precision the same model is around 16 GB and leaves effectively nothing, so eight-bit is the working configuration here rather than a compromise.

8B at INT8 ~9 of 16 GB

Count the slots,
then count the watts.

The last row is the whole page in one number: how many boards a 1,000 W GPU budget supports. It is the only column where this card wins, and it is the reason to rent one.

RTX A4000 RTX A5000 RTX 3090 RTX 4090
VRAM 16 GB GDDR6 ECC 24 GB GDDR6 ECC 24 GB GDDR6X 24 GB GDDR6X
Board power 140 W 230 W 350 W 450 W
Memory bandwidth 448 GB/s 768 GB/s 936 GB/s 1,008 GB/s
CUDA cores 6,144 8,192 10,496 16,384
FP16 / BF16 tensor dense 76.7 TFLOPS 111.1 TFLOPS 142 TFLOPS 165.2 TFLOPS
Boards per 1,000 W GPU budget 7 4 2 2

specs from each board's NVIDIA datasheet · the last row is 1,000 W divided by board power, rounded down, and ignores everything else in the machine

Six servers, eight cards,
all of them booked.

This is the only page in the pro tier with a sample worth quoting, and it is still six machines. The prices below are what those six hosts were asking, not a rate CLORE sets or endorses.

6-SERVER SAMPLE

On-demand asking prices

$0.049 / GPU-hour median
low of $0.014 · measured at 18:05 UTC on 21 Sep 2026
  • Per-GPU figures, derived from whole-server daily prices
  • Every card was rented at the instant of the snapshot
  • Six hosts is a thin sample; treat the median as an indication
  • No bare-metal A4000 contract exists, so this is the only route
See today's listings

How the order types differ

2 ways to book
spot and on-demand, both metered by the minute
  • On-demand holds the machine until you release it
  • Spot costs less and an on-demand order can displace it
  • Cancel inside the first ten minutes and no creation fee applies
  • Settlement is computed in CLORE; pay in BTC, USDT, USDC or CLORE
Read the renter docs
Settled in CLORE, paid in
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Rent the chassis,
not just the card.

On a board whose selling point is density, the interesting question is how many of them are in the machine you are booking. These four steps are ordered around that.

01 / SHAPE

Decide if your work splits

Four separate jobs at 16 GB each is where this card wins. One job needing 20 GB is where it loses completely, and no amount of cards fixes that, because there is no NVLink and no memory pooling.

02 / COUNT

Filter on GPU count, not just model

Eight cards across six servers on 21 Sep 2026 means some machines carry more than one. If you want parallel jobs on one box, the card count is the filter that matters.

$ clore rent --gpu "RTX A4000"
03 / QUANTISE

Plan for eight bits, not sixteen

An 8B model at full precision leaves nothing behind in 16 GB. At eight bits it uses around 9 GB and the rest is yours for the KV cache. Decide this before you pick a container image.

04 / WATCH

Check the clocks early

Single-slot boards packed together share their neighbour's exhaust. You cannot see the chassis you rented, so if the job runs for days, confirm the clocks hold during the first hour.

Questions hosts and renters ask.

What does ECC on a rented GPU actually protect me from on a multi-day run?

A bit in memory flipping without anything telling you. On a board without error correction that shows up much later as a weight that is slightly wrong, an epoch with an unexplained loss spike, or a rendered frame nobody catches. The expensive version is not a crash, it is a job that completes and is quietly incorrect after three days of compute you have already paid for. ECC corrects most single-bit errors outright and turns the rest into a visible failure.

Why would I pick four A4000s over one A6000 with the same 64 GB total?

You would not, if the thing you are running needs more than 16 GB in one allocation. Four boards give you 4 by 16 GB, not 64, and no framework changes that. What four boards give you instead is four independent jobs, four times the aggregate bandwidth for embarrassingly parallel work, and a failure that costs a quarter of your capacity instead of all of it. Choose by shape of workload, not by the sum on the sticker.

Does the single-slot blower design cost me clock speed under sustained load?

Being single-slot is not itself the problem; the card is specified at 140 W and the cooler is built to shed that. What costs you clocks is four of them packed together, where each board's intake is whatever the board beside it just exhausted. On a rented machine you cannot see the chassis, so if a long job depends on stable clocks, watch them for the first hour rather than assuming the datasheet applies to someone else's build.

Is 448 GB/s a hard ceiling for 8B-class serving?

It is a real ceiling and worth doing the arithmetic on. Token generation reads the resident weights once per token, so an 8B model at eight bits, around 9 GB, cannot be read more than about fifty times a second at 448 GB/s no matter what else you optimise. That is the physics, and it sits well below what a 936 GB/s consumer board manages. If throughput per card is your metric, this is the wrong card. If it is jobs per chassis, it is not.

The A4000 has no NVLink. How do I scale past 16 GB on one of these?

Across PCIe, with the framework splitting the model. Tensor and pipeline parallelism in vLLM or DeepSpeed need an interconnect, not specifically NVLink, and PCIe qualifies. The cost lands hardest on gradient-heavy training where data crosses the link every step, and lightest on pipeline-parallel inference where one activation tensor moves per stage. If a peer bridge is genuinely required, the A5000 and A6000 have connectors and this board does not.

All six listed A4000 servers were rented. Is that demand or scarcity?

It is a single instant, and it cannot tell the two apart. Six servers carrying eight cards, none free at 18:05 UTC on 21 September 2026, is a snapshot rather than a utilisation rate, and treating it as one would be the most common mistake made with this kind of data. What it does tell you practically is that the A4000 shelf is small enough to empty, so check the marketplace before planning a job around finding one.

Three jobs that fit,
and one that does not.

Memory budgets use the standard precision arithmetic: about 2 GB per billion parameters at FP16, 1 GB at eight bits, 0.5 GB at four, plus 10 to 30 per cent for the KV cache and activations. Speed depends on the runtime, so what is quoted here is what fits.

Llama 3 8B at eight bits
The working configuration
~9 GB weights, ~7 GB left

At full precision the same model is around 16 GB and leaves nothing for the KV cache. Quantising is not a compromise on this board, it is the only way the model has room to serve anything.

Read the guide →
SDXL at 1024 by 1024
Small batches, long queues
Fits, with ECC on

Diffusion at this resolution is comfortable inside 16 GB at modest batch sizes. Unattended runs over thousands of images are the case where corrected memory stops being theoretical.

Read the guide →
Anything past 16 GB
The hard stop
4 × 16 GB, never 64

No NVLink connector, no memory pooling. Four boards in one chassis are four separate 16 GB spaces. A model larger than one of them has to be sharded by the framework over PCIe, or run somewhere else.

Read the guide →

How many boards fit
a 1,000 watt budget.

The last two columns are the point of this page. Everything else is context for why a slower card can still be the right rental: seven of these run on the power two 4090s need, and each one is a separate rentable GPU. Follow a row for that card.

GPU
VRAM
Board power
Bandwidth
NVLink
Per 1,000 W
Summed, not pooled
RTX A4000 / this page
16 GB ECC
140 W
448 GB/s
none
7 boards
112 GB
RTX A5000
24 GB ECC
230 W
768 GB/s
112.5 GB/s
4 boards
96 GB
RTX A6000
48 GB ECC
300 W
768 GB/s
112.5 GB/s
3 boards
144 GB
RTX 3090
24 GB
350 W
936 GB/s
112.5 GB/s
2 boards
48 GB
RTX 4090
24 GB
450 W
1,008 GB/s
none
2 boards
48 GB

Seven guides that respect
a sixteen-gigabyte ceiling.

Container images and commands on the docs site. These were picked because none of them needs more memory than this board has, which rules out a great deal of what the internet suggests running on a GPU.

Other Workloads
Blender + Cycles GPU
Renders that finish while nobody is looking, on corrected memory.
Image Generation
A1111 WebUI on CLORE.AI
Diffusion with a batch size the 16 GB budget will actually allow.
Video Processing
FFmpeg + NVENC
Transcoding leans on fixed-function silicon, not on the 448 GB/s.
Language Models
Ollama on CLORE.AI
The shortest path to an 8B model at eight bits on one card.
Training
Jupyter for ML training
Exploratory work, where four cheap boards beat one expensive one.
Computer Vision
YOLOv8 detection
A small model that runs all day, which is this card's natural shape.
Advanced
CLORE API integration
Useful when the shelf keeps emptying: poll instead of refreshing.
See all guides →

More memory, more speed,
or more watts.

RTX A5000
24 GB ECC · 230 W · 768 GB/s · NVLink
Same ECC, more of it →
RTX 3090
24 GB · 936 GB/s · 350 W · no ECC · 219 listed
Twice the bandwidth →
RTX A6000
48 GB ECC · 300 W · holds a 70B at four bits
When 16 GB is the wall →

Small board.
Small shelf.

Eight A4000 cards existed across six listed servers on 21 Sep 2026, and all eight were rented. A shelf that small empties and refills quickly, so open the marketplace and see what is there rather than planning a week of work around finding one.