The RTX A6000 carries 48 GB of GDDR6 with ECC, and that number is the whole argument for the card. Llama 3 70B quantised to INT4 is about 40 GB of weights, so it loads on one board instead of two. What 48 GB does not buy is speed: the A6000 reads memory at 768 GB/s, exactly the figure the 24 GB A5000 posts. One A6000 server was listed on the per-minute marketplace on 21 Sep 2026 and it was rented at that moment; dedicated capacity is sold as bare metal from 8 cards on a 30-day term.
Every workload below is here for one reason: it does not fit in 24 GB and it does not need HBM. If your model already fits in 24 GB, an A5000 does the same job at the same 768 GB/s.
A 70-billion-parameter model quantised to four bits is about 40 GB of weights. It fits here and it does not fit on a 24 GB card, which is the entire reason this page exists. What is left over, roughly 8 GB, has to cover the KV cache, activations and the CUDA context, so plan for short contexts and modest concurrency rather than a busy public endpoint.
Large-scene Blender and Omniverse work where displacement, instancing and volume caches all have to stay resident. The moment a scene spills into host memory the render stops being bound by the GPU, and 48 GB is the cheapest way to stop that happening. ECC earns its keep on renders measured in hours.
At INT8 a 32B model is around 32 GB of weights, roughly 35 GB resident once the KV cache and activations are counted. That leaves a useful budget instead of the sliver a 70B INT4 load leaves behind. This is the size where 48 GB stops being a bare fit and starts being comfortable, and where batching actually helps.
Four cards a renter genuinely chooses between once the model outgrows 24 GB. Read the bandwidth row before the VRAM row: it is the column that explains why an A100 costs what it costs.
specs from the NVIDIA datasheet for each board · dense tensor figures, not 2:4 sparse
On the per-minute marketplace every host sets their own price, so there is no platform rate to quote and the supply of this particular card is thin. The bare-metal product is the opposite: a fixed published ladder, with a floor on quantity and term.
On a card this thin on the marketplace, the order of operations matters. Work out the VRAM budget before you go shopping, because it decides whether you want one A6000 or something else entirely.
Weights at INT4 cost about 0.5 GB per billion parameters, INT8 about 1 GB, FP16 about 2 GB. Add 10 to 30 per cent for the KV cache and activations, then check the total against 48.
Filter the marketplace on RTX A6000. There was one listed server on 21 Sep 2026 and it was rented, so treat availability as something to verify, not assume.
Pick a container, get SSH and a Jupyter endpoint. Nothing about the A6000 needs special handling beyond a CUDA build that targets Ampere.
Two A6000s are two separate 48 GB spaces joined by NVLink, useful for tensor or pipeline parallel. For one tenant that needs more than 48 GB in a single allocation, look at an 80 GB card instead.
Roughly 8 GB, minus whatever the runtime holds for activations and the CUDA context. That is enough for short prompts at low concurrency and it runs out quickly as you add either one. If the plan is a busy endpoint with long contexts, budget for a second card rather than assuming the headroom stretches. The 40 GB figure comes from the usual INT4 rule of thumb, about 0.5 GB per billion parameters.
On single-stream token generation, which is bound by how fast the card reads the weights once per token. A 40 GB INT4 model at 768 GB/s has a floor on tokens per second that spare VRAM does not move. Capacity decides whether the model loads at all, bandwidth decides how fast it answers. The A6000 wins the first question and ties the A5000 on the second.
When you want NVLink and you do not need FP8. The A6000 has an NVLink connector and the RTX 6000 Ada does not, so a paired-card setup keeps a 112.5 GB/s peer link instead of falling back to PCIe. The Ada card reads memory faster, 960 GB/s against 768, and adds FP8 tensor cores. If your stack is INT4 or INT8 and your scaling story is two boards talking to each other, the Ampere part is the sensible one.
No. NVLink is a peer-to-peer link between two boards, not a memory controller that merges them. Two A6000s are two 48 GB address spaces with a 112.5 GB/s path between them. Frameworks use that path for tensor or pipeline parallelism, so a model larger than 48 GB can be split across the pair, but no single allocation ever sees 96 GB. Anything promising a 96 GB unified pool on this card is wrong.
Scenes where geometry, textures and volumes have to be resident at once: full-resolution displacement, dense instancing, and volumetric caches a 24 GB board cannot hold without falling back to host memory. Once a Cycles or Omniverse scene spills to system RAM, render time stops being a function of the GPU. ECC matters here too, because a render that runs for hours has hours of exposure to a single-bit error.
One server carrying one A6000 card was listed on the per-minute marketplace when the snapshot was taken at 18:05 UTC on 21 September 2026, and it was rented at that moment. That is a thin market and it moves, so the marketplace listing page is the only honest answer for any given day. If you need a guaranteed block of A6000s, the bare-metal product starts at 8 cards on a 30-day term.
Weight footprints below use the standard precision arithmetic: about 0.5 GB per billion parameters at INT4, 1 GB at INT8, 2 GB at FP16, plus 10 to 30 per cent for the KV cache and activations. Throughput depends on your runtime and prompt shape, so we quote capacity, which does not.
The reason the card exists. It is a genuine fit and a tight one, so keep contexts short and concurrency low until you have measured your own KV usage.
Read the guide →Half the headroom problem disappears at this size. Enough KV budget to batch properly, which is where an A6000 starts behaving like a serving card rather than a demo.
Read the guide →No quantisation to hide behind here: either the scene fits in VRAM or the render falls off a cliff. Long unattended renders are also the clearest case for error-corrected memory.
Read the guide →Same headline capacity, four different answers on bandwidth, FP8 and whether a pair of them can talk over NVLink. Bare-metal figures are the published customer-facing rate per GPU-hour at the 30-day minimum. Follow a row to that card's page.
Each of these is on the docs site with the container and the commands. They are ordered the way an A6000 job usually goes: load something large, measure what is left, then decide whether one card was the right call.
Earn up to ~$630/mo per card on Clore.ai, listing per minute or supplying bare-metal contracts. The host page walks through both routes and what renters look for in a 48 GB listing.
Per-minute A6000 supply was one server on 21 Sep 2026 and it changes daily, so the marketplace is the only place with today's answer. For a block you can count on, the bare-metal configurator prices 8 cards or more on a 30-day minimum.