Log in Rent RTX A6000
RTX A6000 · 48 GB GDDR6 ECC · capacity before bandwidth

48 GB on one card.
70B at INT4,
with tensor parallel.

The RTX A6000 carries 48 GB of GDDR6 with ECC, and that number is the whole argument for the card. Llama 3 70B quantised to INT4 is about 40 GB of weights, so it loads on one board instead of two. What 48 GB does not buy is speed: the A6000 reads memory at 768 GB/s, exactly the figure the 24 GB A5000 posts. One A6000 server was listed on the per-minute marketplace on 21 Sep 2026 and it was rented at that moment; dedicated capacity is sold as bare metal from 8 cards on a 30-day term.

●48 GB GDDR6 with ECC ●768 GB/s, same as the A5000 ●NVLink pairs, not a 96 GB pool ●No MIG on this card
SPECIFICATION SHEET
RTX A6000 · Ampere GA102
48 GB · 300 W NVIDIA datasheet
ArchitectureAmpere · GA102
Memory48 GB GDDR6 · ECC
Memory bandwidth768 GB/s
CUDA cores10,752
FP16 / BF16 tensor · dense154.8 TFLOPS
FP16 / BF16 tensor · 2:4 sparse309.7 TFLOPS
FP8 tensor coresnone
MIG partitioningnot supported
NVLink112.5 GB/s · pairs only
Board power300 W
Holds at INT4Llama 3 70B · ~40 GB
Will not holdLlama 3 70B at FP16
BARE METAL · 30 D $0.95/GPU-h
90 D $0.83/GPU-h
180 D $0.75/GPU-h
360 D AND UP $0.67/GPU-h
48GB
GDDR6 with ECC on a single board
768GB/s
Memory bandwidth, identical to the 24 GB A5000
1
A6000 server listed per-minute as of 21 Sep 2026, rented
8
Minimum GPU count on a bare-metal A6000 contract

Forty gigabytes of weights,
eight gigabytes of headroom.

Every workload below is here for one reason: it does not fit in 24 GB and it does not need HBM. If your model already fits in 24 GB, an A5000 does the same job at the same 768 GB/s.

Llama 3 70B at INT4, on one board

A 70-billion-parameter model quantised to four bits is about 40 GB of weights. It fits here and it does not fit on a 24 GB card, which is the entire reason this page exists. What is left over, roughly 8 GB, has to cover the KV cache, activations and the CUDA context, so plan for short contexts and modest concurrency rather than a busy public endpoint.

Weights at INT4 ~40 of 48 GB

Scenes a 24 GB board has to page

Large-scene Blender and Omniverse work where displacement, instancing and volume caches all have to stay resident. The moment a scene spills into host memory the render stops being bound by the GPU, and 48 GB is the cheapest way to stop that happening. ECC earns its keep on renders measured in hours.

Resident scene budget 48 GB with ECC

Qwen2.5 32B with real headroom

At INT8 a 32B model is around 32 GB of weights, roughly 35 GB resident once the KV cache and activations are counted. That leaves a useful budget instead of the sliver a 70B INT4 load leaves behind. This is the size where 48 GB stops being a bare fit and starts being comfortable, and where batching actually helps.

Qwen2.5 32B at INT8 ~35 GB resident

Where 48 GB sits
between 24 and 80.

Four cards a renter genuinely chooses between once the model outgrows 24 GB. Read the bandwidth row before the VRAM row: it is the column that explains why an A100 costs what it costs.

RTX A6000 RTX A5000 NVIDIA A40 A100 80GB
VRAM 48 GB GDDR6 ECC 24 GB GDDR6 ECC 48 GB GDDR6 ECC 80 GB HBM2e
Memory bandwidth 768 GB/s 768 GB/s 696 GB/s 1,935 GB/s
CUDA cores 10,752 8,192 10,752 6,912
FP16 / BF16 tensor dense 154.8 TFLOPS 111.1 TFLOPS 149.7 TFLOPS 312 TFLOPS
MIG partitioning no no no up to 7
Board power 300 W 230 W 300 W 400 W

specs from the NVIDIA datasheet for each board · dense tensor figures, not 2:4 sparse

Two routes to an A6000.
Only one has a published rate.

On the per-minute marketplace every host sets their own price, so there is no platform rate to quote and the supply of this particular card is thin. The bare-metal product is the opposite: a fixed published ladder, with a floor on quantity and term.

Per-minute marketplace

1 server listed
1 card · 0 free at 18:05 UTC on 21 Sep 2026
  • The host sets the price, not CLORE
  • Billed by the minute, stop whenever
  • A spot order can be displaced by an on-demand one
  • Supply moves daily, so check before you plan around it
Open the marketplace
PUBLISHED RATE

Bare metal

$0.67 / GPU-hour
at a 360-day term · $0.95 at the 30-day minimum
  • Minimum 8 A6000 cards per contract
  • Minimum 30-day term, price falls as the term grows
  • Offered in the USA and France
  • Rates above are the figures the configurator returns
Configure a bare-metal block
Renters pay in
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Size the model first,
then go looking for the card.

On a card this thin on the marketplace, the order of operations matters. Work out the VRAM budget before you go shopping, because it decides whether you want one A6000 or something else entirely.

01 / BUDGET

Do the VRAM arithmetic

Weights at INT4 cost about 0.5 GB per billion parameters, INT8 about 1 GB, FP16 about 2 GB. Add 10 to 30 per cent for the KV cache and activations, then check the total against 48.

02 / CHECK SUPPLY

Look before you plan

Filter the marketplace on RTX A6000. There was one listed server on 21 Sep 2026 and it was rented, so treat availability as something to verify, not assume.

$ clore rent --gpu "RTX A6000"
03 / DEPLOY

Bring your own image

Pick a container, get SSH and a Jupyter endpoint. Nothing about the A6000 needs special handling beyond a CUDA build that targets Ampere.

04 / SCALE

Decide what two cards mean

Two A6000s are two separate 48 GB spaces joined by NVLink, useful for tensor or pipeline parallel. For one tenant that needs more than 48 GB in a single allocation, look at an 80 GB card instead.

Six questions about 48 GB.

Llama 3 70B at INT4 is about 40 GB. What does that leave for the KV cache on 48 GB?

Roughly 8 GB, minus whatever the runtime holds for activations and the CUDA context. That is enough for short prompts at low concurrency and it runs out quickly as you add either one. If the plan is a busy endpoint with long contexts, budget for a second card rather than assuming the headroom stretches. The 40 GB figure comes from the usual INT4 rule of thumb, about 0.5 GB per billion parameters.

The A6000 and A5000 both run at 768 GB/s. When does that make 48 GB feel slow?

On single-stream token generation, which is bound by how fast the card reads the weights once per token. A 40 GB INT4 model at 768 GB/s has a floor on tokens per second that spare VRAM does not move. Capacity decides whether the model loads at all, bandwidth decides how fast it answers. The A6000 wins the first question and ties the A5000 on the second.

A6000 versus RTX 6000 Ada: same 48 GB, so when is the older card the right buy?

When you want NVLink and you do not need FP8. The A6000 has an NVLink connector and the RTX 6000 Ada does not, so a paired-card setup keeps a 112.5 GB/s peer link instead of falling back to PCIe. The Ada card reads memory faster, 960 GB/s against 768, and adds FP8 tensor cores. If your stack is INT4 or INT8 and your scaling story is two boards talking to each other, the Ampere part is the sensible one.

Does an NVLink pair give me 96 GB?

No. NVLink is a peer-to-peer link between two boards, not a memory controller that merges them. Two A6000s are two 48 GB address spaces with a 112.5 GB/s path between them. Frameworks use that path for tensor or pipeline parallelism, so a model larger than 48 GB can be split across the pair, but no single allocation ever sees 96 GB. Anything promising a 96 GB unified pool on this card is wrong.

What large-scene rendering work justifies 48 GB that 24 GB cannot do at all?

Scenes where geometry, textures and volumes have to be resident at once: full-resolution displacement, dense instancing, and volumetric caches a 24 GB board cannot hold without falling back to host memory. Once a Cycles or Omniverse scene spills to system RAM, render time stops being a function of the GPU. ECC matters here too, because a render that runs for hours has hours of exposure to a single-bit error.

How many A6000s can I actually rent by the minute today?

One server carrying one A6000 card was listed on the per-minute marketplace when the snapshot was taken at 18:05 UTC on 21 September 2026, and it was rented at that moment. That is a thin market and it moves, so the marketplace listing page is the only honest answer for any given day. If you need a guaranteed block of A6000s, the bare-metal product starts at 8 cards on a 30-day term.

Three memory budgets
against 48 GB.

Weight footprints below use the standard precision arithmetic: about 0.5 GB per billion parameters at INT4, 1 GB at INT8, 2 GB at FP16, plus 10 to 30 per cent for the KV cache and activations. Throughput depends on your runtime and prompt shape, so we quote capacity, which does not.

Llama 3 70B at INT4
vLLM, one card
~40 GB weights, ~8 GB left

The reason the card exists. It is a genuine fit and a tight one, so keep contexts short and concurrency low until you have measured your own KV usage.

Read the guide →
Qwen2.5 32B at INT8
Quantised serving, one card
~32 GB weights, ~16 GB left

Half the headroom problem disappears at this size. Enough KV budget to batch properly, which is where an A6000 starts behaving like a serving card rather than a demo.

Read the guide →
Cycles scenes over 24 GB
Blender on CUDA or OptiX
48 GB resident, ECC on

No quantisation to hide behind here: either the scene fits in VRAM or the render falls off a cliff. Long unattended renders are also the clearest case for error-corrected memory.

Read the guide →

Four boards hold 48 GB.
They are not interchangeable.

Same headline capacity, four different answers on bandwidth, FP8 and whether a pair of them can talk over NVLink. Bare-metal figures are the published customer-facing rate per GPU-hour at the 30-day minimum. Follow a row to that card's page.

GPU
Architecture
Bandwidth
FP8 cores
NVLink
Board power
Bare metal, 30 d
RTX A6000 / this page
Ampere GA102
768 GB/s
no
112.5 GB/s
300 W
$0.95
RTX 6000 Ada
Ada AD102
960 GB/s
yes
none
300 W
$1.66
NVIDIA A40
Ampere GA102
696 GB/s
no
112.5 GB/s
300 W
$0.75
NVIDIA L40S
Ada AD102
864 GB/s
yes
none
350 W
$1.40

Guides worth reading
before you book 48 GB.

Each of these is on the docs site with the container and the commands. They are ordered the way an A6000 job usually goes: load something large, measure what is left, then decide whether one card was the right call.

Other Workloads
Blender + Cycles GPU
The case where 48 GB either holds the scene or the render collapses.
Training
LLM fine-tuning
Four-bit LoRA is how a 70B base model becomes trainable at this capacity.
Language Models
vLLM serving
Paged attention is what makes the leftover 8 GB stretch further than it should.
Video Generation
Hunyuan Video
A video stack whose text encoder, transformer and VAE all want to stay resident.
Training
DeepSpeed multi-GPU training
Sharding is the honest answer when one 48 GB card is not enough.
Language Models
Llama 3.3 on CLORE.AI
Another 70B at INT4, which lands in the same ~40 GB bracket.
Advanced
Multi-GPU setup
What an NVLink bridge between two A6000s does, and what it does not.
See all guides →

One step down, one across,
one up.

RTX A5000
24 GB ECC · same 768 GB/s · 230 W
Half the VRAM, none of the speed loss →
RTX 6000 Ada
48 GB ECC · 960 GB/s · FP8 · no NVLink
The newer 48 GB board →
A100 80GB
80 GB HBM2e · 1,935 GB/s · MIG up to 7
When 48 GB is the constraint →

Check the listing page
before you commit.

Per-minute A6000 supply was one server on 21 Sep 2026 and it changes daily, so the marketplace is the only place with today's answer. For a block you can count on, the bare-metal configurator prices 8 cards or more on a 30-day minimum.