Log in RTX 6000 Ada listings
RTX 6000 Ada · 960 GB/s · FP8 · and no NVLink connector

960 GB/s and FP8
across 48 GB of ECC.
With NVLink.

The generational break from the A6000 runs in both directions, and pages that treat the two as interchangeable get it wrong. What the Ada part gains: 960 GB/s against 768, fourth-generation tensor cores with native FP8, and about 70 per cent more CUDA cores in the same 300 W. What it loses: the NVLink connector, so two of these talk over PCIe and nothing else. One server carrying two cards was listed on the per-minute marketplace on 21 Sep 2026 and it was rented; dedicated blocks are sold as bare metal from 8 cards on a 30-day term.

●48 GB GDDR6 with ECC ●FP8 tensor cores, natively ●No NVLink connector ●No MIG on this card
960GB/s
The fastest 48 GB GDDR board CLORE lists
18,176
CUDA cores, inside the same 300 W as an A6000
0
NVLink connectors on this board, unlike the A6000
1
Server listed per-minute as of 21 Sep 2026, 2 cards, rented

Same 48 GB as the A6000.
Different silicon underneath.

Capacity is the one number where these two boards tie, so it is the least interesting thing about either. Everything below is a place where the Ada generation does something the Ampere one cannot, or where it stops short.

FP8 with hardware behind it

Eight-bit floating point costs about 1 GB per billion parameters and, on this board, runs on tensor cores designed for it. The A6000 has no FP8 units, so the closest thing there is INT8 with its own calibration story. If your serving stack already speaks FP8, this is the cheapest board on the platform that speaks it back with 48 GB of room.

32B at FP8 ~32 of 48 GB

Bandwidth you can feel on generation

Token generation re-reads the whole resident model for every token, so memory speed sets the floor. Against the A6000's 768 GB/s this board reads 1.25 times faster, which is the honest ceiling on what that column buys. It is a real difference on decode and no difference at all on anything compute-bound.

Against 768 GB/s 1.25× on reads

Multi-card plans need rewriting

This is the constraint people discover late. There is no NVLink connector on the board, so a two-card job runs its collectives over PCIe. Pipeline-parallel inference barely notices. Gradient-heavy training notices a great deal. Decide which of those you are doing before you decide how many cards to rent.

Peer interconnect PCIe only

The generational break,
column by column.

Two 48 GB boards and two 24 GB ones, chosen so the Ada and Ampere rows sit next to each other. The NVLink row is the one that surprises people, and the FP8 row is the one that justifies the price.

RTX 6000 Ada RTX A6000 NVIDIA L40S RTX 4090
Architecture Ada AD102 Ampere GA102 Ada AD102 Ada AD102
VRAM 48 GB GDDR6 ECC 48 GB GDDR6 ECC 48 GB GDDR6 ECC 24 GB GDDR6X
Memory bandwidth 960 GB/s 768 GB/s 864 GB/s 1,008 GB/s
FP8 tensor cores yes no yes yes
NVLink connector none 112.5 GB/s none none
FP16 / BF16 tensor dense 182.1 TFLOPS 154.8 TFLOPS 181 TFLOPS 165.2 TFLOPS

dense tensor throughput from each board's NVIDIA datasheet · 2:4 sparse figures are double and are not mixed in here

The configurator above
is the price. This is the supply.

Two things are true at once here. The bare-metal ladder for this card is published and fixed, and it is the widget at the top of the page. The per-minute side was a single listing at the snapshot, asking $0.208 per GPU-hour, which is one host's price and a fair place to start.

Per-minute marketplace

1 server, 2 cards
0 free at 18:05 UTC on 21 Sep 2026
  • The host sets the price, not CLORE
  • A 2-card server, so tensor-parallel work has somewhere to go
  • A spot order can be displaced by an on-demand one
  • One per-minute listing at the snapshot, $0.208 per GPU-hour, and bare metal from $1.30
Open the marketplace
PUBLISHED RATE

Bare metal

$1.30 / GPU-hour
at a 360-day term · $1.66 at the 30-day minimum
  • $1.66 at 30 days, $1.44 at 90, $1.36 at 180, $1.30 from 360
  • Minimum 8 RTX 6000 Ada cards per contract
  • Offered in the USA, France and Japan
  • Roughly 75 per cent above the A6000 ladder, for FP8 and 960 GB/s
Configure a bare-metal block
Renters pay in
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Settle the precision question
before anything else.

On this board the interesting decisions are about number formats and interconnect, not about clicking rent. Both of them change what you should book, so they go first.

01 / PRECISION

Check whether your stack speaks FP8

If it does, you are buying something the A6000 cannot do and about 1 GB per billion parameters is your weight budget. If it does not, you are buying bandwidth and CUDA cores, which is a much smaller gap.

02 / TOPOLOGY

Decide how many cards, knowing there is no NVLink

Collectives run over PCIe here. Pipeline-parallel inference copes; gradient-heavy training pays for it. If you need a peer link, you are looking at a different board.

03 / SUPPLY

Look at what exists

One 2-card server on 21 Sep 2026, rented. Filter and see what today holds, or price the bare-metal ladder if you need certainty rather than luck.

$ clore rent --gpu "RTX 6000 Ada"
04 / MEASURE

Watch the clocks for an hour

18,176 cores inside 300 W means the power limit binds before the cores do, and sustained behaviour depends on a chassis you do not own. Check it before you commit a long run.

Questions hosts and renters ask.

The RTX 6000 Ada has no NVLink. How do I scale past one card?

Over PCIe, with the framework doing the work. Tensor and pipeline parallelism in vLLM, DeepSpeed or FSDP do not require NVLink, they require a working interconnect, and PCIe is one. What you give up is peer bandwidth, which hurts most in training where gradients cross the link every step, and least in pipeline-parallel inference where the handoff is one activation tensor per stage. If your plan is gradient-heavy multi-card training, the A6000 and its 112.5 GB/s bridge is the Ampere-era answer.

What does FP8 on Ada give me on 48 GB that the A6000's INT8 does not?

A narrower weight format with silicon behind it. Ada's fourth-generation tensor cores execute FP8 natively, so a model in eight-bit floating point costs roughly 1 GB per billion parameters on a path the hardware was designed for. The A6000 has no FP8 units at all, so the comparable saving there comes from INT8, which needs its own calibration and behaves differently on outlier activations. On a 48 GB board the practical effect is the same headline capacity with more of it left after the weights land.

960 GB/s versus the A6000's 768: where does that show up in serving latency?

In the token-by-token part of generation, which re-reads the weights from memory for every token produced. A 32 GB resident model means 32 GB of reads per token, and at 960 GB/s that is a shorter trip than at 768. The ratio, about 1.25, is the honest upper bound on what bandwidth alone can give you. Prompt processing and anything compute-bound will not track it, so treat 25 per cent as a ceiling on one component, not a promise about your end-to-end latency.

Is 18,176 CUDA cores at 300 W realistic under sustained load, or does it throttle?

The board is specified at 300 W and the power limit, not the core count, is the binding constraint. Whether it holds clocks over a long run is a property of the chassis: inlet temperature, airflow, and whether the card is sandwiched against another. On a rented machine you do not control any of that, so if clock stability matters to your job, watch the clocks for the first hour rather than trusting a datasheet. Worth noting: this is the same 300 W envelope as the A6000 with about 70 per cent more CUDA cores inside it.

Qwen2.5 32B at FP8 is about 32 GB. What concurrency does 48 GB support at that size?

At roughly 1 GB per billion parameters, a 32B model at FP8 is about 32 GB, leaving something near 16 GB for the KV cache, activations and context. That is a materially better position than a 70B four-bit load on the same board, which leaves around 8 GB. How many concurrent sequences that becomes depends on your context length and attention implementation, so measure it: the count falls roughly in proportion to tokens of context per request, and a paged KV cache stretches it considerably further than a naive one.

Is this card worth 75 per cent more per GPU-hour than an A6000 on bare metal?

Look at what the difference buys. Bare-metal RTX 6000 Ada is $1.66 per GPU-hour at the 30-day minimum against $0.95 for the A6000, and $1.30 against $0.67 at 360 days. For that you get 960 GB/s instead of 768, FP8 tensor cores, and about 70 per cent more CUDA cores, and you lose the NVLink connector. If your workload is FP8 inference, the newer board does something the older one physically cannot. If it is four-bit inference or rendering, you may be paying for headroom you will not use.

What FP8 changes,
and what it does not.

Weight footprints use the standard precision arithmetic: about 1 GB per billion parameters at FP8 and INT8, 0.5 GB at INT4, 2 GB at FP16, plus 10 to 30 per cent for the KV cache and activations. Throughput belongs to your runtime, so the figures quoted here are capacities.

Qwen2.5 32B at FP8
Native eight-bit float, one card
~32 GB weights, ~16 GB left

The size this board was built for. Enough spare capacity to batch seriously, on a numeric format the tensor cores execute directly rather than emulate.

Read the guide →
Llama 3.3 70B needs four bits here
INT4, not FP8
~70 GB at FP8, ~40 GB at INT4

Worth stating plainly because it is widely got wrong: a 70B model in eight-bit float is about 70 GB and does not fit this card. Four-bit does, with roughly 8 GB to spare.

Read the guide →
Diffusion and path tracing
Flux, SDXL, Cycles
48 GB resident, 300 W

Image and render work leans on raw cores rather than precision tricks, and this is where the 70 per cent CUDA-core advantage over the A6000 shows up most directly.

Read the guide →

What a term commitment
does to the price.

Every pro-tier board CLORE offers as bare metal, at the four term lengths the configurator prices. Rates are per GPU-hour, customer-facing, and taken from the live configurator on 21 Sep 2026. Beyond 360 days the price stops falling.

GPU
30 days
90 days
180 days
360 days +
Min GPUs
Regions
RTX 6000 Ada / this page
$1.66
$1.44
$1.36
$1.30
8
US / FR / JP
RTX A6000
$0.95
$0.83
$0.75
$0.67
8
US / FR
NVIDIA A40
$0.75
$0.65
$0.59
$0.53
8
US / FR / JP
NVIDIA L40S
$1.40
$1.21
$1.10
$1.04
8
US / JP / SI
A100 40GB
$2.28
$1.98
$1.81
$1.54
8
US / JP / SI

Seven guides,
written for this generation.

Container, commands and caveats for each, on the docs site. Several of them care whether your card has FP8 units, which is exactly the question this board answers differently from the A6000.

Other Workloads
Blender + Cycles GPU
Where 18,176 cores show up more plainly than in any language model.
Training
LLM fine-tuning
Adapter training, which stays inside one board and avoids the PCIe question.
Language Models
vLLM serving
Start here if your quantised weights are eight-bit float rather than integer.
Video Generation
Wan Video
Video generation, which wants every gigabyte the board has and then some.
Language Models
Llama 3.3 on CLORE.AI
A 70B model, so plan on four-bit weights rather than eight on 48 GB.
Image Generation
Flux.1 on CLORE.AI
Diffusion at full precision, with 48 GB removing the usual offload dance.
Advanced
Multi-GPU setup
Read the NCCL part and skip the NVLink part, because this board has none.
See all guides →

The three boards people
weigh against this one.

RTX A6000
48 GB ECC · 768 GB/s · no FP8 · has NVLink
The previous generation →
RTX 5090
32 GB GDDR7 · consumer board, no ECC
Faster, smaller, cheaper →
A100 80GB
80 GB HBM2e · 1,935 GB/s · MIG · no FP8
More memory, older tensor cores →

FP8 or four bits.
Decide, then book.

An eight-bit model at 32B leaves 16 GB of room on this board; a 70B needs four-bit weights and leaves 8 GB. That one choice decides whether the FP8 units are why you are here. Per-minute supply was one server on 21 Sep 2026, so check the listing page, or price a bare-metal block if you need capacity you can plan around.