The generational break from the A6000 runs in both directions, and pages that treat the two as interchangeable get it wrong. What the Ada part gains: 960 GB/s against 768, fourth-generation tensor cores with native FP8, and about 70 per cent more CUDA cores in the same 300 W. What it loses: the NVLink connector, so two of these talk over PCIe and nothing else. One server carrying two cards was listed on the per-minute marketplace on 21 Sep 2026 and it was rented; dedicated blocks are sold as bare metal from 8 cards on a 30-day term.
Capacity is the one number where these two boards tie, so it is the least interesting thing about either. Everything below is a place where the Ada generation does something the Ampere one cannot, or where it stops short.
Eight-bit floating point costs about 1 GB per billion parameters and, on this board, runs on tensor cores designed for it. The A6000 has no FP8 units, so the closest thing there is INT8 with its own calibration story. If your serving stack already speaks FP8, this is the cheapest board on the platform that speaks it back with 48 GB of room.
Token generation re-reads the whole resident model for every token, so memory speed sets the floor. Against the A6000's 768 GB/s this board reads 1.25 times faster, which is the honest ceiling on what that column buys. It is a real difference on decode and no difference at all on anything compute-bound.
This is the constraint people discover late. There is no NVLink connector on the board, so a two-card job runs its collectives over PCIe. Pipeline-parallel inference barely notices. Gradient-heavy training notices a great deal. Decide which of those you are doing before you decide how many cards to rent.
Two 48 GB boards and two 24 GB ones, chosen so the Ada and Ampere rows sit next to each other. The NVLink row is the one that surprises people, and the FP8 row is the one that justifies the price.
dense tensor throughput from each board's NVIDIA datasheet · 2:4 sparse figures are double and are not mixed in here
Two things are true at once here. The bare-metal ladder for this card is published and fixed, and it is the widget at the top of the page. The per-minute side was a single listing at the snapshot, asking $0.208 per GPU-hour, which is one host's price and a fair place to start.
On this board the interesting decisions are about number formats and interconnect, not about clicking rent. Both of them change what you should book, so they go first.
If it does, you are buying something the A6000 cannot do and about 1 GB per billion parameters is your weight budget. If it does not, you are buying bandwidth and CUDA cores, which is a much smaller gap.
Collectives run over PCIe here. Pipeline-parallel inference copes; gradient-heavy training pays for it. If you need a peer link, you are looking at a different board.
One 2-card server on 21 Sep 2026, rented. Filter and see what today holds, or price the bare-metal ladder if you need certainty rather than luck.
18,176 cores inside 300 W means the power limit binds before the cores do, and sustained behaviour depends on a chassis you do not own. Check it before you commit a long run.
Over PCIe, with the framework doing the work. Tensor and pipeline parallelism in vLLM, DeepSpeed or FSDP do not require NVLink, they require a working interconnect, and PCIe is one. What you give up is peer bandwidth, which hurts most in training where gradients cross the link every step, and least in pipeline-parallel inference where the handoff is one activation tensor per stage. If your plan is gradient-heavy multi-card training, the A6000 and its 112.5 GB/s bridge is the Ampere-era answer.
A narrower weight format with silicon behind it. Ada's fourth-generation tensor cores execute FP8 natively, so a model in eight-bit floating point costs roughly 1 GB per billion parameters on a path the hardware was designed for. The A6000 has no FP8 units at all, so the comparable saving there comes from INT8, which needs its own calibration and behaves differently on outlier activations. On a 48 GB board the practical effect is the same headline capacity with more of it left after the weights land.
In the token-by-token part of generation, which re-reads the weights from memory for every token produced. A 32 GB resident model means 32 GB of reads per token, and at 960 GB/s that is a shorter trip than at 768. The ratio, about 1.25, is the honest upper bound on what bandwidth alone can give you. Prompt processing and anything compute-bound will not track it, so treat 25 per cent as a ceiling on one component, not a promise about your end-to-end latency.
The board is specified at 300 W and the power limit, not the core count, is the binding constraint. Whether it holds clocks over a long run is a property of the chassis: inlet temperature, airflow, and whether the card is sandwiched against another. On a rented machine you do not control any of that, so if clock stability matters to your job, watch the clocks for the first hour rather than trusting a datasheet. Worth noting: this is the same 300 W envelope as the A6000 with about 70 per cent more CUDA cores inside it.
At roughly 1 GB per billion parameters, a 32B model at FP8 is about 32 GB, leaving something near 16 GB for the KV cache, activations and context. That is a materially better position than a 70B four-bit load on the same board, which leaves around 8 GB. How many concurrent sequences that becomes depends on your context length and attention implementation, so measure it: the count falls roughly in proportion to tokens of context per request, and a paged KV cache stretches it considerably further than a naive one.
Look at what the difference buys. Bare-metal RTX 6000 Ada is $1.66 per GPU-hour at the 30-day minimum against $0.95 for the A6000, and $1.30 against $0.67 at 360 days. For that you get 960 GB/s instead of 768, FP8 tensor cores, and about 70 per cent more CUDA cores, and you lose the NVLink connector. If your workload is FP8 inference, the newer board does something the older one physically cannot. If it is four-bit inference or rendering, you may be paying for headroom you will not use.
Weight footprints use the standard precision arithmetic: about 1 GB per billion parameters at FP8 and INT8, 0.5 GB at INT4, 2 GB at FP16, plus 10 to 30 per cent for the KV cache and activations. Throughput belongs to your runtime, so the figures quoted here are capacities.
The size this board was built for. Enough spare capacity to batch seriously, on a numeric format the tensor cores execute directly rather than emulate.
Read the guide →Worth stating plainly because it is widely got wrong: a 70B model in eight-bit float is about 70 GB and does not fit this card. Four-bit does, with roughly 8 GB to spare.
Read the guide →Image and render work leans on raw cores rather than precision tricks, and this is where the 70 per cent CUDA-core advantage over the A6000 shows up most directly.
Read the guide →Every pro-tier board CLORE offers as bare metal, at the four term lengths the configurator prices. Rates are per GPU-hour, customer-facing, and taken from the live configurator on 21 Sep 2026. Beyond 360 days the price stops falling.
Container, commands and caveats for each, on the docs site. Several of them care whether your card has FP8 units, which is exactly the question this board answers differently from the A6000.
On Clore.ai it can earn up to ~$1,100/mo per card, from per-minute listings to bare-metal contracts. The host page covers setup, pricing and the contract route.
An eight-bit model at 32B leaves 16 GB of room on this board; a 70B needs four-bit weights and leaves 8 GB. That one choice decides whether the FP8 units are why you are here. Per-minute supply was one server on 21 Sep 2026, so check the listing page, or price a bare-metal block if you need capacity you can plan around.