Log in A40 rate card
NVIDIA A40 · 48 GB ECC · bare metal · from $0.53/GPU-hour, blocks from 8 GPUs

Cooled by whatever
chassis you put an A40 in.
It has its own fan.

The NVIDIA A40 is available on Clore.ai as bare metal from $0.53 per GPU-hour, in blocks from 8 GPUs on terms from 30 days, hosted in the USA, France or Japan and payable in BTC, USDT or USDC. Each card is the same Ampere GA102 silicon as the RTX A6000, with 48 GB of GDDR6 and ECC, 696 GB/s of bandwidth and NVLink in pairs, built without a fan of its own to sit in a chassis that supplies the airflow.

●Blocks from 8 GPUs, terms from 30 days ●USA · France · Japan ●Pay in BTC, USDT or USDC ●Ampere GA102, 10,752 CUDA cores
$0.53
Per GPU-hour on bare metal, at the 360-day term
8GPUs
Smallest block you can order
30days
Shortest term, at $0.75 per GPU-hour
3regions
USA, France and Japan to deploy in

The same chip,
a different machine.

Everything that distinguishes the A40 from its workstation twin is physical rather than computational: how it is cooled, how it is connected, and where that lets it live.

A card that cannot cool itself

The A40 has no fan. It is designed for a chassis that forces air across it, which is why it appears in rack servers and not in towers. That single property decides who can host one, and it is the reason A40 supply behaves differently from GeForce supply on this marketplace.

Board power 300 W passive

NVLink, and what it is not

Two bridged A40s talk to each other at roughly 112.5 GB/s, which beats PCIe for tensor and pipeline parallelism. They do not become one 96 GB card. Splitting a model larger than 48 GB across the pair is your framework's job; the interconnect only makes it cheaper.

NVLink topology Pairs only

Ampere capacity, Ampere limits

Forty-eight gigabytes is enough for a 70B model at INT4 or a 32B model at INT8 with cache to spare. What is not on the menu is FP8, which arrived with Ada, and MIG, which this card does not have despite being a datacenter part.

FP8 and MIG Neither

Four Ampere cards,
one of them different.

The A40 against its workstation twin, the smaller A5000, and the HBM part people reach for when bandwidth is the constraint. Tensor figures are dense throughout.

NVIDIA A40 RTX A6000 RTX A5000 A100 40GB
Architecture Ampere GA102 Ampere GA102 Ampere GA102 Ampere GA100
VRAM 48 GB GDDR6 ECC 48 GB GDDR6 ECC 24 GB GDDR6 ECC 40 GB HBM2e
Memory bandwidth 696 GB/s 768 GB/s 768 GB/s 1,555 GB/s
Board power 300 W 300 W 230 W 400 W
FP16 tensor (dense) 149.7 TFLOPS 154.8 TFLOPS 111.1 TFLOPS 312 TFLOPS
NVLink pairs, 112.5 GB/s pairs, 112.5 GB/s pairs, 112.5 GB/s 600 GB/s on SXM4

none of these NVLink links pools memory · only the A100 in this table supports MIG

The rate falls
as the term grows.

A40 capacity is sold on fixed terms from eight GPUs up, and the only variable that moves the price is how long you commit for. Below are the two ends of that ladder. There is no per-minute column here because nothing with an A40 in it was listed on 21 September 2026, and inventing a starting price for capacity that was not there is exactly the habit this page was rewritten to remove.

Shortest term

$0.75 / GPU-hour, 30 days
The floor on commitment, not on price
  • Thirty days is the shortest A40 contract on offer
  • Eight GPUs is the smallest block at any term
  • Still under the RTX A6000, which starts at $0.95 for the same 48 GB
  • Shorter than a month is a marketplace question, and today it has no answer
Check the marketplace anyway
29% CHEAPER PER HOUR

Longest term

$0.53 / GPU-hour, 360 days
And unchanged at 720, 1,080, 1,440 and 1,800 days
  • The ladder in between: $0.65 at 90 days, $0.59 at 180
  • Past a year the rate stops moving, so longer buys certainty rather than discount
  • Hosted in the USA, France or Japan
  • Settled in Bitcoin, USDT or USDC
Open the configurator
Pay with
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Deciding between the twins.

The A40 and the A6000 do the same arithmetic. Which one you want comes down to four questions, in this order.

01 / CAPACITY

Does it fit in 48 GB?

A 70B model at INT4 is about 40 GB and fits with cache to spare. At FP16 it is roughly 140 GB and does not fit on one card, or on a bridged pair.

02 / BANDWIDTH

Is 696 GB/s enough?

If generation speed is your constraint rather than capacity, the extra memory bandwidth of an A100 will do more for you than the extra gigabytes here.

03 / DURATION

Days or months?

With no per-minute listings for this card, short jobs point at the A6000 or the L40S. A month or more is what the A40 rate card is built for.

04 / ORDER

Size it and pay

Eight GPUs is the floor and thirty days the shortest term. Bitcoin, USDT and USDC settle bare-metal orders.

Six questions about the A40.

A40 against RTX A6000: same silicon, same 48 GB. What is actually different?

Cooling and bandwidth, and that is close to the whole list. Both are Ampere GA102 with 10,752 CUDA cores and 48 GB of GDDR6 with ECC at 300 W. The A6000 runs 768 GB/s, the A40 696, so the A40 gives up about nine percent of the memory bandwidth. In exchange it is built as a passively cooled datacenter card that expects chassis airflow rather than carrying its own. On CLORE the practical difference is price: bare-metal A40 runs $0.53 to $0.75 per GPU-hour against $0.67 to $0.95 for the A6000.

The A40 is passively cooled. What does that mean for the machines you are renting?

It means they are rack servers, not desktops. A card with no fan of its own needs a chassis that pushes air through it, which rules out the tower under someone's desk and limits A40 supply to operators with real server hardware. For a renter that is mostly good news, since the population skews toward datacenter deployments. It also explains the availability picture on this page: fewer people can host this card than can host a GeForce part.

Does the A40 support MIG?

It does not, and the surprise is the point worth dwelling on: this is a datacenter card, and datacenter cards are exactly where people expect to find the feature. Ampere shipped it on two products, the A100 and the A30. The A40 is neither. It runs GA102, the die it shares with the RTX A6000, and GA102 was never given partitioning. NVIDIA's supported-GPU list, checked on 21 September 2026, is the reference. So an A40 rental is always the whole card, and a workload that genuinely needs isolated instances belongs on an A100 node.

The A40 came out of virtualization. What does that give you on a single-tenant rental?

On a single-tenant rental, nothing you can point at, and it is worth saying so rather than dressing it up. The features that make this card attractive to a virtualization platform are about carving one GPU between many desktops, and a rented server hands you the whole card. What does carry over is the design intent behind it: ECC memory, passive cooling and a duty cycle meant for machines that stay on, which is a reasonable thing to want from hardware you are renting for a month.

696 GB/s over 48 GB: which 70B quantizations are bandwidth-bound?

All of the ones that fit, which is the honest answer. At roughly 0.5 GB per billion parameters a 70B model at INT4 is about 40 GB, leaving 8 GB for the KV cache. Dividing 696 GB/s by 40 GB gives about 17 full weight passes per second as a ceiling before the cache and kernel overheads take their cut. FP8 is not an option here at all, since Ampere has no FP8 tensor cores, and FP16 at roughly 140 GB is far outside 48 GB. If bandwidth is your binding constraint rather than capacity, an A100 at 1,555 GB/s is the card that fixes it.

The A40 has NVLink. Does a bridged pair give you 96 GB?

No, and this is the most common misreading of the spec. NVLink on the A40 is a peer-to-peer link between two cards at roughly 112.5 GB/s, useful for tensor and pipeline parallelism because it moves data between cards faster than PCIe. It does not merge two 48 GB cards into one 96 GB address space, and no NVLink implementation on any of these Ampere parts does. A model that needs more than 48 GB has to be split across the two cards by your framework, not by the interconnect.

Three sums off the datasheet.

Capacity from parameter counts, ceilings from 696 GB/s, and one card about what the interconnect does not do. None of these are benchmark results.

The big model that fits
Llama 3.3 70B at INT4
≈40 GB of 48, 8 GB for cache

Half a gigabyte per billion parameters gets a 70B model onto one A40. The same model at FP16 is roughly 140 GB and is not a single-card proposition anywhere in this tier.

Read the guide →
What 696 GB/s allows
Bandwidth divided by 40 GB of weights
≈17 passes per second, at best

Every generated token reads the weights once, so this is the arithmetic ceiling on a 70B INT4 deployment before cache traffic and kernel overheads take their share.

Read the guide →
Two cards, not one big one
NVLink between a bridged pair
112.5 GB/s, still 48 GB each

The link is faster than PCIe for parallel work across cards. It is not a memory pool, so a model over 48 GB still has to be partitioned by the framework.

Read the guide →

Four cards with 48 GB.

Same capacity, four different answers on bandwidth, interconnect and low-precision support. The last column is what a block of eight costs per GPU-hour on a contract.

GPU
VRAM
Bandwidth (GB/s)
NVLink
FP8
Board power (W)
Bare metal $/GPU-hr
NVIDIA A40 / this page
48 GB GDDR6 ECC
696
pairs, 112.5
no
300
$0.53–$0.75
RTX A6000
48 GB GDDR6 ECC
768
pairs, 112.5
no
300
$0.67–$0.95
RTX 6000 Ada
48 GB GDDR6 ECC
960
none
yes
300
$1.30–$1.66
NVIDIA L40S
48 GB GDDR6 ECC
864
none
yes
350
$1.04–$1.40

NVLink here is a card-to-card link, never a memory pool · none of the four supports MIG · rates are the 30-day to 360-day band for a minimum of 8 GPUs

Seven guides for a 48 GB Ampere card.

Walkthroughs on docs.clore.ai. Everything here assumes capacity rather than bandwidth is your advantage, and none of it assumes FP8.

Other Workloads
Blender + Cycles GPU
Production rendering with Cycles on CUDA/OptiX.
Training
LLM fine-tuning
LoRA / QLoRA fine-tuning workflow.
Image Generation
ComfyUI on CLORE.AI
Node-based pipeline for SDXL, Flux, and SD3.
Video Generation
Stable Video Diffusion
Stability's image-to-video model.
Training
DreamBooth training
Fine-tune SDXL on your subject with DreamBooth.
Language Models
text-gen WebUI
The oobabooga WebUI for chat, RAG, and agents.
Advanced
Multi-GPU setup
Configure NVLink, NCCL, and distributed training.
See all guides →

The neighbours.

RTX A6000
48 GB · 768 GB/s · has its own fan
The same chip, unrestricted →
RTX 6000 Ada
48 GB · 960 GB/s · FP8, no NVLink
The Ada generation of this idea →
A100 40GB
40 GB HBM2e · 1,555 GB/s · MIG
Less memory, far more bandwidth →

Forty-eight gigabytes,
on a thirty-day floor.

The A40 is the cheapest 48 GB card CLORE offers on contract, at $0.53 to $0.75 per GPU-hour depending on the term. Every number on this page is dated and sourced.