Log in Rent an A100 40GB
A100 40GB · MIG-capable 6 servers listed, as of 21 Sep 2026

Rent an A100 40GB
and slice it seven ways.
One tenant per card.

This is the cheapest card on CLORE that partitions in hardware. MIG carves one A100 into as many as seven instances of roughly 5 GB, each with its own multiprocessors and its own path to memory, so seven tenants share a board without sharing a failure domain. 40 GB of HBM2e at 1,555 GB/s, 6,912 CUDA cores, 312 dense BF16 TFLOPS, NVLink 3 between SXM4 boards. Ampere, so BF16 and TF32 but no FP8. Billed per minute, paid in BTC, CLORE, USDT or USDC.

●Billed per minute ●SSH, Docker and Jupyter ●Spot and on-demand ●Bare metal from 8 GPUs
7×
MIG instances per card, about 5 GB each
40GB
HBM2e at 1,555 GB/s
6
Servers listed per minute, 21 Sep 2026 (10 cards)
$0.417/GPU-h
Median on-demand at that snapshot

One board,
seven tenants.

Multi-Instance GPU is the reason to pick this card over anything cheaper. NVIDIA supports it on the A100, A30, H100, H200, B200 and GB200, and on the RTX PRO Blackwell parts. Nowhere else. On CLORE, the A100 40GB is the lowest-priced way to get it.

Partitions that are real, not scheduled

A MIG instance owns its multiprocessors and its own slice of HBM2e. Two tenants on the same board cannot contend for bandwidth, and a crash in one instance leaves the rest running. Time-slicing and MPS share the same engines and give you neither guarantee, which is why multi-tenant platforms reach for this card specifically.

MIG instances per board up to 7

Small models at density

A five-gigabyte slice is enough for Llama 3.2 1B or Qwen2.5 1.5B at FP16, or an 8B quantised to INT4 with a short context. Seven of those on one board is a different economic shape from seven separate GPUs, and the isolation is what lets you sell it as capacity rather than as best effort.

Per-slice memory about 5 GB

Or keep the card whole

Partitioning is optional. Undivided, the board runs BF16 and TF32 training at 312 dense TFLOPS, and two SXM4 boards over NVLink 3 give you 80 GB of sharded capacity for a Qwen2.5 32B fine-tune at FP16. What it will not do is hold a 70B at INT4 with usable context, which is the 80 GB card's job.

NVLink 3 between boards 600 GB/s

Against the other cards
that can be partitioned.

MIG exists on a short list of datacenter parts. Here is what each of the rentable ones costs in memory, bandwidth and supply, with an unpartitionable consumer flagship for scale. Listing counts are the CLORE marketplace on 21 September 2026.

A100 40GB A100 80GB H100 80GB SXM5 RTX 4090
Architecture Ampere GA100 Ampere GA100 Hopper GH100 Ada Lovelace
MIG instances up to 7 up to 7 up to 7 none
VRAM 40 GB HBM2e 80 GB HBM2e 80 GB HBM3 24 GB GDDR6X
Memory bandwidth 1,555 GB/s 1,935 GB/s 3,350 GB/s 1,008 GB/s
FP16 / BF16 (dense) 312 TFLOPS 312 TFLOPS 989 TFLOPS 165 TFLOPS
FP8 tensor cores no no yes yes
Servers listed 21 Sep 2026 6 0 4 243

specs from the NVIDIA A100 datasheet and the NVIDIA MIG supported-GPUs list · listing counts from the CLORE marketplace, 21 Sep 2026

What the 6 listings
were charging.

Hosts set their own rates, so these are a snapshot, not a tariff. Figures below are per GPU-hour across the 6 A100 40GB servers listed on 21 September 2026, derived from whole-server daily prices. Live prices are on the marketplace.

Spot

$0.333 / GPU-hour
median of 6 listings · cheapest seen $0.167
  • Cheapest way onto the card
  • Billed per minute
  • An on-demand renter can take the machine back
  • Suits checkpointed fine-tunes and batch jobs
Browse spot listings
NOT PREEMPTIBLE

On-demand

$0.417 / GPU-hour
median of 6 listings · cheapest seen $0.208
  • Yours until you release it
  • No preemption
  • Billed per minute
  • Suits a MIG-partitioned serving box
Rent on demand

Need capacity you can plan around? CLORE also sells A100 bare metal from 8 GPUs on a 30-day minimum term, $1.54 to $2.28 per GPU-hour by contract length, in the USA, Japan and Slovenia. The configurator's A100 SKU is the 80 GB SXM part. Configure bare metal →

Pay with
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

From listing to MIG slices.

Partitioning is a host-side setting on the machine you rent, so the order of operations matters: take the card first, then decide how to cut it.

01 / FILTER

Find a board

Filter the marketplace by A100 40GB. Supply is thin, so also sort by reliability and check the GPU count if you want more than one board in the same machine.

02 / RENT

Pick spot or on-demand

Spot is cheaper and preemptible; on-demand is yours until you stop it. Choose a CUDA image or bring your own.

marketplace filter → A100 40GB
03 / CONNECT

SSH or Jupyter

You get an endpoint, an SSH key and Jupyter on port 8888. Confirm the card with nvidia-smi before you commit a long run.

04 / PARTITION

Cut it, or don't

MIG profiles are configured through nvidia-smi mig and need the instances created before workloads attach. Undivided is fine too. Billing rounds to the minute either way.

Questions about MIG and 40 GB.

Seven MIG instances of about 5 GB each. Which models actually fit in one slice?

A 1g.5gb-class slice holds roughly 5 GB, so plan for small weights: Llama 3.2 1B at FP16 (about 2 GB), Qwen2.5 1.5B at FP16 (about 3 GB), or Llama 3 8B quantised to INT4 (about 4 GB) with a short context window. Anything at 7B or 8B in FP16 needs 16 GB or more and will not start in a single slice. If you need bigger models per tenant, use fewer, larger slices (the profiles scale up to the whole 40 GB card) or move to the 80 GB part where a slice is about 10 GB.

Does a MIG instance give me hardware isolation, or just a scheduling quota?

Hardware. A MIG instance gets its own streaming multiprocessors, its own slice of HBM2e and its own path to memory, so one tenant cannot steal bandwidth or cache from another and a crash in one instance does not take the others down. That is the difference between MIG and time-slicing or MPS, which share the same engines and only divide attention. It is also why partitioning is fixed at configuration time rather than negotiated per job.

Llama 3 70B at INT4 is about 40 GB. Why does the 40 GB card not really work for it?

Because 40 GB of weights on a 40 GB card leaves nothing for the KV cache, the activations or the CUDA context. In practice you get a model that loads and then runs out of memory at the first long prompt or the second concurrent request. Treat 40 GB as the ceiling for weights plus overhead, not for weights alone. The A100 80GB holds the same quantised 70B with real context headroom, and two 40 GB boards over NVLink give you 80 GB if you are willing to shard.

Ampere has no FP8. What is the A100 actual serving format in 2026?

BF16 or TF32 for anything training-shaped, and INT8 or INT4 weight quantisation for serving. FP8 tensor cores arrived with Hopper, so an FP8 checkpoint or an FP8 kernel path will either refuse to build or silently fall back on this card. The practical consequence is that you compare an A100 against an H100 on memory and bandwidth, not on the FP8 headline numbers, and you pick quantisation schemes that Ampere actually executes.

This card runs at 1,555 GB/s against the 80 GB part 1,935 GB/s. When does that 25 percent matter?

On decode. Token generation reads the whole weight set for every token, so throughput tracks memory bandwidth almost linearly and a 25 percent gap shows up as roughly a 25 percent difference in tokens per second on the same model. It matters far less on prefill, on training steps that are compute-bound, and on anything small enough to sit in cache. If your workload is batch fine-tuning rather than interactive serving, the cheaper board is usually the better buy.

How many A100 40GB servers can I actually rent today?

At the 21 September 2026 snapshot the per-minute marketplace carried 6 servers with 10 A100 40GB cards between them, and 4 of those servers were not rented at that instant. That is a small pool, so treat it as opportunistic capacity rather than something to plan a quarter around. For capacity you can plan around, the alternative is bare metal: CLORE's bare-metal A100 is the 80 GB SXM part, sold in blocks of at least 8 GPUs on a 30-day minimum term at $1.54 to $2.28 per GPU-hour depending on length, in the USA, Japan and Slovenia. Current listings and prices are always on the marketplace.

Three shapes of A100 40GB job.

Specs are from the NVIDIA A100 datasheet and the NVIDIA MIG supported-GPUs list. Throughput depends on your model, batch and context, so benchmark before you budget.

Multi-tenant serving
MIG profiles + Triton or vLLM per slice
7 instances × ~5 GB

One process per slice, one model per tenant, no noisy neighbour. Size the models to the slice, not the board.

Read the guide →
Fine-tuning on a whole board
LoRA or QLoRA + BF16 + Flash Attention 2
312 dense BF16 TFLOPS

40 GB carries a 7B or 8B adapter run comfortably once optimiser state is sharded or quantised. BF16 and TF32 both execute natively on Ampere.

Read the guide →
Two boards over NVLink
torchrun or DeepSpeed across SXM4 pairs
600 GB/s NVLink 3

Two 40 GB boards give 80 GB of sharded capacity, which is what a Qwen2.5 32B at FP16 needs. Sharding is not pooling: the model has to be split.

Read the guide →

What a slice costs, card by card.

Supply and bare-metal terms as of 21 September 2026. Bare-metal prices are per GPU-hour and vary with contract length; every bare-metal SKU here starts at 8 GPUs and 30 days. For planned A100 capacity, that SKU is the 80 GB SXM part.

GPU
VRAM
MIG
Mem BW (GB/s)
BF16 dense TFLOPS
FP8
Listed 21 Sep 2026
Bare metal $/GPU-h
A100 40GB / this page
40 GB HBM2e
7 × ~5 GB
1,555
312
—
6 servers
80 GB SXM: $1.54–$2.28
A100 80GB
80 GB HBM2e
7 × ~10 GB
1,935
312
—
0 servers
$1.54–$2.28
H100 80GB SXM5
80 GB HBM3
up to 7
3,350
989
yes
4 servers
$1.81–$2.62
Tesla V100
32 GB HBM2
—
900
125
—
186 servers
$0.36–$0.48
RTX 4090
24 GB GDDR6X
—
1,008
165
yes
243 servers
$0.40–$0.94

Guides that assume Ampere.

Every one of these runs on BF16 or INT8 rather than FP8, which is the constraint that matters on this card. Start with the serving guide if you plan to partition.

Training
LLM fine-tuning
LoRA / QLoRA fine-tuning workflow.
Training
DeepSpeed multi-GPU training
ZeRO-2/3 training across multiple cards.
Language Models
vLLM serving
High-throughput LLM serving with PagedAttention.
Training
HF Transformers training
Train and fine-tune with the Trainer API.
Language Models
Llama 3.3 on CLORE.AI
Run Meta's flagship Llama-3.3 on your rented card.
Language Models
Qwen 2.5
Alibaba's Qwen 2.5 family.
Advanced
Multi-GPU setup
Configure NVLink, NCCL, and distributed training.
See all guides →

Where to go from here.

A100 80GB
Same compute, 80 GB, ~10 GB per slice
Compare →
H100 80GB
MIG plus FP8 and 3,350 GB/s
Compare →
Tesla V100
No MIG, but 1,389 cards listed
Compare →

Seven tenants,
one board.

Six A100 40GB servers were listed at the last snapshot and four of them were free. Check what is on the marketplace now, or configure a bare-metal block if you need the capacity to still be there next month.