This is the cheapest card on CLORE that partitions in hardware. MIG carves one A100 into as many as seven instances of roughly 5 GB, each with its own multiprocessors and its own path to memory, so seven tenants share a board without sharing a failure domain. 40 GB of HBM2e at 1,555 GB/s, 6,912 CUDA cores, 312 dense BF16 TFLOPS, NVLink 3 between SXM4 boards. Ampere, so BF16 and TF32 but no FP8. Billed per minute, paid in BTC, CLORE, USDT or USDC.
Multi-Instance GPU is the reason to pick this card over anything cheaper. NVIDIA supports it on the A100, A30, H100, H200, B200 and GB200, and on the RTX PRO Blackwell parts. Nowhere else. On CLORE, the A100 40GB is the lowest-priced way to get it.
A MIG instance owns its multiprocessors and its own slice of HBM2e. Two tenants on the same board cannot contend for bandwidth, and a crash in one instance leaves the rest running. Time-slicing and MPS share the same engines and give you neither guarantee, which is why multi-tenant platforms reach for this card specifically.
A five-gigabyte slice is enough for Llama 3.2 1B or Qwen2.5 1.5B at FP16, or an 8B quantised to INT4 with a short context. Seven of those on one board is a different economic shape from seven separate GPUs, and the isolation is what lets you sell it as capacity rather than as best effort.
Partitioning is optional. Undivided, the board runs BF16 and TF32 training at 312 dense TFLOPS, and two SXM4 boards over NVLink 3 give you 80 GB of sharded capacity for a Qwen2.5 32B fine-tune at FP16. What it will not do is hold a 70B at INT4 with usable context, which is the 80 GB card's job.
MIG exists on a short list of datacenter parts. Here is what each of the rentable ones costs in memory, bandwidth and supply, with an unpartitionable consumer flagship for scale. Listing counts are the CLORE marketplace on 21 September 2026.
specs from the NVIDIA A100 datasheet and the NVIDIA MIG supported-GPUs list · listing counts from the CLORE marketplace, 21 Sep 2026
Hosts set their own rates, so these are a snapshot, not a tariff. Figures below are per GPU-hour across the 6 A100 40GB servers listed on 21 September 2026, derived from whole-server daily prices. Live prices are on the marketplace.
Need capacity you can plan around? CLORE also sells A100 bare metal from 8 GPUs on a 30-day minimum term, $1.54 to $2.28 per GPU-hour by contract length, in the USA, Japan and Slovenia. The configurator's A100 SKU is the 80 GB SXM part. Configure bare metal →
Partitioning is a host-side setting on the machine you rent, so the order of operations matters: take the card first, then decide how to cut it.
Filter the marketplace by A100 40GB. Supply is thin, so also sort by reliability and check the GPU count if you want more than one board in the same machine.
Spot is cheaper and preemptible; on-demand is yours until you stop it. Choose a CUDA image or bring your own.
You get an endpoint, an SSH key and Jupyter on port 8888. Confirm the card with nvidia-smi before you commit a long run.
MIG profiles are configured through nvidia-smi mig and need the instances created before workloads attach. Undivided is fine too. Billing rounds to the minute either way.
A 1g.5gb-class slice holds roughly 5 GB, so plan for small weights: Llama 3.2 1B at FP16 (about 2 GB), Qwen2.5 1.5B at FP16 (about 3 GB), or Llama 3 8B quantised to INT4 (about 4 GB) with a short context window. Anything at 7B or 8B in FP16 needs 16 GB or more and will not start in a single slice. If you need bigger models per tenant, use fewer, larger slices (the profiles scale up to the whole 40 GB card) or move to the 80 GB part where a slice is about 10 GB.
Hardware. A MIG instance gets its own streaming multiprocessors, its own slice of HBM2e and its own path to memory, so one tenant cannot steal bandwidth or cache from another and a crash in one instance does not take the others down. That is the difference between MIG and time-slicing or MPS, which share the same engines and only divide attention. It is also why partitioning is fixed at configuration time rather than negotiated per job.
Because 40 GB of weights on a 40 GB card leaves nothing for the KV cache, the activations or the CUDA context. In practice you get a model that loads and then runs out of memory at the first long prompt or the second concurrent request. Treat 40 GB as the ceiling for weights plus overhead, not for weights alone. The A100 80GB holds the same quantised 70B with real context headroom, and two 40 GB boards over NVLink give you 80 GB if you are willing to shard.
BF16 or TF32 for anything training-shaped, and INT8 or INT4 weight quantisation for serving. FP8 tensor cores arrived with Hopper, so an FP8 checkpoint or an FP8 kernel path will either refuse to build or silently fall back on this card. The practical consequence is that you compare an A100 against an H100 on memory and bandwidth, not on the FP8 headline numbers, and you pick quantisation schemes that Ampere actually executes.
On decode. Token generation reads the whole weight set for every token, so throughput tracks memory bandwidth almost linearly and a 25 percent gap shows up as roughly a 25 percent difference in tokens per second on the same model. It matters far less on prefill, on training steps that are compute-bound, and on anything small enough to sit in cache. If your workload is batch fine-tuning rather than interactive serving, the cheaper board is usually the better buy.
At the 21 September 2026 snapshot the per-minute marketplace carried 6 servers with 10 A100 40GB cards between them, and 4 of those servers were not rented at that instant. That is a small pool, so treat it as opportunistic capacity rather than something to plan a quarter around. For capacity you can plan around, the alternative is bare metal: CLORE's bare-metal A100 is the 80 GB SXM part, sold in blocks of at least 8 GPUs on a 30-day minimum term at $1.54 to $2.28 per GPU-hour depending on length, in the USA, Japan and Slovenia. Current listings and prices are always on the marketplace.
Specs are from the NVIDIA A100 datasheet and the NVIDIA MIG supported-GPUs list. Throughput depends on your model, batch and context, so benchmark before you budget.
One process per slice, one model per tenant, no noisy neighbour. Size the models to the slice, not the board.
Read the guide →40 GB carries a 7B or 8B adapter run comfortably once optimiser state is sharded or quantised. BF16 and TF32 both execute natively on Ampere.
Read the guide →Two 40 GB boards give 80 GB of sharded capacity, which is what a Qwen2.5 32B at FP16 needs. Sharding is not pooling: the model has to be split.
Read the guide →Supply and bare-metal terms as of 21 September 2026. Bare-metal prices are per GPU-hour and vary with contract length; every bare-metal SKU here starts at 8 GPUs and 30 days. For planned A100 capacity, that SKU is the 80 GB SXM part.
Every one of these runs on BF16 or INT8 rather than FP8, which is the constraint that matters on this card. Start with the serving guide if you plan to partition.
Earn up to ~$920/mo per card, paid for every rented minute in BTC, CLORE, USDT or USDC. MIG puts your board in front of renters who need isolated instances.
Six A100 40GB servers were listed at the last snapshot and four of them were free. Check what is on the marketplace now, or configure a bare-metal block if you need the capacity to still be there next month.