Log in Bare-metal A10
NVIDIA A10 · 24 GB · bare metal · from $0.43/GPU-hour in the USA, France and Japan

The A10 is not
an A10G.
It is an A100.

The NVIDIA A10 is available on Clore.ai as bare metal: dedicated servers from $0.43 per GPU-hour on the longest term, in blocks from 8 GPUs with terms starting at 30 days, deployed in the USA, France or Japan and paid in BTC, USDT or USDC. The A10 family is the hardware hyperscalers reach for when they build an inference instance type, which is why people arrive here carrying a number from one. Those cloud instances run the sibling A10G, so an A10 gets you close rather than identical: 24 GB of GDDR6 at 600 GB/s on a single-slot, passive card.

●600 GB/s of memory bandwidth ●Blocks from 8 GPUs, 30-day terms ●USA, France, Japan ●Pay in BTC, USDT or USDC
$0.43/GPU-h
Lowest bare-metal rate, on a 360-day term
8GPUs
Smallest block, on terms from 30 days
3regions
USA, France and Japan
24GB
GDDR6 per card, single-slot and passive

Three things people
get wrong about it.

Almost every mistake made about this card comes from one of three sources: a name collision, a MIG claim that was never true, and an assumption that Ampere and Ada behave the same at low precision.

The name is doing a lot of work

A10, A10G and A100 are three different parts. Everything quoted here comes from NVIDIA's A10 datasheet and describes only the A10: GA102 silicon, 9,216 CUDA cores, 24 GB of GDDR6, 600 GB/s, 150 W. Before you carry a benchmark over from somewhere else, confirm which of the three it was run on.

Architecture Ampere GA102

No MIG. Not four instances, not any

Four-way MIG partitioning on an A10 is a claim you will find in plenty of places, and it is wrong in all of them. The capability does not exist on this part and never has. MIG belongs to the A100, A30, H100, H200 and B200 families and the RTX PRO Blackwell parts, verified against NVIDIA's own supported-GPU list on 21 September 2026.

MIG instances None

Ampere stops at INT8

FP8 tensor cores arrived with Ada and Hopper, so the A10 does not have them. In capacity terms INT8 puts you in the same place, about 1 GB per billion parameters, but a serving stack built on the Ada FP8 path will fall back to something else here. Plan quantization around INT8 and INT4.

FP8 tensor cores Absent

Twenty-four gigabytes,
four ways.

The A10 next to the three cards it is most often weighed against. The bottom row is the one worth reading twice.

NVIDIA A10 NVIDIA L4 RTX A5000 A100 40GB
Architecture Ampere GA102 Ada AD104 Ampere GA102 Ampere GA100
VRAM 24 GB GDDR6 24 GB GDDR6 24 GB GDDR6 ECC 40 GB HBM2e
Memory bandwidth 600 GB/s 300 GB/s 768 GB/s 1,555 GB/s
Board power 150 W 72 W 230 W 400 W
FP16 tensor (dense) 125 TFLOPS 60.5 TFLOPS 111.1 TFLOPS 312 TFLOPS
MIG partitions none none none up to 7

MIG column checked against the NVIDIA MIG supported-GPU list on 21 Sep 2026 · tensor figures are dense

Priced as a contract,
not as a listing.

Zero A10 listings on 21 September 2026, so nothing on this page quotes an hourly marketplace rate for the card. The contract route was available, and its rate card is public.

Per-minute marketplace

0 servers listed
Snapshot taken 21 Sep 2026
  • No server carrying an A10 was on the marketplace at that moment
  • So there is no minimum, no median, and no starting price to quote
  • Supply comes from independent hosts and can change on any day
  • The marketplace is the only thing that knows what is online now
Look anyway
THE ROUTE THAT EXISTS

Bare metal

$0.43 – $0.64 / GPU-hour
Minimum 8 GPUs, minimum 30-day term
  • $0.64 per GPU-hour on a 30-day term
  • $0.54 at 90 days, $0.50 at 180, $0.43 from 360 days on
  • Deployed in the USA, France or Japan
  • Payable in Bitcoin, USDT or USDC
Configure a block
Pay with
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

How an A10 block is ordered.

With nothing on the per-minute marketplace, the path runs through the configurator rather than through a listing.

01 / CONFIRM

Check it is the A10 you want

24 GB, 600 GB/s, 150 W, no FP8, no MIG. If the requirement came from a benchmark on another part with a similar name, revisit it before ordering.

02 / SIZE

Pick a count and a term

Eight GPUs is the floor and thirty days is the shortest term. Longer terms step the rate down to $0.43 per GPU-hour at 360 days.

03 / PLACE

Choose a region

A10 blocks are offered in the USA, France and Japan. The configurator on this page quotes each one.

04 / PAY

Settle from your balance

Bare-metal orders settle from your Bitcoin balance or your USDT or USDC stablecoin balance.

Six questions, one of them a correction.

If you came here from an AWS g5 benchmark, what does an A10 actually reproduce?

Close to it, not the same as it, and the gap is worth understanding before you order. The A10 family is the hardware the large clouds standardised on for inference instance types, which is why this is such a common route onto the page. The catch is that a g5 instance carries the A10G, a sibling part, and the A100 is a third product sharing nothing but three characters. What CLORE offers, and what every figure on this page describes, is the A10 from NVIDIA's A10 datasheet: 24 GB of GDDR6, 600 GB/s, 9,216 CUDA cores, 150 W, no FP8 and no MIG. Treat a hyperscaler number as a neighbouring data point rather than a target, and confirm which part produced it.

Does the A10 support MIG?

No, however often you read otherwise. MIG exists on the A100, A30, H100, H200 and B200 families and on the RTX PRO Blackwell parts, checked against NVIDIA's MIG supported-GPU list on 21 September 2026. The A10 is not on it. If the plan depends on handing tenants hardware-isolated slices of one card, the A100 is the smallest part that does it, with up to seven instances.

A10 against L4: 600 GB/s at 150 W versus 300 GB/s at 72 W. Which one wins?

Neither, until you say what you are optimising. Per watt the L4 is ahead on every axis. Per card the A10 has twice the bandwidth, which is the number that caps token generation, and roughly twice the dense FP16 tensor throughput at 125 TFLOPS against 60.5. On bare-metal rates the A10 runs $0.43 to $0.64 per GPU-hour against $0.41 to $0.52 for the L4, so the A10 costs slightly more for meaningfully more throughput. Where the L4 pulls ahead on capability rather than efficiency is FP8, which Ampere does not have at all.

Ampere has no FP8. What does that cost against an Ada L4 on the same 24 GB?

Memory, mostly. Without FP8 the practical quantized formats on an A10 are INT8 and INT4, so an 8B model takes about 8 GB at INT8 or 4 GB at INT4 rather than benefiting from an FP8 path with its own accumulation behaviour. In capacity terms INT8 and FP8 land in the same place, roughly 1 GB per billion parameters. What you give up is the tensor-core path Ada added for that format and whatever a given serving stack has built on top of it. On a card whose real constraint is 600 GB/s of bandwidth, that matters less than it would on a bigger part.

Llama 3 8B fits in 24 GB at FP16. What concurrency does 600 GB/s sustain?

Start from the ceiling. 8B at FP16 is about 16 GB of weights, and 600 GB/s divided by 16 GB is roughly 37 full weight passes per second before anything else competes for the bus. That leaves about 8 GB for the KV cache, which is what actually limits how many requests you can hold at once. Drop to INT8 and the weights halve to about 8 GB: the ceiling roughly doubles and the cache budget triples. On this card the quantization decision is a concurrency decision.

How do I rent an A10 on Clore.ai?

As bare metal. A10s on Clore.ai are sold as bare-metal blocks, and the rate card is the price: from $0.43 per GPU-hour. The per-minute marketplace carried no A10 servers on 21 September 2026, so the contract route is the way in: a minimum of 8 A10 GPUs on a minimum 30-day term, $0.64 per GPU-hour at 30 days falling to $0.43 from 360 days, deployed in the USA, France or Japan. If you want a single card for an afternoon instead, the L4 and the RTX 4090 both had listings on the same day.

Three numbers, worked out.

Capacity and bandwidth arithmetic off the A10 datasheet rather than benchmark runs. Read them as ceilings your own measurements will sit under.

An 8B model at full precision
Llama 3 8B at FP16
≈16 GB of the 24 GB gone

Two gigabytes per billion parameters leaves roughly 8 GB for the KV cache, which is the budget that decides how many requests you can hold open at once.

Read the guide →
The same model, quantized
Llama 3 8B at INT8
≈8 GB, and triple the cache

Ampere has INT8 but not FP8, so this is the low-precision path here. Halving the weights roughly triples what is left for concurrency.

Read the guide →
Vision work on 150 watts
Detection and video pipelines
Single-slot, passive, 150 W

The A10 fits standard server airflow without a taller cooler, which is why it turns up in mixed inference and virtual-workstation roles rather than in training racks.

Read the guide →

Four cards with 24 GB.

What separates them is bandwidth, watts and whether the low-precision path exists. The last column is the marketplace snapshot of 21 September 2026 and explains why this page routes to bare metal.

GPU
VRAM
Bandwidth (GB/s)
Board power (W)
FP8
MIG
FP16 dense (TFLOPS)
Listed 21 Sep 2026
NVIDIA A10 / this page
24 GB GDDR6
600
150
no
none
125
0 servers
NVIDIA L4
24 GB GDDR6
300
72
yes
none
60.5
1 server
RTX A5000
24 GB GDDR6 ECC
768
230
no
none
111.1
1 server
GeForce RTX 4090
24 GB GDDR6X
1,008
450
yes
none
165.2
243 servers

not one of these four supports MIG · listing counts are a single snapshot, not an average

Seven guides sized for 24 GB.

Walkthroughs on docs.clore.ai. Nothing here assumes FP8, because this card does not have it, and nothing here assumes a partitioned GPU.

Language Models
vLLM serving
High-throughput LLM serving with PagedAttention.
Language Models
llama.cpp server
GGUF quantized inference with HTTP/OpenAI-compatible API.
Audio Voice
Whisper transcription
OpenAI Whisper-large for speech-to-text.
Computer Vision
YOLOv8 detection
Real-time object detection with YOLOv8.
Vision Models
Florence-2
Microsoft's compact vision-language model.
Video Processing
FFmpeg + NVENC
Hardware-accelerated video transcoding.
Advanced
CLORE API integration
Programmatic order creation via the public API.
See all guides →

Cards that had listings that day.

NVIDIA L4
24 GB · 72 W · 1 server listed
Half the bandwidth, a fifth the watts →
Tesla T4
16 GB · 70 W · 13 servers listed
Older, smaller, actually available →
NVIDIA L40S
48 GB · 864 GB/s · FP8
When 24 GB is the wrong size →

Bare-metal A10s
from $0.43/GPU-hour.

Eight A10 GPUs is the smallest block, thirty days the shortest term, and $0.43 per GPU-hour the floor at a year. Every figure on this page is dated and sourced.