The NVIDIA A10 is available on Clore.ai as bare metal: dedicated servers from $0.43 per GPU-hour on the longest term, in blocks from 8 GPUs with terms starting at 30 days, deployed in the USA, France or Japan and paid in BTC, USDT or USDC. The A10 family is the hardware hyperscalers reach for when they build an inference instance type, which is why people arrive here carrying a number from one. Those cloud instances run the sibling A10G, so an A10 gets you close rather than identical: 24 GB of GDDR6 at 600 GB/s on a single-slot, passive card.
Almost every mistake made about this card comes from one of three sources: a name collision, a MIG claim that was never true, and an assumption that Ampere and Ada behave the same at low precision.
A10, A10G and A100 are three different parts. Everything quoted here comes from NVIDIA's A10 datasheet and describes only the A10: GA102 silicon, 9,216 CUDA cores, 24 GB of GDDR6, 600 GB/s, 150 W. Before you carry a benchmark over from somewhere else, confirm which of the three it was run on.
Four-way MIG partitioning on an A10 is a claim you will find in plenty of places, and it is wrong in all of them. The capability does not exist on this part and never has. MIG belongs to the A100, A30, H100, H200 and B200 families and the RTX PRO Blackwell parts, verified against NVIDIA's own supported-GPU list on 21 September 2026.
FP8 tensor cores arrived with Ada and Hopper, so the A10 does not have them. In capacity terms INT8 puts you in the same place, about 1 GB per billion parameters, but a serving stack built on the Ada FP8 path will fall back to something else here. Plan quantization around INT8 and INT4.
The A10 next to the three cards it is most often weighed against. The bottom row is the one worth reading twice.
MIG column checked against the NVIDIA MIG supported-GPU list on 21 Sep 2026 · tensor figures are dense
Zero A10 listings on 21 September 2026, so nothing on this page quotes an hourly marketplace rate for the card. The contract route was available, and its rate card is public.
With nothing on the per-minute marketplace, the path runs through the configurator rather than through a listing.
24 GB, 600 GB/s, 150 W, no FP8, no MIG. If the requirement came from a benchmark on another part with a similar name, revisit it before ordering.
Eight GPUs is the floor and thirty days is the shortest term. Longer terms step the rate down to $0.43 per GPU-hour at 360 days.
A10 blocks are offered in the USA, France and Japan. The configurator on this page quotes each one.
Bare-metal orders settle from your Bitcoin balance or your USDT or USDC stablecoin balance.
Close to it, not the same as it, and the gap is worth understanding before you order. The A10 family is the hardware the large clouds standardised on for inference instance types, which is why this is such a common route onto the page. The catch is that a g5 instance carries the A10G, a sibling part, and the A100 is a third product sharing nothing but three characters. What CLORE offers, and what every figure on this page describes, is the A10 from NVIDIA's A10 datasheet: 24 GB of GDDR6, 600 GB/s, 9,216 CUDA cores, 150 W, no FP8 and no MIG. Treat a hyperscaler number as a neighbouring data point rather than a target, and confirm which part produced it.
No, however often you read otherwise. MIG exists on the A100, A30, H100, H200 and B200 families and on the RTX PRO Blackwell parts, checked against NVIDIA's MIG supported-GPU list on 21 September 2026. The A10 is not on it. If the plan depends on handing tenants hardware-isolated slices of one card, the A100 is the smallest part that does it, with up to seven instances.
Neither, until you say what you are optimising. Per watt the L4 is ahead on every axis. Per card the A10 has twice the bandwidth, which is the number that caps token generation, and roughly twice the dense FP16 tensor throughput at 125 TFLOPS against 60.5. On bare-metal rates the A10 runs $0.43 to $0.64 per GPU-hour against $0.41 to $0.52 for the L4, so the A10 costs slightly more for meaningfully more throughput. Where the L4 pulls ahead on capability rather than efficiency is FP8, which Ampere does not have at all.
Memory, mostly. Without FP8 the practical quantized formats on an A10 are INT8 and INT4, so an 8B model takes about 8 GB at INT8 or 4 GB at INT4 rather than benefiting from an FP8 path with its own accumulation behaviour. In capacity terms INT8 and FP8 land in the same place, roughly 1 GB per billion parameters. What you give up is the tensor-core path Ada added for that format and whatever a given serving stack has built on top of it. On a card whose real constraint is 600 GB/s of bandwidth, that matters less than it would on a bigger part.
Start from the ceiling. 8B at FP16 is about 16 GB of weights, and 600 GB/s divided by 16 GB is roughly 37 full weight passes per second before anything else competes for the bus. That leaves about 8 GB for the KV cache, which is what actually limits how many requests you can hold at once. Drop to INT8 and the weights halve to about 8 GB: the ceiling roughly doubles and the cache budget triples. On this card the quantization decision is a concurrency decision.
As bare metal. A10s on Clore.ai are sold as bare-metal blocks, and the rate card is the price: from $0.43 per GPU-hour. The per-minute marketplace carried no A10 servers on 21 September 2026, so the contract route is the way in: a minimum of 8 A10 GPUs on a minimum 30-day term, $0.64 per GPU-hour at 30 days falling to $0.43 from 360 days, deployed in the USA, France or Japan. If you want a single card for an afternoon instead, the L4 and the RTX 4090 both had listings on the same day.
Capacity and bandwidth arithmetic off the A10 datasheet rather than benchmark runs. Read them as ceilings your own measurements will sit under.
Two gigabytes per billion parameters leaves roughly 8 GB for the KV cache, which is the budget that decides how many requests you can hold open at once.
Read the guide →Ampere has INT8 but not FP8, so this is the low-precision path here. Halving the weights roughly triples what is left for concurrency.
Read the guide →The A10 fits standard server airflow without a taller cooler, which is why it turns up in mixed inference and virtual-workstation roles rather than in training racks.
Read the guide →What separates them is bandwidth, watts and whether the low-precision path exists. The last column is the marketplace snapshot of 21 September 2026 and explains why this page routes to bare metal.
not one of these four supports MIG · listing counts are a single snapshot, not an average
Walkthroughs on docs.clore.ai. Nothing here assumes FP8, because this card does not have it, and nothing here assumes a partitioned GPU.
Hosts who supply A10s on bare-metal contracts earn up to ~$420/mo per card, and single servers can be listed per minute at their own price.
Eight A10 GPUs is the smallest block, thirty days the shortest term, and $0.43 per GPU-hour the floor at a year. Every figure on this page is dated and sourced.