The NVIDIA L40S is a datacenter card that made two unusual choices. It carries FP8 tensor cores without HBM, so 48 GB of GDDR6 with ECC sits behind 864 GB/s rather than the 1,935 an A100 80GB gets. And it keeps its ray-tracing hardware, which the H100 dropped, so it is the one card here that can render and serve. What it does not have is NVLink or MIG. Page facts come from the NVIDIA L40S datasheet and a marketplace snapshot of 21 September 2026.
The L40S is not a budget H100 and not an oversized workstation card. It is 48 GB of ECC memory with FP8 arithmetic and graphics hardware attached, on a bus that is deliberately not HBM. Each of those choices costs something and buys something.
Ada's fourth-generation tensor cores give the L40S FP8, which no Ampere card on this site has, including the A100. It pays for that with bandwidth: 864 GB/s against the 1,935 of an A100 80GB. Compute-bound prefill favours the L40S. Bandwidth-bound decoding does not.
NVIDIA stripped the graphics pipeline out of the H100 and left it in the L40S. That is why this card turns up in render farms and virtual-production work as well as in inference racks, and why a single box can do both jobs on different shifts.
No MIG, so no hardware-isolated tenants on one card. No NVLink, so cards talk over PCIe and tightly coupled multi-GPU training is off the table. And 48 GB stops well short of a 70B model at FP16, which needs roughly 140 GB.
Deliberately so. Against the other 48 GB Ada card and the two HBM parts it competes with, the L40S trades bandwidth for price and keeps a graphics pipeline nobody else in this table has.
H100 figures are the SXM5 part · the PCIe H100 has 14,592 CUDA cores and roughly 2,000 GB/s
One L40S server was on the per-minute marketplace on 21 September 2026 and it was busy. A single occupied listing sets no price level, so the only rates on this page are the contract ones, which do not depend on any host being free.
Start from the model, not from the card. On 48 GB the precision decides whether the thing fits at all, and the bandwidth decides how fast it runs once it does.
Weights are about 2 GB per billion parameters at FP16, 1 GB at FP8 and 0.5 GB at INT4. Add a tenth to a third again for the KV cache before you decide 48 GB is enough.
A single listing on the marketplace suits an experiment. Anything that has to be there next month is a bare-metal contract, which starts at eight GPUs.
Marketplace orders run a container image of your choosing and bill per minute for as long as the order is up.
Bitcoin, CLORE and the USDT or USDC stablecoin balance are all accepted for marketplace and bare-metal orders alike.
When your model fits in 48 GB and your arithmetic is FP8. The L40S has FP8 tensor cores and Ampere does not, so on an A100 the same weights have to run at INT8 or FP16 and take more room. Where the A100 wins is memory: 80 GB against 48, and 1,935 GB/s against 864, more than twice the bandwidth that sets the token-generation ceiling. It also partitions into up to seven MIG instances, which the L40S cannot do at all. The short version is that the L40S is the better card for one model that fits and the A100 for one that does not, or for splitting a card between tenants.
Two things. Without MIG you cannot hand different tenants hardware-isolated slices of the same card, so multi-tenancy has to be containers and software limits. Without NVLink, multi-GPU work falls back to PCIe between cards, which rules the L40S out of tightly coupled training setups that assume a fast interconnect and makes tensor parallelism across cards expensive. Neither limit touches the common case, which is one model served from one card. Both are worth knowing before you plan around eight of them behaving as one.
It makes it possible, which the datacenter alternatives do not. The L40S retains ray-tracing hardware that an H100 does not carry, so a box that renders during one shift and serves a model during the next can be one card rather than two. Whether that is the right buy depends on the mix: if the render half is occasional, a cheaper 48 GB Ampere card such as the A40 covers it, and if the inference half is the whole job, the money is better spent on bandwidth.
INT4. At roughly 0.5 GB per billion parameters a 70B model lands near 40 GB, which leaves about 8 GB for the KV cache inside 48 GB of GDDR6 with ECC. FP8 would want about 70 GB and does not fit, and FP16 at around 140 GB is not close. Bandwidth then sets the pace: 864 GB/s across 40 GB of weights is about 21 full passes per second as a ceiling, before the cache and kernel overheads take their share. A 32B model at FP8 is the more comfortable configuration if you want concurrency rather than the largest model that will squeeze in.
The data format is the same idea and the machine around it is not. The L40S gets FP8 from fourth-generation Ada tensor cores, while FP8 and the Transformer Engine arrived with Hopper on the H100. The sizes tell the rest: 181 dense FP16 tensor TFLOPS against 989.4, and 864 GB/s of GDDR6 against 3,350 GB/s of HBM3 on the SXM5 part. Treat the L40S as a card that can run FP8 kernels, not as a cheaper H100.
Very little, in either direction. A sample of one server cannot tell you a utilization rate, and CLORE's own guidance is not to publish a single snapshot as though it were an average. What it does tell you is practical. If you needed L40S capacity on that day, the per-minute marketplace would not have given it to you, and the bare-metal route with its eight-GPU minimum is the one that was actually available.
Capacity from the parameter count, ceilings from the 864 GB/s bus. These are arithmetic on the datasheet rather than benchmark runs, so read them as the number your measurements will fall short of.
At roughly 0.5 GB per billion parameters the weights fit with about 8 GB spare for the KV cache. FP8 would need around 70 GB and does not fit at all.
Read the guide →A 32B model at FP8 leaves a third of the card for the cache, which is what buys concurrency. This is the shape the L40S is genuinely good at.
Read the guide →The L40S keeps the ray-tracing hardware NVIDIA removed from Hopper, so a render queue and a served endpoint can share one machine on different shifts.
Read the guide →The comparison this card is always dragged into. Bandwidth is where it loses, FP8 and price are where it answers. Tensor figures are dense, never with sparsity.
H100 bandwidth is the SXM5 figure · bare-metal rates are the 30-day to 360-day range for a minimum block of 8 GPUs
Walkthroughs on docs.clore.ai chosen for the two jobs this card was built for: serving a large quantized model, and generative media that needs capacity more than bandwidth.
List them on Clore.ai and earn up to ~$930/mo per card, per minute or on a bare-metal contract. Hosts price their own servers and are paid for every minute an order runs.
Whether an L40S is free on the marketplace changes by the hour. The bare-metal rate card does not: eight GPUs, thirty days, and a price that falls as the term lengthens.