The NVIDIA L4 is the lowest-power 24 GB card on this site: single-slot, low-profile, passively cooled, 72 W. It is the card you can leave serving an endpoint around the clock. Its honest weakness is the other side of that budget, 300 GB/s of memory bandwidth, the lowest of any 24 GB card on this site. Everything below is either from NVIDIA's L4 datasheet or from a marketplace snapshot taken on 21 September 2026.
Three things the L4 does that no other 24 GB card on this site does at the same time: it fits a single low-profile slot, it runs on 72 W, and it carries both FP8 tensor cores and NVIDIA's video encode and decode engines.
The A10 needs 150 W for the same 24 GB. The L40S needs 350 W for 48 GB. A GeForce RTX 4090 needs 450 W for 24 GB. Passive cooling and a low-profile single slot mean the L4 goes into chassis that have no room left for a taller card and no headroom left in the power budget.
Weights take roughly 2 GB per billion parameters at FP16, 1 GB at FP8 or INT8 and 0.5 GB at INT4, plus a tenth to a third again for the KV cache and activations. On 24 GB that puts Llama 3 8B, Mistral 7B and the smaller Qwen2.5 sizes comfortably inside the card at FP8, with room left for concurrency.
The L4 carries dedicated encode and decode engines alongside the tensor cores, which is what NVIDIA designed it around. A pipeline that decodes a stream, runs a detector or a captioning model over the frames and re-encodes the result runs end to end on one card instead of a GPU plus a transcode box.
Not the H100. The L4 sits next to the T4 it replaces and the A10 it is usually weighed against, with a GeForce RTX 4090 on the end for scale. Every tensor figure below is the dense number; none of them are with-sparsity figures.
figures from the NVIDIA L4, T4, A10 and GeForce RTX 4090 datasheets · no prices in this table
L4 supply on the per-minute marketplace is thin, and this page is not going to pretend otherwise. The volume route is bare metal, where the card is offered in three regions on fixed terms.
Which one you take depends on whether you want a card for an afternoon or a block of them for a quarter.
Search for L4 and see what hosts have listed today. Supply on this model changes, so the snapshot on this page is a reference point, not a promise.
Choose a container image for the listing, or bring your own, and the order starts billing per minute from the moment it runs.
If you need more than a listing or two, the configurator takes a GPU count from 8 upward and a term from 30 days upward and quotes the rate.
Marketplace orders accept Bitcoin, CLORE, USDT or USDC; bare-metal orders settle in Bitcoin, USDT or USDC.
On an on-demand order, usually yes. Partial GPU rental lets you take some of the GPUs on a server whose cards are all the same model, and the price scales with the share you take, so one card on a four-card L4 server costs a quarter of the server's rate. Spot orders still take the whole machine, and a host can switch partial rental off for a server, so check the listing before you plan around it.
Token generation reads the whole weight set once per token, so bandwidth sets the ceiling. Take 8 GB of FP8 weights: 300 GB/s divides into that about 37 times per second, 600 GB/s about 75. That is an arithmetic ceiling and nothing more, because real decoding also moves the KV cache and never reaches peak bandwidth, but the ratio is the honest shape of it. The A10 is the faster decoder of the two 24 GB cards. The L4 answers with less than half the board power, FP8 tensor cores the Ampere A10 does not have, and a single-slot card.
No. The video engines are separate silicon from the tensor cores, so a text-only endpoint neither uses them nor loses anything by ignoring them. They matter when the same rented box also ingests or re-encodes video, which is the workload NVIDIA built this card around. A pipeline that decodes a stream, runs a vision model over the frames and writes the result back fits on one L4 instead of a GPU plus a separate transcode machine.
No, and on a 24 GB card the absence costs less than it sounds. Partitioning exists to carve a large GPU into several smaller ones, which presumes memory to spare; an L4 holding an 8B model at FP8 has already spent a third of itself and behaves as a single-tenant device anyway. The parts that do have the feature are a short list: A100 and A30 on Ampere, H100 and H200 on Hopper, B200 and GB200 on Blackwell, and the RTX PRO Blackwell cards, per NVIDIA's supported-GPU list as checked on 21 September 2026. Multiple containers against one L4 remain fine, they are simply isolated by software rather than by silicon.
The hardware half is settled: the L4 carries fourth-generation Ada tensor cores with FP8, which is the main capability it has that the Ampere A10 does not. The software half moves faster than a static page can track, so check the serving guide on docs.clore.ai rather than a date-stamped claim here. What is safe to plan around is the memory arithmetic: at FP8 a model takes roughly 1 GB per billion parameters, so 8B weights sit near 8 GB and leave most of the 24 GB for the KV cache.
Llama 3 70B at any precision. Even at INT4 the weights land near 40 GB against 24 GB of GDDR6, so that model wants an L40S, an A40 or two cards. MIG partitioning does not fit either. The third category is anything latency-critical whose bottleneck is memory bandwidth, because 300 GB/s is the lowest figure of any 24 GB card on this site.
The figures below are arithmetic on the datasheet, not benchmark results: capacity from the model size, ceilings from the bandwidth. Treat them as the upper bound your own measurements will sit under.
At about 1 GB per billion parameters, an 8B model uses a third of the card. The rest is KV cache, which is what concurrency actually costs.
Read the guide →Every generated token reads the weights once. This is the hard ceiling before the KV cache and kernel overheads take their share.
Read the guide →Encode and decode run on dedicated blocks, so a transcode job and a vision model share a card without fighting over the tensor cores.
Read the guide →The four cards CLORE groups as inference parts, ranked by nothing. Dense FP16 tensor figures divided by board power, and the bare-metal rate each one carries. None of the four supports MIG.
the L4 does not lead this table on TFLOPS per watt. It leads on gigabytes per watt, and it is the only card here with FP8 under 150 W.
Walkthroughs on docs.clore.ai, chosen for workloads that fit inside 24 GB and do not need the bandwidth of a bigger card.
Supply them on a bare-metal contract and earn up to ~$350/mo per card, or list them per minute at a price you set and get paid in BTC, CLORE, USDT or USDC.
The marketplace shows what hosts have online at this moment, which is a different question from what was online on 21 September 2026. The bare-metal configurator quotes a term price for any count from eight upward.