Nothing on CLORE rents cheaper. Across the 13 T4 servers listed on 21 September 2026 the median was $0.018 per GPU-hour on demand, the lowest of any card on the platform and well under half the next one up. What you are buying is 2018 Turing: 16 GB of GDDR6 at 320 GB/s, 2,560 CUDA cores, 70 W in a single passive slot, and no FP8, no BF16, no MIG and no NVLink. For a long queue of small jobs that is an excellent trade. For anything that has to be fast, it is not.
The T4 wins on cost per finished item and loses on time to finish any single one. Every workload below is shaped the first way: many independent jobs, none of them urgent, none of them large.
At the 21 September 2026 snapshot the median T4 listing was $0.018 per GPU-hour. The next cheapest listed cards were an RTX A4000 at $0.049 and a Tesla V100 at $0.065, then an L4 at $0.083, an RTX 4070 at $0.104, an RTX 3090 at $0.155 and an RTX 3070 at $0.167. That gap is not a promotion; it is what a 70 W card with no modern data types is worth in a market that mostly wants them.
Turing has first-generation RT cores and 65 dense FP16 TFLOPS, and diffusion is compute-bound. SDXL at full resolution is slow enough here to be impractical, and cheap hours do not help when you need twenty of them. Send that work to a card with the arithmetic to finish it.
Sixteen gigabytes is enough to load Llama 3 8B at INT4, run a real CUDA test suite, or transcode through NVENC. At these rates a runner that actually has a GPU attached costs less per month than most people assume it does.
Every card here is picked for cost rather than for speed, and each gives up something different to get there. All TFLOPS figures are dense, never with sparsity.
specs from the NVIDIA T4, L4, V100 and GeForce datasheets · rates are the median per-GPU hourly price across CLORE listings on 21 Sep 2026
These are two different numbers and it matters which one you plan against. The medians below cover all 13 servers listed on 21 September 2026; 11 of them were already rented. Hosts set their own rates and the marketplace is the live source.
CLORE also sells T4 bare metal, from 8 GPUs on a 30-day minimum term: $0.40 per GPU-hour at 30 days, $0.34 at 90, $0.31 at 180 and $0.26 from 360, in the USA and Japan. That is more than ten times the per-minute median, and it buys capacity that is still there tomorrow. Configure bare metal →
Nothing about this card rewards ceremony. The only step worth thinking about is the first one, where you decide whether the job is queue-shaped.
Many small independent jobs with no one waiting: rent the T4. One large model or an interactive endpoint: rent something else and stop reading here.
Filter by Tesla T4 and check the price on the listing rather than the headline median, because the two differ on this card.
You get an endpoint, an SSH key and Jupyter on port 8888. Sixteen gigabytes goes fast, so size the model and the batch before you scale the queue.
A cheap hour is only cheap if the job finishes in it. Time one unit of work, multiply, and compare that against a faster card before you commit the batch.
It rules out every data type introduced after Turing, which means no BF16 and no FP8, and it rules out kernels compiled only for Ampere and newer, which includes FlashAttention. It also rules out MIG and NVLink, because the T4 has neither. What still works is the large body of inference that only ever needed FP16 and INT8: speech recognition, object detection, embedding and reranking, video transcode through NVENC, and small language models quantised down to four bits. Eight years is old for a GPU, but none of those workloads have moved on.
FP32, FP16 and INT8 in hardware, plus whatever INT4 scheme your runtime implements on top of INT8 arithmetic. In practice that means GGUF and other weight-quantised formats work, FP16 checkpoints load directly, and anything published in BF16 has to be cast on the way in, which is usually fine for inference and unreliable for training. FP8 checkpoints are simply not executable here. The narrower FP16 exponent range is also why nobody should be training on this card.
Work you would otherwise run on a CPU, in volume, where the job is embarrassingly parallel and latency does not matter: batch transcription, bulk image classification and detection, embedding a corpus for retrieval, and hardware video transcoding. It is also the cheapest way to keep a GPU attached to a continuous-integration runner so that CUDA tests actually execute. The unifying trait is that the model is small and the queue is long.
For throughput-shaped work, often yes, because the cards pack densely, rent at the lowest rate on the platform, and a queue of independent jobs scales linearly across them. For anything that has to be fast rather than merely finished, no, because splitting one model across many small cards costs you interconnect the T4 does not have. The honest rule is that many T4s beat one big card when the work is a queue, and lose badly when the work is a single large model.
When the job is compute-bound and you are paying for wall-clock time rather than for a result. Diffusion is the clearest case: SDXL at full resolution is slow enough on this card to be impractical, and a job that takes many times longer at a fraction of the price can easily be the worse deal rather than the better one. The same applies to anything interactive, where a user is waiting. Multiply the rate by the hours you expect, not by the hours a faster card would need.
More, and this is the part a median hides. Across all 13 servers listed on 21 September 2026 the median was $0.018 per GPU-hour on demand and the median spot bid was $0.014, with the cheapest seen at $0.015 on demand and a $0.014 spot bid. But only 2 servers were unrented, and those were priced at a $0.094 median on demand, with a $0.083 median spot bid. Treat the low figures as what the market clears at and the higher ones as what you can actually book today. Live prices are always on the marketplace.
Memory figures are weights only, before activations. Throughput depends on your batch, your input length and your runtime, so time one unit of work on the card before you size the job.
A backlog of audio, a card that costs almost nothing per hour, and nobody waiting on any single file. This is the shape the T4 was built for.
Read the guide →Turing added INT8 tensor cores, which is the one modern-ish feature this card does have. Small vision models quantise cleanly and run well inside 16 GB.
Read the guide →Create an order, run one batch, destroy it. At the cheapest rate on the platform, a scripted box that lives for ten minutes costs a rounding error.
Read the guide →Listing counts and medians are the CLORE marketplace on 21 September 2026. None of these four cards supports MIG. Bare-metal rates are per GPU-hour and move with contract length, from 8 GPUs and 30 days upward.
Small models, quantised weights and batch pipelines. Anything here that mentions FP8 or BF16 will need an INT8 path on this card.
A block of eight on a bare-metal contract earns up to ~$270/mo per card, and a single T4 can be listed by the minute at whatever rate you choose.
Time one unit of work on a T4, multiply it out, and compare. If the answer still favours the cheapest card on the platform, the listings are one click away.