Log in T4 listings
Tesla T4 · 16 GB GDDR6 · 70 W 13 servers listed, as of 21 Sep 2026

The Tesla T4
is the price floor.
And a modern GPU.

Nothing on CLORE rents cheaper. Across the 13 T4 servers listed on 21 September 2026 the median was $0.018 per GPU-hour on demand, the lowest of any card on the platform and well under half the next one up. What you are buying is 2018 Turing: 16 GB of GDDR6 at 320 GB/s, 2,560 CUDA cores, 70 W in a single passive slot, and no FP8, no BF16, no MIG and no NVLink. For a long queue of small jobs that is an excellent trade. For anything that has to be fast, it is not.

●Billed per minute ●SSH, Docker and Jupyter ●Turing, FP16 and INT8 only ●Bare metal from 8 GPUs
MARKETPLACE SNAPSHOT Tesla T4 · per-minute listings 21 SEP 2026
SERVERS LISTED
13
▬ 43 cards between them
UNRENTED
2 of 13
▬ at that instant
FREE CARDS
8 of 43
▬ on those 2 servers
SERVERS RENTED 11 / 13
CARDS FREE 8 / 43
SPOT BID VS ON-DEMAND $0.014 / $0.018
MEDIAN ON-DEMAND
$0.018/GPU-h
MEDIAN SPOT BID
$0.014/GPU-h
UNRENTED MEDIAN
$0.094/GPU-h
BARE METAL
$0.26–$0.40/GPU-h
$0.018/GPU-h
Median on-demand across 13 listings, 21 Sep 2026
16GB
GDDR6 at 320 GB/s per card
70W
Single-slot, passively cooled
2018
Turing TU104, launched September that year

Long queues,
small models.

The T4 wins on cost per finished item and loses on time to finish any single one. Every workload below is shaped the first way: many independent jobs, none of them urgent, none of them large.

Cheaper than anything else listed

At the 21 September 2026 snapshot the median T4 listing was $0.018 per GPU-hour. The next cheapest listed cards were an RTX A4000 at $0.049 and a Tesla V100 at $0.065, then an L4 at $0.083, an RTX 4070 at $0.104, an RTX 3090 at $0.155 and an RTX 3070 at $0.167. That gap is not a promotion; it is what a 70 W card with no modern data types is worth in a market that mostly wants them.

Median rate against an A4000 $0.018 vs $0.049

Not image or video generation

Turing has first-generation RT cores and 65 dense FP16 TFLOPS, and diffusion is compute-bound. SDXL at full resolution is slow enough here to be impractical, and cheap hours do not help when you need twenty of them. Send that work to a card with the arithmetic to finish it.

FP16 dense 65 TFLOPS

A GPU on your CI runner

Sixteen gigabytes is enough to load Llama 3 8B at INT4, run a real CUDA test suite, or transcode through NVENC. At these rates a runner that actually has a GPU attached costs less per month than most people assume it does.

Llama 3 8B at INT4 ~4 GB of 16

Four cheap cards,
four different compromises.

Every card here is picked for cost rather than for speed, and each gives up something different to get there. All TFLOPS figures are dense, never with sparsity.

Tesla T4 NVIDIA L4 RTX 3070 Tesla V100
Architecture Turing TU104 Ada AD104 Ampere GA104 Volta GV100
CUDA cores 2,560 7,424 5,888 5,120
VRAM 16 GB GDDR6 24 GB GDDR6 8 GB GDDR6 32 GB HBM2
Memory bandwidth 320 GB/s 300 GB/s 448 GB/s 900 GB/s
FP16 / BF16 (dense) 65 TFLOPS 61 TFLOPS 81 TFLOPS 125 TFLOPS
BF16 / FP8 no / no yes / yes yes / no no / no
Board power 70 W 72 W 220 W 300 W
Median on-demand 21 Sep 2026 $0.018 /GPU-h $0.083 $0.167 $0.065

specs from the NVIDIA T4, L4, V100 and GeForce datasheets · rates are the median per-GPU hourly price across CLORE listings on 21 Sep 2026

The median and
the bookable price.

These are two different numbers and it matters which one you plan against. The medians below cover all 13 servers listed on 21 September 2026; 11 of them were already rented. Hosts set their own rates and the marketplace is the live source.

Spot

$0.014 / GPU-hour
median spot bid across 13 listings · unrented ones sat at a $0.083 median bid
  • The cheapest rate on the platform
  • Billed per minute
  • An on-demand renter can take the machine back
  • Suits a queue you can restart from anywhere
Browse spot listings
NOT PREEMPTIBLE

On-demand

$0.018 / GPU-hour
median of 13 listings · unrented ones sat at a $0.094 median
  • Yours until you release it
  • No preemption
  • Two servers were free at the snapshot, and both were priced well above the median
  • Billed per minute
Rent on demand

CLORE also sells T4 bare metal, from 8 GPUs on a 30-day minimum term: $0.40 per GPU-hour at 30 days, $0.34 at 90, $0.31 at 180 and $0.26 from 360, in the USA and Japan. That is more than ten times the per-minute median, and it buys capacity that is still there tomorrow. Configure bare metal →

Pay with
Bitcoin on-chain
CLORE native token
USDT / USDC ERC-20 · BEP-20

Renting the floor.

Nothing about this card rewards ceremony. The only step worth thinking about is the first one, where you decide whether the job is queue-shaped.

01 / CHECK

Is this a queue or a deadline?

Many small independent jobs with no one waiting: rent the T4. One large model or an interactive endpoint: rent something else and stop reading here.

02 / RENT

Filter and take one

Filter by Tesla T4 and check the price on the listing rather than the headline median, because the two differ on this card.

marketplace filter → Tesla T4
03 / CONNECT

SSH or Jupyter

You get an endpoint, an SSH key and Jupyter on port 8888. Sixteen gigabytes goes fast, so size the model and the batch before you scale the queue.

04 / MEASURE

Cost per item, not per hour

A cheap hour is only cheap if the job finishes in it. Time one unit of work, multiply, and compare that against a faster card before you commit the batch.

Questions about an eight-year-old card.

The T4 is 2018 Turing. What does that rule out in 2026, and what still works fine?

It rules out every data type introduced after Turing, which means no BF16 and no FP8, and it rules out kernels compiled only for Ampere and newer, which includes FlashAttention. It also rules out MIG and NVLink, because the T4 has neither. What still works is the large body of inference that only ever needed FP16 and INT8: speech recognition, object detection, embedding and reranking, video transcode through NVENC, and small language models quantised down to four bits. Eight years is old for a GPU, but none of those workloads have moved on.

No BF16 and no FP8 on a T4. Which model formats does that leave me?

FP32, FP16 and INT8 in hardware, plus whatever INT4 scheme your runtime implements on top of INT8 arithmetic. In practice that means GGUF and other weight-quantised formats work, FP16 checkpoints load directly, and anything published in BF16 has to be cast on the way in, which is usually fine for inference and unreliable for training. FP8 checkpoints are simply not executable here. The narrower FP16 exponent range is also why nobody should be training on this card.

At a $0.018 median per GPU-hour, what is the T4 genuinely the cheapest way to do?

Work you would otherwise run on a CPU, in volume, where the job is embarrassingly parallel and latency does not matter: batch transcription, bulk image classification and detection, embedding a corpus for retrieval, and hardware video transcoding. It is also the cheapest way to keep a GPU attached to a continuous-integration runner so that CUDA tests actually execute. The unifying trait is that the model is small and the queue is long.

Single-slot and passive. Do T4 fleets beat fewer, bigger cards per job?

For throughput-shaped work, often yes, because the cards pack densely, rent at the lowest rate on the platform, and a queue of independent jobs scales linearly across them. For anything that has to be fast rather than merely finished, no, because splitting one model across many small cards costs you interconnect the T4 does not have. The honest rule is that many T4s beat one big card when the work is a queue, and lose badly when the work is a single large model.

When do 2,560 CUDA cores make a job slower than it is cheap?

When the job is compute-bound and you are paying for wall-clock time rather than for a result. Diffusion is the clearest case: SDXL at full resolution is slow enough on this card to be impractical, and a job that takes many times longer at a fraction of the price can easily be the worse deal rather than the better one. The same applies to anything interactive, where a user is waiting. Multiply the rate by the hours you expect, not by the hours a faster card would need.

The cheapest T4 listings are all rented. What do the free ones cost?

More, and this is the part a median hides. Across all 13 servers listed on 21 September 2026 the median was $0.018 per GPU-hour on demand and the median spot bid was $0.014, with the cheapest seen at $0.015 on demand and a $0.014 spot bid. But only 2 servers were unrented, and those were priced at a $0.094 median on demand, with a $0.083 median spot bid. Treat the low figures as what the market clears at and the higher ones as what you can actually book today. Live prices are always on the marketplace.

Three queues worth a T4.

Memory figures are weights only, before activations. Throughput depends on your batch, your input length and your runtime, so time one unit of work on the card before you size the job.

Offline transcription
Whisper, FP16 or INT8 weights
16 GB, no urgency

A backlog of audio, a card that costs almost nothing per hour, and nobody waiting on any single file. This is the shape the T4 was built for.

Read the guide →
Bulk detection and classification
YOLOv8 with TensorRT, INT8
INT8 in hardware

Turing added INT8 tensor cores, which is the one modern-ish feature this card does have. Small vision models quantise cleanly and run well inside 16 GB.

Read the guide →
A GPU behind an API call
Programmatic orders, short-lived boxes
Billed by the minute

Create an order, run one batch, destroy it. At the cheapest rate on the platform, a scripted box that lives for ten minutes costs a rounding error.

Read the guide →

What the low-power cards cost to rent.

Listing counts and medians are the CLORE marketplace on 21 September 2026. None of these four cards supports MIG. Bare-metal rates are per GPU-hour and move with contract length, from 8 GPUs and 30 days upward.

GPU
VRAM
TDP (W)
Mem BW (GB/s)
FP8
Servers listed
Median on-demand
Bare metal $/GPU-h
Tesla T4 / this page
16 GB GDDR6
70
320
—
13
$0.018
$0.26–$0.40
NVIDIA L4
24 GB GDDR6
72
300
yes
1
$0.083
$0.41–$0.52
NVIDIA A10
24 GB GDDR6
150
600
—
0
—
$0.43–$0.64
RTX 3070
8 GB GDDR6
220
448
—
291
$0.167
not offered

Guides that fit in 16 GB.

Small models, quantised weights and batch pipelines. Anything here that mentions FP8 or BF16 will need an INT8 path on this card.

Language Models
Ollama on CLORE.AI
One-command LLM inference for Llama, Mistral, Phi.
Audio Voice
Whisper transcription
OpenAI Whisper-large for speech-to-text.
Computer Vision
YOLOv8 detection
Real-time object detection with YOLOv8.
Video Processing
FFmpeg + NVENC
Hardware-accelerated video transcoding.
Vision Models
Florence-2
Microsoft's compact vision-language model.
Language Models
Phi-4
Microsoft Phi-4 small-but-capable model.
Advanced
CLORE API integration
Programmatic order creation via the public API.
See all guides →

When the floor is too low.

NVIDIA L4
24 GB and FP8, still at 72 W
Compare →
Tesla V100
32 GB of HBM2 at 900 GB/s
Compare →
RTX 4090
For the diffusion jobs a T4 cannot finish
Compare →

Cheap hours only help
if the job finishes.

Time one unit of work on a T4, multiply it out, and compare. If the answer still favours the cheapest card on the platform, the listings are one click away.