NVIDIA T4 - 16GB usable

70W, single-slot, 16GB, and it runs off slot power in almost any server. Only 320 GB/s, so it favours always-on inference over speed.

Specifications

BrandNVIDIA
ModelT4
Usable VRAM16GB
ArchitectureTuring
CUDA / Stream Processors2,560
Memory Bandwidth320 GB/s
TDP70W
FP32 TFLOPS8.1

Current Offers

Used from 690

Prices last updated:

GPUDojo is reader-supported. When you buy through links on our site, we may earn an affiliate commission.

Price History

  • eBay690at high
Mar 13Apr 17May 22Jul 3Aug 7Sep 4600690

For AI / LLM Use

Good for 14B models. 30B requires aggressive quantization. Slower generation, usable but not snappy. Datacenter card with no display output, may need aftermarket cooling.

What Models Can It Run?

  • 14B Q6_K, 30B Q3_K (tight)
  • 14B Q4_K_M, 7B full precision
  • 7B Q6_K, 14B Q3_K (tight)
  • 7B Q4_K_M only

Estimated Performance

Generation: ~24 tokens/sec

Prefill: ~145 tokens/sec

Recommended Quantisations

  • Q4_K_M for 14B models
  • Q6_K for 7B-8B models
  • Q8 for 7B if VRAM allows

Pros & Cons

Pros

  • Only 70W TDP: power efficient
  • Turing architecture: broad CUDA software support

Cons

  • 16GB usable VRAM: may need quantization for 30B+ models
  • Moderate memory bandwidth: not the fastest for inference
  • No display output: headless only
  • May need aftermarket cooling solution

Community Verdict

No community reviews yet for the T4. Know a good review? Let us know.