NVIDIA Tesla P4 - 8GB usable

75W, half-height, 8GB, and it runs off slot power alone. Pascal at 192 GB/s: an inference sidecar, not a main GPU.

Specifications

BrandNVIDIA
ModelTesla P4
Usable VRAM8GB
ArchitecturePascal
CUDA / Stream Processors2,560
Memory Bandwidth192 GB/s
TDP75W
FP32 TFLOPS5.5

Current Offers

Used from 189

Prices last updated:

GPUDojo is reader-supported. When you buy through links on our site, we may earn an affiliate commission.

Price History

  • eBay189mid-range
Aug 21Aug 28Sep 1Sep 4189189

For AI / LLM Use

Limited VRAM restricts you to 7B quantized models. Slower generation, usable but not snappy. Older architecture may have limited software support (check CUDA compatibility). Datacenter card with no display output, may need aftermarket cooling.

What Models Can It Run?

  • 7B Q6_K, 14B Q3_K (tight)
  • 7B Q4_K_M only

Estimated Performance

Generation: ~14 tokens/sec

Prefill: ~98 tokens/sec

Recommended Quantisations

  • Q4_K_M for 7B models
  • Q3_K for larger experiments

Pros & Cons

Pros

  • Only 75W TDP: power efficient

Cons

  • Only 8GB usable VRAM: limited to small models
  • Low memory bandwidth: slower token generation
  • Older Pascal architecture: verify current CUDA support
  • No display output: headless only
  • May need aftermarket cooling solution

Community Verdict

No community reviews yet for the Tesla P4. Know a good review? Let us know.