NVIDIA Tesla P4 - 8GB usable
75W, half-height, 8GB, and it runs off slot power alone. Pascal at 192 GB/s: an inference sidecar, not a main GPU.
Specifications
| Brand | NVIDIA |
|---|---|
| Model | Tesla P4 |
| Usable VRAM | 8GB |
| Architecture | Pascal |
| CUDA / Stream Processors | 2,560 |
| Memory Bandwidth | 192 GB/s |
| TDP | 75W |
| FP32 TFLOPS | 5.5 |
Current Offers
Used from £117
Prices last updated:
GPUDojo is reader-supported. When you buy through links on our site, we may earn an affiliate commission.
Price History
- eBay£117at high
For AI / LLM Use
Limited VRAM restricts you to 7B quantized models. Slower generation, usable but not snappy. Older architecture may have limited software support (check CUDA compatibility). Datacenter card with no display output, may need aftermarket cooling.
What Models Can It Run?
- 7B Q6_K, 14B Q3_K (tight)
- 7B Q4_K_M only
Estimated Performance
Generation: ~14 tokens/sec
Prefill: ~98 tokens/sec
Recommended Quantisations
- Q4_K_M for 7B models
- Q3_K for larger experiments
Pros & Cons
Pros
- Only 75W TDP: power efficient
Cons
- Only 8GB usable VRAM: limited to small models
- Low memory bandwidth: slower token generation
- Older Pascal architecture: verify current CUDA support
- No display output: headless only
- May need aftermarket cooling solution
Community Verdict
No community reviews yet for the Tesla P4. Know a good review? Let us know.