NVIDIA Tesla P100 - 16GB usable
Fast HBM2 bandwidth on a budget, but 16GB capacity and Pascal-era support make it a specialist choice.
Specifications
| Brand | NVIDIA |
|---|---|
| Model | Tesla P100 |
| Usable VRAM | 16GB |
| Architecture | Pascal |
| CUDA / Stream Processors | 3,584 |
| Memory Bandwidth | 732 GB/s |
| TDP | 250W |
| FP32 TFLOPS | 9.3 |
Current Offers
Used from €149
Prices last updated:
GPUDojo is reader-supported. When you buy through links on our site, we may earn an affiliate commission.
Price History
- eBay€149at high
For AI / LLM Use
Good for 14B models. 30B requires aggressive quantization. Older architecture may have limited software support (check CUDA compatibility). Datacenter card with no display output, may need aftermarket cooling.
What Models Can It Run?
- 14B Q6_K, 30B Q3_K (tight)
- 14B Q4_K_M, 7B full precision
- 7B Q6_K, 14B Q3_K (tight)
- 7B Q4_K_M only
Estimated Performance
Generation: ~55 tokens/sec
Prefill: ~166 tokens/sec
Recommended Quantisations
- Q4_K_M for 14B models
- Q6_K for 7B-8B models
- Q8 for 7B if VRAM allows
Pros & Cons
Pros
- High-bandwidth HBM2 memory for its used price
- Often faster than a P40 when the workload fits in 16GB
Cons
- 16GB usable VRAM: may need quantization for 30B+ models
- Moderate memory bandwidth: not the fastest for inference
- Older Pascal architecture: verify current CUDA support
- No display output: headless only
- May need aftermarket cooling solution
- Only 16GB, so many 30B-class configurations will not fit
- Passive cooling and no display output
Community Verdict
- r/LocalLLaMA
16GB HBM2 on a budget. Pascal architecture limits software support but handles 14B Q4 models.
Source