NVIDIA Tesla K80 - 12GB usable
Avoid for modern local AI: its advertised 24GB is split across two independent 12GB GPUs.
Specifications
| Brand | NVIDIA |
|---|---|
| Model | Tesla K80 |
| Usable VRAM | 12GB (24GB split across the board) |
| Architecture | Kepler |
| CUDA / Stream Processors | 4,992 |
| Memory Bandwidth | 480 GB/s |
| TDP | 300W |
| FP32 TFLOPS | 8.7 |
Current Offers
Used from €199
Prices last updated:
GPUDojo is reader-supported. When you buy through links on our site, we may earn an affiliate commission.
Price History
- eBay€199at high
For AI / LLM Use
Entry-level for local AI. Handles 7B-8B models well. Older architecture may have limited software support (check CUDA compatibility). Datacenter card with no display output, may need aftermarket cooling.
What Models Can It Run?
- 14B Q4_K_M, 7B full precision
- 7B Q6_K, 14B Q3_K (tight)
- 7B Q4_K_M only
Estimated Performance
Generation: ~36 tokens/sec
Prefill: ~155 tokens/sec
Recommended Quantisations
- Q4_K_M for 14B (tight fit)
- Q6_K or Q8 for 7B models
Pros & Cons
Pros
- Very low acquisition cost
Cons
- 12GB usable VRAM: may need quantization for 30B+ models
- 24GB is split across multiple GPUs; one model sees 12GB per GPU
- Moderate memory bandwidth: not the fastest for inference
- 300W TDP: high power draw
- Older Kepler architecture: verify current CUDA support
- No display output: headless only
- May need aftermarket cooling solution
- Only 12GB is usable by one model without explicit sharding
- Kepler software support is obsolete for many current stacks
- Dual-GPU 300W board with substantial cooling requirements
Community Verdict
- r/LocalLLaMA
Avoid. Dual-GPU means 12GB per die, Kepler lacks modern CUDA support, and it draws 300W. Buy a P40 instead.
Source