NVIDIA Tesla P40 - 24GB usable

The best low-cost route to 24GB if you can solve passive cooling and do not need display output.

Specifications

BrandNVIDIA
ModelTesla P40
Usable VRAM24GB
ArchitecturePascal
CUDA / Stream Processors3,840
Memory Bandwidth347 GB/s
TDP250W
FP32 TFLOPS12

Current Offers

Used from £675

Prices last updated:

GPUDojo is reader-supported. When you buy through links on our site, we may earn an affiliate commission.

Buying Guidance

Prioritise tested cards from established sellers and budget for a fan duct, power adapter and adequate case airflow. Compare the delivered cost, not the bare-card price.

Software: Pascal works well with llama.cpp and many CUDA inference stacks, but newer kernels and features increasingly favour Turing or later.

Price History

  • eBay£675at high
Aug 21Aug 28Sep 1Sep 4£219£675

For AI / LLM Use

Solid choice for 30B models and comfortable 14B inference. Slower generation, usable but not snappy. Older architecture may have limited software support (check CUDA compatibility). Datacenter card with no display output, may need aftermarket cooling.

What Models Can It Run?

  • 30B Q4_K_M, 14B full precision, 70B Q2 (tight)
  • 14B Q6_K, 30B Q3_K (tight)
  • 14B Q4_K_M, 7B full precision
  • 7B Q6_K, 14B Q3_K (tight)
  • 7B Q4_K_M only

Estimated Performance

Generation: ~26 tokens/sec

Prefill: ~214 tokens/sec

Recommended Quantisations

  • Q4_K_M recommended for 30B models
  • Q6_K or Q8 for 14B and below
  • Full precision for 7B

Pros & Cons

Pros

  • 24GB usable VRAM: handles large models
  • 24GB of genuinely usable VRAM on one GPU
  • Large community of local-AI cooling and setup guides

Cons

  • Moderate memory bandwidth: not the fastest for inference
  • Older Pascal architecture: verify current CUDA support
  • No display output: headless only
  • May need aftermarket cooling solution
  • Passive server cooling requires a deliberate desktop modification
  • No display output and no native fast FP16 path

Community Verdict

  • r/LocalLLaMA

    Budget legend. Needs a blower cooler mod and no display output, but 24GB at a low used price remains hard to beat.

    Source