NVIDIA A40 - 48GB usable

48GB of Ampere in a passive shell, an A6000 without the blower. Needs server front-to-back airflow, not a desktop case.

Specifications

BrandNVIDIA
ModelA40
Usable VRAM48GB
ArchitectureAmpere
CUDA / Stream Processors10,752
Memory Bandwidth696 GB/s
TDP300W
FP32 TFLOPS37.4

Current Offers

Used from $4600New from $5699

Prices last updated:

GPUDojo is reader-supported. When you buy through links on our site, we may earn an affiliate commission.

Price History

  • eBay$4600at high
  • Amazon$5699current
  • Newegg$8000mid-range
Jul 31Aug 7Aug 14Aug 28Sep 1Sep 4$8000$8000$4600$4600$5699

For AI / LLM Use

Best for running 70B+ models on a single GPU. Datacenter card with no display output, may need aftermarket cooling.

What Models Can It Run?

  • 70B Q4_K_M, 30B full precision
  • 30B Q6_K, 70B Q2_K
  • 30B Q4_K_M, 14B full precision, 70B Q2 (tight)
  • 14B Q6_K, 30B Q3_K (tight)
  • 14B Q4_K_M, 7B full precision
  • 7B Q6_K, 14B Q3_K (tight)
  • 7B Q4_K_M only

Estimated Performance

Generation: ~52 tokens/sec

Prefill: ~668 tokens/sec

Recommended Quantisations

  • Q6_K or Q8 for best quality
  • Q4_K_M for larger models
  • Full precision for 30B and below

Pros & Cons

Pros

  • 48GB usable VRAM: handles large models
  • Ampere architecture: broad CUDA software support

Cons

  • Moderate memory bandwidth: not the fastest for inference
  • 300W TDP: high power draw
  • No display output: headless only
  • May need aftermarket cooling solution

Community Verdict

No community reviews yet for the A40. Know a good review? Let us know.