NVIDIA A40 - 48GB usable
48GB of Ampere in a passive shell, an A6000 without the blower. Needs server front-to-back airflow, not a desktop case.
Specifications
| Brand | NVIDIA |
|---|---|
| Model | A40 |
| Usable VRAM | 48GB |
| Architecture | Ampere |
| CUDA / Stream Processors | 10,752 |
| Memory Bandwidth | 696 GB/s |
| TDP | 300W |
| FP32 TFLOPS | 37.4 |
Current Offers
GPUDojo is reader-supported. When you buy through links on our site, we may earn an affiliate commission.
Price History
Price tracking started: chart will appear after the next snapshot.
For AI / LLM Use
Best for running 70B+ models on a single GPU. Datacenter card with no display output, may need aftermarket cooling.
What Models Can It Run?
- 70B Q4_K_M, 30B full precision
- 30B Q6_K, 70B Q2_K
- 30B Q4_K_M, 14B full precision, 70B Q2 (tight)
- 14B Q6_K, 30B Q3_K (tight)
- 14B Q4_K_M, 7B full precision
- 7B Q6_K, 14B Q3_K (tight)
- 7B Q4_K_M only
Estimated Performance
Generation: ~52 tokens/sec
Prefill: ~668 tokens/sec
Recommended Quantisations
- Q6_K or Q8 for best quality
- Q4_K_M for larger models
- Full precision for 30B and below
Pros & Cons
Pros
- 48GB usable VRAM: handles large models
- Ampere architecture: broad CUDA software support
Cons
- Moderate memory bandwidth: not the fastest for inference
- 300W TDP: high power draw
- No display output: headless only
- May need aftermarket cooling solution
Community Verdict
No community reviews yet for the A40. Know a good review? Let us know.