NVIDIA A30 - 24GB usable

24GB of HBM2 at 933 GB/s, so it generates faster than its FP32 figure suggests. Passive and 165W: a server part, not a drop-in.

Specifications

BrandNVIDIA
ModelA30
Usable VRAM24GB
ArchitectureAmpere
CUDA / Stream Processors3,584
Memory Bandwidth933 GB/s
TDP165W
FP32 TFLOPS10.3

Current Offers

Used from $2599New from $3890

Prices last updated:

GPUDojo is reader-supported. When you buy through links on our site, we may earn an affiliate commission.

Price History

  • eBay$2599mid-range
  • Amazon$3890current
  • Newegg$3944at high
Jul 31Aug 7Aug 14Aug 28Sep 1Sep 4$2799$3944$2599$2599$3890

For AI / LLM Use

Solid choice for 30B models and comfortable 14B inference. Datacenter card with no display output, may need aftermarket cooling.

What Models Can It Run?

  • 30B Q4_K_M, 14B full precision, 70B Q2 (tight)
  • 14B Q6_K, 30B Q3_K (tight)
  • 14B Q4_K_M, 7B full precision
  • 7B Q6_K, 14B Q3_K (tight)
  • 7B Q4_K_M only

Estimated Performance

Generation: ~70 tokens/sec

Prefill: ~184 tokens/sec

Recommended Quantisations

  • Q4_K_M recommended for 30B models
  • Q6_K or Q8 for 14B and below
  • Full precision for 7B

Pros & Cons

Pros

  • 24GB usable VRAM: handles large models
  • High memory bandwidth for fast generation
  • Ampere architecture: broad CUDA software support

Cons

  • No display output: headless only
  • May need aftermarket cooling solution

Community Verdict

No community reviews yet for the A30. Know a good review? Let us know.