NVIDIA L4 - 24GB usable

Specifications

BrandNVIDIA
ModelL4
Usable VRAM24GB
ArchitectureAda
CUDA / Stream Processors7,424
Memory Bandwidth300 GB/s
TDP72W
FP32 TFLOPS30

Current Offers

Used from $3699New from $3395

Prices last updated:

GPUDojo is reader-supported. When you buy through links on our site, we may earn an affiliate commission.

Price History

  • eBay$3699at high
  • Amazon$3395current
  • Newegg$3599at high
Aug 21Aug 28Sep 1Sep 4$2750$3599$3500$3699$3395

For AI / LLM Use

Solid choice for 30B models and comfortable 14B inference. Slower generation, usable but not snappy. Datacenter card with no display output, may need aftermarket cooling.

What Models Can It Run?

  • 30B Q4_K_M, 14B full precision, 70B Q2 (tight)
  • 14B Q6_K, 30B Q3_K (tight)
  • 14B Q4_K_M, 7B full precision
  • 7B Q6_K, 14B Q3_K (tight)
  • 7B Q4_K_M only

Estimated Performance

Generation: ~23 tokens/sec

Prefill: ~536 tokens/sec

Recommended Quantisations

  • Q4_K_M recommended for 30B models
  • Q6_K or Q8 for 14B and below
  • Full precision for 7B

Pros & Cons

Pros

  • 24GB usable VRAM: handles large models
  • Only 72W TDP: power efficient
  • Ada architecture: broad CUDA software support

Cons

  • Moderate memory bandwidth: not the fastest for inference
  • No display output: headless only
  • May need aftermarket cooling solution

Community Verdict

No community reviews yet for the L4. Know a good review? Let us know.