Tesla M40 24GB in 2026: Still Worth Buying for AI?

The Tesla M40 24GB used to have one clear argument: it was the cheapest way to get 24GB of VRAM, at $70-90 on eBay. That argument has weakened considerably. The M40 now starts around $229, while the 16GB Tesla P100 sits near $80 with two and a half times the memory bandwidth, and the Tesla P40, a genuinely better card in every respect, is only $10 more.

That does not make the M40 useless. It makes it a narrower proposition than it was: 24GB of capacity, on 2015 Maxwell silicon, with no native FP16 and a passive cooler you will have to solve. Whether that is worth $229 depends entirely on whether you need the capacity more than the speed. Here are the numbers to decide with.

Tesla M40 24GB: Key Specs
GPU ArchitectureMaxwell (GM200) - 2015
CUDA Cores3,072
VRAM24GB GDDR5
Memory Bandwidth288 GB/s
FP32 Performance7.0 TFLOPS
FP16 PerformanceNone (no native FP16)
TDP250W
CoolingPassive (requires server airflow)
Compute Capability5.2
Lowest Used Price$229 (live, September 2026)

The Case For the M40

The M40's advantage is capacity: 24GB of VRAM for $229, on a card that modern software still supports. This lets you:

The key word is capacity. If you need to fit a model larger than 16GB and cannot stretch to a P40 at $239, the M40 is the cheapest card that will hold it. If your models fit in 16GB, the P100 at $80 is both cheaper and dramatically faster: see the comparison below.

Pros

  • 24GB of VRAM at $229, the cheapest card that holds a 20B+ model
  • Full CUDA 12 support (compute capability 5.2)
  • Modern PyTorch/llama.cpp works fine
  • 24GB unified VRAM (not split like K80)
  • Abundant supply on eBay
  • Can run 20B+ models at Q4

Cons

  • Slow - 288 GB/s bandwidth limits tok/s
  • No FP16 support (no native half-precision)
  • Passive cooling only - overheats in desktop
  • 250W power draw
  • No display output
  • No video encode/decode

Real-World Performance

Here's what you can realistically expect running models on the M40 with llama.cpp:

Tesla M40 24GB - Estimated tok/s (llama.cpp, Q4_K_M)

Llama 3 8B (Q4) ~30-35 tok/s
Llama 3 8B (Q8) ~18-22 tok/s
Llama 3 14B (Q4) ~18-22 tok/s
Qwen 2.5 32B (Q4) ~10-14 tok/s
Llama 3 8B (FP16) ~15-18 tok/s
Deepseek Coder 33B (Q3) ~8-10 tok/s

These speeds are usable but not fast. For comparison, a P40 is about 20-40% faster across the board, and an RTX 3090 is 3-4x faster. If you're accustomed to ChatGPT's response speed, the M40 will feel noticeably slower for larger models.

The Cooling Problem

This Is the M40's Biggest Challenge

The Tesla M40, like all data center GPUs of its era, has no fans. It's a completely passive card designed for server chassis with front-to-back forced airflow. In a standard desktop PC, it will thermal throttle within minutes and may shut down to prevent damage.

You must add aftermarket cooling. Options include:

Temperature Targets

  • Under 85C: Safe for sustained operation
  • 85-95C: Thermal throttling begins, performance degrades
  • Above 95C: Risk of shutdown, potential long-term damage

Monitor with nvidia-smi -l 1 during first use to check temperatures under load.

M40 vs P40: Is the Extra $10 Worth It?

Spec M40 24GB ($229) P40 24GB ($239) Difference
Architecture Maxwell Pascal 1 generation newer
Bandwidth 288 GB/s 347 GB/s +20%
VRAM 24GB GDDR5 24GB GDDR5X Same capacity, faster memory
INT8 Inference No Yes P40 has INT8 support
tok/s (8B Q4) ~32 ~45 ~40% faster
tok/s (32B Q4) ~12 ~16 ~33% faster
Power 250W 250W Same
Cooling Passive Passive Both need aftermarket
Software support Good Better P40 has wider kernel support

The P40 is 20-40% faster across the board, adds INT8 inference, and has wider kernel support. Whether $10 buys enough of that to matter is a judgement about your own patience. The M40 is not unusable, it is just slower at everything.

M40 vs P100: The Comparison That Changed

This is the one worth thinking hardest about, because it inverted. When this article was first written the M40 was $70-90 and the P100 cost more. Today the M40 starts at $229 and the P100 starts at $80.

Spec M40 24GB ($229) P100 16GB ($80) What it means
VRAM 24GB GDDR5 16GB HBM2 M40 fits larger models
Bandwidth 288 GB/s 732 GB/s P100 is ~2.5x faster at generation
Native FP16 No Yes Matters for training and some runtimes
Architecture Maxwell (2015) Pascal (2016) Both CUDA 12 capable
Cooling Passive Passive Both need aftermarket

The decision is capacity against speed, and the cheaper card is now the faster one. If your models fit inside 16GB, the P100 is faster and costs less, so there is no argument for the M40. If you need more than 16GB, the M40 is still the cheapest card that holds it, and 24GB of slow VRAM beats 16GB you cannot fit into. Work out your model size first; the answer follows from it.

Verdict

A capacity card, not a value card

The Tesla M40 24GB at $229 is a functioning 24GB GPU that modern software still supports. It runs 20B+ parameter models, and for batch work where you set a job going and come back later, its slowness costs you little.

What it no longer is, is the obvious budget pick. At $80 the P100 is cheaper and far faster if 16GB is enough for you, and at $239 the P40 gives you the same 24GB with real speed. The M40 now occupies a narrow gap: you need more than 16GB, and you cannot or will not pay $239. That is a real position to be in, but make sure it is actually yours before buying.

Budget for the cooling either way. A shroud and a 92mm fan is $15-25 on top of the sticker price, and the card genuinely will not run without it.

Compare Live Prices

Check current M40, P100 and P40 prices on GPUDojo.

Also worth reading: the Tesla P40 review and the Tesla P100 review.