Tesla M40 24GB in 2026: Still Worth Buying for AI?
The Tesla M40 24GB used to have one clear argument: it was the cheapest way to get 24GB of VRAM, at $70-90 on eBay. That argument has weakened considerably. The M40 now starts around $229, while the 16GB Tesla P100 sits near $80 with two and a half times the memory bandwidth, and the Tesla P40, a genuinely better card in every respect, is only $10 more.
That does not make the M40 useless. It makes it a narrower proposition than it was: 24GB of capacity, on 2015 Maxwell silicon, with no native FP16 and a passive cooler you will have to solve. Whether that is worth $229 depends entirely on whether you need the capacity more than the speed. Here are the numbers to decide with.
| GPU Architecture | Maxwell (GM200) - 2015 |
|---|---|
| CUDA Cores | 3,072 |
| VRAM | 24GB GDDR5 |
| Memory Bandwidth | 288 GB/s |
| FP32 Performance | 7.0 TFLOPS |
| FP16 Performance | None (no native FP16) |
| TDP | 250W |
| Cooling | Passive (requires server airflow) |
| Compute Capability | 5.2 |
| Lowest Used Price | $229 (live, September 2026) |
The Case For the M40
The M40's advantage is capacity: 24GB of VRAM for $229, on a card that modern software still supports. This lets you:
- Run 20B+ parameter models at Q4 that won't fit on 16GB cards
- Run 7B-14B models at higher quantization (Q8, even FP16 for 7B) for better quality
- Use larger context windows without running out of VRAM
- Experiment with fine-tuning small models (LoRA on 7B fits in 24GB)
The key word is capacity. If you need to fit a model larger than 16GB and cannot stretch to a P40 at $239, the M40 is the cheapest card that will hold it. If your models fit in 16GB, the P100 at $80 is both cheaper and dramatically faster: see the comparison below.
Pros
- 24GB of VRAM at $229, the cheapest card that holds a 20B+ model
- Full CUDA 12 support (compute capability 5.2)
- Modern PyTorch/llama.cpp works fine
- 24GB unified VRAM (not split like K80)
- Abundant supply on eBay
- Can run 20B+ models at Q4
Cons
- Slow - 288 GB/s bandwidth limits tok/s
- No FP16 support (no native half-precision)
- Passive cooling only - overheats in desktop
- 250W power draw
- No display output
- No video encode/decode
Real-World Performance
Here's what you can realistically expect running models on the M40 with llama.cpp:
Tesla M40 24GB - Estimated tok/s (llama.cpp, Q4_K_M)
These speeds are usable but not fast. For comparison, a P40 is about 20-40% faster across the board, and an RTX 3090 is 3-4x faster. If you're accustomed to ChatGPT's response speed, the M40 will feel noticeably slower for larger models.
The Cooling Problem
This Is the M40's Biggest Challenge
The Tesla M40, like all data center GPUs of its era, has no fans. It's a completely passive card designed for server chassis with front-to-back forced airflow. In a standard desktop PC, it will thermal throttle within minutes and may shut down to prevent damage.
You must add aftermarket cooling. Options include:
- 3D-printed fan shroud ($5-10 in materials, or $15-25 printed for you) - Search "Tesla M40 fan shroud" on Printables/Thingiverse. Mounts a 92mm or dual 80mm fans onto the heatsink. This is the best solution.
- Zip-tied fans (free if you have spare fans) - Zip-tie a 92mm Noctua or Arctic fan directly to the card's heatsink. Ugly but effective.
- High-airflow case - A case with strong front-to-back airflow (like a server chassis) can work without modifications, but only if you have genuinely strong fans.
- Open test bench - Running the card on an open bench with a desk fan pointed at it. Works for testing but not a long-term solution.
Temperature Targets
- Under 85C: Safe for sustained operation
- 85-95C: Thermal throttling begins, performance degrades
- Above 95C: Risk of shutdown, potential long-term damage
Monitor with nvidia-smi -l 1 during first use to check temperatures under load.
M40 vs P40: Is the Extra $10 Worth It?
| Spec | M40 24GB ($229) | P40 24GB ($239) | Difference |
|---|---|---|---|
| Architecture | Maxwell | Pascal | 1 generation newer |
| Bandwidth | 288 GB/s | 347 GB/s | +20% |
| VRAM | 24GB GDDR5 | 24GB GDDR5X | Same capacity, faster memory |
| INT8 Inference | No | Yes | P40 has INT8 support |
| tok/s (8B Q4) | ~32 | ~45 | ~40% faster |
| tok/s (32B Q4) | ~12 | ~16 | ~33% faster |
| Power | 250W | 250W | Same |
| Cooling | Passive | Passive | Both need aftermarket |
| Software support | Good | Better | P40 has wider kernel support |
The P40 is 20-40% faster across the board, adds INT8 inference, and has wider kernel support. Whether $10 buys enough of that to matter is a judgement about your own patience. The M40 is not unusable, it is just slower at everything.
M40 vs P100: The Comparison That Changed
This is the one worth thinking hardest about, because it inverted. When this article was first written the M40 was $70-90 and the P100 cost more. Today the M40 starts at $229 and the P100 starts at $80.
| Spec | M40 24GB ($229) | P100 16GB ($80) | What it means |
|---|---|---|---|
| VRAM | 24GB GDDR5 | 16GB HBM2 | M40 fits larger models |
| Bandwidth | 288 GB/s | 732 GB/s | P100 is ~2.5x faster at generation |
| Native FP16 | No | Yes | Matters for training and some runtimes |
| Architecture | Maxwell (2015) | Pascal (2016) | Both CUDA 12 capable |
| Cooling | Passive | Passive | Both need aftermarket |
The decision is capacity against speed, and the cheaper card is now the faster one. If your models fit inside 16GB, the P100 is faster and costs less, so there is no argument for the M40. If you need more than 16GB, the M40 is still the cheapest card that holds it, and 24GB of slow VRAM beats 16GB you cannot fit into. Work out your model size first; the answer follows from it.
Verdict
A capacity card, not a value card
The Tesla M40 24GB at $229 is a functioning 24GB GPU that modern software still supports. It runs 20B+ parameter models, and for batch work where you set a job going and come back later, its slowness costs you little.
What it no longer is, is the obvious budget pick. At $80 the P100 is cheaper and far faster if 16GB is enough for you, and at $239 the P40 gives you the same 24GB with real speed. The M40 now occupies a narrow gap: you need more than 16GB, and you cannot or will not pay $239. That is a real position to be in, but make sure it is actually yours before buying.
Budget for the cooling either way. A shroud and a 92mm fan is $15-25 on top of the sticker price, and the card genuinely will not run without it.
Compare Live Prices
Check current M40, P100 and P40 prices on GPUDojo.
Also worth reading: the Tesla P40 review and the Tesla P100 review.