Skip to main content

Local LLM Hardware Recommendations

Detailed model-by-model hardware sizing, quantization advice, throughput tables, cooling and PSU guidance, and cloud GPU rental fallback for self-hosting plugins-pro/paid/ai with GPU acceleration.

Companion to the Local LLM Hardware Guide. That page covers the high-level decision (local vs API). This page is the per-model hardware matrix you need when buying or renting GPUs.

All tokens-per-second numbers below are batch-size 1, single user, with the listed quantization. Throughput drops 30-50% under multi-user load; bump up a tier for shared deployments.

Hardware Tiers

Tier 1. Entry, 8 GB VRAM

  • GPUs: RTX 3060 12 GB, RTX 4060 8 GB, RTX 4060 Ti 16 GB
  • Apple Silicon equivalent: M1/M2 base 8-16 GB unified memory
  • Cost: $300-500 used / new
  • Runs well: 7B-13B models at Q4-Q5
  • Don’t try: anything 30B+, quantization loss kills quality
  • GPUs: RTX 4090 24 GB, RTX 3090 24 GB, RTX A4000 16 GB
  • Apple Silicon equivalent: M3 Pro 18-36 GB, M3 Max 64 GB
  • Cost: $1,000-2,000 used / new
  • Runs well: 30B at Q4, 70B at Q4 (tight, 2-bit weights), Mixtral 8x7B at Q5
  • Sweet spot: 95% of self-host workloads land here

Tier 3. Pro, 48 GB+ VRAM

  • GPUs: RTX A6000 48 GB, dual RTX 4090 (NVLink), H100 80 GB, dual A6000
  • Apple Silicon equivalent: M2 Ultra 192 GB, M3 Ultra 192 GB
  • Cost: $4,500-30,000+
  • Runs well: 70B at FP16, 405B at Q4 (multi-GPU), Mixtral 8x22B at Q5, multi-modal vision models
  • When to buy: regulated data, sustained 50M+ token/month workloads, on-prem requirements

Per-Model Recommendations

Model Parameters Recommended VRAM Quantization Tokens/sec Tier
Llama 3.1 8B 8B 6-8 GB Q4_K_M 30-60 (4090) / 8-15 (M2 base) 1
Llama 3.1 70B 70B 24-48 GB Q4_K_M 15-25 (4090) / 8-12 (M3 Max) 2-3
Llama 3.1 405B 405B 220+ GB Q4_K_M 5-10 (8x A100 / M2 Ultra) 3
Mistral 7B 7B 6-8 GB Q4_K_M 40-70 (4090) 1
Mixtral 8x7B 47B (13B active) 24-32 GB Q4_K_M 30-50 (4090) 2
Mixtral 8x22B 141B (39B active) 80-100 GB Q4_K_M 15-25 (A100) / 10-15 (M2 Ultra) 3
Qwen 2.5 7B 7B 6-8 GB Q4_K_M 35-65 (4090) 1
Qwen 2.5 32B 32B 20-24 GB Q4_K_M 18-28 (4090) 2
Qwen 2.5 72B 72B 40-48 GB Q4_K_M 12-20 (A6000) 3
DeepSeek-V3 671B (37B active) 380+ GB Q4_K_M 10-15 (8x H100) 3+
Phi-4 14B 14B 10-12 GB Q4_K_M 25-45 (4090) 1-2

Quantization key: Q4_K_M is the standard 4-bit GGUF format with mixed precision on critical layers. FP16 doubles VRAM needs but improves coherence on edge cases. Drop to Q3_K_S if you must squeeze a tier larger model into limited VRAM, but expect quality loss.

CPU vs GPU vs Apple Silicon

Path Best for Throughput floor Throughput ceiling
CPU only (DDR5) 7B emergency fallback 1-3 tok/s 8 tok/s
Consumer NVIDIA (4060-4090) Single user, mixed sizes 8 tok/s 70 tok/s
Datacenter NVIDIA (A100, H100) Multi-user batch serving 30 tok/s 200+ tok/s
Apple Silicon (M3 Max, M2 Ultra) Quiet desk, large models in unified memory 6 tok/s 30 tok/s
AMD (RX 7900 XTX, MI300X) ROCm-capable workloads 10 tok/s 80 tok/s

Apple Silicon punches above its weight on 70B+ thanks to unified memory, one M2 Ultra holds models that need a multi-GPU rig on NVIDIA, at the cost of lower peak throughput.

Cooling, PSU, Chassis

For multi-GPU rigs:

  • PSU: 1000-1600W 80+ Platinum. RTX 4090 peaks near 600W. Dual 4090 needs 1600W minimum with margin.
  • Cooling: front-to-back airflow, minimum 3 intake fans. GPU temps above 80°C throttle inference. Open-frame mining chassis are fine for home rigs.
  • PCIe lanes: prefer x16/x16 over x8/x8 for multi-GPU model sharding. Threadripper or Xeon W gives you the lanes; consumer Intel/AMD desktops don’t.
  • RAM: 64 GB system RAM minimum for 70B+ model loading. Slow DDR4 is fine; the model lives on the GPU.
  • Storage: NVMe SSD with 100+ GB free per model. Llama 3.1 405B at Q4 is ~230 GB on disk.

Single-GPU rigs in a tower case work without modification on a quality 850W PSU.

Cloud GPU Rental Fallback

Before buying, validate the workload on rented hardware. Three reputable options:

Provider A100 40GB H100 80GB Notes
RunPod $1.19/hr $2.69/hr Spot pricing 30-50% cheaper, hourly billing
Lambda Labs $1.29/hr $2.49/hr Reserved instances cheaper, US-only
Vast.ai $0.40-1.00/hr $1.50-2.50/hr Marketplace, variable reliability

Rough monthly costs at 100% utilization: A100 ~$870-940, H100 ~$1,800-1,950. Compare against amortized hardware cost (RTX 4090 used $1,200 / 36 months ≈ $33 + $50 power = $83/mo) before committing.

Rent for: workload sizing, occasional bursts, regulated data with verified-private cloud. Buy for: predictable steady load, air-gapped requirements, multi-year horizon.

Reference Benchmarks

Verify on your own hardware before sizing, model releases and inference engine optimizations move the numbers monthly.

See Also