DeepSeek Hardware Requirements: GPU, VRAM and RAM

DeepSeek-V4.1-Flash is a 510 GB checkpoint that needs about eight 80 GB GPUs; DeepSeek-V4-Pro is 865 GB and needs eight 141 GB GPUs.

A Mixture-of-Experts model must hold every expert in memory, so active parameters affect speed, not the VRAM bill.

V4.1-Flash keeps its global KV cache at 890 bytes per token, so a full 1M-token context adds under 1 GB.

Laptops and single GPUs can run only the R1 distilled models, for example the 32B distill on a 24 GB card at 4-bit.

Self-hosting rarely beats the API on price: an eight-GPU rental needs roughly 174 billion tokens a month to break even.

Read the full post

Read the full post