AI Tools

DeepSeek Hardware Requirements: GPU, VRAM and RAM

By · Wed Oct 07 2026 · 9 min read · 0 views

View as a Web Story

AI Tools#local llm#deepseek#gpu#VRAM#Self-hosting

GPU server rack with a table of DeepSeek VRAM requirements

Running the current DeepSeek V4 models yourself takes a multi-GPU server, not a gaming PC. The official DeepSeek-V4.1-Flash checkpoint is 510 GB, which needs about eight 80 GB GPUs. The larger DeepSeek-V4-Pro checkpoint is 865 GB and needs eight 141 GB GPUs or sixteen 80 GB ones. A laptop or a single consumer GPU can only run the smaller R1 distilled models from 2025, which this guide covers separately. Every file size and parameter count below comes from the model pages on Hugging Face, checked on October 7, 2026.

DeepSeek hardware requirements at a glance

DeepSeek is a Chinese AI lab that publishes its models on Hugging Face under the MIT licence. Which hardware you need depends on which model you want. This table gives the memory each one needs just to hold its weights.

Model Total parameters Active per token Checkpoint size Smallest practical GPU setup
DeepSeek-V4-Pro 1.6T 49B 865 GB 8 x 141 GB GPUs (H200 class)
DeepSeek-V4.1-Flash 552B 8B prefill, 16B decode 510 GB 8 x 80 GB GPUs (H100 class)
DeepSeek-V4-Flash (June) 284B 13B 160 GB 4 x 80 GB GPUs
DeepSeek-R1 / V3-class (older) 671B 37B about 670 GB in FP8 8 x 141 GB GPUs
R1-Distill-Llama-70B 70.6B dense 141 GB in BF16 1 x 48 GB GPU at 4-bit
R1-Distill-Qwen-32B 32.8B dense 65.5 GB in BF16 1 x 24 GB GPU at 4-bit
R1-Distill-Qwen-14B 14.8B dense 29.5 GB in BF16 1 x 12 GB GPU at 4-bit
R1-Distill-Qwen-7B 7.6B dense 15.2 GB in BF16 1 x 8 GB GPU at 4-bit

Sources: the Hugging Face pages for DeepSeek-V4.1-Flash, DeepSeek-V4-Pro and DeepSeek-V4-Flash, plus the R1-Distill-Qwen-32B card. The GPU setups are our sizing from those file sizes. They are estimates, not DeepSeek recommendations, and the 4-bit rows assume a community quantisation.

How much VRAM does DeepSeek V4 need?

DeepSeek-V4.1-Flash needs at least 510 GB of GPU memory for weights alone, and DeepSeek-V4-Pro needs at least 865 GB. These are the sizes of the official checkpoints, which store the Mixture-of-Experts layers in 4-bit floating point and most other layers in 8-bit. You cannot load fewer weights, because a Mixture-of-Experts model must hold every expert in memory even though each token only uses a few of them.

Here is the arithmetic for common GPU configurations. Memory figures are the cards' nominal capacity, and real usable memory is lower.

Setup Total GPU memory V4.1-Flash (510 GB) V4-Pro (865 GB)
4 x H100 80 GB 320 GB No No
8 x H100 80 GB 640 GB Yes, 130 GB spare No
16 x H100 80 GB 1,280 GB Yes Yes, 415 GB spare
4 x H200 141 GB 564 GB Yes, only 54 GB spare No
8 x H200 141 GB 1,128 GB Yes Yes, 263 GB spare
8 x 192 GB GPUs (MI300X or B200 class) 1,536 GB Yes Yes

Read the "spare" column with care. The spare memory has to hold the key-value cache, activations and framework overhead. A 54 GB margin on four H200 cards is thin for production traffic, while 130 GB on eight H100 cards is comfortable. DeepSeek's own reference code for V4.1-Flash converts the weights into one checkpoint per tensor-parallel rank and runs on eight ranks in its example, per the inference README. That is a reference implementation, not a production server.

Active parameters do not shrink the memory bill. They set speed. V4.1-Flash activates only 8B parameters per token during prefill and 16B during decode, which makes each token cheap to compute once the weights are loaded.

How much memory does a long context need?

V4.1-Flash keeps its global KV cache at 890 bytes per token, so a full 1M-token context adds under 1 GB. DeepSeek reports this in the model card and says the figure is roughly one quarter of the previous Flash model. At 890 bytes per token the maths is simple:

Context length Global KV cache for one sequence
32K tokens about 29 MB
128K tokens about 117 MB
1M tokens about 0.9 GB

This changes how you size a server. On older dense models, context length and concurrency decide how many GPUs you need. On V4.1-Flash the weights dominate, so a box that fits the 510 GB checkpoint with some headroom can serve many long-context requests at once. The figure covers the global cache. The model also keeps a short sliding-window state and uses a replay mechanism, so treat the number as a floor and measure real usage under load.

Can you run DeepSeek on a laptop or one GPU?

Not the V4 models. A 510 GB checkpoint does not fit in a 512 GB unified-memory machine once the cache and the operating system are counted, and no single consumer GPU comes close. What you can run is the family of R1 distilled models that DeepSeek released in early 2025. They are dense models from 7.6B to 70.6B parameters, published under the MIT licence.

Advertisement

The formula is weights times bytes per parameter, plus runtime overhead. At 4 bits a parameter takes half a byte, and at 8 bits it takes one byte. The table below adds a flat 20% for the KV cache and framework overhead at a moderate context length, and your real figure will move with context size and engine.

Model Parameters 4-bit weights 4-bit with 20% overhead 8-bit with 20% overhead Fits on
R1-Distill-Qwen-7B 7.6B 3.8 GB 4.6 GB 9.1 GB 8 GB GPU at 4-bit
R1-Distill-Qwen-14B 14.8B 7.4 GB 8.9 GB 17.8 GB 12 GB GPU at 4-bit
R1-Distill-Qwen-32B 32.8B 16.4 GB 19.7 GB 39.4 GB 24 GB GPU at 4-bit
R1-Distill-Llama-70B 70.6B 35.3 GB 42.4 GB 84.7 GB 48 GB GPU, or two 24 GB GPUs at 4-bit

The model card shows DeepSeek's own example serving the 32B distill with vLLM across two GPUs, using the command vllm serve deepseek-ai/DeepSeek-R1-Distill-Qwen-32B --tensor-parallel-size 2 --max-model-len 32768. Quantised builds from the community are what let one 24 GB card hold it. Quality drops as you quantise, so test the build on your own prompts.

A distilled model is not the same as the full model. The distills are smaller dense models trained on outputs from the large R1 model, and they are weaker than DeepSeek's current API models. If your goal is simply to use DeepSeek, calling the API costs far less than buying hardware. See what the DeepSeek API costs and when it bills peak rates.

How much RAM and disk do you need?

DeepSeek publishes no system-RAM requirement for V4, so follow your serving engine's guidance and keep the host well above the GPU's staging needs. Plan for roughly two copies of the checkpoint on disk when you use DeepSeek's conversion script. The disk figure follows from the reference workflow, which converts the Hugging Face weights into a separate per-rank checkpoint. That is our reading of the convert step, so confirm it with a test run before you buy storage.

Item V4.1-Flash V4-Pro
Download size 510 GB 865 GB
Disk with one converted copy (estimate) about 1 TB about 1.7 TB
Download time at 1 Gbps about 68 minutes about 115 minutes
Download time at 10 Gbps about 7 minutes about 12 minutes

Use fast NVMe storage. Loading hundreds of gigabytes from a slow disk can take longer than the download.

What should you check before you buy hardware?

Check three things before spending money: whether your engine supports the model, whether the numerical formats run well on your GPU, and what your real traffic looks like.

  1. Engine support. DeepSeek publishes a reference implementation and a Rust toolkit called deepseek-recipe for prompt formatting and response parsing. Production engines such as vLLM may lag a new architecture, so confirm support for the exact model and version before you commit.
  2. Numeric formats. The official V4 weights use FP4 for experts and FP8 for most other layers. Older GPUs may lack fast native kernels for those formats, which could slow serving down. Verify throughput on a rented node first.
  3. Sampling settings. DeepSeek recommends temperature 1.0, top_p of 0.95 or 1.0, and a max_tokens of at least 256K for V4.1-Flash. Long reasoning outputs are normal and keep requests on the GPU longer.
  4. Licence. The model cards list the MIT licence. Read the licence file in the repository yourself before commercial use.

When does self-hosting beat the API?

Self-hosting rarely beats the DeepSeek API on price. It wins on data control, custom serving and very high volume. Here is the break-even maths with an assumed rental rate. We assume $2.50 per GPU-hour for an H100-class card, which you should replace with a current quote.

Item Value
Eight GPUs at $2.50 per GPU-hour $20 per hour
Always-on for a 730-hour month $14,600
API cost for a 100M-input, 10M-output agent workload, off-peak $9.24
Workloads needed to match the rental bill about 1,580
Tokens per month at that point about 174 billion

The API example uses the deepseek-flash off-peak rates of $0.003 for cache hits, $0.15 for cache misses and $0.60 per million output tokens. A rented eight-GPU node has to process roughly 174 billion tokens a month before it undercuts the API on these assumptions, and that ignores engineering time and idle hours. Self-host when your data cannot leave your network, when you need a fixed latency guarantee, or when your volume is above that scale. Otherwise the API is cheaper and simpler. The same logic appears in why the cheapest AI API is not the cheapest to run.

If privacy is your reason, also read whether the free AI API tier trains on your data before you decide the hosted route is unsafe.

Is the API a better fit than your own servers?

For most teams, yes. See how the hosted models compare with the closed alternatives in DeepSeek vs OpenAI and DeepSeek vs Claude, which include monthly cost tables for the same agent workload used above.

Should you buy a GPU for DeepSeek now?

For local experiments with distilled models, a single 24 GB card is the practical ceiling and covers the 32B distill at 4-bit. For the V4 family, rent GPUs by the hour instead of buying, because the hardware cost is far beyond a workstation and the model lineup changes every few weeks. If you are weighing a purchase anyway, our guides on whether to buy a GPU now or wait out 2026 and whether RAM prices are cooling enough to buy now cover timing.

Sources

DeepSeek-V4.1-Flash on Hugging Face, DeepSeek-V4-Pro on Hugging Face, DeepSeek-V4-Flash on Hugging Face, V4.1-Flash inference README, DeepSeek-R1-Distill-Qwen-32B, DeepSeek-R1-Distill-Llama-70B, deepseek-recipe on GitHub, DeepSeek V4.1-Flash release notes, DeepSeek API pricing, NVIDIA H200, NVIDIA H100, AMD Instinct MI300X.

Advertisement

FAQ

What are the hardware requirements for DeepSeek V4?

DeepSeek-V4.1-Flash needs at least 510 GB of GPU memory for its weights, so eight 80 GB GPUs work. DeepSeek-V4-Pro needs at least 865 GB, so eight 141 GB GPUs or sixteen 80 GB GPUs. These are checkpoint sizes; the cache and overhead need extra memory.

How much VRAM does DeepSeek R1 need?

The full 671B R1 model needs about 670 GB in FP8, which means eight 141 GB GPUs. The R1 distilled models are far smaller: the 7B needs about 5 GB at 4-bit, the 14B about 9 GB, the 32B about 20 GB and the 70B about 42 GB, including 20% overhead.

Can I run DeepSeek on my laptop or a single GPU?

You can run the R1 distilled models, not the V4 models. A 24 GB GPU holds the 32B distill at 4-bit, and an 8 GB GPU holds the 7B distill. The V4 checkpoints are 510 GB and 865 GB, which no consumer GPU can hold.

Do active parameters reduce DeepSeek's memory needs?

No. DeepSeek V4 models are Mixture-of-Experts, so every expert must be loaded even though only a few run per token. V4.1-Flash activates 8B parameters during prefill and 16B during decode, which makes it fast, but the full 510 GB must still be in memory.

Is it cheaper to self-host DeepSeek than use the API?

Usually not. An eight-GPU node at an assumed $2.50 per GPU-hour costs about $14,600 a month, and the API would need roughly 174 billion tokens a month at deepseek-flash off-peak rates to cost the same. Self-hosting pays off for data control or very high volume.

How much memory does DeepSeek's long context use?

DeepSeek-V4.1-Flash stores its global KV cache at 890 bytes per token, so a 1M-token context adds about 0.9 GB per sequence. That is roughly a quarter of the previous Flash model, and it makes the weights, not the context, the main memory cost.

Comments

Loading…

Sign in to join the conversation.

Related posts

What Is a Proxy on Janitor AI? How It Works

What Is a Proxy on Janitor AI? How It Works

A proxy on Janitor AI is a connection that lets the site use an outside language model, such as DeepSeek, instead of its built-in one. It is not a VPN and it does not hide your traffic. Janitor AI's

Wed Oct 07 2026 · 8 min read · 0 views

AI Tools

Janitor AI Suspended From Gemini? What to Check

Janitor AI Suspended From Gemini? What to Check

Short answer: if Gemini stopped working in Janitor AI, the cause is usually one of three things: a rate-limit or quota error (429), a content filter block, or an API key that stopped working. Only the

Wed Oct 07 2026 · 7 min read · 0 views

AI Tools

Is Janitor AI Down? How to Check and Fix Errors

Is Janitor AI Down? How to Check and Fix Errors

Short answer: to check whether Janitor AI is down, open its official status page at status.janitorai.com, which Janitor AI's own help centre points to. On October 7, 2026 at 16:58 UTC it showed All

Wed Oct 07 2026 · 5 min read · 0 views

AI Tools