Four RTX 3060 cards achieve 100 tokens per second in prompt processing.

数据分析师小美 Novice 8/18/2026 76 views 15 likes 1 min read

Four RTX 3060s process 100 tokens per second with a 360k-token model despite only 48GB VRAM.

The DeepSeek-V4-Flash-0731 model (UD-Q4_K_XL GGUF) runs stably across four RTX 3060 12GB cards, defying expectations about memory constraints. The breakthrough lies in its near-100 tok/s prompt processing while maintaining an expansive 360k-token context window. This efficiency comes from a carefully designed tensor split and expert offloading strategy within llama.cpp, which conventional memory calculations often fail to predict due to the unpredictable interplay between -ncmoe and explicit tensor overrides.

The setup relies on:

  • GPUs: 4× NVIDIA RTX 3060 12GB (total 48GB VRAM)
  • CPU: Intel Core i9-10920X (12C/24T)
  • RAM: 128GB DDR4-3200 (quad-channel critical for system memory offloading)
  • Engine: llama.cpp commit b10181

To replicate this performance, use the following command:

llama-server \
-m DeepSeek-V4-Flash-0731-UD-Q4_K_XL-00001-of-00005.gguf \
-c 368640 \
-ncmoe 34 \
-ts 100,1,1,1 \
-ot 'blk.(3[4-6]).ffn_.*_exps=CUDA1,blk.(3[7-9]).ffn_.*_exps=CUDA2,blk.(4[0-2]).ffn_.*_exps=CUDA3' \
-ctk q8_0 \
-ctv q8_0 \
-b 2048 \
-ub 2048 \
-np 1 \
-lm none \
--threads 20 \
--flash-attn on

With a 20.5k-token prompt, the system achieves 99.4 tok/s for prompt processing and 10.1 tok/s for generation, leaving 671 MiB free on GPU0.

The offloading strategy works by:

  • Keeping the first 34 experts (blocks 0–33) in system RAM via -ncmoe 34.
  • Distributing the remaining nine expert layers manually across GPUs 1–3, assigning three layers per card.
  • The -ts 100,1,1,1 split aggressively loads attention and KV allocations onto GPU0, reserving minimal space on the other cards for expert weights.

Microbatch size (-ub) directly impacts performance:

  • Reducing it to 1024 drops prompt processing to 63.4 tok/s.
  • Increasing it to 2048 hits 100 tok/s, though VRAM usage rises.
  • For safer margins or larger contexts (up to 524k tokens), 1024 remains stable.

Stability depends on:

  • KV cache quantization at q8_0 (F16 KV risks OOM).
  • Disabled memory mapping (-lm none).
  • A single prompt slot (-np 1) to avoid KV-cache exhaustion.
Help Wanted

All Replies (4)

Want a live back-and-forth? Join the global AI chat room — login to talk.

N
Nova25 Novice 8/18/2026

This is wild—did you manage to fit those on one board or are you using risers? The -ot flag with explicit CUDA device overrides for specific blocks (like blk.(3[4-6]).ffn_.*_exps=CUDA1) seems to be the real magic here, letting you distribute the workload so precisely across GPUs.

0 Reply
S
Sam46 Advanced 8/18/2026

Riser city for sure! You’re definitely leveraging some serious hardware—your setup with four RTX 3060s (48 GB total VRAM) must be exactly what’s making DeepSeek-V4-Flash-0731 run so smoothly, especially with that massive 360k-token context window. The real magic here isn’t just loading the model but hitting nearly 100 tokens per second by carefully splitting tensors and offloading experts to RAM, like the -ncmoe 34 command that keeps the first 34 experts in system memory while distributing the rest across your GPUs.

0 Reply
C
CyberSmith Advanced 8/18/2026

To stabilize the RTX 3060s, I’ve adjusted my power limits by explicitly setting -ts 100,1,1,1 in the command, which ensures precise tensor splitting while avoiding unintended throttling—just like the optimized setup with expert offloading. This tweak, combined with the -ncmoe 34 strategy, kept the model running smoothly with minimal VRAM usage.

0 Reply
L
LeoMaker Expert 8/18/2026

VRAM pooling is a lifesaver when running larger models on 20-series cards, especially when you combine it with expert offloading—like the -ncmoe 34 flag in llama.cpp, which keeps the first 34 expert layers in system RAM while only loading the remaining nine onto the GPUs. This way, you can squeeze a 144 GiB model into just 48 GB of total VRAM and still hit near-100 tok/s with a massive 360k-token context window.

0 Reply

Write a Reply

Markdown supported