Standard -ngl vs --cpu-moe: Offloading Mixture of Experts in llama.cpp
How modern MoE architectures like Qwen 3.6 35B A3B and Gemma 4 26B differ from dense models, what lives inside each layer, and why --cpu-moe saves speed for MoE models, when compared to standard -ngl.
If you have tried running modern Mixture of Experts (MoE) models like Qwen 3.6 35B A3B, or Gemma 4 26B A4B on a 12 GB or 16 GB graphics card, you have probably run into this problem.
At a 4-bit quant (Q4_K_M), Qwen 3.6 35B sits around 21 GB, and Gemma 4 26B takes about 16 GB. Neither fits cleanly onto a 12 GB RTX 3060 or a 16 GB card once you factor in KV cache and overhead.
So you treat it like any normal dense model: you open your terminal, set -ngl 24 in llama.cpp to offload whatever fits on your GPU, and let the remaining layers spill over to system RAM.
Then you test a prompt, and the generation speed collapses. A model that should easily run at 30 to 45 tokens per second crawls along at 3.
To understand why this happens, and why flags like --cpu-moe and --n-cpu-moe exist, you have to look at what actually lives inside an MoE layer and how memory moves during inference.
How MoE models differ from dense models
In a standard dense model like Qwen 3.8 27B, every parameter is active on every single token. If a model has 27 billion parameters, your GPU computes matrix math for all 27 billion parameters on every forward pass.
MoE models decouple the total parameters on disk from the active parameters in compute.
An MoE model has two numbers that matter:
- •Total parameters: Every weight stored on disk that must be loaded into memory.
- •Active parameters: The tiny subset of weights that actually execute for any given token.
Look at some models:
- •Qwen 3.6 35B A3B: 35 billion total parameters, but only ~3.4 billion active parameters per token.
- •Gemma 4 26B A4B: 25.2 billion total parameters, but only ~3.8 billion active parameters per token.
- •DeepSeek V4 Flash: 284 billion total parameters, but only routes to 6 out of 256 experts per token.
Your system RAM and VRAM must hold the total parameters. But your compute silicon only has to process the active parameters.
This is why a 35B MoE model can generate text at the speed of a tiny 3B model. But the moment you start offloading layers naively between GPU and CPU, that advantage evaporates.
What a transformer layer contains: Dense vs. Expert weights
To see why standard offloading breaks down on MoE models, look at what sits inside a single transformer block. Every weight inside the layer falls strictly into one of two buckets:
1. The Dense Weights (Always Active)
These weights run on every single token without exception:
- •Self-Attention Projections: The Query, Key, Value, and Output matrices (
q_proj,k_proj,v_proj,o_proj), along with the layer's KV Cache. - •Normalization: RMSNorm / LayerNorm scaling tensors.
- •The Router Gate: A small dense matrix that inspects each incoming token and calculates which experts should handle it.
Why Dense weights must stay on the GPU: Self-attention is strictly memory-bandwidth bound. Every generated token reads through the full past KV cache history. On modern graphics cards, this happens at GDDR6 or HBM bandwidth (500 to 1,000+ GB/s). On system RAM, it drops to dual-channel DDR speeds (50 to 90 GB/s). If your dense weights and KV cache spill into system RAM, your entire forward pass slows to a crawl.
The good news? Dense weights are small. They typically account for only 8% to 15% of an MoE layer's total size.
2. The Expert Weights (Conditionally Active)
Instead of a single feed-forward network (FFN), an MoE layer contains a massive swarm of independent expert FFNs:
- •In Qwen 3.6 35B: 256 independent sets of
gate_proj,up_proj, anddown_projmatrices per layer. - •In Gemma 4 26B: 128 independent sets per layer.
Good News: The expert bank accounts for 85% to 92% of the layer's weight bytes.
For any given token:
- •In Qwen 3.6 35B: only 8 out of 256 experts activate. The other 248 experts (96.8% of the expert weights) sit completely untouched.
- •In Gemma 4 26B A4B: only 8 out of 128 experts activate per token (routing to ~3.8B active parameters). The other 120 experts (93.8% of the expert weights) sit idle.
This distinction is the key to memory offloading: Dense weights are small but bandwidth-critical (they must stay in VRAM). Expert weights are huge but mostly idle (they can tolerate living in system RAM).
What happens in -ngl vs --cpu-moe vs --n-cpu-moe
Once you see the divide between Dense weights and Expert weights, the behavior of each llama.cpp flag becomes immediately obvious.
1. Standard -ngl (Blind Horizontal Slicing)
The default -ngl flag was designed for dense models. It has no concept of Dense vs. Expert weights—it simply slices the model horizontally from top to bottom.
If you have a 40-layer model like Qwen 3.6 35B and set -ngl 25:
- •The top 25 layers live 100% on your GPU.
- •The bottom 15 layers live 100% in CPU system RAM.
Here is the fatal flaw: For those 15 layers on the CPU, you dragged the Dense weights (Self-Attention and KV Cache) out of VRAM and into system RAM.
Now, on every single token:
- 1.Self-attention for those 15 layers runs across slow DDR memory bandwidth (10x slower than GPU).
- 2.Activations must cross the PCIe bus from GPU to CPU, wait for the CPU to compute attention and experts, and then travel back across PCIe to the GPU.
Moving the Dense weights to CPU is what destroys your token generation speed.
2. --cpu-moe (Architecture-Aware Splitting)
The --cpu-moe flag respects the Dense vs. Expert boundary:
- •It protects 100% of Dense weights:
-nglis automatically forced to 99, keeping all Self-Attention projections, all KV Caches, all normalization weights, and all router gates pinned in fast GPU VRAM across every layer. - •It offloads only the Expert weights: The bulky 85% to 92% expert FFN bank is diverted to system RAM.
Because no Dense weights ever leave the GPU, self-attention always runs at full GPU memory bandwidth with zero CPU slowdown. The GPU computes attention, the router selects the active experts, and only the required active expert math executes through system memory.
You save 15 to 25 GB of VRAM without sacrificing attention speed.
3. --n-cpu-moe (Partial Expert Offload)
--cpu-moe moves all expert layers to RAM. But what if you have an RTX 4070 Ti (16 GB) or an RTX 4090 (24 GB), and you have spare VRAM sitting empty after loading the Dense weights?
That is where --n-cpu-moe comes in.
It gives you fine-grained control over the Expert weights:
- •Dense weights stay 100% in GPU VRAM across every single layer.
- •Expert weights are divided: Only the bottom
Nlayers have their expert banks in system RAM; the remaining layers keep their expert banks accelerated in GPU VRAM.
This gives you a precision dial to pack your remaining VRAM with as many expert weights as fit, without ever spilling Dense attention weights to the CPU.
Comparison: Dense vs. Expert memory placement
| Setting | Dense Weights (Attention & KV) | Expert Weights (FFNs) | PCIe Bus Overhead | Real Speed Impact |
|---|---|---|---|---|
-ngl (Partial) |
Split between GPU and CPU RAM | Split between GPU and CPU RAM | Full activations every token | Severe (70%–90% drop) |
--cpu-moe |
100% in GPU VRAM | 100% in System RAM | Only active expert passes | Moderate (20%–40% drop) |
--n-cpu-moe |
100% in GPU VRAM | Bottom N in RAM, rest in VRAM | Only offloaded expert passes | Low to moderate |
Which one should you choose and when?
Once you think in terms of Dense vs. Expert weights, the decision is simple:
Scenario A: You are running a Dense model (Llama 3.1 8B, Qwen 2.5 14B)
- •Use standard
-ngl. - •Dense models contain only Dense weights and have no separate expert banks. Flags like
--cpu-moedo nothing here.
Scenario B: Your VRAM can fit both Dense weights AND all Expert weights
- •Use
-ngl 99. - •For example, running Gemma 4 26B (Q4_K_M, ~15 GB) on a 24 GB RTX 4090. If your GPU can hold everything, keep all weights on the GPU for maximum speed.
Scenario C: Your VRAM can fit Dense weights, but NOT all Expert weights
- •Never use standard partial
-nglalone. That drags Dense weights to the CPU. - •If you have spare VRAM beyond Dense weights: Use
--n-cpu-moe. For example, running Qwen 3.6 35B A3B on a 16 GB card: set-ngl 99 --n-cpu-moe 10. That keeps all Dense weights and 30 layers of experts on the GPU, offloading only the last 10 expert layers to RAM. - •If your VRAM is tight: Use
-ngl 99 --cpu-moe. For example, running Qwen 3.6 35B or Gemma 4 26B on an 8 GB or 12 GB card. This guarantees all Dense attention weights stay accelerated in VRAM while system RAM absorbs the expert bulk.
Example commands
Qwen 3.6 35B A3B on a 12 GB or 16 GB card (Full Expert Offload):
llama-cli -m Qwen3.6-35B-A3B.Q4_K_M.gguf \
-ngl 99 \
--cpu-moe \
-c 8192
Qwen 3.6 35B A3B on a 16 GB or 24 GB card (Partial Expert Offload):
llama-cli -m Qwen3.6-35B-A3B.Q4_K_M.gguf \
-ngl 99 \
--n-cpu-moe 10 \
-c 16384
Gemma 4 26B A4B on an 8 GB or 12 GB card:
llama-cli -m Gemma-4-26B-A4B.Q4_K_M.gguf \
-ngl 99 \
--cpu-moe \
-c 8192
If you want to know the exact number of layers to pass for --n-cpu-moe on your specific card, you can test your hardware in the WhichLLMModel Simulator to calculate the exact VRAM fit.
Written by Zubair Tahir
Founder & Lead Developer
Building independent, empirical evaluation tools to eliminate model decision fatigue for AI engineers and developers.