Skip to content

Speculative decoding (MTP)

Speculative decoding drafts several tokens per step using a separate assistant or the model’s built-in prediction head. The target model verifies the drafts in parallel before emitting accepted tokens. Configure it through the MTP (Multi-Token Prediction) API.

Terminal window
mistralrs run -m google/gemma-4-E4B-it --quant 8 \
--mtp-model google/gemma-4-E4B-it-assistant \
--mtp-n-predict 6

--mtp-model accepts a Hugging Face id or a local path. The same flags work with mistralrs serve. See run and serve flag references.

--mtp-n-predict N fixes the number of draft tokens per step. Without it, the engine tunes depth using draft acceptance, step time, and batch size, periodically trying other depths as the workload changes. Built-in Qwen heads choose among 2, 3, 4, and 6 drafts. Gemma 4 uses the same candidates, capped by num_assistant_tokens in the assistant’s generation_config.json (6 if absent). DFlash uses the policy described below.

Mode Target models Assistant model Status
MTP Gemma 4 Gemma 4 assistant checkpoints (--mtp-model) Supported with paged attention
MTP Qwen3.5, Qwen3.8 Built into the checkpoint (--mtp) Supported with paged attention
MTP Qwen3.8-Flash-Next Built into the checkpoint (--mtp) Safetensors/UQFF with paged attention
DFlash / DFlash 2 Qwen3.5, Qwen3.8 DFlash draft checkpoints (--mtp-model) Supported with paged attention

Qwen3.5 and Qwen3.8 checkpoints ship a multi-token prediction head (mtp.* weights) sharing the target’s embeddings and lm_head. Pass --mtp to load it; no separate assistant model is needed:

Terminal window
mistralrs serve -m Qwen/Qwen3.8-27B --quant 4 --mtp

The head adds one attention layer’s paged KV cache. Safetensors and UQFF checkpoints must contain the mtp.* weights; GGUF loading does not support built-in MTP.

Multi-token verification can use different quantized matrix and recurrent kernels from ordinary decoding. Log probabilities and greedy token choices can differ, so token-for-token equivalence is not guaranteed.

Built-in Qwen MTP recomputes cached prompt prefixes, which can increase time to first token for repeated prompts. This restriction does not apply to DFlash.

Qwen3.8-Flash-Next uses the same built-in-head flags:

Terminal window
mistralrs serve -m Qwen/Qwen3.8-Flash-Next --isq q4k --mtp

Flash-Next supports only its built-in head. External assistants and DFlash are unsupported. See model-family notes for memory requirements.

DFlash drafters are small block-diffusion models that propose a whole block of tokens in one pass, conditioned on hidden states tapped from the target’s intermediate layers. They reach much higher acceptance than the built-in single-token MTP head. Pass a DFlash checkpoint (a HF id or local path) as --mtp-model; mistral.rs detects DFlash vs DFlash 2 from the checkpoint’s architectures and loads the path selector and dynamic convolutions when present:

Terminal window
mistralrs serve -m Qwen/Qwen3.8-27B --quant 4 \
--mtp-model incoai/Qwen3.8-27B-DFlash2

--mtp-draft-sampling defaults to auto. A DFlash 2 checkpoint uses probabilistic drafting when its CUDA candidate selector is supported, while DFlash 1 and unsupported devices use greedy drafting. Pass --mtp-draft-sampling greedy to force greedy drafting. Passing probabilistic requires a supported DFlash 2 selector and fails during loading when that path is unavailable.

The maximum automatic draft depth is min(block_size - 1, 7). During serving, mistral.rs may use a shallower depth so the recurrent checkpoint layout fits the configured sequence capacity while preserving memory for the requested paged KV cache. Pass --mtp-n-predict to request a fixed depth instead. Automatic CUDA serving also adapts the active depth to live load: up to 8 live DFlash contexts use the full reserved depth, while larger loads use up to 3 drafts. Set MISTRALRS_DFLASH_ADAPTIVE=0 (or false) to tune depth from measured throughput instead. That tuner also runs when adaptation by live context count is unavailable. It chooses among 3, 5, and 7 drafts, capped by the reserved depth. The drafter’s projection weights are requantized to the target’s in-situ quantization type at load (override with the MISTRALRS_DFLASH_ISQ environment variable, e.g. q6k or none for bf16).

Acceptance is strongly content-dependent: incoai/Qwen3.8-27B-DFlash2 accepts ~7 of 7 drafts on math/code and much less on thinking-mode reasoning traces, which are out of its training distribution.

Gemma 4 assistant checkpoints are MTP drafters for Gemma 4 target models. See the google/gemma-4-E4B-it-assistant model card for the upstream checkpoint. A downloaded checkout works too:

Terminal window
mistralrs run -m google/gemma-4-E4B-it --quant 8 \
--mtp-model ./gemma-4-E4B-it-assistant \
--mtp-n-predict 6

Non-paged KV-cache MTP is intentionally disabled for now, which is why paged attention is required (see the note at the top).

The target and assistant configs must match where the implementation requires it, including vocabulary size and target hidden size. Mismatches fail at load, before generation starts.

Throughput gain depends on how many proposed tokens the target accepts and on the cost of the target verification pass.

MTP supports batched generation and constrained decoding.

MTP is configured at launch time only (CLI flags, Runner(...), or the model builder). There is no per-request HTTP field to toggle it; load the server with the assistant attached and every request uses it.