Speculative decoding (MTP)
Speculative decoding drafts several tokens per step using a separate assistant or the model’s built-in prediction head. The target model verifies the drafts in parallel before emitting accepted tokens. Configure it through the MTP (Multi-Token Prediction) API.
mistralrs run -m google/gemma-4-E4B-it --quant 8 \ --mtp-model google/gemma-4-E4B-it-assistant \ --mtp-n-predict 6--mtp-model accepts a Hugging Face id or a local path. The same flags work with mistralrs serve. See run and serve flag references.
from mistralrs import Runner, Which
runner = Runner( which=Which.MultimodalPlain(model_id="google/gemma-4-E4B-it"), in_situ_quant="8", mtp_model="google/gemma-4-E4B-it-assistant", mtp_n_predict=6,)Gemma 4 loads via Which.MultimodalPlain and uses a separate assistant checkpoint.
let model = mistralrs::ModelBuilder::new("google/gemma-4-E4B-it") .with_mtp_model("google/gemma-4-E4B-it-assistant", Some(6)) .build() .await?;You can also pass MtpConfig::new("google/gemma-4-E4B-it-assistant", Some(6)) to with_mtp_config. Use MtpConfig::builtin(None) for a built-in head with automatic draft depth. These methods are available on the text, multimodal, and auto-detecting model builders.
--mtp-n-predict N fixes the number of draft tokens per step. Without it, the engine tunes depth using draft acceptance, step time, and batch size, periodically trying other depths as the workload changes. Built-in Qwen heads choose among 2, 3, 4, and 6 drafts. Gemma 4 uses the same candidates, capped by num_assistant_tokens in the assistant’s generation_config.json (6 if absent). DFlash uses the policy described below.
Supported models
Section titled “Supported models”| Mode | Target models | Assistant model | Status |
|---|---|---|---|
| MTP | Gemma 4 | Gemma 4 assistant checkpoints (--mtp-model) |
Supported with paged attention |
| MTP | Qwen3.5, Qwen3.8 | Built into the checkpoint (--mtp) |
Supported with paged attention |
| MTP | Qwen3.8-Flash-Next | Built into the checkpoint (--mtp) |
Safetensors/UQFF with paged attention |
| DFlash / DFlash 2 | Qwen3.5, Qwen3.8 | DFlash draft checkpoints (--mtp-model) |
Supported with paged attention |
Qwen3.5 / Qwen3.8
Section titled “Qwen3.5 / Qwen3.8”Qwen3.5 and Qwen3.8 checkpoints ship a multi-token prediction head (mtp.* weights) sharing the target’s embeddings and lm_head. Pass --mtp to load it; no separate assistant model is needed:
mistralrs serve -m Qwen/Qwen3.8-27B --quant 4 --mtpThe head adds one attention layer’s paged KV cache. Safetensors and UQFF checkpoints must contain the mtp.* weights; GGUF loading does not support built-in MTP.
Multi-token verification can use different quantized matrix and recurrent kernels from ordinary decoding. Log probabilities and greedy token choices can differ, so token-for-token equivalence is not guaranteed.
Built-in Qwen MTP recomputes cached prompt prefixes, which can increase time to first token for repeated prompts. This restriction does not apply to DFlash.
Qwen3.8-Flash-Next uses the same built-in-head flags:
mistralrs serve -m Qwen/Qwen3.8-Flash-Next --isq q4k --mtpFlash-Next supports only its built-in head. External assistants and DFlash are unsupported. See model-family notes for memory requirements.
DFlash and DFlash 2 (Qwen3.5 / Qwen3.8)
Section titled “DFlash and DFlash 2 (Qwen3.5 / Qwen3.8)”DFlash drafters are small block-diffusion models that propose a
whole block of tokens in one pass, conditioned on hidden states tapped from the target’s
intermediate layers. They reach much higher acceptance than the built-in single-token MTP head.
Pass a DFlash checkpoint (a HF id or local path) as --mtp-model; mistral.rs detects DFlash vs
DFlash 2 from the checkpoint’s architectures and loads the path selector and dynamic convolutions
when present:
mistralrs serve -m Qwen/Qwen3.8-27B --quant 4 \ --mtp-model incoai/Qwen3.8-27B-DFlash2--mtp-draft-sampling defaults to auto. A DFlash 2 checkpoint uses probabilistic drafting when
its CUDA candidate selector is supported, while DFlash 1 and unsupported devices use greedy
drafting. Pass --mtp-draft-sampling greedy to force greedy drafting. Passing probabilistic
requires a supported DFlash 2 selector and fails during loading when that path is unavailable.
The maximum automatic draft depth is min(block_size - 1, 7). During serving, mistral.rs may use a
shallower depth so the recurrent checkpoint layout fits the configured sequence capacity while
preserving memory for the requested paged KV cache. Pass --mtp-n-predict to request a fixed depth
instead. Automatic CUDA serving also adapts the active depth to live load: up to 8 live DFlash
contexts use the full reserved depth, while larger loads use up to 3 drafts. Set
MISTRALRS_DFLASH_ADAPTIVE=0 (or false) to tune depth from measured throughput instead.
That tuner also runs when adaptation by live context count is unavailable. It chooses among
3, 5, and 7 drafts, capped by the reserved depth. The drafter’s projection
weights are requantized to the target’s in-situ quantization type at load (override with the
MISTRALRS_DFLASH_ISQ environment variable, e.g. q6k or none for bf16).
Acceptance is strongly content-dependent: incoai/Qwen3.8-27B-DFlash2 accepts ~7 of 7 drafts on math/code and much less on thinking-mode reasoning traces, which are out of its training distribution.
Gemma 4
Section titled “Gemma 4”Gemma 4 assistant checkpoints are MTP drafters for Gemma 4 target models. See the google/gemma-4-E4B-it-assistant model card for the upstream checkpoint. A downloaded checkout works too:
mistralrs run -m google/gemma-4-E4B-it --quant 8 \ --mtp-model ./gemma-4-E4B-it-assistant \ --mtp-n-predict 6Non-paged KV-cache MTP is intentionally disabled for now, which is why paged attention is required (see the note at the top).
The target and assistant configs must match where the implementation requires it, including vocabulary size and target hidden size. Mismatches fail at load, before generation starts.
Throughput gain depends on how many proposed tokens the target accepts and on the cost of the target verification pass.
MTP supports batched generation and constrained decoding.
MTP is configured at launch time only (CLI flags, Runner(...), or the model builder). There is no per-request HTTP field to toggle it; load the server with the assistant attached and every request uses it.