Skip to content

Model family notes

Most models need nothing beyond mistralrs run -m <id> (see Run models). This page collects the per-family exceptions. The full architecture inventory is in the supported models reference.

reasoning_effort accepts off, low, medium, high, and xhigh; none is an alias for off. If no effort is set, mistral.rs leaves it unspecified and enables thinking. An explicit positive effort enables thinking, while off disables it. The legacy enable_thinking boolean remains supported, but contradictory explicit values return an error.

These controls are template inputs, not sampling parameters. mistral.rs passes an explicit effort under both reasoning_effort and the template compatibility name reasoning_strength. Each model’s template decides whether it distinguishes every positive tier.

Qwen3 and SmolLM3 are hybrid reasoning models; their chat templates enable thinking by default. Toggle it per request. Inline /think and /no_think prompt tags work everywhere; the --thinking flag and the enable_thinking field do the same without editing user text (true forces on and false forces off).

Terminal window
mistralrs run --thinking false -m Qwen/Qwen3-4B

--thinking and --reasoning-effort apply to both one-shot and interactive use. Prompt tags also work inline:

How many rs are in blueberry? /no_think
Are you sure? /think

Qwen3 also publishes FP8 pre-quantized checkpoints; pass the FP8 model ID directly when you want those weights instead of runtime ISQ (in-situ quantization). For example, Qwen3.8 can serve published block-FP8 weights and an FP8 paged KV cache together:

Terminal window
mistralrs serve -m Qwen/Qwen3.8-27B-FP8 --pa-cache-type f8e4m3 --max-seqs 16

The checkpoint’s dynamic activation scaling and modules_to_not_convert declarations are applied automatically; --quant is not needed for a pre-quantized FP8 repository.

MoE (Mixture of Experts) families (DeepSeek V2/V3, GLM-4.7, GLM-4.7-Flash, Phi 3.5 MoE, Qwen3 MoE, Qwen3-VL MoE, Qwen3.5/Qwen3.6 MoE, LFM2/LFM2.5 MoE) support MoQE (Mixture of Quantized Experts): quantizing only the routed experts, which dominate memory, while leaving the rest of the model alone. Enable it when applying ISQ with --isq-organization moqe:

Terminal window
mistralrs run --isq 4 --isq-organization moqe -m Qwen/Qwen3-30B-A3B

MoQE also applies when generating UQFF with mistralrs quantize. It does not rewrite an already-quantized GGUF or UQFF selected by --quant; those weights load as published.

In the Python SDK, pass organization=IsqOrganization.MoQE inside Which.Plain(...) or Which.MultimodalPlain(...). Expect small output differences between quantization levels: router decisions are sensitive to numerical noise.

MLA models (DeepSeek V2/V3, GLM-4.7-Flash)

Section titled “MLA models (DeepSeek V2/V3, GLM-4.7-Flash)”

DeepSeek V2, DeepSeek V3 (including non-distill R1, which uses the V3 architecture), and GLM-4.7-Flash use MLA (Multi-head Latent Attention). The KV cache stores a low-dimensional latent instead of full K/V, so the cache footprint is substantially smaller than standard attention at the same context length.

On CUDA (Unix builds), a specialized MLA decode kernel is used when all of the following hold:

  • single-token decode (no attention mask, sequence length 1);
  • paged attention enabled;
  • FlashInfer (NVIDIA’s attention-kernel library) paged metadata available.

A parallel fast path covers prefill with prefix caching (paged attention on, CUDA device). Otherwise the generic attention path reconstructs the latent per step.

MISTRALRS_NO_MLA=1 forces the generic path; use it when debugging suspected MLA kernel issues, and try --paged-attn off as a sanity check for unexpected paged-attention behavior. Background: the DeepSeek V2 paper.

GPT-OSS experts are stored pre-quantized in MXFP4 (4-bit microscaling float), and its attention uses per-head sinks. Load it without a quantization flag first:

Terminal window
mistralrs run -m openai/gpt-oss-20b

ISQ applies only to the attention layers (and lm_head); the expert weights are already quantized.

Qwen3 Next mixes Gated Delta Network (linear attention) layers with full softmax attention, so its cost profile at long contexts differs from a pure softmax model. Qwen3-Coder-Next checkpoints use the same loader.

Qwen3.5 and Qwen3.6 dense and MoE checkpoints are supported. Qwen3.6 shares the Qwen3.5 HF model_type, so it loads through the same Qwen3_5 and Qwen3_5Moe paths.

Use nested Hugging Face config overrides to enable YaRN without editing the downloaded checkpoint. This example extends a native 262,144-token Qwen3.5 model to a 1,010,000-token serving limit:

Terminal window
mistralrs serve -m Qwen/Qwen3.5-27B \
--hf-overrides '{"text_config":{"rope_parameters":{"rope_type":"yarn","factor":4.0,"original_max_position_embeddings":262144}}}' \
--max-model-len 1010000

Overrides merge recursively, so the checkpoint’s rope_theta, partial_rotary_factor, and MRoPE sections remain intact. --max-model-len only sets the prompt-plus-output limit; it does not enable RoPE scaling and cannot exceed the context length supported by the final config. Static YaRN can reduce quality on shorter inputs, so enable it only when the extended context is needed.

Qwen3.8-Flash-Next (HF qwen4_exp) adds block-sparse attention (QSA), hyper-connections, and a hashed n-gram embedding (PLE) to the Qwen3.5 MoE hybrid. The checkpoint is roughly 330 GB, including a 100 GB bf16 PLE table. Use ISQ to reduce memory use:

Terminal window
mistralrs serve -m Qwen/Qwen3.8-Flash-Next --isq q4k
  • On CUDA, the safetensors PLE table is quantized separately at startup: 8-bit if it fits in device memory, otherwise 4-bit (about 29 GB). This reads the whole table once and adds to startup time. If neither fits, rows are read from the memory-mapped checkpoint; CUDA decode graphs then require pageable memory access (ATS/HMM). CPU and Metal gather mapped rows on the host.
  • QSA preserves full causal attention through 2,051 tokens. Longer contexts use QSA over the paged KV cache: each query selects 512 four-token blocks plus the incomplete tail, for at most 2,051 tokens. QSA also needs a small per-token indexer cache.
  • Paged attention is CUDA-only for this model. On other devices, use --paged-attn off. FP8 KV caches and tensor parallelism are not supported.
  • Add --mtp for built-in speculative decoding with safetensors or UQFF checkpoints containing mtp.* weights. Draft depth adapts among 2, 3, 4, and 6 tokens; --mtp-n-predict N fixes it. MTP needs extra weights, one attention layer’s KV cache, and recurrent verification state.
  • Built-in MTP recomputes cached prompt prefixes, so repeated prompts can take longer to prefill. GGUF loading does not support MTP. External MTP and DFlash drafters are also unsupported for this model.
  • qwen4exp GGUF files load without config.json. --quant selects a variant and its vision projector for image and video input. On CUDA, the IQ4_NL PLE table is copied into device memory without requantization if it fits. Otherwise, rows are gathered on the host and CUDA decode graphs are disabled.
Terminal window
mistralrs run -m unsloth/Qwen3.8-Flash-Next-GGUF --quant 4

On CUDA devices with unified memory, such as DGX Spark, the default memory budget is available system RAM minus 1 GiB. The KV cache is capped to the model context length. Use --pa-context-len N for a smaller cache to leave room for other applications or MTP; use --max-model-len N to limit prompt-plus-output length. MISTRALRS_IGPU_MEMORY_FRACTION overrides the memory budget.

IBM Granite 4.0 checkpoints (e.g. ibm-granite/granite-4.0-micro) mix Mamba-2 recurrent layers with attention layers. They load through auto-detection like any other text model.

LiquidAI LFM2 and LFM2.5 dense and MoE checkpoints are supported. LFM2-VL and LFM2.5-VL are supported for image input.

The official repos ship chat_template.jinja, including the Liquid tool-call format. It is auto-detected like other model-native tool-call formats.

Gemma repos are gated: accept the license on the Hugging Face model page, then authenticate with mistralrs login.

Gemma 4 accepts image, audio, and video parts mixed in one message, and enforces its tool-call format through constrained decoding by default; see tool calling.

Muse Glimmer loads through normal model auto-detection:

Terminal window
mistralrs run -m meta-models/Muse-Glimmer-30B

The native safetensors checkpoint accepts text, images, and video; it does not accept audio. Video is represented as timestamped frame groups, and the model card notes that the model was trained primarily for image rather than video understanding.

The model’s ATEM chat template is detected automatically. Tool calling supports automatic, required, named, and parallel calls through the same HTTP, Python, Rust, and agent APIs as other tool-capable models. Muse reads the reasoning_strength template variable, so mistral.rs maps the public reasoning_effort control to it. The published Muse tiers are low, medium, high, and xhigh; omission uses the template’s high default.

ISQ, UQFF, paged attention, prefix caching, tensor parallelism, and text-backbone LoRA use the standard multimodal model paths. For GGUF-specific requirements and its image-only projector limitation, see GGUF compatibility.

MatFormer-trained models encode multiple model sizes in one checkpoint; the desired slice is selected at load time with two values:

  • matformer_config_path: path to the slice config file (CSV or JSON) shipped with the model card.
  • matformer_slice_name: the named slice within that file.

Without these, the default (full) configuration loads. Gemma 3n (google/gemma-3n-E4B-it) is the MatFormer model in the supported list; the bundled matformer_configs/gemma3n.csv contains the full E4B configuration, the official E2B slice, and intermediate E1.96B-E3.79B slices:

Terminal window
mistralrs run -m google/gemma-3n-E4B-it \
--matformer-config-path matformer_configs/gemma3n.csv \
--matformer-slice-name "Config for E2.49B (block-level)"

The same slice selection is available on every surface:

  • CLI: --matformer-config-path / --matformer-slice-name on run, serve, and bench.
  • TOML configs: matformer_config_path / matformer_slice_name.
  • Python: matformer_config_path / matformer_slice_name on the Which selectors.
  • Rust SDK: with_matformer_config_path / with_matformer_slice_name on the model builders.

Use the full configuration for quality and smaller slices for constrained devices.

Mistral Small 3 checkpoints can do tool calling, but some repos do not ship the right chat template. Use the bundled one:

Terminal window
mistralrs serve --quant 4 \
--jinja-explicit chat_templates/mistral_small_tool_call.jinja \
-m mistralai/Mistral-Small-3.2-24B-Instruct-2506

Mistral-backed LLaVA checkpoints work with the default template. Vicuna-backed checkpoints need the Vicuna template:

Terminal window
mistralrs run -m llava-hf/llava-v1.6-vicuna-7b-hf \
-c chat_templates/vicuna.json --image photo.jpg -i "Describe this image"

Per-request video frame-sampling overrides are not exposed. In multi-turn conversations reusing prefix cache entries, pixel inputs are narrowed per turn by grid count, not image count.

Llama 4 Scout supports up to 10M tokens of context. Using the full window requires paged attention with a large memory budget, generally with multi-GPU tensor parallelism.

For most multimodal models the text backbone holds most of the parameters, so device mapping and topology apply mainly to the text portion; the vision, audio, or video encoder stays on its supported device path.

Phi 3.5 Vision works best with a single image; multiple images are resized together. Phi 4 Multimodal accepts audio and image parts in the same message. Phi 3.5 MoE has 16 experts and routes each token to 2 of them; it benefits from MoQE.

Block-diffusion generation has its own page: block-diffusion models.

File an issue on GitHub, with a reproducer when possible.