Skip to content

mistralrs run

Run model in interactive mode, or one-shot mode with -i

mistralrs run [OPTIONS] [COMMAND]
Option Default Description
-m, --model-id <MODEL_ID> Hugging Face model ID or local model directory; optional when -f names local files
-t, --tokenizer <TOKENIZER> Path to local tokenizer.json file
-a, --arch <ARCH> Model architecture (auto-detected if not specified)
--dtype <DTYPE> auto Model data type
--hf-overrides <HF_OVERRIDES> Recursively merged JSON overrides for the Hugging Face model config
--max-model-len <MAX_MODEL_LEN> Runtime model context length
--format <FORMAT> Model format: plain (safetensors), GGUF, or GGML. Auto-detected from -f when not specified. Possible values: plain, gguf, ggml.
-f, --quantized-file <QUANTIZED_FILE> GGUF/GGML filename(s); the suffix selects the format (semicolon-separated for multiple)
--mmproj <MMPROJ> GGUF projector override; auto-selected when unambiguous (semicolon-separated for multiple)
--tok-model-id <TOK_MODEL_ID> Optional model ID overriding configuration, tokenizer, and processor assets for a quantized model
--gqa <GQA> 1 GQA value for GGML models
--enable-lora false Enable dynamic LoRA without preloading an adapter. Supports compatible text and multimodal language models, including GGUF. Vision, audio, and projector adapters are unsupported
--lora <ALIAS=SOURCE|JSON> Preload a language-model LoRA adapter as ALIAS=SOURCE. Supports compatible text and multimodal language models, including GGUF. Remote adapters use revision main. May be repeated. Vision, audio, and projector adapters are unsupported
--lora-max-adapters <LORA_MAX_ADAPTERS> 16 Maximum loaded LoRA aliases and, independently, resident adapter generations
--lora-max-rank <LORA_MAX_RANK> 256 Maximum rank accepted for a LoRA adapter
--lora-max-bytes <BYTES> 8589934592 Maximum memory used by loaded adapters
--legacy-lora <SOURCE> Static LoRA adapter source for GGML or a Phi3 GGUF model
--legacy-lora-order <LEGACY_LORA_ORDER> Ordering JSON file for a legacy raw GGUF or GGML LoRA adapter
--xlora <XLORA> X-LoRA adapter model ID
--xlora-order <XLORA_ORDER> X-LoRA ordering JSON file
--tgt-non-granular-index <TGT_NON_GRANULAR_INDEX> Target non-granular index for X-LoRA
--quant <QUANT> Quantization target. Inference commands select a matching GGUF or UQFF artifact when available. Source checkpoints without a matching UQFF use in-situ quantization. tune evaluates the requested level instead of selecting an artifact. Accepts numeric levels (2, 3, 4, 5, 6, 8) or supported quantization names
--isq <IN_SITU_QUANT> In-situ quantization target. Accepts numeric levels (2, 3, 4, 5, 6, 8) or raw quant names (q4k, q8_0, etc.). Supports compatible GGUF sources
--from-uqff <FROM_UQFF> UQFF artifact to load. Accepts a filename, numeric quantization level (2, 3, 4, 5, 6, 8), or quantization type (q4k, afq8, etc.). Report-declared artifacts and conventional shard names expand to all of their shards. Use semicolons only to list disjoint shards manually
--isq-organization <ISQ_ORGANIZATION> ISQ organization strategy: default or moqe
--imatrix <IMATRIX> imatrix file for enhanced quantization
--calibration-file <CALIBRATION_FILE> Calibration file for imatrix generation
--cpu false Force CPU-only execution
-n, --device-layers <DEVICE_LAYERS> Device layer mapping (format: ORD:NUM;… e.g., “0:10;1:20”) Omit for automatic device mapping
--topology <TOPOLOGY> Topology YAML file for device mapping
--hf-cache <HF_CACHE> Custom Hugging Face cache directory
--max-seq-len <MAX_SEQ_LEN> 4096 Max sequence length for automatic device mapping
--max-batch-size <MAX_BATCH_SIZE> 1 Max batch size for automatic device mapping
--paged-attn <MODE> auto PagedAttention mode - auto: enabled on CUDA, disabled on Metal/CPU (default) - on: force enable (fails if unsupported) - off: force disable. Possible values: auto, on, off.
--pa-context-len <CONTEXT_LEN> Allocate KV cache for this many tokens. Defaults to a memory budget on dedicated GPUs, capped to the model context length on CUDA unified-memory devices
--pa-memory-mb <MEMORY_MB> GPU memory to allocate in MBs (alternative to context-len)
--pa-memory-fraction <MEMORY_FRACTION> GPU memory utilization fraction 0.0-1.0 (alternative to context-len/memory-mb)
--pa-block-size <BLOCK_SIZE> Tokens per block (default: 32 on CUDA)
--pa-cache-type <CACHE_TYPE> auto KV cache quantization type
--encoder-cache-memory-mb <ENCODER_CACHE_MEMORY_MB> Maximum logical tensor memory retained by the multimodal encoder cache, in MiB
--max-edge <MAX_EDGE> Maximum edge length for image resizing (aspect ratio preserved)
--max-num-images <MAX_NUM_IMAGES> Maximum number of images per request
--max-image-length <MAX_IMAGE_LENGTH> Maximum image dimension for device mapping
--max-seqs <MAX_SEQS> 32 Maximum concurrent sequences
--max-num-batched-tokens <MAX_NUM_BATCHED_TOKENS> 4096 Maximum tokens processed in one paged-attention scheduler step
--max-prefill-chunk-tokens <MAX_PREFILL_CHUNK_TOKENS> 512 CUDA prompt-token quantum used while decode is resident and for recurrent prefill batching
--max-decode-steps-before-prefill <MAX_DECODE_STEPS_BEFORE_PREFILL> 8 Maximum decode steps before a waiting prefill batch is admitted
--no-kv-cache false Disable KV cache entirely
--prefix-cache-n <PREFIX_CACHE_N> 16 Number of prefix caches to hold (0 to disable)
-c, --chat-template <CHAT_TEMPLATE> Custom chat template file (.json or .jinja)
-j, --jinja-explicit <JINJA_EXPLICIT> Explicit JINJA template override
--matformer-config-path <MATFORMER_CONFIG_PATH> Path to a MatFormer config (CSV/JSON describing available slices). See model card
--matformer-slice-name <MATFORMER_SLICE_NAME> MatFormer slice to load (must match a slice name in the config file)
--mtp false Enable MTP speculative decoding with the head built into the model checkpoint
--mtp-model <MTP_MODEL> MTP assistant model id or path
--mtp-n-predict <MTP_N_PREDICT> Fixed MTP draft tokens per step; omit to adapt the depth automatically
--mtp-draft-sampling <MTP_DRAFT_SAMPLING> auto MTP draft sampling policy. Auto uses probabilistic DFlash2 drafting when supported. Possible values: auto, greedy, probabilistic.
--mcp-config <MCP_CONFIG> Path to an MCP client configuration JSON. Also reads MCP_CONFIG_PATH if unset
--agent false Build a local agent: enables web search, Python code execution, and shell execution, runs the agentic tool loop with a per-session temp workdir. Equivalent to passing --enable-search --enable-code-execution --enable-shell together
--enable-search false Enable web search (requires embedding model)
--search-embedding-model <SEARCH_EMBEDDING_MODEL> Search embedding model to use. Requires --enable-search or --agent. Possible values: embedding-gemma.
--enable-code-execution false Enable Python code execution tool (WARNING: allows arbitrary code execution)
--enable-shell false Enable shell execution tool (WARNING: allows arbitrary command execution)
--code-exec-python <CODE_EXEC_PYTHON> Python interpreter path for code execution. Requires code execution to be on (via --enable-code-execution or --agent). Defaults to python3
--code-exec-timeout <CODE_EXEC_TIMEOUT> Code execution timeout in seconds (default: 60). Requires code execution to be on
--code-exec-workdir <CODE_EXEC_WORKDIR> Working directory for code execution. Defaults to a temp dir; use “.” for cwd. Requires code execution to be on
--shell-path <SHELL_PATH> Shell executable path. Requires shell execution to be on. Defaults to /bin/sh
--shell-timeout <SHELL_TIMEOUT> Shell execution timeout in seconds (default: 600). Requires shell execution to be on
--shell-workdir <SHELL_WORKDIR> Root directory for per-session shell working directories. Defaults to temp dirs
--skills-dir <SKILLS_DIR> Directory for uploaded OpenAI-compatible Skills. Defaults to the system temp directory
--agent-permission <PERMISSION> auto Agent action permission mode. Possible values: auto, ask, deny.
--sandbox <MODE> auto Sandbox mode. Possible values: auto, on, off.
--sandbox-profile <PROFILE> Sandbox policy profile. Possible values: restricted, developer.
--sb-max-memory-mb <MEMORY_MB> Per-session memory cap in MiB (default: 2048)
--sb-max-cpu-secs <CPU_SECS> Per-session CPU time cap in seconds (default: 600). Raised to at least enabled code/shell timeouts
--sb-max-procs <PROCS> Per-session process/thread cap (default: 64)
--sandbox-network <NETWORK> Network access permitted to the sandboxed session. Possible values: none, loopback, full.
--thinking <THINKING> Control thinking mode for models that support it. Use –thinking or –thinking true to force on, –thinking false to force off. If both reasoning controls are omitted, effort is unspecified and thinking is enabled. Possible values: true, false.
--reasoning-effort <REASONING_EFFORT> Set reasoning effort without changing the model’s sampling parameters. Values are off, low, medium, high, or xhigh. “none” is an alias for off
-i, --input <INPUT> One-shot text prompt. When provided, sends a single request and exits instead of entering interactive mode. Combine with –image, –video, or –audio for multimodal requests
--image <IMAGE> Image URL(s) or file path(s) to include in the request (requires -i). Can be specified multiple times: –image img1.jpg –image img2.png
--video <VIDEO> Video URL(s) or file path(s) to include in the request (requires -i). Can be specified multiple times: –video vid1.mp4 –video vid2.webm
--audio <AUDIO> Audio URL(s) or file path(s) to include in the request (requires -i). Can be specified multiple times: –audio audio1.wav –audio audio2.mp3
--adapter <ADAPTER> LoRA adapter alias to use for requests. Omit to run the base model

Auto-detect model type (recommended)

mistralrs run auto [OPTIONS] --model-id <MODEL_ID>
Option Default Description
-m, --model-id <MODEL_ID> required Hugging Face model ID or local path to model directory
-t, --tokenizer <TOKENIZER> Path to local tokenizer.json file
-a, --arch <ARCH> Model architecture (auto-detected if not specified)
--dtype <DTYPE> auto Model data type
--hf-overrides <HF_OVERRIDES> Recursively merged JSON overrides for the Hugging Face model config
--max-model-len <MAX_MODEL_LEN> Runtime model context length
--format <FORMAT> Model format: plain (safetensors), GGUF, or GGML. Auto-detected from -f when not specified. Possible values: plain, gguf, ggml.
-f, --quantized-file <QUANTIZED_FILE> GGUF/GGML filename(s); the suffix selects the format (semicolon-separated for multiple)
--mmproj <MMPROJ> GGUF projector override; auto-selected when unambiguous (semicolon-separated for multiple)
--tok-model-id <TOK_MODEL_ID> Optional model ID overriding configuration, tokenizer, and processor assets for a quantized model
--gqa <GQA> 1 GQA value for GGML models
--enable-lora false Enable dynamic LoRA without preloading an adapter. Supports compatible text and multimodal language models, including GGUF. Vision, audio, and projector adapters are unsupported
--lora <ALIAS=SOURCE|JSON> Preload a language-model LoRA adapter as ALIAS=SOURCE. Supports compatible text and multimodal language models, including GGUF. Remote adapters use revision main. May be repeated. Vision, audio, and projector adapters are unsupported
--lora-max-adapters <LORA_MAX_ADAPTERS> 16 Maximum loaded LoRA aliases and, independently, resident adapter generations
--lora-max-rank <LORA_MAX_RANK> 256 Maximum rank accepted for a LoRA adapter
--lora-max-bytes <BYTES> 8589934592 Maximum memory used by loaded adapters
--legacy-lora <SOURCE> Static LoRA adapter source for GGML or a Phi3 GGUF model
--legacy-lora-order <LEGACY_LORA_ORDER> Ordering JSON file for a legacy raw GGUF or GGML LoRA adapter
--xlora <XLORA> X-LoRA adapter model ID
--xlora-order <XLORA_ORDER> X-LoRA ordering JSON file
--tgt-non-granular-index <TGT_NON_GRANULAR_INDEX> Target non-granular index for X-LoRA
--quant <QUANT> Quantization target. Inference commands select a matching GGUF or UQFF artifact when available. Source checkpoints without a matching UQFF use in-situ quantization. tune evaluates the requested level instead of selecting an artifact. Accepts numeric levels (2, 3, 4, 5, 6, 8) or supported quantization names
--isq <IN_SITU_QUANT> In-situ quantization target. Accepts numeric levels (2, 3, 4, 5, 6, 8) or raw quant names (q4k, q8_0, etc.). Supports compatible GGUF sources
--from-uqff <FROM_UQFF> UQFF artifact to load. Accepts a filename, numeric quantization level (2, 3, 4, 5, 6, 8), or quantization type (q4k, afq8, etc.). Report-declared artifacts and conventional shard names expand to all of their shards. Use semicolons only to list disjoint shards manually
--isq-organization <ISQ_ORGANIZATION> ISQ organization strategy: default or moqe
--imatrix <IMATRIX> imatrix file for enhanced quantization
--calibration-file <CALIBRATION_FILE> Calibration file for imatrix generation
--cpu false Force CPU-only execution
-n, --device-layers <DEVICE_LAYERS> Device layer mapping (format: ORD:NUM;… e.g., “0:10;1:20”) Omit for automatic device mapping
--topology <TOPOLOGY> Topology YAML file for device mapping
--hf-cache <HF_CACHE> Custom Hugging Face cache directory
--max-seq-len <MAX_SEQ_LEN> 4096 Max sequence length for automatic device mapping
--max-batch-size <MAX_BATCH_SIZE> 1 Max batch size for automatic device mapping
--paged-attn <MODE> auto PagedAttention mode - auto: enabled on CUDA, disabled on Metal/CPU (default) - on: force enable (fails if unsupported) - off: force disable. Possible values: auto, on, off.
--pa-context-len <CONTEXT_LEN> Allocate KV cache for this many tokens. Defaults to a memory budget on dedicated GPUs, capped to the model context length on CUDA unified-memory devices
--pa-memory-mb <MEMORY_MB> GPU memory to allocate in MBs (alternative to context-len)
--pa-memory-fraction <MEMORY_FRACTION> GPU memory utilization fraction 0.0-1.0 (alternative to context-len/memory-mb)
--pa-block-size <BLOCK_SIZE> Tokens per block (default: 32 on CUDA)
--pa-cache-type <CACHE_TYPE> auto KV cache quantization type
--encoder-cache-memory-mb <ENCODER_CACHE_MEMORY_MB> Maximum logical tensor memory retained by the multimodal encoder cache, in MiB
--max-edge <MAX_EDGE> Maximum edge length for image resizing (aspect ratio preserved)
--max-num-images <MAX_NUM_IMAGES> Maximum number of images per request
--max-image-length <MAX_IMAGE_LENGTH> Maximum image dimension for device mapping

Text generation model with explicit configuration

mistralrs run text [OPTIONS] --model-id <MODEL_ID>
Option Default Description
-m, --model-id <MODEL_ID> required Hugging Face model ID or local path to model directory
-t, --tokenizer <TOKENIZER> Path to local tokenizer.json file
-a, --arch <ARCH> Model architecture (auto-detected if not specified)
--dtype <DTYPE> auto Model data type
--hf-overrides <HF_OVERRIDES> Recursively merged JSON overrides for the Hugging Face model config
--max-model-len <MAX_MODEL_LEN> Runtime model context length
--format <FORMAT> Model format: plain (safetensors), GGUF, or GGML. Auto-detected from -f when not specified. Possible values: plain, gguf, ggml.
-f, --quantized-file <QUANTIZED_FILE> GGUF/GGML filename(s); the suffix selects the format (semicolon-separated for multiple)
--mmproj <MMPROJ> GGUF projector override; auto-selected when unambiguous (semicolon-separated for multiple)
--tok-model-id <TOK_MODEL_ID> Optional model ID overriding configuration, tokenizer, and processor assets for a quantized model
--gqa <GQA> 1 GQA value for GGML models
--enable-lora false Enable dynamic LoRA without preloading an adapter. Supports compatible text and multimodal language models, including GGUF. Vision, audio, and projector adapters are unsupported
--lora <ALIAS=SOURCE|JSON> Preload a language-model LoRA adapter as ALIAS=SOURCE. Supports compatible text and multimodal language models, including GGUF. Remote adapters use revision main. May be repeated. Vision, audio, and projector adapters are unsupported
--lora-max-adapters <LORA_MAX_ADAPTERS> 16 Maximum loaded LoRA aliases and, independently, resident adapter generations
--lora-max-rank <LORA_MAX_RANK> 256 Maximum rank accepted for a LoRA adapter
--lora-max-bytes <BYTES> 8589934592 Maximum memory used by loaded adapters
--legacy-lora <SOURCE> Static LoRA adapter source for GGML or a Phi3 GGUF model
--legacy-lora-order <LEGACY_LORA_ORDER> Ordering JSON file for a legacy raw GGUF or GGML LoRA adapter
--xlora <XLORA> X-LoRA adapter model ID
--xlora-order <XLORA_ORDER> X-LoRA ordering JSON file
--tgt-non-granular-index <TGT_NON_GRANULAR_INDEX> Target non-granular index for X-LoRA
--quant <QUANT> Quantization target. Inference commands select a matching GGUF or UQFF artifact when available. Source checkpoints without a matching UQFF use in-situ quantization. tune evaluates the requested level instead of selecting an artifact. Accepts numeric levels (2, 3, 4, 5, 6, 8) or supported quantization names
--isq <IN_SITU_QUANT> In-situ quantization target. Accepts numeric levels (2, 3, 4, 5, 6, 8) or raw quant names (q4k, q8_0, etc.). Supports compatible GGUF sources
--from-uqff <FROM_UQFF> UQFF artifact to load. Accepts a filename, numeric quantization level (2, 3, 4, 5, 6, 8), or quantization type (q4k, afq8, etc.). Report-declared artifacts and conventional shard names expand to all of their shards. Use semicolons only to list disjoint shards manually
--isq-organization <ISQ_ORGANIZATION> ISQ organization strategy: default or moqe
--imatrix <IMATRIX> imatrix file for enhanced quantization
--calibration-file <CALIBRATION_FILE> Calibration file for imatrix generation
--cpu false Force CPU-only execution
-n, --device-layers <DEVICE_LAYERS> Device layer mapping (format: ORD:NUM;… e.g., “0:10;1:20”) Omit for automatic device mapping
--topology <TOPOLOGY> Topology YAML file for device mapping
--hf-cache <HF_CACHE> Custom Hugging Face cache directory
--max-seq-len <MAX_SEQ_LEN> 4096 Max sequence length for automatic device mapping
--max-batch-size <MAX_BATCH_SIZE> 1 Max batch size for automatic device mapping
--paged-attn <MODE> auto PagedAttention mode - auto: enabled on CUDA, disabled on Metal/CPU (default) - on: force enable (fails if unsupported) - off: force disable. Possible values: auto, on, off.
--pa-context-len <CONTEXT_LEN> Allocate KV cache for this many tokens. Defaults to a memory budget on dedicated GPUs, capped to the model context length on CUDA unified-memory devices
--pa-memory-mb <MEMORY_MB> GPU memory to allocate in MBs (alternative to context-len)
--pa-memory-fraction <MEMORY_FRACTION> GPU memory utilization fraction 0.0-1.0 (alternative to context-len/memory-mb)
--pa-block-size <BLOCK_SIZE> Tokens per block (default: 32 on CUDA)
--pa-cache-type <CACHE_TYPE> auto KV cache quantization type

Multimodal model

mistralrs run multimodal [OPTIONS] --model-id <MODEL_ID>
Option Default Description
-m, --model-id <MODEL_ID> required Hugging Face model ID or local path to model directory
-t, --tokenizer <TOKENIZER> Path to local tokenizer.json file
-a, --arch <ARCH> Model architecture (auto-detected if not specified)
--dtype <DTYPE> auto Model data type
--hf-overrides <HF_OVERRIDES> Recursively merged JSON overrides for the Hugging Face model config
--max-model-len <MAX_MODEL_LEN> Runtime model context length
--format <FORMAT> Model format: plain (safetensors), GGUF, or GGML. Auto-detected from -f when not specified. Possible values: plain, gguf, ggml.
-f, --quantized-file <QUANTIZED_FILE> GGUF/GGML filename(s); the suffix selects the format (semicolon-separated for multiple)
--mmproj <MMPROJ> GGUF projector override; auto-selected when unambiguous (semicolon-separated for multiple)
--tok-model-id <TOK_MODEL_ID> Optional model ID overriding configuration, tokenizer, and processor assets for a quantized model
--gqa <GQA> 1 GQA value for GGML models
--enable-lora false Enable dynamic LoRA for the language model without preloading an adapter. Vision, audio, and projector adapters are unsupported
--lora <ALIAS=SOURCE|JSON> Preload a language-model LoRA adapter as ALIAS=SOURCE. Remote adapters use revision main. May be repeated. Vision, audio, and projector adapters are unsupported
--lora-max-adapters <LORA_MAX_ADAPTERS> 16 Maximum loaded LoRA aliases and, independently, resident adapter generations
--lora-max-rank <LORA_MAX_RANK> 256 Maximum rank accepted for a LoRA adapter
--lora-max-bytes <BYTES> 8589934592 Maximum memory used by loaded adapters
--quant <QUANT> Quantization target. Inference commands select a matching GGUF or UQFF artifact when available. Source checkpoints without a matching UQFF use in-situ quantization. tune evaluates the requested level instead of selecting an artifact. Accepts numeric levels (2, 3, 4, 5, 6, 8) or supported quantization names
--isq <IN_SITU_QUANT> In-situ quantization target. Accepts numeric levels (2, 3, 4, 5, 6, 8) or raw quant names (q4k, q8_0, etc.). Supports compatible GGUF sources
--from-uqff <FROM_UQFF> UQFF artifact to load. Accepts a filename, numeric quantization level (2, 3, 4, 5, 6, 8), or quantization type (q4k, afq8, etc.). Report-declared artifacts and conventional shard names expand to all of their shards. Use semicolons only to list disjoint shards manually
--isq-organization <ISQ_ORGANIZATION> ISQ organization strategy: default or moqe
--imatrix <IMATRIX> imatrix file for enhanced quantization
--calibration-file <CALIBRATION_FILE> Calibration file for imatrix generation
--cpu false Force CPU-only execution
-n, --device-layers <DEVICE_LAYERS> Device layer mapping (format: ORD:NUM;… e.g., “0:10;1:20”) Omit for automatic device mapping
--topology <TOPOLOGY> Topology YAML file for device mapping
--hf-cache <HF_CACHE> Custom Hugging Face cache directory
--max-seq-len <MAX_SEQ_LEN> 4096 Max sequence length for automatic device mapping
--max-batch-size <MAX_BATCH_SIZE> 1 Max batch size for automatic device mapping
--paged-attn <MODE> auto PagedAttention mode - auto: enabled on CUDA, disabled on Metal/CPU (default) - on: force enable (fails if unsupported) - off: force disable. Possible values: auto, on, off.
--pa-context-len <CONTEXT_LEN> Allocate KV cache for this many tokens. Defaults to a memory budget on dedicated GPUs, capped to the model context length on CUDA unified-memory devices
--pa-memory-mb <MEMORY_MB> GPU memory to allocate in MBs (alternative to context-len)
--pa-memory-fraction <MEMORY_FRACTION> GPU memory utilization fraction 0.0-1.0 (alternative to context-len/memory-mb)
--pa-block-size <BLOCK_SIZE> Tokens per block (default: 32 on CUDA)
--pa-cache-type <CACHE_TYPE> auto KV cache quantization type
--encoder-cache-memory-mb <ENCODER_CACHE_MEMORY_MB> Maximum logical tensor memory retained by the multimodal encoder cache, in MiB
--max-edge <MAX_EDGE> Maximum edge length for image resizing (aspect ratio preserved)
--max-num-images <MAX_NUM_IMAGES> Maximum number of images per request
--max-image-length <MAX_IMAGE_LENGTH> Maximum image dimension for device mapping

Image generation model (diffusion)

mistralrs run diffusion [OPTIONS] --model-id <MODEL_ID>
Option Default Description
-m, --model-id <MODEL_ID> required Hugging Face model ID or local path to model directory
-t, --tokenizer <TOKENIZER> Path to local tokenizer.json file
-a, --arch <ARCH> Model architecture (auto-detected if not specified)
--dtype <DTYPE> auto Model data type
--hf-overrides <HF_OVERRIDES> Recursively merged JSON overrides for the Hugging Face model config
--max-model-len <MAX_MODEL_LEN> Runtime model context length
--cpu false Force CPU-only execution
-n, --device-layers <DEVICE_LAYERS> Device layer mapping (format: ORD:NUM;… e.g., “0:10;1:20”) Omit for automatic device mapping
--topology <TOPOLOGY> Topology YAML file for device mapping
--hf-cache <HF_CACHE> Custom Hugging Face cache directory
--max-seq-len <MAX_SEQ_LEN> 4096 Max sequence length for automatic device mapping
--max-batch-size <MAX_BATCH_SIZE> 1 Max batch size for automatic device mapping

Speech synthesis model

mistralrs run speech [OPTIONS] --model-id <MODEL_ID>
Option Default Description
-m, --model-id <MODEL_ID> required Hugging Face model ID or local path to model directory
-t, --tokenizer <TOKENIZER> Path to local tokenizer.json file
-a, --arch <ARCH> Model architecture (auto-detected if not specified)
--dtype <DTYPE> auto Model data type
--hf-overrides <HF_OVERRIDES> Recursively merged JSON overrides for the Hugging Face model config
--max-model-len <MAX_MODEL_LEN> Runtime model context length
--cpu false Force CPU-only execution
-n, --device-layers <DEVICE_LAYERS> Device layer mapping (format: ORD:NUM;… e.g., “0:10;1:20”) Omit for automatic device mapping
--topology <TOPOLOGY> Topology YAML file for device mapping
--hf-cache <HF_CACHE> Custom Hugging Face cache directory
--max-seq-len <MAX_SEQ_LEN> 4096 Max sequence length for automatic device mapping
--max-batch-size <MAX_BATCH_SIZE> 1 Max batch size for automatic device mapping

Embedding model

mistralrs run embedding [OPTIONS] --model-id <MODEL_ID>
Option Default Description
-m, --model-id <MODEL_ID> required Hugging Face model ID or local path to model directory
-t, --tokenizer <TOKENIZER> Path to local tokenizer.json file
-a, --arch <ARCH> Model architecture (auto-detected if not specified)
--dtype <DTYPE> auto Model data type
--hf-overrides <HF_OVERRIDES> Recursively merged JSON overrides for the Hugging Face model config
--max-model-len <MAX_MODEL_LEN> Runtime model context length
--format <FORMAT> Model format: plain (safetensors), GGUF, or GGML. Auto-detected from -f when not specified. Possible values: plain, gguf, ggml.
-f, --quantized-file <QUANTIZED_FILE> GGUF/GGML filename(s); the suffix selects the format (semicolon-separated for multiple)
--mmproj <MMPROJ> GGUF projector override; auto-selected when unambiguous (semicolon-separated for multiple)
--tok-model-id <TOK_MODEL_ID> Optional model ID overriding configuration, tokenizer, and processor assets for a quantized model
--gqa <GQA> 1 GQA value for GGML models
--quant <QUANT> Quantization target. Inference commands select a matching GGUF or UQFF artifact when available. Source checkpoints without a matching UQFF use in-situ quantization. tune evaluates the requested level instead of selecting an artifact. Accepts numeric levels (2, 3, 4, 5, 6, 8) or supported quantization names
--isq <IN_SITU_QUANT> In-situ quantization target. Accepts numeric levels (2, 3, 4, 5, 6, 8) or raw quant names (q4k, q8_0, etc.). Supports compatible GGUF sources
--from-uqff <FROM_UQFF> UQFF artifact to load. Accepts a filename, numeric quantization level (2, 3, 4, 5, 6, 8), or quantization type (q4k, afq8, etc.). Report-declared artifacts and conventional shard names expand to all of their shards. Use semicolons only to list disjoint shards manually
--isq-organization <ISQ_ORGANIZATION> ISQ organization strategy: default or moqe
--imatrix <IMATRIX> imatrix file for enhanced quantization
--calibration-file <CALIBRATION_FILE> Calibration file for imatrix generation
--cpu false Force CPU-only execution
-n, --device-layers <DEVICE_LAYERS> Device layer mapping (format: ORD:NUM;… e.g., “0:10;1:20”) Omit for automatic device mapping
--topology <TOPOLOGY> Topology YAML file for device mapping
--hf-cache <HF_CACHE> Custom Hugging Face cache directory
--max-seq-len <MAX_SEQ_LEN> 4096 Max sequence length for automatic device mapping
--max-batch-size <MAX_BATCH_SIZE> 1 Max batch size for automatic device mapping
--paged-attn <MODE> auto PagedAttention mode - auto: enabled on CUDA, disabled on Metal/CPU (default) - on: force enable (fails if unsupported) - off: force disable. Possible values: auto, on, off.
--pa-context-len <CONTEXT_LEN> Allocate KV cache for this many tokens. Defaults to a memory budget on dedicated GPUs, capped to the model context length on CUDA unified-memory devices
--pa-memory-mb <MEMORY_MB> GPU memory to allocate in MBs (alternative to context-len)
--pa-memory-fraction <MEMORY_FRACTION> GPU memory utilization fraction 0.0-1.0 (alternative to context-len/memory-mb)
--pa-block-size <BLOCK_SIZE> Tokens per block (default: 32 on CUDA)
--pa-cache-type <CACHE_TYPE> auto KV cache quantization type