diff --git a/docs/cookbook/diffusion/MiniMax/MiniMax-H3.mdx b/docs/cookbook/diffusion/MiniMax/MiniMax-H3.mdx index 187cfeba6..b8225dd0a 100644 --- a/docs/cookbook/diffusion/MiniMax/MiniMax-H3.mdx +++ b/docs/cookbook/diffusion/MiniMax/MiniMax-H3.mdx @@ -47,11 +47,13 @@ Choose H3 when synchronized audiovisual output or reference-driven generation ma | --- | --- | --- | | Text to video and audio | `t2va` | Text prompt only | | First/last frame to video and audio | `fl2va` | First frame, last frame, or both | -| Reference to video and audio | `ref2va` | Image, video, and audio references | +| Reference to video and audio | `ref2va` | Image, video, and audio references, optionally combined with first/last keyframes | Video-to-video (V2V) is a supported `ref2va` use case, not a fourth task value. Run the `Ref2VA` partition and provide a video reference in -`conditions`. +`conditions`. A hybrid `ref2va` request may also include the same ordered +first/last keyframes accepted by `fl2va`, but it must still contain at least one +reference condition. Use the selected Hub's root model ID: `MiniMaxAI/MiniMax-H3` on Hugging Face or `MiniMax/MiniMax-H3` on ModelScope. Select the checkpoint variant with @@ -70,6 +72,48 @@ The builder accepts legal custom GPU counts and topologies, marking them Static H3 head or partition violations disable Copy before they reach `sglang serve`. +### Checkpoint and adapter formats + +Start with the command emitted by the builder. Every row below is an overlay on +the same native SGLang pipeline; component repositories contribute their own +config and weights, while weight files retain the base component config. + +| Scope | Format or variant | Add to the base command | Contract | +| --- | --- | --- | --- | +| Full model | Official mixed BF16/FP32 | `--model-variant fl2va` or `--model-variant ref2va` | Lossless reference and consistency GT path. | +| DiT | Official Diffusers component layout | `--component-paths.transformer MiniMaxAI/MiniMax-H3/transformer` (`fl2va`) or `.../transformer_ref` (`ref2va`) | Loads the official component through the native SGLang graph; no Diffusers runtime fallback. | +| DiT | [AdaLN-pruned Diffusers component](https://huggingface.co/multimodalart/MiniMax-H3-Pruned) | `--component-paths.transformer multimodalart/MiniMax-H3-Pruned/transformer` or `.../transformer_ref` | Approximate curve-AdaLN architecture; its config and basis metadata are loaded natively. | +| DiT | [Full or AdaLN-pruned BF16 safetensors](https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/diffusion_models) | `--component-weights-paths.transformer OWNER/REPO/path/FILE.safetensors` | Weight-only override. Match the FL2VA/Ref2VA partition; pruned or dtype-converted exports are approximate. | +| DiT | Comfy FP8 or self-describing MXFP8 safetensors | `--component-weights-paths.transformer OWNER/REPO/path/FILE.safetensors` | Per-layer metadata selects static/dynamic FP8 or MXFP8 automatically. | +| DiT | ConvRot INT8, W4A8, W4A4, or mixed W4A4+INT8 safetensors ([INT8](https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/diffusion_models), [W4A8](https://huggingface.co/Winnougan/MiniMax-H3-INT4_Convrot_ComfyUI), [W4A4](https://huggingface.co/Merserk/MiniMax-H3-INT4-ConvRot)) | `--component-weights-paths.transformer OWNER/REPO/path/FILE.safetensors` | Auto-detected; requires `comfy-kitchen`. TP must preserve each file's ConvRot group boundaries. | +| DiT | [NVFP4, optionally mixed with INT8 or FP8](https://huggingface.co/Abiray/Minimax-H3-nvfp4-INT4-INT8-Convrot) | `--component-weights-paths.transformer OWNER/REPO/path/FILE.safetensors` | Auto-detected; NVFP4 execution requires NVIDIA compute capability 10.0+. | +| DiT | [AutoRound W4A16 component](https://huggingface.co/Ar4ikov/MiniMax-H3-transformer-W4A16-RTN) | `--component-paths.transformer Ar4ikov/MiniMax-H3-transformer-W4A16-RTN` | Self-describing Diffusers component; SGLang reuses the SRT GPTQ/Marlin backend. The linked export is FL2VA. | +| DiT | GGUF, full or AdaLN-pruned ([full](https://huggingface.co/leejet/MiniMax-H3-GGUF), [pruned](https://huggingface.co/unsloth/MiniMax-H3-GGUF)) | `--component-weights-paths.transformer OWNER/REPO/FILE.gguf` | CUDA capacity path; aligned TP and layerwise offload are supported, FSDP and LoRA are not. | +| DiT | Online FP8 or ConvRot INT8 | `--component-quantizations.transformer fp8` or `... kitchen_int8` | Quantizes the official weights during load; this is not a pre-quantized checkpoint format. | +| Text encoder | [Serialized FP8 component](https://huggingface.co/Qwen/Qwen3-VL-32B-Instruct-FP8) | `--component-paths.text_encoder Qwen/Qwen3-VL-32B-Instruct-FP8` | Only eligible language-model linears use FP8; embeddings, norms, and the vision tower keep their declared precision. | +| Text encoder | ConvRot INT8, W4A8, or W4A4 safetensors ([INT8](https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/text_encoders), [W4A8](https://huggingface.co/Winnougan/MiniMax-H3-INT4_Convrot_ComfyUI), [W4A4](https://huggingface.co/Merserk/MiniMax-H3-INT4-ConvRot)) | `--component-weights-paths.text_encoder OWNER/REPO/path/FILE.safetensors` | Auto-detected; requires `comfy-kitchen`. Unmarked vision and embedding tensors keep their declared precision. | +| Text encoder | [NVFP4-AWQ](https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/text_encoders) or [Quanto qint8](https://huggingface.co/DeepBeepMeep/MiniMax-H3/tree/main/Qwen3-VL-32B-Instruct) safetensors | `--component-weights-paths.text_encoder OWNER/REPO/path/FILE.safetensors` | Memory-oriented formats: compressed storage is restored, then each active matrix uses BF16/FP16 compute. | +| Text encoder | [GGUF Qwen3-VL](https://huggingface.co/DeepBeepMeep/MiniMax-H3/tree/main/Qwen3-VL-32B-Instruct) | `--component-weights-paths.text_encoder OWNER/REPO/FILE.gguf` | CUDA capacity path with encoder TP/layerwise support; encoder FSDP is not supported. | +| Text encoder | [Compact Qwen3-VL 4B/8B + ClipProj](https://huggingface.co/NicoLab28/ClipProj-MiniMax-H3) | `--component-paths.text_encoder ENCODER_REPO --component-paths.conditioning_projection PROJECTION.safetensors` | Approximate conditioning replacement. A separate weight-only override may quantize the selected small encoder. | +| Adapter | Native fused or Diffusers/PEFT LoRA | `--lora-path OWNER/REPO [--lora-weight-name FILE] --lora-merge-mode auto` | Layout and alpha metadata are normalized at load time; see [LoRA recipes](#5-lora-recipes). | + +For H3, the registered component names are `transformer`, `text_encoder`, +`video_vae`, and `audio_vae`. The shorter `--transformer-weights-path` and +`--text-encoder-path` aliases remain supported. `conditioning_projection` is an +H3 text-encoder sidecar key, not a standalone model component. Plain +video/audio VAE safetensors can use `--component-weights-paths.video_vae` or +`--component-weights-paths.audio_vae`, but SGLang does not currently advertise +a native quantized H3 VAE format. + +Pre-quantized files are self-describing: do not combine those rows with +`--quantization` or `--component-quantizations.*`. Only byte-identical official +full weights—whether loaded from the original model, Diffusers component, or a +weight-only layout—belong to the consistency GT. Pruned, quantized, +compact-encoder, and LoRA routes are outside that baseline and the audited +`quality="high"` contract. Packed and per-layer mixed formats reject FSDP unless +their row says otherwise; see +[Quantization](/docs/sglang-diffusion/quantization) for backend-wide constraints. + For a four-card H200 host, keep the full BF16/FP32 model resident by default. The model fits without FSDP, so this path avoids the per-block parameter all-gathers of the memory-oriented FSDP profile: @@ -184,144 +228,15 @@ baseline; everything else stays identical. GPU peak stays about 18 GB either way because streaming offload is set by the offload buffers and VAE decode, not the weight dtype. -### Pre-quantized GGUF transformer +The first launch resolves every selected source through the normal Hub path. If +a repository requires authentication, export a Hugging Face token in the +server environment; no manual pre-download is required. -Use `--transformer-weights-path` to replace only the DiT with a GGUF file; the -base repository continues to provide the text encoder, VAEs, scheduler, and -tokenizers. Do not also pass `--quantization gguf`. - -```bash 1×RTX 5090 Q4_K_M -sglang serve \ - --model-path MiniMaxAI/MiniMax-H3 \ - --model-variant fl2va \ - --transformer-weights-path \ - leejet/MiniMax-H3-GGUF/minimax_h3_fl2va-Q4_K_M.gguf \ - --attention-backend fa \ - --performance-mode memory \ - --layerwise-offload-components dit,text_encoder \ - --dit-offload-prefetch-size 1 \ - --dit-layerwise-resident-layers 0 \ - --enable-torch-compile false \ - --port 30010 -``` - -The loader also recognizes pruned checkpoints that replace the timestep MLP -with `adaln_t_table`. Repositories containing both FL2VA and Ref2VA variants -need a full file reference: - -```bash Pruned FL2VA Q4_K -sglang serve \ - --model-path MiniMaxAI/MiniMax-H3 \ - --model-variant fl2va \ - --transformer-weights-path \ - unsloth/MiniMax-H3-GGUF/minimax_h3_fl2va_pruned-Q4_K.gguf \ - --performance-mode memory \ - --layerwise-offload-components dit,text_encoder \ - --enable-torch-compile false \ - --port 30010 -``` - -The linear adapter reuses SRT's GGUF type definitions and CUDA dequantization, -then runs the native GEMM. SRT's fused MMVQ/MMQ kernels target the low-token LLM -regime and are slower at diffusion sequence lengths. TP is supported when each -row-parallel input shard remains GGML-block aligned; incompatible degrees fail -during model construction. FSDP, LoRA merging, and the separate MiniMax-H3 -AdaLN cache flags are not compatible with packed GGUF weights. - -The first launch downloads the model through the selected Hub. If the Hugging -Face repository requires authentication, export a Hugging Face token in the -server environment. - -For MiniMax-H3, `--performance-mode speed` deliberately keeps the DiT eager. The current `torch.compile` path changes the model's numerical output, so it is not enabled implicitly by any recommended lossless preset. An explicit `--enable-torch-compile true` remains available for controlled experiments, but it should not be used to generate consistency ground truth. - -### AdaLN-pruned safetensors transformers - -[Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3/tree/main/diffusion_models) -publishes smaller DiT-only checkpoints that replace the original AdaLN branches -with an interpolated curve table. Select one file explicitly; the base model -still supplies the text encoder and VAEs. - -```bash Command -sglang serve \ - --model-path MiniMaxAI/MiniMax-H3 \ - --model-variant fl2va \ - --transformer-weights-path \ - Comfy-Org/MiniMax-H3/diffusion_models/minimax_h3_fl2va_pruned_bf16.safetensors \ - --num-gpus 4 \ - --tp-size 2 \ - --ulysses-degree 2 \ - --performance-mode speed \ - --port 30010 -``` - -The pruned checkpoint is approximate and is therefore rejected by -`quality="high"`, which remains limited to the audited official BF16 DiT. - -The Comfy `pruned_fp8_scaled` FL2VA and Ref2VA files are also supported. Their -per-layer markers are detected automatically; do not add `--quantization`: - -```bash Command -sglang serve \ - --model-path MiniMaxAI/MiniMax-H3 \ - --model-variant fl2va \ - --transformer-weights-path \ - Comfy-Org/MiniMax-H3/diffusion_models/minimax_h3_fl2va_pruned_fp8_scaled.safetensors \ - --num-gpus 4 \ - --tp-size 2 \ - --ulysses-degree 2 \ - --performance-mode speed \ - --port 30010 -``` - -SGLang uses its native static-activation FP8 linear path when the checkpoint -stores input scales, and automatically uses dynamic activation scaling for -Comfy FP8 exports that omit them. The checkpoint above marks `fc2` for -full-precision matrix multiplication, so -SGLang retains its FP8 storage but materializes and scales one compute-dtype -`fc2` matrix for each call. This preserves the checkpoint's mixed execution -contract and low resident weight memory, but that part is slower than a fully -quantized FP8 GEMM. TP, Ulysses/Ring sequence parallelism, and -component/layerwise offload are supported; FSDP inference is rejected. - -The `pruned_int8_convrot` files use the same override path and are detected -automatically. Install `comfy-kitchen`, replace the FP8 filename above with -`minimax_h3_fl2va_pruned_int8_convrot.safetensors`, and still omit -`--quantization`. SGLang loads their serialized INT8 weights and row scales into -the fused ConvRot kernel without requantizing them. TP1/2/4 and sequence or -layerwise offload are supported; TP8 violates the checkpoint's 256-element -ConvRot group boundary, and FSDP is rejected. - -Self-describing community MXFP8 files use the same -`--transformer-weights-path ` form and need no -quantization flag. SGLang reads their global format metadata, keeps unmarked H3 -layers in their original dtype, and reuses SRT's MXFP8 dense kernels. The -selected SRT backend must support MXFP8 on the target GPU; FSDP is rejected for -this mixed per-layer layout. - -W4A8 ConvRot DiT files use the same flagless flow. With -`comfy-kitchen>=0.2.27`, pass a file such as -`starsfriday/MiniMax-H3-w4a8/minimax_h3_fl2va_pruned_w4a8_mixed.safetensors` -to `--transformer-weights-path`; SGLang reads the serialized per-layer metadata -and packed INT4 tensors automatically. Do not add `--quantization`. TP remains -subject to each row-parallel shard preserving the checkpoint's ConvRot group -boundary, and FSDP is rejected. - -W4A4 ConvRot files are also detected from their layer metadata. Install a -current `comfy-kitchen`, then pass a full or pruned FL2VA / Ref2VA file such as -`Merserk/MiniMax-H3-INT4-ConvRot/minimax_h3_fl2va_pruned_int4_convrot.safetensors` -to `--transformer-weights-path`. The packed weights stay INT4 and the runtime -honors each layer's activation mode; omit `--quantization`. -Mixed exports use the same command: SGLang dispatches each marked layer to its -serialized W4A4 or INT8 ConvRot kernel instead of applying one global method. - -Comfy NVFP4 DiTs reuse SGLang's ModelOpt NVFP4 backend, which requires CUDA -compute capability 10.0 or newer. Pass a pruned FL2VA / Ref2VA file such as -`Abiray/Minimax-H3-nvfp4-INT4-INT8-Convrot/MiniMax_H3_FL2VA_pruned_nvfp4.safetensors` -to `--transformer-weights-path` and omit `--quantization`. SGLang infers the -packed group size and Comfy scale layout from the checkpoint; FSDP is rejected. -Mixed files may mark selected linears as `int8_tensorwise` or dynamic/static -FP8; SGLang dispatches those layers to their serialized Kitchen INT8 or native -FP8 path automatically. +For MiniMax-H3, `--performance-mode speed` deliberately keeps the DiT eager. +The current `torch.compile` path changes the model's numerical output, so no +recommended lossless preset enables it implicitly. An explicit +`--enable-torch-compile true` remains available for controlled experiments, but +do not use it to generate consistency ground truth. ### Advanced: precomputed AdaLN cache @@ -707,11 +622,11 @@ the adapter author explicitly provides Ref2VA-compatible weights. Also avoid stacking a distilled adapter with `quality: "high"`: both alter denoising, and that combination has not been quality-validated. - -LoRAs trained for a pruned or structurally modified ComfyUI graph are not -automatically compatible with the native H3 weights. Use only adapters whose -architecture and target modules match the full native H3 checkpoint. - +AdaLN-pruned Diffusers components that publish `adaln_basis` and `adaln_mean` +can also consume a LoRA trained against the released full-width AdaLN modules: +SGLang projects those adapter factors onto the pruned coordinates at load time. +A structurally modified checkpoint without that metadata still fails closed, +and packed GGUF weights remain incompatible with LoRA. ## 6. Sampling and output controls @@ -1016,90 +931,11 @@ the attention algorithm. Neither is a consistency ground-truth mode. The BF16 path is unchanged when `comfy-kitchen` is not installed. -The Qwen3-VL text encoder can be replaced independently of the DiT. To reduce -its resident memory, point the text-encoder component at the serialized FP8 -checkpoint used in validation: - -```bash Command -sglang serve \ - --model-path MiniMaxAI/MiniMax-H3 \ - --model-variant fl2va \ - --component-paths.text_encoder Qwen/Qwen3-VL-32B-Instruct-FP8 \ - --num-gpus 4 \ - --tp-size 2 \ - --ulysses-degree 2 \ - --performance-mode speed \ - --port 30010 -``` - -`--text-encoder-path` is accepted as a shorter alias. No separate quantization -flag is required: SGLang reads the checkpoint's `quantization_config` and -fails closed if the native encoder does not support that format. The language -linear layers use FP8 while embeddings, normalization, and the vision tower -remain BF16. This is an approximate serve-time choice and is incompatible with -the strict `quality="high"` deployment contract. - -The mixed W4A8 community encoder follows the same component-path contract: - -```bash Overlay ---component-paths.text_encoder \ - Winnougan/MiniMax-H3-INT4_Convrot_ComfyUI/qwen3vl_32b_minimax_h3-w4a8_convrot.safetensors -``` - -Install `comfy-kitchen>=0.2.27` and omit `--quantization`. SGLang automatically -loads its W4A8 language linears and tensorwise INT8 embedding; the unmarked -vision tower remains BF16. - -W4A4 Qwen3-VL files use the same overlay, for example: - -```bash Overlay ---component-paths.text_encoder \ - Merserk/MiniMax-H3-INT4-ConvRot/qwen3vl_32b_minimax_h3_int4_convrot.safetensors -``` - -This checkpoint keeps its unmarked embedding and vision tower in their source -precision; no explicit component quantization option is needed. - -The official Comfy NVFP4-AWQ encoder uses the same flagless overlay: - -```bash Overlay ---component-paths.text_encoder \ - Comfy-Org/MiniMax-H3/text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors -``` - -SGLang auto-detects its row-wise INT8 embedding, NVFP4 language linears, and -AWQ input pre-scales. The weights stay compressed at rest; each active linear -is dequantized for a BF16/FP16 matrix multiplication, so this path primarily -reduces resident memory rather than encoder latency. Do not add a component or -transformer quantization option. - -The same component option accepts a self-describing Quanto qint8 file without -an additional quantization flag: - -```bash Overlay ---component-paths.text_encoder \ - DeepBeepMeep/MiniMax-H3/Qwen3-VL-32B-Instruct/Qwen3-VL-32B-Instruct-layer50_quanto_bf16_int8.safetensors -``` - -This variant keeps the declared language and vision linear weights in qint8 -storage, then dequantizes only the active matrix for BF16/FP16 linear math. Use -it as a memory option, not as an INT8 throughput claim. The embedded Quanto map -is the selector; adding `--quantization` would describe a different operation. - -Qwen3-VL GGUF files use the same component option: - -```bash Overlay ---component-paths.text_encoder \ - DeepBeepMeep/MiniMax-H3/Qwen3-VL-32B-Instruct/qwen3vl-32B-MiniMax-H3-Q4_K_M.gguf -``` - -The file selects its own Q4/Q6 layout; do not add `--quantization`. SGLang keeps -the aligned vocabulary table and language linears in packed GGML storage. The -smaller Comfy vision tower, whose flattened blocks can cross row and TP shard -boundaries, is restored one matrix at a time to BF16 during loading. -This CUDA-only memory option supports encoder TP and layerwise placement, but -not encoder FSDP. Like the Quanto variant, it is approximate and is not a -throughput claim. +Pre-quantized and compact H3 text encoders are listed once in +[Checkpoint and adapter formats](#checkpoint-and-adapter-formats). They use +component-local paths and never inherit the DiT's `--quantization` setting. +SGLang reads their metadata before constructing the native Qwen3-VL encoder and +fails closed when the selected format, projection, or topology is incompatible. diff --git a/docs/docs/sglang-diffusion/api/cli.mdx b/docs/docs/sglang-diffusion/api/cli.mdx index 878fd6212..95b428747 100644 --- a/docs/docs/sglang-diffusion/api/cli.mdx +++ b/docs/docs/sglang-diffusion/api/cli.mdx @@ -134,13 +134,12 @@ pipeline's registered module name: | Quantize an unquantized component online | `--component-quantizations. {METHOD}` | `---quantization {METHOD}` | Apply a method supported by that component's native loader | | Keep selected component layers unquantized | `--component-quantization-ignored-layers. {PATTERN...}` | None | Pass component-local ignored-layer patterns to its online quantizer | -For example, pair a replacement text-encoder configuration with a separate +For example, pair any replacement text-encoder configuration with a separate single-file checkpoint as follows: ```bash ---component-paths.text_encoder Qwen/Qwen3-VL-4B-Instruct \ ---component-weights-paths.text_encoder \ - Comfy-Org/Krea-2/text_encoders/qwen3vl_4b_fp8_scaled.safetensors +--component-paths.text_encoder COMPONENT_REPO_OR_DIRECTORY \ +--component-weights-paths.text_encoder WEIGHTS_FILE_OR_REPO_FILE ``` The transformer-specific `--transformer-weights-path` spelling remains @@ -162,7 +161,10 @@ while quantized materialization is capability-based. Native auxiliary loaders whose current materializer expects plain state dicts reject unsupported quantization metadata before model construction. See [Quantized Component Repositories](../quantization#quantized-component-repositories) -for the current component matrix. +for the current component matrix. A model cookbook is the source of truth for +published, model-specific checkpoint examples; for example, all H3 sources and +their exact overlays are kept in one +[MiniMax-H3 compatibility table](/cookbook/diffusion/MiniMax/MiniMax-H3#checkpoint-and-adapter-formats). Component overrides accept a local component directory, a standalone Hub repository, or a Hub component subfolder written as `owner/repo/subfolder`. diff --git a/docs/docs/sglang-diffusion/compatibility_matrix.mdx b/docs/docs/sglang-diffusion/compatibility_matrix.mdx index 2a825570e..c31487257 100644 --- a/docs/docs/sglang-diffusion/compatibility_matrix.mdx +++ b/docs/docs/sglang-diffusion/compatibility_matrix.mdx @@ -133,7 +133,7 @@ Rows are grouped when a family shares the same runtime path or optimization supp MiniMax-H3 -
MiniMaxAI/MiniMax-H3
+
MiniMaxAI/MiniMax-H3Checkpoint formats
T2VA / FL2VA / Ref2VA, 768p at 24 fps with synchronized audio Cache-DiTSageOnline FP8GGUF diff --git a/docs/docs/sglang-diffusion/quantization.mdx b/docs/docs/sglang-diffusion/quantization.mdx index 27ee163a9..b526a0059 100644 --- a/docs/docs/sglang-diffusion/quantization.mdx +++ b/docs/docs/sglang-diffusion/quantization.mdx @@ -2,7 +2,7 @@ title: "Quantization" tag: "approx" metatags: - description: "Configure quantized component repositories, transformer and native text-encoder checkpoint quantization, plus Quant-VideoGen causal KV-cache quantization in SGLang-Diffusion." + description: "Configure quantized component repositories, component-local online quantization, and Quant-VideoGen causal KV-cache quantization in SGLang-Diffusion." --- SGLang-Diffusion treats component path selection and quantized checkpoint @@ -13,28 +13,27 @@ precision are resolved independently. ## Quick Reference -Use these paths: +Choose the option from the source you have, not from the precision name: -- `--model-path`: the base or original model -- `--component-paths.` / `---path`: replace a component from `model_index.json` or a native registered module with an independent repo/local directory; for transformer and native encoder loaders, an explicit weight filename retains the base component config -- `--transformer-path`: a quantized transformers-style transformer component directory that already contains its own `config.json` -- `--transformer-weights-path`: replacement transformer weights in safetensors - format (file, directory, or Hub repository/file) or a supported GGUF file - (local or Hub) -- `--quantization`: override the quantization method used by the transformer loader -- `--quantization-ignored-layers`: transformer layer name patterns to keep unquantized during online quantization (e.g. `attention.to_`) -- `--component-quantizations.`: explicitly apply supported online quantization to an unquantized component -- `--component-quantization-ignored-layers.`: component-local layer patterns to keep unquantized during that online quantization -- `--component-paths.text_encoder`: replace a native text encoder with a checkpoint whose `quantization_config` is auto-detected -- `--text-encoder-path`: shorter alias for `--component-paths.text_encoder` -- `--kv-cache-quant`: compress completed causal KV-cache chunks for supported realtime models +| Source or operation | CLI | Result | +| --- | --- | --- | +| Complete component repo/directory | `--component-paths. ` | Replaces that component's config and weights. | +| Weight file/directory for an existing component | `--component-weights-paths. ` | Keeps the base component config and replaces only its weights. | +| Online quantization of unquantized weights | `--component-quantizations. ` | Constructs a supported quantized implementation during load; use `--component-quantization-ignored-layers.` for exclusions. | +| Primary DiT convenience spellings | `--transformer-path`, `--transformer-weights-path`, `--quantization` | Aliases for pipelines with one primary DiT; they remain supported but are not generic component names. | +| Causal KV-cache compression | `--kv-cache-quant` | Quantizes completed runtime cache chunks, not checkpoint weights. | + +`--model-path` always selects the base model. The component key must be a real +key from its `model_index.json` or native pipeline registry; examples include +`transformer`, `transformer_2`, `text_encoder`, `video_vae`, and `audio_vae`, +but no one name is universal. Recommended example for pre-quantized checkpoints: ```bash sglang generate \ --model-path black-forest-labs/FLUX.2-dev \ - --transformer-weights-path black-forest-labs/FLUX.2-dev-NVFP4 \ + --component-weights-paths.transformer black-forest-labs/FLUX.2-dev-NVFP4 \ --prompt "a curious pikachu" ``` @@ -43,23 +42,21 @@ For quantized transformers-style transformer component folders: ```bash sglang generate \ --model-path /path/to/base-model \ - --transformer-path /path/to/quantized-transformer \ + --component-paths.transformer /path/to/quantized-transformer \ --prompt "A Logo With Bold Large Text: SGL Diffusion" ``` NOTE: Some model-specific integrations also accept a quantized repo or local directory directly as `--model-path`, but that is a compatibility path. If a -repo contains multiple candidate checkpoints, pass -`--transformer-weights-path` explicitly. +repo contains multiple candidate checkpoints, select an exact component or +weight source explicitly. -MiniMax-H3 is a verified example for Comfy safetensors with per-layer metadata, -including `pruned_fp8_scaled`, serialized ConvRot formats, and the official -NVFP4-AWQ Qwen3-VL encoder. Other Comfy FP8 exports are also auto-detected: the -presence of an input scale selects static activation scaling, while its absence -selects dynamic scaling. Pass one selected DiT or component file by local path, -`owner/repo/path/file.safetensors`, or direct Hugging Face file URL; do not -combine it with an explicit quantization option. Its GGUF usage is documented in -the [MiniMax-H3 cookbook](/cookbook/diffusion/MiniMax/MiniMax-H3#pre-quantized-gguf-transformer). +Pre-quantized files are self-describing. Pass one selected file by local path, +`owner/repo/path/file.safetensors`, or direct Hugging Face file URL and do not +add an online quantization option. The +[MiniMax-H3 compatibility table](/cookbook/diffusion/MiniMax/MiniMax-H3#checkpoint-and-adapter-formats) +is the canonical list of H3 component formats and published example sources; +this page defines the backend families and their shared constraints. ## Quantized Component Repositories @@ -109,7 +106,7 @@ backend. quant_family checkpoint form - canonical CLI + checkpoint selector supported models extra dependency platform / notes @@ -119,7 +116,7 @@ backend. fp8 / mxfp4 (online quantization) Unquantized checkpoint (offline via AMD Quark coming soon) - --quantization {fp8,mxfp4} + --component-quantizations.<component> {fp8,mxfp4} Z-Image-Turbo (validated), others likely work. More support coming soon. MXFP4: aiter on ROCm MXFP4 requires ROCm and MI350+ (gfx95x). Weights quantized at load time, activations quantized to fp8 / mxfp4 dynamically. @@ -127,7 +124,7 @@ backend. kitchen_int8 (online quantization) Unquantized BF16/FP16 checkpoint - --quantization kitchen_int8 + --component-quantizations.<component> kitchen_int8 MiniMax-H3 (validated on 1× RTX 4090 24 GB) comfy-kitchen Data-free INT8 ConvRot at load time via comfy_kitchen.int8_linear. Layers whose input dim is not divisible by the group size stay in BF16. Requires Turing+ (SM75). @@ -135,7 +132,7 @@ backend. fp8 (offline quantization) Quantized transformer component folder, or safetensors with quantization_config metadata - --transformer-path or --transformer-weights-path + --component-paths.<component> or --component-weights-paths.<component> Native DiTs whose linear layers support the selected FP8 method; validate quality per model None Component-folder and single-file flows are both supported @@ -143,7 +140,7 @@ backend. modelopt-fp8 Converted ModelOpt FP8 transformer directory or repo with config.json - --transformer-path + --component-paths.<component> FLUX.1, FLUX.2, Wan2.2, HunyuanVideo, Qwen Image, Qwen Image Edit None Serialized config stays quant_method=modelopt with quant_algo=FP8; dit_layerwise_offload is supported and dit_cpu_offload stays disabled @@ -151,7 +148,7 @@ backend. auto-round W4A16 Transformer component repo with a self-describing quantization_config and auto_round:auto_gptq packing - --transformer-path + --component-paths.<component> Native dense DiTs with compatible component parameter mappings; MiniMax-H3 Diffusers components are supported None Auto-detected; reuses the SRT GPTQ/Marlin backend. No --quantization flag is needed; use TP/sequence parallelism rather than FSDP. @@ -159,7 +156,7 @@ backend. modelopt-nvfp4 Mixed transformer directory/repo with config.json, raw or Comfy layer-marked NVFP4 safetensors, or full ModelOpt Diffusers repo - --transformer-path for mixed overrides; --transformer-weights-path for raw exports; --model-path for full repos + --component-paths.<component> for component repos; --component-weights-paths.<component> for raw files; --model-path for full repos FLUX.1, FLUX.2, Wan2.2, Qwen Image, Qwen Image 2512, Qwen Image Edit, Qwen Image Edit 2511, MiniMax-H3 None Mixed override repos keep the base model separate; full Qwen Image exports can be loaded directly as --model-path; raw exports such as black-forest-labs/FLUX.2-dev-NVFP4 use the weights-path flow. Comfy markers select NVFP4 plus INT8 or FP8 companion linears automatically; omit --quantization. @@ -167,7 +164,7 @@ backend. gguf One selected GGUF DiT or native encoder file - --transformer-weights-path, or an explicit encoder file through --component-paths.<component> + --component-weights-paths.<component> MiniMax-H3 original or pruned FL2VA / Ref2VA DiTs and its Qwen3-VL text encoder None CUDA only and auto-detected; supports standard and K-quant GGML types. TP requires aligned GGML blocks; component/layerwise offload is supported and FSDP is not. @@ -175,7 +172,7 @@ backend. comfy-fp8 One selected Comfy safetensors file with per-layer comfy_quant metadata - --transformer-weights-path + --component-weights-paths.<component> MiniMax-H3 pruned FL2VA / Ref2VA DiTs None CUDA; auto-detected; TP, sequence parallelism, and component/layerwise offload are supported, while FSDP is not. Checkpoint-marked fc2 layers retain FP8 storage and use compute-dtype matmul. @@ -183,7 +180,7 @@ backend. comfy-int8-convrot One selected safetensors file with per-layer int8_tensorwise and ConvRot metadata - --transformer-weights-path, or an explicit weight file through --component-paths.<component> + --component-weights-paths.<component> Native DiTs and encoders whose parameter mappings preserve each marked linear; MiniMax-H3 DiT and Qwen3-VL encoder checkpoints have validated tensor contracts comfy-kitchen CUDA; auto-detected; uses the fused Kitchen INT8 kernel and validates weight/scale layout before model construction. TP requires every row-parallel input shard to preserve the checkpoint's ConvRot group boundary. The H3 256-group DiT supports TP1/2/4, not TP8; its Qwen3-VL encoder keeps TP8 by replicating only incompatible row projections. Offload is supported; FSDP is not. @@ -191,7 +188,7 @@ backend. mxfp8 Self-describing serialized weights, or BF16/FP16 weights for online quantization - --transformer-weights-path, or --quantization mxfp8 for online quantization + --component-weights-paths.<component>, or --component-quantizations.<component> mxfp8 online Native diffusion linears; mixed MiniMax-H3 checkpoints are selected per layer SRT's platform MXFP8 backend Serialized metadata is auto-detected. NVIDIA and ROCm reuse SRT's dense MXFP8 kernels; Ascend keeps its native online path. Mixed per-layer checkpoints do not support FSDP. @@ -199,7 +196,7 @@ backend. comfy-w4a8-convrot Safetensors with serialized asym_w4a8_int8 layer metadata and packed weights - --transformer-weights-path or --component-paths.text_encoder + --component-weights-paths.<component> MiniMax-H3 FL2VA / Ref2VA DiTs and its native Qwen3-VL encoder comfy-kitchen>=0.2.27 Auto-detected; omit --quantization. Requires SM80+ and validates packed weights, group/channel scales, and optional codebooks before model construction. Mixed encoder files may keep their embedding tensorwise INT8. TP must preserve ConvRot group boundaries; offload is supported and FSDP is not. @@ -207,7 +204,7 @@ backend. comfy-w4a4-convrot Safetensors with serialized convrot_w4a4 metadata, optionally mixed with int8_tensorwise layers - --transformer-weights-path or --component-paths.text_encoder + --component-weights-paths.<component> Native DiTs and encoders with matching parameter mappings; MiniMax-H3 FL2VA / Ref2VA DiTs and Qwen3-VL encoder layouts are recognized comfy-kitchen Auto-detected; omit --quantization. Each layer dispatches to its serialized W4A4 or INT8 ConvRot kernel. CUDA requires SM75+; TP must preserve each format's quantization and ConvRot group boundaries. Offload is supported and FSDP is not. @@ -215,7 +212,7 @@ backend. comfy-nvfp4-full-precision Safetensors with serialized nvfp4 and optional row-wise int8_tensorwise layer metadata - --component-paths.text_encoder + --component-weights-paths.text_encoder MiniMax-H3 native Qwen3-VL encoder None Auto-detected; omit explicit quantization. Preserves packed storage, high-nibble-first weights, swizzled block scales, and AWQ input pre-scales. Each active NVFP4 matrix is dequantized for BF16/FP16 compute, so this is a memory path rather than a native FP4 speed path. @@ -223,7 +220,7 @@ backend. quanto-int8 One native encoder safetensors file with an embedded Quanto quantization map - An explicit weight file through --component-paths.<component> + --component-weights-paths.<component> Native encoders whose mapped linear layers consume every declared qint8 entry; MiniMax-H3's Qwen3-VL encoder is supported None Auto-detected weight-only qint8 storage. Each active matrix is dequantized to the compute dtype for the ordinary linear operation, so this reduces stored/resident weight memory rather than promising INT8 GEMM speed. TP and offload are supported; FSDP is not. @@ -252,14 +249,6 @@ backend. None Currently only compatible with the Ascend NPU family and supports mxfp8, mxfp4, w8a8, and w4a4 - - gguf - A single community .gguf holding a transformer or native encoder - --transformer-weights-path or --component-paths.<component> - MiniMax-H3 fl2va DiT (original and pruned AdaLN curve) and Qwen3-VL text encoder - None - CUDA only, no FSDP. TP requires GGML-aligned shard boundaries. Packed weights reduce checkpoint and resident-weight memory; they are dequantized per use, so this is not a speed claim. See GGUF. - @@ -374,11 +363,14 @@ Apply FP8 quantization to a supported unquantized DiT checkpoint: ```bash sglang generate \ --model-path Tongyi-MAI/Z-Image-Turbo \ - --quantization fp8 \ + --component-quantizations.transformer fp8 \ --prompt "a beautiful sunset" \ --save-output ``` +For a pipeline whose primary DiT is named `transformer`, the shorter +`--quantization fp8` spelling is equivalent. + MiniMax-H3 supports this path while preserving its required FP32 patch, timestep, and output projections. See the [MiniMax-H3 cookbook](/cookbook/diffusion/MiniMax/MiniMax-H3#7-feature-contracts-and-advanced-recipes) @@ -391,7 +383,7 @@ MXFP4 provides aggressive 4-bit compression with online quantization. **Note: Re ```bash sglang generate \ --model-path Tongyi-MAI/Z-Image-Turbo \ - --quantization mxfp4 \ + --component-quantizations.transformer mxfp4 \ --prompt "a beautiful sunset" \ --save-output ``` @@ -399,13 +391,14 @@ sglang generate \ ### Kitchen INT8 -Serialized Comfy ConvRot INT8 DiTs use `--transformer-weights-path`; compatible -native encoders use an explicit file through `--component-paths.` or -its component alias. Both are auto-detected from per-layer markers and load -INT8 weights and row scales directly; omit `--quantization`. +Serialized Comfy ConvRot INT8 DiTs and compatible native encoders use +`--component-weights-paths.`. Both are auto-detected from per-layer +markers and load INT8 weights and row scales directly; omit online +quantization. `--transformer-weights-path` remains the primary-DiT convenience +spelling. -For a BF16 checkpoint, `--quantization kitchen_int8` instead performs online -quantization after loading: +For a BF16 checkpoint, `--component-quantizations.transformer kitchen_int8` +instead performs online quantization after loading: `kitchen_int8` quantizes DiT linear weights online from the stock BF16 checkpoint. Forward uses the fused `comfy_kitchen.int8_linear` op (rotation, @@ -420,7 +413,7 @@ pip install comfy-kitchen sglang generate \ --model-path MiniMaxAI/MiniMax-H3 \ --model-variant fl2va \ - --quantization kitchen_int8 \ + --component-quantizations.transformer kitchen_int8 \ --attention-backend fa \ --performance-mode memory \ --layerwise-offload-components dit,text_encoder \ @@ -692,8 +685,8 @@ sglang generate \ ### Notes -- `--transformer-path` is the canonical flag for converted ModelOpt FP8 - transformer component repos or directories that already carry `config.json`. +- `--transformer-path` is the primary-DiT convenience spelling for converted + ModelOpt FP8 component repos or directories that already carry `config.json`. - If the override repo or local directory contains its own `config.json`, SGLang reads the quantization config from that override instead of relying on the base model config. @@ -840,16 +833,17 @@ them ends up narrower. No extra install: `gguf` is already a core SGLang dependency. -`--model-path` stays the base model; `--transformer-weights-path` takes the -GGUF. A local path, `owner/repo/file.gguf`, or `owner/repo:QUANT_TYPE` all work. -The quant-type shorthand is accepted only when exactly one repository file -matches it; otherwise SGLang lists the candidates and asks for a full path. +`--model-path` stays the base model; +`--component-weights-paths.` takes the GGUF. A local path, +`owner/repo/file.gguf`, or `owner/repo:QUANT_TYPE` all work. The quant-type +shorthand is accepted only when exactly one repository file matches it; +otherwise SGLang lists the candidates and asks for a full path. ```bash sglang serve \ --model-path MiniMaxAI/MiniMax-H3 \ --model-variant fl2va \ - --transformer-weights-path \ + --component-weights-paths.transformer \ leejet/MiniMax-H3-GGUF/minimax_h3_fl2va-Q4_K_M.gguf \ --num-gpus 1 \ --attention-backend fa \ @@ -861,25 +855,17 @@ sglang serve \ --port 30010 ``` -Note that `--quantization gguf` is not the selector — the quantization is read -from the file itself, so passing the file is what enables the path. +Here `transformer` is H3's registered DiT component name. The +`--transformer-weights-path` alias produces the same result. Note that +`--quantization gguf` is not the selector: the quantization is read from the +file itself, so passing the file is what enables the path. MiniMax-H3 GGUF checkpoints may use either the original timestep MLP or the pruned AdaLN curve architecture. Repositories often contain both FL2VA and Ref2VA files, so use the full Hub file reference instead of an ambiguous -`owner/repo:QUANT_TYPE` selector: - -```bash -sglang serve \ - --model-path MiniMaxAI/MiniMax-H3 \ - --model-variant fl2va \ - --transformer-weights-path \ - unsloth/MiniMax-H3-GGUF/minimax_h3_fl2va_pruned-Q4_K.gguf \ - --performance-mode memory \ - --layerwise-offload-components dit,text_encoder \ - --enable-torch-compile false \ - --port 30010 -``` +`owner/repo:QUANT_TYPE` selector. The +[H3 compatibility table](/cookbook/diffusion/MiniMax/MiniMax-H3#checkpoint-and-adapter-formats) +links both published layouts without duplicating their commands here. The pruned architecture keeps its sampled curve and reduced AdaLN projections in FP32, matching the published checkpoint implementation. This precision @@ -1093,9 +1079,9 @@ sglang generate \ ### Notes -- `--transformer-weights-path` is the canonical flag for Nunchaku checkpoints. - Older config names such as `quantized_model_path` are treated as - compatibility aliases. +- `--transformer-weights-path` is the primary-DiT convenience spelling used by + Nunchaku examples. Older config names such as `quantized_model_path` are + treated as compatibility aliases. - Auto-detection only happens when the checkpoint basename matches `svdq-(int4|fp4)_r{rank}`. - The CLI values are `int4` and `nvfp4`. In filenames, the NVFP4 variant is