Add HunyuanVideo ModelOpt FP8 diffusion support (#23199)
This commit is contained in:
@@ -43,21 +43,21 @@ backend.
|
||||
| quant_family | checkpoint form | canonical CLI | supported models | extra dependency | platform / notes |
|
||||
|-------------------|--------------------------------------------------------------------------------------------|------------------------------------------------------------------------|-----------------------------------------|---------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| `fp8` | Quantized transformer component folder, or safetensors with `quantization_config` metadata | `--transformer-path` or `--transformer-weights-path` | ALL | None | Component-folder and single-file flows are both supported |
|
||||
| `modelopt-fp8` | Converted ModelOpt FP8 transformer directory or repo with `config.json` | `--transformer-path` | FLUX.1, FLUX.2, Wan2.2, Qwen Image, Qwen Image Edit | None | Serialized config stays `quant_method=modelopt` with `quant_algo=FP8`; `dit_layerwise_offload` is supported and `dit_cpu_offload` stays disabled |
|
||||
| `modelopt-fp8` | Converted ModelOpt FP8 transformer directory or repo with `config.json` | `--transformer-path` | FLUX.1, FLUX.2, Wan2.2, HunyuanVideo, Qwen Image, Qwen Image Edit | None | Serialized config stays `quant_method=modelopt` with `quant_algo=FP8`; `dit_layerwise_offload` is supported and `dit_cpu_offload` stays disabled |
|
||||
| `modelopt-nvfp4` | Mixed transformer directory/repo with `config.json`, or raw NVFP4 safetensors export/repo | `--transformer-path` for mixed overrides; `--transformer-weights-path` for raw exports | FLUX.1, FLUX.2, Wan2.2 | None | Mixed override repos keep the base model separate; raw exports such as `black-forest-labs/FLUX.2-dev-NVFP4` still use the weights-path flow |
|
||||
| `nunchaku-svdq` | Pre-quantized Nunchaku transformer weights, usually named `svdq-{int4\|fp4}_r{rank}-...` | `--transformer-weights-path` | Model-specific support such as Qwen-Image, FLUX, and Z-Image | `nunchaku` | SGLang can infer precision and rank from the filename and supports both `int4` and `nvfp4` |
|
||||
| `msmodelslim` | Pre-quantized msmodelslim transformer weights | `--model-path` | Wan2.2 family | None | Currently only compatible with the Ascend NPU family and supports both `w8a8` and `w4a4` |
|
||||
|
||||
## Validated ModelOpt Checkpoints
|
||||
|
||||
This section is the canonical support matrix for the diffusion ModelOpt
|
||||
This section is the canonical support matrix for the nine diffusion ModelOpt
|
||||
checkpoints currently wired up in SGLang docs and validation coverage.
|
||||
|
||||
Published checkpoints keep the serialized quantization config as
|
||||
`quant_method=modelopt`; the FP8 vs NVFP4 split below is a documentation label
|
||||
derived from `quant_algo`.
|
||||
|
||||
Seven of the eight repos live under `lmsys/*`. The FLUX.2 NVFP4 entry keeps the
|
||||
Eight of the nine repos live under `lmsys/*`. The FLUX.2 NVFP4 entry keeps the
|
||||
official `black-forest-labs/FLUX.2-dev-NVFP4` repo.
|
||||
|
||||
| Quant Algo | Base Model | Preferred CLI | HF Repo | Current Scope | Notes |
|
||||
@@ -65,13 +65,14 @@ official `black-forest-labs/FLUX.2-dev-NVFP4` repo.
|
||||
| `FP8` | `black-forest-labs/FLUX.1-dev` | `--transformer-path` | `lmsys/flux1-dev-modelopt-fp8-sglang-transformer` | single-transformer override, deterministic latent/image comparison, H100 benchmark, torch-profiler trace | SGLang converter keeps a validated BF16 fallback set for modulation and FF projection layers; use `--model-id FLUX.1-dev` for local mirrors |
|
||||
| `FP8` | `black-forest-labs/FLUX.2-dev` | `--transformer-path` | `lmsys/flux2-dev-modelopt-fp8-sglang-transformer` | single-transformer override load and generation path | published SGLang-ready transformer override |
|
||||
| `FP8` | `Wan-AI/Wan2.2-T2V-A14B-Diffusers` | `--transformer-path` | `lmsys/wan22-t2v-a14b-modelopt-fp8-sglang-transformer` | primary `transformer` quantized, `transformer_2` kept BF16 | primary-transformer-only path; keep `transformer_2` on the base checkpoint, and do not describe this as dual-transformer full-model FP8 unless that path is validated separately |
|
||||
| `FP8` | `hunyuanvideo-community/HunyuanVideo` | `--transformer-path` | `lmsys/hunyuanvideo-modelopt-fp8-sglang-transformer` | single-transformer override, BF16-vs-FP8 video comparison, H100 benchmark, torch-profiler trace | HunyuanVideo uses different ModelOpt/diffusers and SGLang runtime module names; the converter maps those names before writing FP8 scale tensors and BF16 fallback ignores |
|
||||
| `FP8` | `Qwen/Qwen-Image` | `--transformer-path` | `lmsys/qwen-image-modelopt-fp8-sglang-transformer` | single-transformer override, BF16-vs-FP8 image comparison, H100 benchmark, torch-profiler trace | shares the Qwen Image FP8 fallback preset; keep `img_in`, `txt_in`, timestep embedder, `norm_out.linear`, `proj_out`, `img_mod`/`txt_mod`, and `img_mlp.net.2` in BF16 |
|
||||
| `FP8` | `Qwen/Qwen-Image-Edit-2511` | `--transformer-path` | `lmsys/qwen-image-edit-modelopt-fp8-sglang-transformer` | TI2I edit smoke, BF16-vs-FP8 image comparison, H100 benchmark | shares `QwenImageTransformer2DModel` with Qwen Image and uses the same Qwen Image FP8 fallback preset |
|
||||
| `FP8` | `Qwen/Qwen-Image-Edit-2511` | `--transformer-path` | `lmsys/qwen-image-edit-modelopt-fp8-sglang-transformer` | TI2I edit path, BF16-vs-FP8 image comparison, H100 benchmark | shares `QwenImageTransformer2DModel` with Qwen Image and uses the same Qwen Image FP8 fallback preset |
|
||||
| `NVFP4` | `black-forest-labs/FLUX.1-dev` | `--transformer-path` | `lmsys/flux1-dev-modelopt-nvfp4-sglang-transformer` | mixed BF16+NVFP4 transformer override, correctness validation, 4x RTX 5090 benchmark, torch-profiler trace | use `build_modelopt_nvfp4_transformer.py`; validated builder keeps selected FLUX.1 modules in BF16 and sets `swap_weight_nibbles=false` |
|
||||
| `NVFP4` | `black-forest-labs/FLUX.2-dev` | `--transformer-weights-path` | `black-forest-labs/FLUX.2-dev-NVFP4` | packed-QKV load path | official raw export repo; validated packed export detection and runtime layout handling |
|
||||
| `NVFP4` | `Wan-AI/Wan2.2-T2V-A14B-Diffusers` | `--transformer-path` | `lmsys/wan22-t2v-a14b-modelopt-nvfp4-sglang-transformer` | primary `transformer` quantized with ModelOpt NVFP4, `transformer_2` kept BF16 | primary-transformer-only path; keep `transformer_2` on the base checkpoint, and current B200/Blackwell bring-up uses `SGLANG_DIFFUSION_FLASHINFER_FP4_GEMM_BACKEND=cudnn` |
|
||||
|
||||
These eight checkpoints are also the intended case set for the B200 diffusion
|
||||
These nine checkpoints are also the intended case set for the B200 diffusion
|
||||
CI job (`multimodal-gen-test-1-b200`).
|
||||
|
||||
## ModelOpt FP8
|
||||
@@ -98,6 +99,15 @@ sglang generate \
|
||||
--save-output
|
||||
```
|
||||
|
||||
```bash
|
||||
sglang generate \
|
||||
--model-path hunyuanvideo-community/HunyuanVideo \
|
||||
--transformer-path lmsys/hunyuanvideo-modelopt-fp8-sglang-transformer \
|
||||
--height 544 --width 960 --num-frames 17 \
|
||||
--prompt "A cinematic shot of a red sports car driving through rain at night" \
|
||||
--save-output
|
||||
```
|
||||
|
||||
```bash
|
||||
sglang generate \
|
||||
--model-path Qwen/Qwen-Image \
|
||||
@@ -131,6 +141,17 @@ sglang generate \
|
||||
- On disk, the quantization config stays `quant_method=modelopt` with
|
||||
`quant_algo=FP8`; the `modelopt-fp8` label in this document is a support
|
||||
family name, not a serialized config key.
|
||||
- `hunyuanvideo-community/HunyuanVideo` uses the `hunyuan-video` converter
|
||||
preset. Use `--model-type hunyuan-video` to force it, or rely on
|
||||
auto-detection from `_class_name=HunyuanVideoTransformer3DModel`.
|
||||
- The validated HunyuanVideo FP8 fallback preset keeps `context_embedder`,
|
||||
`x_embedder.proj`, timestep/guidance/text embedder linear layers,
|
||||
`norm_out.linear`, `proj_out`, double-block modulation linear layers, and
|
||||
single-block modulation linear layers in BF16.
|
||||
- HunyuanVideo ModelOpt exports use diffusers module names that do not match
|
||||
SGLang runtime module names for fused QKV and fused QKV+MLP layers. The
|
||||
converter maps the names before selecting scale tensors and before writing
|
||||
the runtime ignore list.
|
||||
- `Qwen/Qwen-Image` and `Qwen/Qwen-Image-Edit-2511` share the `qwen-image`
|
||||
converter preset. Use `--model-type qwen-image` to force it, or rely on
|
||||
auto-detection from `_class_name=QwenImageTransformer2DModel`.
|
||||
|
||||
Reference in New Issue
Block a user