docs(diffusion): update skills for MiniMax-H3 (#33282)
This commit is contained in:
@@ -7,7 +7,7 @@ description: Use when adding a new diffusion model or Diffusers pipeline to SGLa
|
||||
|
||||
Use this skill when adding a new diffusion model or pipeline variant to `sglang.multimodal_gen`.
|
||||
|
||||
## Two Pipeline Styles
|
||||
## Three Pipeline Styles
|
||||
|
||||
### Style A: Hybrid Monolithic Pipeline (Recommended)
|
||||
|
||||
@@ -33,11 +33,32 @@ This style is appropriate when:
|
||||
|
||||
See existing Modular examples: `QwenImagePipeline` (uses `add_standard_t2i_stages`), `FluxPipeline`, `WanPipeline`, `SanaPipeline`, `StableDiffusion3Pipeline`, and `ZImagePipeline`.
|
||||
|
||||
### Style C: Native Task-Contract Pipeline
|
||||
|
||||
Use this only when one checkpoint exposes multiple tightly coupled modalities
|
||||
or request profiles that cannot be represented safely by generic image/video
|
||||
sampling fields. MiniMax-H3 is the reference: it selects FL2VA or Ref2VA
|
||||
weights from one root model ID, validates canonical `task` / `conditions` /
|
||||
`target` requests before queueing, packs text/video/audio tokens into one
|
||||
denoise sequence, and returns synchronized video plus audio.
|
||||
|
||||
This style still uses `ComposedPipelineBase`, but owns a model-specific chain
|
||||
under `stages/model_specific_stages/<model>/`. Keep request validation, media
|
||||
materialization, packed-sequence construction, per-modality encode/decode, and
|
||||
presentation as explicit stages. Do not force coupled state into the standard
|
||||
`DenoisingStage` / `DecodingStage` contract just to resemble a simpler model.
|
||||
|
||||
Choose this style only with source evidence that the public API, scheduler, or
|
||||
joint latent state needs it. Preserve one canonical request object from API
|
||||
admission through offline generation and server execution so the two entry
|
||||
points cannot silently diverge.
|
||||
|
||||
### How to Choose
|
||||
|
||||
| Situation | Recommended Style |
|
||||
|-----------|-------------------|
|
||||
| Model has unique/complex pre-processing (VLM captioning, AR token generation, custom latent packing, etc.) | **Hybrid** — consolidate into a BeforeDenoisingStage |
|
||||
| Model jointly denoises multiple modalities or exposes partitioned task contracts from one root checkpoint | **Native task contract** — use MiniMax-H3 as the reference and keep model-specific stages explicit |
|
||||
| Model fits neatly into standard text-to-image or text+image-to-image pattern | **Modular** — use `add_standard_t2i_stages()` / `add_standard_ti2i_stages()` |
|
||||
| Porting a Diffusers pipeline with many custom steps | **Hybrid** — copy the `__call__` logic into a single stage |
|
||||
| Adding a variant of an existing model that shares most logic | **Modular** — reuse existing stages, customize via PipelineConfig callbacks |
|
||||
@@ -94,7 +115,7 @@ Once you have the reference code, study it thoroughly:
|
||||
**Before creating any new files, check whether an existing pipeline or stage can be reused or extended.** Only create new pipelines/stages when the existing ones would require extensive modifications or when no similar implementation exists.
|
||||
|
||||
Specifically:
|
||||
1. **Compare the new model's architecture against existing pipelines** before creating files. Current native families include LTX-2/2.3, HunyuanVideo/FastHunyuan, Wan/FastWan/TurboWan/LingBot World, MOVA, FLUX/FLUX.2/Klein, Z-Image, Qwen-Image/edit/layered, GLM-Image, SD3, Hunyuan3D, Helios, Cosmos3, SANA/SANA-WM, FireRed, ERNIE-Image, JoyAI, and Ideogram4. If the new model shares most of its structure with an existing one (e.g., same text encoders, similar latent format, compatible denoising loop), prefer:
|
||||
1. **Compare the new model's architecture against existing pipelines** before creating files. Current native families include MiniMax-H3, LTX-2/2.3, HunyuanVideo/FastHunyuan, Wan/FastWan/TurboWan/LingBot World, MOVA, FLUX/FLUX.2/Klein, Z-Image, Qwen-Image/edit/layered, GLM-Image, SD3, Hunyuan3D, Helios, Cosmos3, SANA/SANA-WM, FireRed, ERNIE-Image, JoyAI, and Ideogram4. If the new model shares most of its structure with an existing one (e.g., same text encoders, similar latent format, compatible denoising loop), prefer:
|
||||
- Adding a new config variant to the existing pipeline rather than creating a new pipeline class
|
||||
- Reusing the existing `BeforeDenoisingStage` with minor parameter differences
|
||||
- Using `add_standard_t2i_stages()` / `add_standard_ti2i_stages()` / `add_standard_ti2v_stages()` if the model fits standard patterns
|
||||
@@ -579,6 +600,12 @@ After implementation, **you must verify that the generated output is not noise**
|
||||
| FireRed/JoyAI image edit | `runtime/pipelines/qwen_image.py`, `runtime/pipelines/joy_image.py` | FireRed reuses Qwen edit-plus config; JoyAI has its own edit pipeline |
|
||||
| Wan | `runtime/pipelines/wan_pipeline.py` | Uses `add_standard_ti2v_stages()` |
|
||||
|
||||
### Native Task-Contract Style (coupled multimodal requests)
|
||||
|
||||
| Model | Pipeline | Request / stage references |
|
||||
|-------|----------|----------------------------|
|
||||
| MiniMax-H3 | `runtime/pipelines/minimax_h3_pipeline.py` | `configs/sample/minimax_h3.py` owns the canonical request fields; `stages/model_specific_stages/minimax_h3/` owns admission, material I/O, packed video/audio/text denoising, separate video/audio VAE work, and synchronized presentation |
|
||||
|
||||
---
|
||||
|
||||
## Checklist
|
||||
@@ -607,6 +634,17 @@ Before submitting, verify:
|
||||
- [ ] **BeforeDenoisingStage** at `stages/model_specific_stages/{model_name}.py`
|
||||
- [ ] `BeforeDenoisingStage.forward()` populates all fields needed by `DenoisingStage`
|
||||
|
||||
**Native task-contract style only:**
|
||||
|
||||
- [ ] Root checkpoint plus variant selection maps to the intended partition;
|
||||
do not require users to discover internal subdirectories
|
||||
- [ ] Offline `generate` and HTTP serving lower through the same validated
|
||||
request contract
|
||||
- [ ] Task, condition role/order, target canvas/time, and output container are
|
||||
rejected early when invalid
|
||||
- [ ] Joint-modality correctness covers every output stream; a valid video is
|
||||
insufficient when the model also generates audio or action data
|
||||
|
||||
## Common Pitfalls
|
||||
|
||||
1. **`batch.sigmas` must be a Python list**, not a numpy array. Use `.tolist()` to convert.
|
||||
|
||||
+13
@@ -33,6 +33,15 @@ B200-only groups are not currently inputs to the component-accuracy selector;
|
||||
add or identify a representative regular GPU case when that coverage is
|
||||
required.
|
||||
|
||||
Larger-topology smoke cases need the same explicit decision even when they are
|
||||
not enrolled in component accuracy. MiniMax-H3's current
|
||||
`MINIMAX_H3_FOUR_GPU_H100_CASES` is the reference: it exercises a real FL2VA
|
||||
request with TP2 + Ulysses2, but deliberately disables component accuracy and
|
||||
pipeline consistency because its native joint video/audio components do not
|
||||
have a directly comparable Diffusers pipeline contract. Pair this GPU smoke
|
||||
case with focused unit tests for request admission, packed-sequence layout,
|
||||
denoise scheduling, media handling, and VAE parallel-mode rejection.
|
||||
|
||||
The component-accuracy harness compares SGLang components against Diffusers/HF
|
||||
reference components. This is stricter than pipeline-level inference. New GPU
|
||||
cases commonly fail here for one of three reasons:
|
||||
@@ -87,6 +96,10 @@ Tests should cover:
|
||||
- single-GPU inference producing non-noise output
|
||||
- multi-GPU inference if TP/SP is supported
|
||||
- relevant unit tests for new math, parsing, scheduling, or loader behavior
|
||||
- every generated modality and delivery contract. For a joint model such as
|
||||
MiniMax-H3, validate the MP4 video stream, synchronized audio stream, frame
|
||||
rate/sample rate, and multi-output grouping; a visually valid frame sequence
|
||||
alone is not sufficient
|
||||
|
||||
For performance data:
|
||||
|
||||
|
||||
+8
-2
@@ -42,10 +42,10 @@ If any benchmark, perf-dump, or `torch.profiler` command prints one of those sig
|
||||
|
||||
## Main Reference
|
||||
|
||||
- [benchmark-and-profile.md](benchmark-and-profile.md) — canonical denoise benchmark, perf dump, and `torch.profiler` workflow; uses checked-in nightly-aligned presets plus current-source extras such as FLUX.2 Klein, Cosmos3, Ideogram4, ERNIE/GLM/SANA image models, FastWan2.2, `LTX-2.3` one-stage/two-stage/HQ, HunyuanVideo, MOVA, Helios, JoyAI/FireRed image edit, and Hunyuan3D shape
|
||||
- [benchmark-and-profile.md](benchmark-and-profile.md) — canonical denoise benchmark, perf dump, and `torch.profiler` workflow; uses checked-in nightly-aligned presets plus current-source extras such as MiniMax-H3 joint video/audio T2VA, FLUX.2 Klein, Cosmos3, Ideogram4, ERNIE/GLM/SANA image models, FastWan2.2, `LTX-2.3` one-stage/two-stage/HQ, HunyuanVideo, MOVA, Helios, JoyAI/FireRed image edit, and Hunyuan3D shape
|
||||
- [existing-fast-paths.md](existing-fast-paths.md) — map bottlenecks to existing fused kernels, packed QKV paths, fused `QK norm + RoPE`, distributed overlap patterns, and open optimization PRs before proposing new code
|
||||
- [scripts/diffusion_skill_env.py](scripts/diffusion_skill_env.py) — preflight helper: repo root discovery via `sglang.__file__`, write-access probe, benchmark/profile output directories, idle GPU selection
|
||||
- [scripts/bench_diffusion_denoise.py](scripts/bench_diffusion_denoise.py) — end-to-end denoise benchmark preset runner via `sglang generate`; supports `--no-torch-compile`, validates nightly preset drift with `--validate-nightly-alignment`, and saves perf dumps by label for `compare_perf.py`
|
||||
- [scripts/bench_diffusion_denoise.py](scripts/bench_diffusion_denoise.py) — end-to-end denoise benchmark preset runner via `sglang generate`; supports `--no-torch-compile`, forces the H3 preset to its eager consistency mode, validates nightly preset drift with `--validate-nightly-alignment`, and saves perf dumps by label for `compare_perf.py`
|
||||
|
||||
## Opportunity Discovery Rule
|
||||
|
||||
@@ -56,6 +56,8 @@ Always rule out these existing families first:
|
||||
- LTX upsampler GroupNorm+SiLU
|
||||
- Z-Image bf16-native Triton RMSNorm scale/tanh-residual modulation
|
||||
- SANA packed self-attention Q/K/V and cross-attention K/V GEMMs
|
||||
- MiniMax-H3 indexed modulation, fused QK norm + RoPE, packed Ulysses QKV,
|
||||
USP relayout, and batched TP AdaLN collectives
|
||||
- fused diffusion `QK norm + RoPE`
|
||||
- LTX2 split RoPE
|
||||
- LTX2 residual-gate add
|
||||
@@ -72,6 +74,10 @@ default benchmark preset invocation unchanged. Either pass the checked-in
|
||||
benchmark helper its no-compile switch or run the equivalent manual command
|
||||
without `--enable-torch-compile`.
|
||||
|
||||
MiniMax-H3 is always an eager consistency case on current main. Use
|
||||
`--model minimax-h3-t2va`; its preset writes the H3 request fields through a
|
||||
generated config and suppresses the helper's global compile default.
|
||||
|
||||
For FLUX-family manual profiling runs with a quantized transformer override:
|
||||
- use `sglang generate` directly
|
||||
- pass the override as `--transformer-path <dir>`
|
||||
|
||||
+117
-1
@@ -37,11 +37,15 @@ python3 "$ENV_PY" check-write-access >/dev/null
|
||||
|
||||
export HF_TOKEN=<your_hf_token> # required for gated repos such as black-forest-labs/FLUX.*
|
||||
export FLASHINFER_DISABLE_VERSION_CHECK=1
|
||||
export CUDA_VISIBLE_DEVICES=$(python3 "$ENV_PY" print-idle-gpus --count 1)
|
||||
# Leave CUDA_VISIBLE_DEVICES unset to let the preset helper select the number
|
||||
# of idle GPUs it requires. For manual runs, set --count to that command's
|
||||
# exact --num-gpus value.
|
||||
|
||||
ASSET_DIR=$(python3 "$ENV_PY" print-assets-dir --mkdir)
|
||||
BENCH_DIR=$(python3 "$ENV_PY" print-output-dir --kind benchmarks --mkdir)
|
||||
PROFILE_DIR=$(python3 "$ENV_PY" print-output-dir --kind profiles --mkdir)
|
||||
CONFIG_DIR="${BENCH_DIR}/generated_configs"
|
||||
mkdir -p "${CONFIG_DIR}"
|
||||
export PROFILE_DIR
|
||||
|
||||
check() {
|
||||
@@ -163,6 +167,17 @@ PYTHONPATH=python python3 "$BENCH_PY" \
|
||||
--output-dir "${BENCH_DIR}"
|
||||
```
|
||||
|
||||
Run the current-source MiniMax-H3 T2VA preset. The helper forces eager mode
|
||||
for this model even when its global compile default is enabled:
|
||||
|
||||
```bash
|
||||
export CUDA_VISIBLE_DEVICES=$(python3 "$ENV_PY" print-idle-gpus --count 4)
|
||||
PYTHONPATH=python python3 "$BENCH_PY" \
|
||||
--model minimax-h3-t2va \
|
||||
--label baseline \
|
||||
--output-dir "${BENCH_DIR}"
|
||||
```
|
||||
|
||||
Run the full preset sweep only when you have enough GPU time for both the
|
||||
nightly-aligned cases and the source-tracked extras:
|
||||
|
||||
@@ -202,6 +217,7 @@ Use the preset categories this way:
|
||||
| `ideogram4-fp8` | `ideogram-ai/ideogram-4-fp8` | Yes: `ideogram4_fp8_t2i_2gpu` | Prompt, 1024x1024, seed 42, 2 GPUs, TP size 2, FlashAttention backend; sampling preset owns steps/guidance |
|
||||
| `cosmos3-super-t2v` | `nvidia/Cosmos3-Super` | Yes: `cosmos3_super_t2v_2gpu` | Prompt, 1280x720, 81 frames, seed 42, 2 GPUs, TP size 2, guardrails disabled for benchmark isolation |
|
||||
| `wan-i2v` | `Wan-AI/Wan2.2-I2V-A14B-Diffusers` | Yes: `wan22_i2v_a14b_720p` | Nightly cat image and motion prompt, 1280x720, 81 frames, 4 GPUs, CFG parallel, Ulysses degree 2, text encoder CPU offload and pinned CPU memory |
|
||||
| `minimax-h3-t2va` | `MiniMaxAI/MiniMax-H3` | No | Current-source H3 FL2VA-partition T2VA baseline: 1344x768 resolved canvas, 5 seconds / 124 frames at 24 fps, 50 joint video-audio steps, 4 GPUs, TP2 + Ulysses2, eager BF16/FP32. The helper writes H3's `task`, `conditions`, `target`, and audio/video flow shifts to a generated config. |
|
||||
| `ltx2` | `Lightricks/LTX-2` | No | Current-source two-stage LTX-2 preset with 2 GPUs, CFG parallel, 768x512, 121 frames |
|
||||
| `qwen-image` | `Qwen/Qwen-Image` | No | Current-source extra covering the base Qwen-Image native path, separate from the nightly `Qwen-Image-2512` case |
|
||||
| `qwen-edit-2509` | `Qwen/Qwen-Image-Edit-2509` | No | Current-source extra for the pre-2511 edit-plus path; uses the cat image, 1024x1024 |
|
||||
@@ -232,6 +248,92 @@ and **best latency tuning**:
|
||||
- do not assume that is the fastest topology
|
||||
- for pure latency tuning, benchmark pure Ulysses too, for example `--ulysses-degree=4 --ring-degree=1` on 4 GPUs, and on 8 GPUs compare pure `--ulysses-degree=8` against `--enable-cfg-parallel --ulysses-degree=4`
|
||||
|
||||
For MiniMax-H3, keep the native contract intact:
|
||||
|
||||
- use the root model ID and select `fl2va` or `ref2va` with
|
||||
`--model-variant`; do not point at a checkpoint subdirectory
|
||||
- use eager BF16/FP32 for consistency ground truth; current H3
|
||||
`torch.compile` changes numerical output
|
||||
- use Ulysses, not Ring, for H3's packed multi-segment attention; CFG parallel
|
||||
is invalid because the released pipeline has one denoising branch
|
||||
- keep the released overlapping tiled video-VAE decode. H3 rejects
|
||||
`spatial`, `spatial_shard`, and patch decode modes after output mismatches
|
||||
|
||||
### Manual command example: MiniMax-H3 T2VA
|
||||
|
||||
Create `${CONFIG_DIR}/minimax-h3-t2va.json` with the model-specific request
|
||||
fields below. The generic width, height, and frame flags are intentionally
|
||||
absent because H3 resolves all three from `target`:
|
||||
|
||||
```json
|
||||
{
|
||||
"task": "t2va",
|
||||
"conditions": [],
|
||||
"target": {
|
||||
"short_edge": 768,
|
||||
"aspect_ratio": "16:9",
|
||||
"duration_seconds": 5.0
|
||||
},
|
||||
"num_inference_steps": 50,
|
||||
"flow_shift": 12.0,
|
||||
"audio_flow_shift": 3.0
|
||||
}
|
||||
```
|
||||
|
||||
Then run the same lossless 4-GPU H100 topology and 5-second shape used by the
|
||||
source-tracked preset:
|
||||
|
||||
```bash
|
||||
sglang generate \
|
||||
--backend=sglang \
|
||||
--model-path=MiniMaxAI/MiniMax-H3 \
|
||||
--model-variant=fl2va \
|
||||
--config="${CONFIG_DIR}/minimax-h3-t2va.json" \
|
||||
--prompt="At night, while their owner sleeps in a bedroom, three cats march in loudly playing tiny brass instruments, then abruptly file out." \
|
||||
--seed=1101 --num-gpus=4 --tp-size=2 --ulysses-degree=2 \
|
||||
--performance-mode=speed --enable-torch-compile=false \
|
||||
--save-output --warmup \
|
||||
--perf-dump-path="${BENCH_DIR}/minimax-h3-t2va-baseline.json"
|
||||
```
|
||||
|
||||
The benchmark helper creates this config automatically. For ModelScope, set
|
||||
`SGLANG_USE_MODELSCOPE=true`, replace the root model ID with
|
||||
`MiniMax/MiniMax-H3`, and keep the selected variant unchanged.
|
||||
When `--output-dir` is provided, the helper places the generated config under
|
||||
that directory's `generated_configs/` subdirectory so the run is self-contained.
|
||||
|
||||
For a serving benchmark, use the driver maintained by the H3 cookbook after
|
||||
launching the corresponding `sglang serve` command:
|
||||
|
||||
```bash
|
||||
python3 -m sglang.multimodal_gen.benchmarks.bench_serving \
|
||||
--host 127.0.0.1 --port 30010 \
|
||||
--model MiniMaxAI/MiniMax-H3 \
|
||||
--dataset vbench --task text-to-video \
|
||||
--num-prompts 1 --max-concurrency 1 \
|
||||
--warmup-requests 1 --warmup-inference-steps 50 \
|
||||
--extra-body '{"task":"t2va","conditions":[],"target":{"short_edge":768,"aspect_ratio":"16:9","duration_seconds":5.0},"seconds":5,"flow_shift":12.0,"audio_flow_shift":3.0}'
|
||||
```
|
||||
|
||||
H3 correctness is joint video/audio correctness. Use eager BF16/FP32 as the
|
||||
only ground truth and keep prompt, seed, target, step count, shifts, partition,
|
||||
and topology fixed. For a lossless kernel/runtime change:
|
||||
|
||||
- compare decoded frames after frame-count and timestamp alignment; report at
|
||||
least frame-wise PSNR/SSIM plus the worst frame, not only an average
|
||||
- extract the 32 kHz stereo audio stream and compare channel order, sample
|
||||
count, waveform error, and a time-aligned log-mel or spectral metric
|
||||
- verify the MP4 contract remains H.264 video at 24 fps plus one AAC stereo
|
||||
audio stream
|
||||
- run the relevant kernel/unit exactness test when replacing an existing H3
|
||||
BF16 fast path. Do not hide a failed exact test behind a permissive
|
||||
end-to-end perceptual threshold
|
||||
|
||||
There is no source-wide universal perceptual threshold for arbitrary H3
|
||||
changes. Record the acceptance bounds before optimization and tighten them for
|
||||
changes that claim to preserve eager math. Approximate Cache-DiT or FP8 runs
|
||||
must be labeled separately and validated for both output modalities.
|
||||
|
||||
### Manual command example: LTX-2 Two-Stage
|
||||
|
||||
```bash
|
||||
@@ -417,6 +519,18 @@ Keep model shape, seed, and GPU topology fixed for every comparison. Save one
|
||||
reference image or video before changing code. If the active task requires
|
||||
`torch.compile` off, add `--no-torch-compile` here too.
|
||||
|
||||
MiniMax-H3 always requires eager mode for consistency ground truth. The
|
||||
`minimax-h3-t2va` helper preset enforces it, and manual H3 profile commands
|
||||
must pass `--enable-torch-compile=false`.
|
||||
|
||||
For H3, one `--profile-all-stages` trace separates text/condition encoding,
|
||||
`MiniMaxH3DenoisingStage`, and the aggregate `MiniMaxH3DecodingStage`. The
|
||||
decoding stage contains both video decode and rank-0 audio decode. If decoding
|
||||
is hot, add temporary `record_function` or NVTX scopes around
|
||||
`video_vae.decode_base` and `_decode_audio` in the H3 decoding stage, then
|
||||
re-run the same all-stage profile. Do not attribute aggregate decoding time to
|
||||
one VAE without those inner scopes.
|
||||
|
||||
### 2. Capture a representative trace
|
||||
|
||||
By default SGLang profiles the denoising stage. The default sampling window is
|
||||
@@ -493,6 +607,8 @@ the known mainline families.
|
||||
| LTX-2 split RoPE appears as a long PyTorch elementwise chain | Check the `apply_ltx2_split_rotary_emb` Triton path and its shape guards |
|
||||
| masked attention spends time packing/unpacking Q/K/V | Check whether fused varlen USP pack/scatter should have engaged |
|
||||
| `all_to_all`, ring attention, or async A2A dominate | Classify against Ulysses, USP, or turbo-layer overlap first |
|
||||
| H3 shows separate indexed gather + scale/shift, QK norm + RoPE, or three Q/K/V Ulysses relayouts | Check H3's indexed-modulation, fused QK-norm+RoPE, packed Ulysses-QKV, and USP relayout guards before writing a new kernel |
|
||||
| H3 TP traces show one AdaLN collective per block | Check the batched TP AdaLN projection/all-gather path in `minimax_h3.py` before attempting communication overlap |
|
||||
| split `fc1 -> gelu -> quant -> fc2.lora_down` on Nunchaku FLUX | Treat as a missing fused GELU MLP path |
|
||||
| attention kernels dominate | Confirm backend, topology, and shape guards before proposing a new kernel |
|
||||
|
||||
|
||||
+31
-1
@@ -18,6 +18,11 @@ framework-specific optimization workflow.
|
||||
- `python/sglang/kernels/ops/diffusion/triton/zimage_native_norm.py`
|
||||
- `python/sglang/kernels/ops/diffusion/triton/rotary.py`
|
||||
- `python/sglang/kernels/ops/diffusion/triton/ltx2_rotary.py`
|
||||
- `python/sglang/kernels/ops/diffusion/triton/indexed_modulation.py`
|
||||
- `python/sglang/kernels/ops/diffusion/triton/ulysses_qkv.py`
|
||||
- `python/sglang/kernels/ops/diffusion/usp_relayout.py`
|
||||
- `python/sglang/multimodal_gen/runtime/layers/usp.py`
|
||||
- `python/sglang/multimodal_gen/runtime/models/dits/minimax_h3.py`
|
||||
- `python/sglang/kernels/ops/diffusion/residual_gate_add.py`
|
||||
- `python/sglang/kernels/jit/csrc/diffusion/residual_gate_add.cuh`
|
||||
- `python/sglang/kernels/ops/diffusion/triton/varlen_pack_pad.py`
|
||||
@@ -106,7 +111,23 @@ framework-specific optimization workflow.
|
||||
- Microbench: `test/registered/kernels/benchmark/diffusion/bench_residual_gate_add.py`.
|
||||
- Workflow rule: if LTX2 traces show repeated elementwise `mul` + `add` ladders around attention or MLP residuals, check whether this existing CUDA path was disabled by shape, dtype, contiguity, or a prior runtime failure before proposing another elementwise fusion.
|
||||
|
||||
9. HunyuanVideo / LTX upsampler GroupNorm + SiLU fusion
|
||||
9. MiniMax-H3 indexed AdaLN modulation and gated residual fusion
|
||||
- Kernels: `indexed_scale_shift_bf16_`, `indexed_gate_bf16_`
|
||||
- Locations: `triton/indexed_modulation.py`, `runtime/models/dits/minimax_h3.py`
|
||||
- Use cases: H3's packed video/audio/text rows select per-token modulation with `combined_indices`; the Triton paths replace `index_select` plus scale/shift or gated residual chains in place.
|
||||
- Constraints: CUDA BF16 H3 tensors, BF16 modulation tensors, contiguous disposable inputs; the gated path also requires contiguous `other`. Unsupported shapes/dtypes retain the eager formula.
|
||||
- Numerical contract: the kernels explicitly reproduce H3's eager BF16 rounding boundaries. Do not replace them with a mathematically equivalent contraction without the H3 consistency check.
|
||||
- Workflow rule: if H3 traces show `index_select` plus elementwise ladders around every block, check dtype, contiguity, and input-reuse eligibility before designing another modulation kernel.
|
||||
|
||||
10. MiniMax-H3 packed Ulysses QKV and output relayout
|
||||
- Kernels: `pack_qkv_destination_major`, `usp_merge_heads`
|
||||
- Locations: `triton/ulysses_qkv.py`, `usp_relayout.py`, `runtime/layers/usp.py`, `runtime/models/dits/minimax_h3.py`
|
||||
- Use cases: one destination-major QKV pack plus one collective replaces three separately prepared Ulysses input exchanges; the output JIT kernel replaces `permute(...).contiguous()` when merging gathered heads.
|
||||
- Constraints: packed QKV fast packing requires CUDA fp16/bf16 Q/K/V with matching dtypes, contiguous head dimension, and eager execution. `usp_merge_heads` requires a nonempty contiguous 5D CUDA fp16/bf16/fp32 tensor and is disabled inside `torch.compile`.
|
||||
- Related transport: 2-rank, peer-accessible CUDA groups can use the existing IPC A2A transport; larger or unsupported groups fall back to the normal collective path.
|
||||
- Workflow rule: if an H3 Ulysses trace has three Q/K/V preparation ladders or a large output `permute + contiguous`, first prove why these existing guards missed.
|
||||
|
||||
11. HunyuanVideo / LTX upsampler GroupNorm + SiLU fusion
|
||||
- Kernel: `triton_group_norm_silu`
|
||||
- Locations: `diffusion/group_norm_silu.py`, `triton/group_norm_silu.py`, `runtime/models/vaes/hunyuanvae.py`, `runtime/models/upsampler/latent_upsampler.py`
|
||||
- Use case: `activation(group_norm(x))` when the activation is non-inplace `nn.SiLU` and the GroupNorm is affine.
|
||||
@@ -166,6 +187,11 @@ framework-specific optimization workflow.
|
||||
- Supported head dims: `64, 128, 256`.
|
||||
- Behavior: `apply_qk_norm_rope` prefers the fused JIT kernel when all guards pass; otherwise it falls back to `apply_qk_norm(...)` plus `apply_flashinfer_rope_qk_inplace(...)`.
|
||||
- Validation: `test/registered/kernels/ops/diffusion/test_qknorm_rope.py`.
|
||||
- MiniMax-H3: the H3 DiT calls `fused_inplace_qknorm_rope` directly for BF16
|
||||
head dim 128 with 96 rotary dims, NeoX layout, and
|
||||
`round_norm_before_rope=True`. This flag is part of H3's eager numerical
|
||||
contract. Compiled execution deliberately falls back to separate eager
|
||||
operations.
|
||||
- Workflow rule: treat LTX2 traces that miss the generic fused path as an enablement/shape-guard issue first, and check the separate LTX2 split-RoPE path before proposing new attention-prep kernels.
|
||||
|
||||
**Nunchaku Fused GELU MLP**
|
||||
@@ -197,6 +223,8 @@ framework-specific optimization workflow.
|
||||
`zimage_rmsnorm_tanh_mul_add` in `zimage.py`, backed by
|
||||
`triton/zimage_native_norm.py`.
|
||||
- HunyuanVideo VAE and LTX upsampler GroupNorm+SiLU: `apply_group_norm_silu` in `hunyuanvae.py` and `latent_upsampler.py`; default-eligible when wrapper guards pass.
|
||||
- MiniMax-H3 indexed modulation: `_modulate_scale_shift` and `_modulate_gate` in `minimax_h3.py`, backed by `triton/indexed_modulation.py`.
|
||||
- MiniMax-H3 Ulysses relayout: `_usp_input_all_to_all_packed_qkv` and `usp_merge_heads` through `runtime/layers/usp.py`.
|
||||
- QK norm: `apply_qk_norm` used in `flux.py`, `flux_2.py`, `qwen_image.py`, `zimage.py`, `wanvideo.py`, `ltx_2.py`, `hunyuanvideo.py`.
|
||||
- QK norm + RoPE: `apply_qk_norm_rope` in `layernorm.py`; use this path when the model wants fused attention prep instead of separate QK norm and RoPE calls.
|
||||
- LTX2 split RoPE: `apply_ltx2_split_rotary_emb` in `ltx_2.py`.
|
||||
@@ -210,6 +238,8 @@ framework-specific optimization workflow.
|
||||
**Existing Overlap / Communication Families**
|
||||
|
||||
- Ulysses / USP attention: treat `all_to_all`, `ring_attn`, and head / sequence reshards as an existing distributed attention family, not a new overlap idea.
|
||||
- MiniMax-H3 TP AdaLN: the DiT stacks every block's TP-local AdaLN projection and performs one batched all-gather before the block loop when `_can_batch_block_adaln()` passes. One all-gather per block indicates that this existing batching path missed.
|
||||
- MiniMax-H3 final projections: H3 removes dead text/padding rows before the final TP column gathers and combines video/audio for the SP row gather. Preserve that ordering when optimizing output communication.
|
||||
- Turbo-layer async all-to-all: `all_to_all_single(..., async_op=True)` plus staged waits already form an existing overlap family in `turbo_layer.py`.
|
||||
- TorchInductor compute / communication reorder: `torch._inductor.config.reorder_for_compute_comm_overlap = True` can already partially overlap compiled denoise traces.
|
||||
- Dual-stream diffusion models: `use_dual_stream = True` in models such as `hunyuan3d.py` is an existing overlap family.
|
||||
|
||||
+40
-3
@@ -249,6 +249,38 @@ MODELS = {
|
||||
],
|
||||
},
|
||||
# Source-tracked extras from current registry / GPU test coverage.
|
||||
# MiniMax-H3 owns its temporal canvas through target.duration_seconds, so
|
||||
# the model-specific sampling fields are passed through --config instead
|
||||
# of generic --width/--height/--num-frames flags.
|
||||
"minimax-h3-t2va": {
|
||||
"path": "MiniMaxAI/MiniMax-H3",
|
||||
"prompt": "At night, while their owner sleeps in a bedroom, three cats march in loudly playing tiny brass instruments, then abruptly file out.",
|
||||
"seed": 1101,
|
||||
"config_overrides": {
|
||||
"task": "t2va",
|
||||
"conditions": [],
|
||||
"target": {
|
||||
"short_edge": 768,
|
||||
"aspect_ratio": "16:9",
|
||||
"duration_seconds": 5.0,
|
||||
},
|
||||
"audio_flow_shift": 3.0,
|
||||
"flow_shift": 12.0,
|
||||
"num_inference_steps": 50,
|
||||
},
|
||||
"extra_args": [
|
||||
"--model-variant=fl2va",
|
||||
"--num-gpus=4",
|
||||
"--tp-size=2",
|
||||
"--ulysses-degree=2",
|
||||
"--performance-mode=speed",
|
||||
"--enable-torch-compile=false",
|
||||
],
|
||||
# H3 eager BF16/FP32 is the consistency ground truth. Current
|
||||
# torch.compile changes numerical output, so never add the global
|
||||
# helper default --enable-torch-compile flag for this preset.
|
||||
"force_eager": True,
|
||||
},
|
||||
"ltx2": {
|
||||
"path": "Lightricks/LTX-2",
|
||||
"prompt": "A cat and a dog baking a cake together in a kitchen.",
|
||||
@@ -730,6 +762,7 @@ def build_sglang_cmd(
|
||||
torch_compile: bool = True,
|
||||
seed: int = 42,
|
||||
save_output: bool = True,
|
||||
artifact_dir: Optional[Path] = None,
|
||||
) -> list[str]:
|
||||
"""
|
||||
Build the `sglang generate` command for the given model.
|
||||
@@ -756,9 +789,12 @@ def build_sglang_cmd(
|
||||
cmd.append(f"--image-path={cfg['image_path']}")
|
||||
|
||||
if "config_overrides" in cfg:
|
||||
config_dir = ensure_dir(
|
||||
get_output_dir("benchmarks", REPO_ROOT) / "generated_configs"
|
||||
config_root = (
|
||||
Path(artifact_dir)
|
||||
if artifact_dir is not None
|
||||
else get_output_dir("benchmarks", REPO_ROOT)
|
||||
)
|
||||
config_dir = ensure_dir(config_root / "generated_configs")
|
||||
config_path = config_dir / f"{model_key}.json"
|
||||
with open(config_path, "w") as f:
|
||||
json.dump(cfg["config_overrides"], f, indent=2, sort_keys=True)
|
||||
@@ -770,7 +806,7 @@ def build_sglang_cmd(
|
||||
cmd.append("--save-output")
|
||||
if warmup:
|
||||
cmd.append("--warmup")
|
||||
if torch_compile:
|
||||
if torch_compile and not cfg.get("force_eager", False):
|
||||
cmd.append("--enable-torch-compile")
|
||||
if perf_dump_path:
|
||||
cmd.extend(["--perf-dump-path", perf_dump_path])
|
||||
@@ -793,6 +829,7 @@ def run_benchmark_once(
|
||||
perf_dump_path=str(perf_path),
|
||||
warmup=warmup,
|
||||
torch_compile=torch_compile,
|
||||
artifact_dir=output_dir,
|
||||
)
|
||||
|
||||
env = os.environ.copy()
|
||||
|
||||
@@ -82,6 +82,22 @@ checkpoint repos; the FLUX.2 NVFP4 raw export remains
|
||||
`black-forest-labs/FLUX.2-dev-NVFP4`. Do not use older `BBuf/*` examples unless
|
||||
you are explicitly testing a historical branch.
|
||||
|
||||
### MiniMax-H3 boundary
|
||||
|
||||
MiniMax-H3 is current-main evidence for the separate online FP8 path, not a
|
||||
validated ModelOpt PTQ/export family. Its verified B200/B300 serving recipe
|
||||
loads the unquantized root checkpoint with `--quantization fp8` and preserves
|
||||
the video/audio patch projections, timestep MLP, and final video/audio heads in
|
||||
FP32. Do not add H3 to the ModelOpt support matrix or run the generic ModelOpt
|
||||
converter until an exact H3 export, loader mapping, accuracy check, and
|
||||
benchmark scope have been validated.
|
||||
|
||||
If the user asks for current H3 online quantization, route the command and
|
||||
quality caveats through `sglang-diffusion-performance` and the MiniMax-H3
|
||||
cookbook. Online FP8 is approximate and must be compared against eager
|
||||
BF16/FP32 for both video and audio; combining it with Cache-DiT compounds two
|
||||
approximations.
|
||||
|
||||
## Related PR Watchlist
|
||||
|
||||
These related SGLang PRs are useful as ModelOpt diffusion support history.
|
||||
|
||||
@@ -64,6 +64,64 @@ These options **trade output quality** for speed or VRAM savings. Results will d
|
||||
|
||||
## Quick Recipes
|
||||
|
||||
### MiniMax-H3 first: lossless joint video/audio
|
||||
|
||||
H3 has a stricter contract than the generic recipes below. Keep its DiT eager
|
||||
for consistency ground truth, use Ulysses rather than Ring, do not enable CFG
|
||||
parallel, and leave the released overlapping tiled video-VAE decode in place.
|
||||
|
||||
Four H200 GPUs can keep the complete BF16/FP32 pipeline resident:
|
||||
|
||||
```bash
|
||||
sglang serve \
|
||||
--model-path MiniMaxAI/MiniMax-H3 \
|
||||
--model-variant fl2va \
|
||||
--num-gpus 4 \
|
||||
--ulysses-degree 4 \
|
||||
--performance-mode speed \
|
||||
--enable-torch-compile false \
|
||||
--port 30010
|
||||
```
|
||||
|
||||
On 4x H100 80 GB, start from the fastest measured lossless resident topology:
|
||||
|
||||
```bash
|
||||
sglang serve \
|
||||
--model-path MiniMaxAI/MiniMax-H3 \
|
||||
--model-variant fl2va \
|
||||
--num-gpus 4 \
|
||||
--tp-size 2 \
|
||||
--ulysses-degree 2 \
|
||||
--performance-mode speed \
|
||||
--enable-torch-compile false \
|
||||
--port 30010
|
||||
```
|
||||
|
||||
On B200/B300, the verified resident sweep uses 8 GPUs with Ulysses8. H3 also
|
||||
has a verified 4x B200 FSDP-capacity path, but FSDP all-gathers are a memory
|
||||
policy rather than the default latency choice. Benchmark the target topology
|
||||
with the H3 driver from `sglang-diffusion-benchmark-profile`.
|
||||
|
||||
Use the FL2VA partition for both `t2va` and `fl2va`; use
|
||||
`--model-variant ref2va` for image/video/audio reference conditioning. The root
|
||||
IDs are `MiniMaxAI/MiniMax-H3` on Hugging Face and `MiniMax/MiniMax-H3` on
|
||||
ModelScope. Do not point `--model-path` at a partition subdirectory.
|
||||
|
||||
Current H3 restrictions:
|
||||
|
||||
- `torch.compile` is opt-in experimentation only because it changes numerical
|
||||
output; it is not a lossless baseline
|
||||
- Ring attention and CFG parallel are incompatible with the packed single
|
||||
denoising branch
|
||||
- SageAttention is rejected for the current packed multi-segment attention
|
||||
- `--vae-config.parallel-decode-mode spatial`, `spatial_shard`, and patch VAE
|
||||
decode are rejected after mismatches; use the default tiled recipe
|
||||
- Breakable CUDA Graph is opt-in and signature-specific; the validated
|
||||
1344x768 Ref2VA capture uses `--bcg-text-buckets 5504`, but it did not show a
|
||||
measured speedup
|
||||
- the `quality=high|medium|low` Cache-DiT profiles and online FP8 are
|
||||
approximate; keep them outside lossless comparisons
|
||||
|
||||
### Maximum speed, video model, multi-GPU, lossless (Wan A14B, 8 GPUs)
|
||||
|
||||
```bash
|
||||
@@ -255,6 +313,7 @@ Use these as first commands to benchmark, not as universal winners.
|
||||
|
||||
| Model family | First performance shape | Starting flags | Notes |
|
||||
|---|---|---|---|
|
||||
| MiniMax-H3 | 1344x768 resolved canvas, 5 seconds / 124 frames at 24 fps, 50 joint video/audio steps | H200: `--num-gpus 4 --ulysses-degree 4 --performance-mode speed --enable-torch-compile false`; H100: TP2 + Ulysses2 | Root ID plus `--model-variant fl2va` for T2VA/FL2VA or `ref2va` for Ref2VA. Ulysses only; no Ring/CFG/SageAttention. Preserve tiled video-VAE decode. Profile joint denoise, video VAE, audio VAE/vocoder, encoder, and collectives separately. |
|
||||
| FLUX.1 / FLUX.2 image | 1024x1024, runtime-default steps/guidance, 1 GPU | `--enable-torch-compile --warmup --dit-layerwise-offload false` | `black-forest-labs/FLUX.*` repos are gated; for FP8/NVFP4 use validated `--transformer-path` or `--transformer-weights-path` flows from the quant skill. |
|
||||
| FLUX.2 Klein / Klein Base | 1024x1024, runtime-default steps/guidance, 1 GPU | `--enable-torch-compile --warmup --dit-layerwise-offload false` | Current registry has `black-forest-labs/FLUX.2-klein-4B`, `FLUX.2-klein-9B`, and base variants. Klein is step-distilled; Klein Base is not. |
|
||||
| Qwen-Image / Qwen-Image-Edit | 1024x1024, runtime-default steps/guidance, 1 GPU | `--enable-torch-compile --warmup`; optionally native `SGLANG_CACHE_DIT_ENABLED=true` | Cache-DiT is lossy. For edit tasks, keep reference image, seed, and output size fixed. |
|
||||
@@ -289,6 +348,6 @@ about whether the work has merged:
|
||||
- **Offload tuning**: after the first request, the runtime logs peak GPU memory and which components could stay resident. Use this to decide which `--*-cpu-offload` flags to disable.
|
||||
- **Backend selection**: `--backend sglang` (default, auto-detected) enables native optimizations (fused kernels, SP, native Cache-DiT env knobs, etc.). `--backend diffusers` falls back to Diffusers pipelines and is the path that accepts `--cache-dit-config` plus diffusers attention backend names.
|
||||
- **Wan2.2-I2V sizing**: explicit `--width/--height` on `Wan2.2-I2V-A14B` control the target area while preserving the condition-image aspect ratio.
|
||||
- **Mainline diffusion fast paths**: before proposing a new kernel or overlap scheme, check `sglang-diffusion-benchmark-profile/existing-fast-paths.md`. It covers GroupNorm+SiLU, Z-Image bf16-native Triton norm modulation, fused diffusion `QK norm + RoPE`, LTX2 split RoPE, LTX2 residual-gate add, varlen USP pack/scatter, packed QKV/NVFP4 expectations, and existing multi-GPU overlap families such as Ulysses / USP and turbo-layer async all-to-all.
|
||||
- **Mainline diffusion fast paths**: before proposing a new kernel or overlap scheme, check `sglang-diffusion-benchmark-profile/existing-fast-paths.md`. It covers H3 indexed modulation, fused QK norm + RoPE, packed Ulysses QKV/USP relayout and batched TP AdaLN, plus GroupNorm+SiLU, Z-Image bf16-native Triton norm modulation, LTX2 split RoPE, LTX2 residual-gate add, varlen USP pack/scatter, packed QKV/NVFP4 expectations, and existing multi-GPU overlap families such as Ulysses / USP and turbo-layer async all-to-all.
|
||||
- **NVFP4 trace interpretation**: on FLUX.2 NVFP4 and Nunchaku-style checkpoints, packed QKV is expected. SGLang intentionally uses fused projection modules such as `to_qkv` / `to_added_qkv` instead of separate `to_q` / `to_k` / `to_v`, so a split-QKV trace usually means the quantized path did not engage rather than a brand new fusion opportunity.
|
||||
- **Hotspot workflow split**: use `sglang-diffusion-benchmark-profile` to prove and classify a slowdown with perf dumps plus `torch.profiler`; hand concrete kernel work off with the perf/profile evidence attached instead of expanding the benchmark skill.
|
||||
|
||||
Reference in New Issue
Block a user