Update SGLang diffusion skills (#27496)

This commit is contained in:
Xiaoyu Zhang
2026-06-08 12:22:45 +08:00
committed by GitHub
parent bf7fb6b925
commit 5bf7dd8e4a
10 changed files with 276 additions and 344 deletions
@@ -64,6 +64,8 @@ See existing Modular examples: `QwenImagePipeline` (uses `add_standard_t2i_stage
| Scheduler implementations | `python/sglang/multimodal_gen/runtime/models/schedulers/` |
| Model/VAE/DiT configs | `python/sglang/multimodal_gen/configs/models/dits/`, `vaes/`, `encoders/` |
| Central registry | `python/sglang/multimodal_gen/registry.py` |
| Model component registry | `python/sglang/multimodal_gen/runtime/models/registry.py` |
| Current support list | `docs/diffusion/compatibility_matrix.md` |
---
@@ -92,7 +94,7 @@ Once you have the reference code, study it thoroughly:
**Before creating any new files, check whether an existing pipeline or stage can be reused or extended.** Only create new pipelines/stages when the existing ones would require extensive modifications or when no similar implementation exists.
Specifically:
1. **Compare the new model's architecture against existing pipelines** (Flux, Wan, Qwen-Image, GLM-Image, HunyuanVideo, LTX, etc.). If the new model shares most of its structure with an existing one (e.g., same text encoders, similar latent format, compatible denoising loop), prefer:
1. **Compare the new model's architecture against existing pipelines** before creating files. Current native families include LTX-2/2.3, HunyuanVideo/FastHunyuan, Wan/FastWan/TurboWan/LingBot World, MOVA, FLUX/FLUX.2/Klein, Z-Image, Qwen-Image/edit/layered, GLM-Image, SD3, Hunyuan3D, Helios, Cosmos3, SANA, FireRed, ERNIE-Image, JoyAI, and Ideogram4. If the new model shares most of its structure with an existing one (e.g., same text encoders, similar latent format, compatible denoising loop), prefer:
- Adding a new config variant to the existing pipeline rather than creating a new pipeline class
- Reusing the existing `BeforeDenoisingStage` with minor parameter differences
- Using `add_standard_t2i_stages()` / `add_standard_ti2i_stages()` / `add_standard_ti2v_stages()` if the model fits standard patterns
@@ -524,7 +526,9 @@ register_configs(
`register_configs()` does not take a `model_family` argument. It registers the
sampling and pipeline config classes, then resolves models by exact
`hf_model_paths` or optional detector predicates.
`hf_model_paths` or optional detector predicates. Prefer exact `hf_model_paths`
for public checkpoints used in docs or tests; use detector predicates only for
families where local mirrors, renamed repos, or generated paths are common.
The `EntryClass` in your pipeline file is automatically discovered by the registry's `_discover_and_register_pipelines()` function -- no additional registration needed for the pipeline class itself.
@@ -552,6 +556,10 @@ After implementation, **you must verify that the generated output is not noise**
|-------|----------|---------------------|----------------|
| GLM-Image | `runtime/pipelines/glm_image.py` | `stages/model_specific_stages/glm_image.py` | `configs/pipeline_configs/glm_image.py` |
| Qwen-Image-Layered | `runtime/pipelines/qwen_image.py` (`QwenImageLayeredPipeline`) | `stages/model_specific_stages/qwen_image_layered.py` | `configs/pipeline_configs/qwen_image.py` (`QwenImageLayeredPipelineConfig`) |
| Cosmos3 | `runtime/pipelines/cosmos3_pipeline.py` | `stages/model_specific_stages/cosmos3.py` | `configs/pipeline_configs/cosmos3.py` |
| ErnieImage | `runtime/pipelines/ernie_image.py` | `runtime/pipelines/ernie_image.py` | `configs/pipeline_configs/ernie_image.py` |
| Hunyuan3D | `runtime/pipelines/hunyuan3d_pipeline.py` | `stages/model_specific_stages/hunyuan3d.py` | `configs/pipeline_configs/hunyuan3d.py` |
| LingBot World realtime | `runtime/pipelines/lingbot_world_causal_dmd_pipeline.py` | `stages/model_specific_stages/lingbot_world/` | `configs/pipeline_configs/lingbot_world.py` |
### Modular Style (when standard stages fit well)
@@ -560,6 +568,10 @@ After implementation, **you must verify that the generated output is not noise**
| Qwen-Image (T2I) | `runtime/pipelines/qwen_image.py` | Uses `add_standard_t2i_stages()` — standard text encoding + latent prep fits this model |
| Qwen-Image-Edit | `runtime/pipelines/qwen_image.py` | Uses `add_standard_ti2i_stages()` — standard image-to-image flow |
| Flux | `runtime/pipelines/flux.py` | Uses `add_standard_t2i_stages()` with custom `prepare_mu` |
| FLUX.2 / FLUX.2 Klein | `runtime/pipelines/flux_2.py`, `flux_2_klein.py` | Reuses FLUX.2 stages; Klein differences live in config and sampling params |
| Z-Image | `runtime/pipelines/zimage_pipeline.py` | Uses standard image pipeline stages plus Z-Image-specific config/model code |
| Ideogram4 | `runtime/pipelines/ideogram.py` | Uses dedicated text encoding and denoising stages while keeping standard latent prep |
| SANA | `runtime/pipelines/sana.py` | Spatial image pipeline; reuse the spatial image config pattern |
| Wan | `runtime/pipelines/wan_pipeline.py` | Uses `add_standard_ti2v_stages()` |
---
@@ -1,145 +0,0 @@
---
name: sglang-diffusion-ako4all-kernel
description: Use when optimizing an existing SGLang diffusion kernel with AKO4ALL, including AKO4ALL repo hygiene, custom microbench setup, ncu-guided iteration, and end-to-end denoise validation. Also use when a sibling AKO4ALL repo must be cloned or refreshed before starting kernel tuning work.
---
# SGLang Diffusion AKO4ALL Kernel
Use this skill to run the full AKO4ALL-based optimization loop for an existing SGLang diffusion kernel.
It is the default implementation path once the benchmark/profile skill has already shown that a hotspot is real and not covered by an existing fast path. This workflow bootstraps a custom AKO harness, benchmarks and profiles the kernel, iterates with `ncu`, ports the best version back to `sglang`, then validates with targeted tests and model-level denoise runs.
This skill assumes a sibling repo layout like:
```text
<base-dir>/
├── sglang/
└── AKO4ALL/
```
If `AKO4ALL/` is missing under the current base directory, clone it first.
## Use This Skill When
- tuning an existing diffusion Triton, CUDA JIT, CuTeDSL, or runtime-integrated kernel in `sglang`
- `sglang-diffusion-benchmark-profile` has already ruled out an existing in-repo fast path or overlap family
- creating a custom AKO4ALL harness for a real diffusion kernel instead of using the default benchmark tasks
- validating that a kernel-level win transfers to Qwen, FLUX, Wan, Hunyuan, MOVA, or other diffusion denoise latency
- preparing PR artifacts such as microbench tables, `ncu` before/after data, and proof image outputs
Do not start here when the bottleneck has not been proven yet.
First use [../sglang-diffusion-benchmark-profile/SKILL.md](../sglang-diffusion-benchmark-profile/SKILL.md) to:
- measure the real denoise regression
- collect the perf dump baseline
- capture one representative `torch.profiler` trace
- rule out existing mainline fast paths
- prove the run stayed on the native SGLang diffusion backend, not a diffusers fallback
Before opening AKO, also read
[../sglang-diffusion-benchmark-profile/existing-fast-paths.md](../sglang-diffusion-benchmark-profile/existing-fast-paths.md).
It records current mainline fusions plus the open PR watchlist for diffusion
kernel, VAE, attention, cache, and scheduling work. If an open PR already covers
the same shape family, use it as prior art or decide whether to rebase/extend it
instead of starting a duplicate kernel.
If a future specialized optimization skill matches the kernel family better than AKO4ALL, hand off there instead. The diagnosis contract stays the same.
## Mandatory AKO4ALL Preflight
Before any AKO work:
1. Run `scripts/ensure_ako4all_clean.sh [base-dir]`.
2. If `<base-dir>/AKO4ALL` does not exist, the script clones it.
3. Do not continue unless `AKO4ALL` is:
- on the upstream default branch, usually `main`
- fully clean with no tracked or untracked local changes
- exactly synced to `upstream/<default-branch>`
4. If the script reports local commits, divergence, or a dirty worktree, stop and clean or re-clone the repo before continuing.
The script creates an `upstream` remote automatically when missing.
By default it uses the existing `origin` URL, or `AKO4ALL_URL` if you need to override the clone source.
## Workflow
### 1. Scope the Kernel
- Identify the exact kernel entry point and runtime call sites in `sglang`.
- Record the target shapes, dtypes, model families, and whether the kernel is on a hot path.
- Reuse existing unit tests and benchmark entry points when they already exist.
- Record whether the hotspot overlaps an open PR from `existing-fast-paths.md`;
if it does, note the PR number in the AKO context and final PR artifacts.
### 2. Bootstrap the AKO Harness
Inside the clean `AKO4ALL` repo:
- read `TASK.md` and `HINTS.md`
- create a custom harness instead of relying on the stock benchmark tasks
- mirror the real SGLang kernel into:
- `input/reference.py`
- `input/<kernel>.py`
- `solution/<kernel>.py`
- `bench/bench_<kernel>.py`
- keep a short context note in `context/` when the kernel has model-specific shape assumptions or perf conclusions
The custom benchmark should:
- cover representative diffusion shapes
- check correctness against the reference kernel
- report aggregate runtime plus per-shape results when useful
### 3. Establish the Baseline
- run the AKO custom microbench before changing the kernel
- capture one representative `ncu` baseline on the hottest meaningful shape
- note whether the bottleneck looks like registers, occupancy, instruction count, launch config, or memory latency
### 4. Iterate in AKO4ALL
- change one idea at a time
- rerun the microbench after every change
- update `ITERATIONS.md` with hypothesis, result, and next step
- prefer simple, explainable wins over clever rewrites that do not transfer
After 3 consecutive no-improvement or regression iterations:
- rerun `ncu`
- re-read `ITERATIONS.md`
- change direction instead of continuing blind sweeps
### 5. Port the Best Version Back to SGLang
- apply the best candidate to the real `sglang` kernel file
- run import or syntax checks and targeted tests first
- keep the AKO `solution/` version aligned with the main-tree version you actually want to keep
### 6. Validate on Real Models
- use the benchmark/profile skill for denoise perf dumps and before/after comparison
- prefer exact local snapshot validation when testing local edits on a GPU box
- run targeted kernel tests first
- run model-level denoise benchmarks with perf dumps
- compare baseline vs optimized runs with `compare_perf.py`
- if the PR needs proof that generation still works, save one real model output image
### 7. Prepare PR Artifacts
At minimum, keep:
- one microbench table
- one denoise-stage table
- one end-to-end table
- one `ncu` before/after pair on the most representative kernel shape
- one generated image when the kernel affects production inference
See [references/ako-loop.md](references/ako-loop.md) for the checklist and common stop rules.
## Operating Rules
- Treat AKO4ALL repo hygiene as a gate, not a suggestion.
- Prefer exact local snapshot validation over hand-wavy “remote tree is close enough”.
- Do not start or justify kernel work from traces collected after
`Falling back to diffusers backend`, `Using diffusers backend`, or
`Loaded diffusers pipeline`; fix backend selection and rerun the
benchmark/profile workflow first.
- Keep model-level validation honest: if microbench improves but denoise does not, do not keep the AKO-only variant in the main code path.
- When writing conclusions, explain the win in terms of measurable causes such as lower registers per thread, higher occupancy, fewer executed instructions, or better scheduler eligibility.
@@ -1,54 +0,0 @@
# AKO Loop Checklist
Use this checklist after `scripts/ensure_ako4all_clean.sh` succeeds.
## Minimum Repo Layout
Inside `AKO4ALL/`, prefer these files for a diffusion kernel task:
- `input/reference.py`
- `input/<kernel>.py`
- `solution/<kernel>.py`
- `bench/bench_<kernel>.py`
- `context/<kernel>_notes.md`
## Baseline Checklist
- Reproduce the current SGLang kernel exactly in AKO first.
- Run the custom microbench before making edits.
- Record one representative `ncu` report on a real hot shape.
- Note the baseline bottleneck in plain language.
## Iteration Discipline
- One optimization idea per iteration.
- Re-benchmark after every code change.
- Log the result in `ITERATIONS.md`.
- Keep the best candidate easy to identify.
Stop a direction early when:
- 3 consecutive iterations do not beat the best runtime
- correctness gets fragile
- AKO-only gains stop transferring to real denoise runs
## Real Validation Gate
Before calling a kernel "done", validate all of:
- syntax or import checks
- targeted unit test or regression test
- kernel or op-level benchmark
- model-level denoise benchmark with perf dumps
- one generated image if the PR needs production proof
## PR Artifact Checklist
Prepare these artifacts:
- microbench table
- denoise-stage table
- end-to-end table
- one `ncu` before or after pair
- one short explanation of why the kernel got faster
- one generated output image when applicable
@@ -1,84 +0,0 @@
#!/usr/bin/env bash
set -euo pipefail
BASE_DIR="${1:-$PWD}"
BASE_DIR="$(cd "$BASE_DIR" && pwd)"
AKO_DIR="${BASE_DIR}/AKO4ALL"
CANONICAL_UPSTREAM_URL="https://github.com/TongmingLAIC/AKO4ALL.git"
UPSTREAM_URL="${AKO4ALL_UPSTREAM_URL:-$CANONICAL_UPSTREAM_URL}"
CLONE_URL="${AKO4ALL_URL:-$UPSTREAM_URL}"
say() {
printf '[ako4all] %s\n' "$*"
}
fail() {
printf '[ako4all] ERROR: %s\n' "$*" >&2
exit 1
}
if [[ ! -d "$AKO_DIR/.git" ]]; then
say "AKO4ALL not found under ${BASE_DIR}; cloning ${CLONE_URL}"
git clone "$CLONE_URL" "$AKO_DIR"
fi
cd "$AKO_DIR"
if ! git remote get-url origin >/dev/null 2>&1; then
fail "AKO4ALL exists but has no origin remote."
fi
if ! git remote get-url upstream >/dev/null 2>&1; then
say "Adding missing upstream remote -> ${UPSTREAM_URL}"
git remote add upstream "$UPSTREAM_URL"
fi
git fetch upstream --prune
git remote set-head upstream -a >/dev/null 2>&1 || true
default_branch="${AKO4ALL_BRANCH:-}"
if [[ -z "$default_branch" ]]; then
if upstream_head="$(git symbolic-ref --quiet --short refs/remotes/upstream/HEAD 2>/dev/null)"; then
default_branch="${upstream_head#upstream/}"
else
default_branch="main"
fi
fi
if [[ -n "$(git status --porcelain)" ]]; then
fail "AKO4ALL worktree is dirty. Clean all local changes before using this skill."
fi
if git show-ref --verify --quiet "refs/heads/${default_branch}"; then
git switch "$default_branch" >/dev/null
else
git switch -c "$default_branch" --track "upstream/${default_branch}" >/dev/null
fi
git fetch upstream --prune
local_head="$(git rev-parse HEAD)"
upstream_head="$(git rev-parse "upstream/${default_branch}")"
if [[ "$local_head" != "$upstream_head" ]]; then
if git merge-base --is-ancestor "$local_head" "$upstream_head"; then
say "Fast-forwarding ${default_branch} to upstream/${default_branch}"
git merge --ff-only "upstream/${default_branch}" >/dev/null
else
fail "Local ${default_branch} diverges from upstream/${default_branch}. Reset or re-clone AKO4ALL before continuing."
fi
fi
if [[ -n "$(git status --porcelain)" ]]; then
fail "AKO4ALL became dirty after sync; stop and inspect the repo."
fi
final_head="$(git rev-parse HEAD)"
expected_head="$(git rev-parse "upstream/${default_branch}")"
if [[ "$final_head" != "$expected_head" ]]; then
fail "AKO4ALL is not exactly at upstream/${default_branch}."
fi
say "Ready: ${AKO_DIR}"
say "Branch: ${default_branch}"
say "Commit: ${final_head}"
@@ -12,7 +12,7 @@ This skill is diagnosis-first. It owns:
- perf dump collection and before/after comparison
- `torch.profiler` trace capture and quick hotspot ranking
- mapping hot kernels back to known fast paths and fusion families
- handing confirmed kernel work to a specialized optimization skill such as [../sglang-diffusion-ako4all-kernel/SKILL.md](../sglang-diffusion-ako4all-kernel/SKILL.md)
- packaging confirmed kernel work with enough evidence for the appropriate kernel, Nsight, or framework-specific optimization workflow
This skill does not own low-level kernel authoring or standalone Nsight workflows.
@@ -42,7 +42,7 @@ If any benchmark, perf-dump, or `torch.profiler` command prints one of those sig
## Main Reference
- [benchmark-and-profile.md](benchmark-and-profile.md) — canonical denoise benchmark, perf dump, and `torch.profiler` workflow; uses checked-in nightly-aligned presets plus skill-only stress recipes such as `LTX-2.3` one-stage/two-stage, HunyuanVideo, MOVA, Helios, JoyAI/FireRed image edit, and Hunyuan3D shape
- [benchmark-and-profile.md](benchmark-and-profile.md) — canonical denoise benchmark, perf dump, and `torch.profiler` workflow; uses checked-in nightly-aligned presets plus current-source extras such as FLUX.2 Klein, Cosmos3, Ideogram4, ERNIE/GLM/SANA image models, FastWan2.2, `LTX-2.3` one-stage/two-stage/HQ, HunyuanVideo, MOVA, Helios, JoyAI/FireRed image edit, and Hunyuan3D shape
- [existing-fast-paths.md](existing-fast-paths.md) — map bottlenecks to existing fused kernels, packed QKV paths, fused `QK norm + RoPE`, distributed overlap patterns, and open optimization PRs before proposing new code
- [scripts/diffusion_skill_env.py](scripts/diffusion_skill_env.py) — preflight helper: repo root discovery via `sglang.__file__`, write-access probe, benchmark/profile output directories, idle GPU selection
- [scripts/bench_diffusion_denoise.py](scripts/bench_diffusion_denoise.py) — end-to-end denoise benchmark preset runner via `sglang generate`; supports `--no-torch-compile`, validates nightly preset drift with `--validate-nightly-alignment`, and saves perf dumps by label for `compare_perf.py`
@@ -53,8 +53,11 @@ Before calling a diffusion hotspot "new", first classify it with `existing-fast-
Always rule out these existing families first:
- HunyuanVideo VAE GroupNorm+SiLU
- LTX upsampler GroupNorm+SiLU
- Z-Image residual-form modulation
- fused diffusion `QK norm + RoPE`
- LTX2 split RoPE
- varlen USP attention pack/scatter
- NVFP4 / Nunchaku packed QKV
- Nunchaku fused GELU MLP
- Ulysses / USP attention overlap
@@ -21,10 +21,10 @@ This guide intentionally stops at:
- hotspot ranking
- mapping hotspots to known fast paths
If the hotspot survives this checklist, hand the work to
`sglang-diffusion-ako4all-kernel` or another specialized kernel-optimization
skill. Do not grow this skill back into a general Nsight or kernel-authoring
guide.
If the hotspot survives this checklist, package the perf dump, profiler trace,
exact command, and shape/topology notes for the appropriate kernel, Nsight, or
framework-specific optimization workflow. Do not grow this skill back into a
general Nsight or kernel-authoring guide.
## Prerequisites
@@ -163,7 +163,8 @@ PYTHONPATH=python python3 "$BENCH_PY" \
--output-dir "${BENCH_DIR}"
```
Run the full preset sweep:
Run the full preset sweep only when you have enough GPU time for both the
nightly-aligned cases and the source-tracked extras:
```bash
PYTHONPATH=python python3 "$BENCH_PY" \
@@ -172,7 +173,21 @@ PYTHONPATH=python python3 "$BENCH_PY" \
--output-dir "${BENCH_DIR}"
```
Nightly-aligned presets come first; skill-only presets stay available after them.
Nightly-aligned presets come first, followed by current-source extras from the
registry / GPU test cases, then broader skill-only stress presets.
Use the preset categories this way:
- **Nightly-aligned**: exact mirrors of
`scripts/ci/utils/diffusion/comparison_configs.json`; use these when the goal
is apples-to-apples comparison with CI / nightly coverage.
- **Current-source extras**: models or request shapes with explicit support
evidence in the current registry, GPU cases, compatibility matrix, pipeline
files, or unit tests, but without a nightly comparison case yet.
- **Skill-only stress / coverage presets**: extra profiling scenarios kept by
this skill to stress a topology, high-resolution path, multi-GPU mode, or
model-specific stage. These may be older than the latest registry additions,
so re-check the active source tree before treating them as support-matrix
commitments.
| Preset | Model | Nightly | Notes |
| --- | --- | --- | --- |
@@ -186,6 +201,19 @@ Nightly-aligned presets come first; skill-only presets stay available after them
| `ltx2` | `Lightricks/LTX-2` | Yes: `ltx2_twostage_t2v` | `LTX2TwoStagePipeline`, 2 GPUs, CFG parallel, 768x512, 121 frames, seed 42 |
| `ltx23-ti2v-two-stage` | `Lightricks/LTX-2.3` | Yes: `ltx2.3_twostage_ti2v_2gpus` | Nightly cat image, motion prompt, `LTX2TwoStagePipeline`, 2 GPUs, `--cfg-parallel-size 2`, 768x512, 121 frames, seed 42 |
| `wan-i2v` | `Wan-AI/Wan2.2-I2V-A14B-Diffusers` | Yes: `wan22_i2v_a14b_720p` | Nightly cat image and motion prompt, 1280x720, 81 frames, 4 GPUs, CFG parallel, Ulysses degree 2, text encoder CPU offload and pinned CPU memory |
| `qwen-image` | `Qwen/Qwen-Image` | No | Current-source extra covering the base Qwen-Image native path, separate from the nightly `Qwen-Image-2512` case |
| `qwen-edit-2509` | `Qwen/Qwen-Image-Edit-2509` | No | Current-source extra for the pre-2511 edit-plus path; uses the cat image, 1024x1024 |
| `zimage-base` | `Tongyi-MAI/Z-Image` | No | Current-source extra for non-turbo Z-Image; keep it separate from `zimage` / `Z-Image-Turbo` |
| `flux2-klein` | `black-forest-labs/FLUX.2-klein-4B` | No | Current-source extra for the distilled FLUX.2 Klein path; gated repo, 1024x1024, DiT layerwise offload disabled |
| `flux2-klein-base` | `black-forest-labs/FLUX.2-klein-base-4B` | No | Current-source extra for the undistilled FLUX.2 Klein Base path; gated repo, 1024x1024, DiT layerwise offload disabled |
| `cosmos3-nano-t2i` | `nvidia/Cosmos3-Nano` | No | Current-source extra for the single-frame Cosmos3 image path; sets `SGLANG_DISABLE_COSMOS3_GUARDRAILS=1` in the helper environment |
| `cosmos3-nano-t2v` | `nvidia/Cosmos3-Nano` | No | Current-source extra for a short Cosmos3 video path; sets `SGLANG_DISABLE_COSMOS3_GUARDRAILS=1` in the helper environment |
| `ideogram4-fp8` | `ideogram-ai/ideogram-4-fp8` | No | Current-source extra matching the native Ideogram 4 FP8 pipeline; do not override steps/guidance directly because the sampling preset owns them |
| `ernie-image-turbo` | `baidu/ERNIE-Image-Turbo` | No | Current-source extra for ERNIE-Image Turbo |
| `glm-image` | `zai-org/GLM-Image` | No | Current-source extra for GLM-Image |
| `sana-1.5-1.6b` | `Efficient-Large-Model/SANA1.5_1.6B_1024px_diffusers` | No | Current-source extra for a SANA native image path |
| `fastwan22-ti2v-5b` | `FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers` | No | Current-source extra matching the FastWan2.2 TI2V registered path |
| `ltx23-hq-two-stage` | `Lightricks/LTX-2.3` | No | Current-source extra for `LTX2TwoStageHQPipeline` with `snapshot` device mode; high-resolution and VRAM-heavy |
| `ltx23-one-stage` | `Lightricks/LTX-2.3` | No | Skill-only extra preset for the native `LTX-2.3` one-stage baseline; 2 GPUs, 768x512, 121 frames, fps 24, 30 steps, guidance 3.0, seed 1234 |
| `ltx23-two-stage` | `Lightricks/LTX-2.3` | No | Skill-only high-resolution stress preset for the native `LTX-2.3` two-stage path; uses `LTX2TwoStagePipeline`, 2 GPUs, 1536x1024, 121 frames, fps 24, 30 steps, guidance 3.0, seed 1234 |
| `ltx23-two-stage-cfg-parallel` | `Lightricks/LTX-2.3` | No | Skill-only high-resolution CFG-parallel stress preset matching `ltx23-two-stage` plus `--cfg-parallel-size 2` |
@@ -461,6 +489,8 @@ the known mainline families.
| `fused_inplace_qknorm_rope` missing, but separate qk norm plus rope show up | Check whether the fused diffusion `QK norm + RoPE` path should have engaged |
| `to_q -> to_k -> to_v` on NVFP4 or Nunchaku FLUX-family checkpoints | Treat as a packed-QKV fast-path miss or checkpoint-format mismatch |
| `fused_norm_tanh_mul_add*` missing on Z-Image | Treat as a missing mainline modulation path, not a new fusion request |
| LTX-2 split RoPE appears as a long PyTorch elementwise chain | Check the `apply_ltx2_split_rotary_emb` Triton path and its shape guards |
| masked attention spends time packing/unpacking Q/K/V | Check whether fused varlen USP pack/scatter should have engaged |
| `all_to_all`, ring attention, or async A2A dominate | Classify against Ulysses, USP, or turbo-layer overlap first |
| split `fc1 -> gelu -> quant -> fc2.lora_down` on Nunchaku FLUX | Treat as a missing fused GELU MLP path |
| attention kernels dominate | Confirm backend, topology, and shape guards before proposing a new kernel |
@@ -475,7 +505,7 @@ Only after the hotspot survives the fast-path checklist:
1. save a baseline perf dump
2. save a representative `torch.profiler` trace
3. note the exact model, shape, dtype, and GPU topology
4. hand the work to `sglang-diffusion-ako4all-kernel` or another future specialized optimization skill
4. hand the work to the appropriate kernel, Nsight, or framework-specific optimization workflow
This skill intentionally stops here. It tells you whether you are looking at:
- a missing existing optimization
@@ -490,4 +520,4 @@ This skill intentionally stops here. It tells you whether you are looking at:
- [ ] one representative `torch.profiler` trace saved
- [ ] hotspot classified against `existing-fast-paths.md`
- [ ] reference image or video checked for correctness
- [ ] any remaining kernel work handed to a specialized optimization skill
- [ ] any remaining kernel work handed off with perf/profile evidence attached
@@ -2,11 +2,13 @@
Use this guide when mapping a diffusion bottleneck to an existing fused path or
distributed overlap pattern in `sglang.multimodal_gen`. Prefer reuse and
configuration first before handing the problem to a specialized kernel-optimization skill.
configuration first before handing the problem to a kernel, Nsight, or
framework-specific optimization workflow.
**Key Files**
- `python/sglang/multimodal_gen/runtime/layers/layernorm.py`
- `python/sglang/multimodal_gen/runtime/layers/elementwise.py`
- `python/sglang/multimodal_gen/runtime/layers/fused_scale_shift_gate.py`
- `python/sglang/multimodal_gen/runtime/layers/rotary_embedding/utils.py`
- `python/sglang/jit_kernel/diffusion/triton/scale_shift.py`
- `python/sglang/jit_kernel/diffusion/group_norm_silu.py`
@@ -14,8 +16,14 @@ configuration first before handing the problem to a specialized kernel-optimizat
- `python/sglang/jit_kernel/diffusion/triton/norm.py`
- `python/sglang/jit_kernel/diffusion/triton/rmsnorm_onepass.py`
- `python/sglang/jit_kernel/diffusion/triton/rotary.py`
- `python/sglang/jit_kernel/diffusion/triton/ltx2_rotary.py`
- `python/sglang/jit_kernel/diffusion/triton/varlen_pack_pad.py`
- `python/sglang/jit_kernel/diffusion/cutedsl/scale_residual_norm_scale_shift.py`
- `python/sglang/jit_kernel/tests/diffusion/test_qwen_image_modulation.py`
- `python/sglang/jit_kernel/tests/diffusion/test_group_norm_silu.py`
- `python/sglang/jit_kernel/tests/diffusion/test_varlen_pack_pad.py`
- `python/sglang/jit_kernel/tests/diffusion/test_varlen_uspattn_equivalence.py`
- `python/sglang/jit_kernel/benchmark/diffusion/bench_qwen_image_modulation.py`
- `python/sglang/jit_kernel/benchmark/diffusion/bench_group_norm_silu.py`
- `python/sglang/jit_kernel/norm.py`
- `python/sglang/multimodal_gen/runtime/platforms/cuda.py`
@@ -24,10 +32,10 @@ configuration first before handing the problem to a specialized kernel-optimizat
**Core Fusion Patterns**
1. Scale/Shift elementwise fusion (AdaLN modulation)
- Kernels: `fuse_scale_shift_kernel`, `fuse_scale_shift_gate_select01_kernel`
- Locations: `elementwise.py`, `layernorm.py`, `qwen_image.py`, `triton/scale_shift.py`
- Use cases: `x * (1 + scale) + shift` and `a * (k + b) + c`
1. Scale/Shift elementwise and gate fusion (AdaLN modulation)
- Kernels: `fuse_scale_shift_kernel`, `fuse_layernorm_scale_shift_gate_select01_kernel`, `fuse_residual_layernorm_scale_shift_gate_select01_kernel`
- Locations: `elementwise.py`, `layernorm.py`, `fused_scale_shift_gate.py`, `qwen_image.py`, `triton/scale_shift.py`
- Use cases: `x * (1 + scale) + shift`, `a * (k + b) + c`, and Qwen-style `(layernorm/residual layernorm) + scale/shift + gate select`.
- Constraints: `x` must be CUDA and contiguous. `scale/shift` support 0D/1D/2D/3D/4D broadcast. 4D `[B, F, 1, C]` requires `L % F == 0`.
- NPU fallback: `scale_shift.py` swaps to `npu_fallback` native path.
- Validation: `python/sglang/jit_kernel/tests/diffusion/test_qwen_image_modulation.py`.
@@ -72,11 +80,18 @@ configuration first before handing the problem to a specialized kernel-optimizat
- NPU fallback: `npu_fallback.apply_rotary_embedding_native`.
- Validation: `python/sglang/jit_kernel/tests/test_rope.py`.
7. HunyuanVideo VAE GroupNorm + SiLU fusion
7. LTX2 split RoPE fusion
- Kernel: `apply_ltx2_split_rotary_emb`
- Locations: `triton/ltx2_rotary.py`, `runtime/models/dits/ltx_2.py`
- Use case: LTX-2 split rotary embedding over `[B, S, num_heads * head_dim]` with separate `cos` and `sin` tensors.
- Constraints: `cos` and `sin` shapes must match `[B, H, S, head_dim / 2]`, and `inner_dim == H * head_dim`.
- Workflow rule: if LTX-2 traces show a large split-RoPE PyTorch chain, check whether the LTX2-specific Triton path was disabled by shape or dtype before proposing a new RoPE kernel.
8. HunyuanVideo / LTX upsampler GroupNorm + SiLU fusion
- Kernel: `triton_group_norm_silu`
- Locations: `diffusion/group_norm_silu.py`, `triton/group_norm_silu.py`, `runtime/models/vaes/hunyuanvae.py`
- Locations: `diffusion/group_norm_silu.py`, `triton/group_norm_silu.py`, `runtime/models/vaes/hunyuanvae.py`, `runtime/models/upsampler/latent_upsampler.py`
- Use case: `activation(group_norm(x))` when the activation is non-inplace `nn.SiLU` and the GroupNorm is affine.
- Enablement: mainline uses `apply_group_norm_silu(...)` in HunyuanVideo VAE paths by default; there is no env toggle. The wrapper dispatches to Triton only when guards pass.
- Enablement: mainline uses `apply_group_norm_silu(...)` in HunyuanVideo VAE paths and LTX latent upsampler paths by default; there is no env toggle. The wrapper dispatches to Triton only when guards pass.
- Constraints: CUDA inference path only; no grad, `x.requires_grad == False`, `nn.GroupNorm`, `nn.SiLU(inplace=False)`, affine norm with weight and bias. Unsupported cases fall back to native `activation(norm(x))`.
- Validation: `python/sglang/jit_kernel/tests/diffusion/test_group_norm_silu.py`.
- Microbench: `python/sglang/jit_kernel/benchmark/diffusion/bench_group_norm_silu.py`.
@@ -99,6 +114,12 @@ configuration first before handing the problem to a specialized kernel-optimizat
- Location: `rotary_embedding/utils.py`
- Behavior: `flashinfer.rope.apply_rope_with_cos_sin_cache_inplace` when available, otherwise Triton RoPE fallback.
4. Varlen USP attention pack/scatter
- Locations: `runtime/layers/attention/layer.py`, `triton/varlen_pack_pad.py`
- Behavior: masked `USPAttention.forward` can gather dense Q/K/V into packed `[total_valid, H, D]` rows with `fused_pack_qkv`, run varlen attention, then scatter back with `fused_scatter_to_padded`.
- Validation: `python/sglang/jit_kernel/tests/diffusion/test_varlen_pack_pad.py` and `test_varlen_uspattn_equivalence.py`.
- Workflow rule: if a masked attention trace spends time in Python/advanced indexing pack or scatter, first check whether this fused varlen path should have engaged.
**QK Norm Optimization**
- Entry point: `apply_qk_norm` in `layernorm.py`.
@@ -125,7 +146,7 @@ configuration first before handing the problem to a specialized kernel-optimizat
- Supported head dims: `64, 128, 256`.
- Behavior: `apply_qk_norm_rope` prefers the fused JIT kernel when all guards pass; otherwise it falls back to `apply_qk_norm(...)` plus `apply_flashinfer_rope_qk_inplace(...)`.
- Validation: `python/sglang/jit_kernel/tests/diffusion/test_qknorm_rope.py`.
- Watchlist: PR #24025 adds LTX2-specific QK norm fusion work. Until merged, treat LTX2 traces that miss the generic fused path as an enablement/shape-guard issue first, not proof that the PR path exists locally.
- Workflow rule: treat LTX2 traces that miss the generic fused path as an enablement/shape-guard issue first, and check the separate LTX2 split-RoPE path before proposing new attention-prep kernels.
**Nunchaku Fused GELU MLP**
@@ -144,11 +165,13 @@ configuration first before handing the problem to a specialized kernel-optimizat
**Common Entry Points in Diffusion Models**
- AdaLN modulation: `LayerNormScaleShift`, `RMSNormScaleShift`, `ScaleResidual*` in `layernorm.py`.
- Qwen-Image gating: `fuse_scale_shift_gate_select01_kernel` in `qwen_image.py`.
- Qwen-Image gating: `fuse_layernorm_scale_shift_gate_select01_kernel` and `fuse_residual_layernorm_scale_shift_gate_select01_kernel` through `fused_scale_shift_gate.py` and `qwen_image.py`.
- Z-Image residual-form modulation: `fused_norm_tanh_mul_add` and `fused_norm_tanh_mul_add_norm_scale` in `zimage.py`.
- HunyuanVideo VAE GroupNorm+SiLU: `apply_group_norm_silu` in `hunyuanvae.py`; default-eligible when wrapper guards pass.
- HunyuanVideo VAE and LTX upsampler GroupNorm+SiLU: `apply_group_norm_silu` in `hunyuanvae.py` and `latent_upsampler.py`; default-eligible when wrapper guards pass.
- QK norm: `apply_qk_norm` used in `flux.py`, `flux_2.py`, `qwen_image.py`, `zimage.py`, `wanvideo.py`, `ltx_2.py`, `hunyuanvideo.py`.
- QK norm + RoPE: `apply_qk_norm_rope` in `layernorm.py`; use this path when the model wants fused attention prep instead of separate QK norm and RoPE calls.
- LTX2 split RoPE: `apply_ltx2_split_rotary_emb` in `ltx_2.py`.
- Varlen USP attention: `fused_pack_qkv` and `fused_scatter_to_padded` in `attention/layer.py`.
- Nunchaku fused GELU MLP: `_fused_gelu_mlp` in `flux.py` for quantized FLUX-family checkpoints.
- NVFP4 / packed QKV attention: `to_qkv`, `to_added_qkv`, and `to_qkv_mlp_proj` in FLUX-family quantized paths.
- RoPE: `_apply_rotary_emb` prefers Triton; Q/K RoPE prefers FlashInfer when present.
@@ -161,11 +184,11 @@ configuration first before handing the problem to a specialized kernel-optimizat
- Dual-stream diffusion models: `use_dual_stream = True` in models such as `hunyuan3d.py` is an existing overlap family.
- Workflow rule: if a hotspot is communication-heavy, rule out these in-repo overlap families before proposing a brand new overlap design.
**Open PR Watchlist**
**Historical PR Watchlist**
As of 2026-05-02, these SGLang PRs were still open. Use them as upstream
direction and prior art, not as current-main behavior. Re-check the PR state
before relying on any file path or flag.
These SGLang PRs are useful as upstream direction and prior art, not as
current-main behavior. Re-check the PR state and the active source tree before
relying on any file path, flag, or claim about whether the work has merged.
- Norm, modulation, and packed projection fusions:
- #24025 LTX2 QK norm fusion.
@@ -211,4 +234,4 @@ before relying on any file path or flag.
- Keep CuTe compile cache keys aligned to `(dtype, ndim, D)`.
- Avoid implicit broadcasts that force hidden `contiguous()` copies.
- Preserve NPU and ROCm fallback paths.
- If none of the families above match, package the evidence from the benchmark/profile skill and hand the kernel work to a specialized optimization skill such as `sglang-diffusion-ako4all-kernel`.
- If none of the families above match, package the evidence from the benchmark/profile skill and hand the kernel work to the appropriate kernel, Nsight, or framework-specific optimization workflow.
@@ -12,7 +12,7 @@ Usage:
# Tag the run for later compare_perf.py usage
python3 python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py --model flux --label tuned
# All 20 preset models
# All preset models
python3 python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py --all
# Show preset order, model path, and nightly mapping
@@ -56,13 +56,18 @@ ASSET_DIR = ensure_dir(get_assets_dir(REPO_ROOT))
NIGHTLY_CONFIG_PATH = (
REPO_ROOT / "scripts" / "ci" / "utils" / "diffusion" / "comparison_configs.json"
)
GATED_MODELS = {"flux", "flux2"}
GATED_MODELS = {
"flux",
"flux2",
"flux2-klein",
"flux2-klein-base",
}
DIFFUSERS_FALLBACK_SIGNALS = (
"falling back to diffusers backend",
"using diffusers backend",
"loaded diffusers pipeline",
)
CATALOG_TABLE_WIDTH = 105
CATALOG_TABLE_WIDTH = 140
RESULTS_TABLE_WIDTH = 105
NIGHTLY_PRESET_ORDER = (
"flux",
@@ -80,7 +85,7 @@ NIGHTLY_PRESET_ORDER = (
# ---------------------------------------------------------------------------
# Model configs — kept in exact sync with benchmark-and-profile.md
# Nightly-aligned presets mirror scripts/ci/utils/diffusion/comparison_configs.json
# first, followed by skill-only extras.
# first, followed by current-source extras and skill-only stress / coverage presets.
# Each entry produces the same `sglang generate` command as shown in that doc.
# ---------------------------------------------------------------------------
MODELS = {
@@ -217,7 +222,137 @@ MODELS = {
"--pin-cpu-memory",
],
},
# 11. Skill-only extra preset
# Source-tracked extras from current registry / GPU test coverage.
"qwen-image": {
"path": "Qwen/Qwen-Image",
"prompt": "A futuristic cyberpunk city at night, neon lights reflecting on wet streets",
"extra_args": [
"--width=1024",
"--height=1024",
],
},
# Requires: <repo>/inputs/diffusion_benchmark/figs/cat.png
"qwen-edit-2509": {
"path": "Qwen/Qwen-Image-Edit-2509",
"prompt": "Make the cat wear a red hat",
"image_path": str(ASSET_DIR / "cat.png"),
"extra_args": [
"--width=1024",
"--height=1024",
],
},
"zimage-base": {
"path": "Tongyi-MAI/Z-Image",
"prompt": "A futuristic cyberpunk city at night, neon lights reflecting on wet streets",
"extra_args": [
"--width=1024",
"--height=1024",
],
},
"flux2-klein": {
"path": "black-forest-labs/FLUX.2-klein-4B",
"prompt": "A futuristic cyberpunk city at night, neon lights reflecting on wet streets",
"extra_args": [
"--width=1024",
"--height=1024",
"--dit-layerwise-offload",
"false",
],
},
"flux2-klein-base": {
"path": "black-forest-labs/FLUX.2-klein-base-4B",
"prompt": "A futuristic cyberpunk city at night, neon lights reflecting on wet streets",
"extra_args": [
"--width=1024",
"--height=1024",
"--dit-layerwise-offload",
"false",
],
},
"cosmos3-nano-t2i": {
"path": "nvidia/Cosmos3-Nano",
"prompt": "A red cube on a white table, product photo.",
"env": {
"SGLANG_DISABLE_COSMOS3_GUARDRAILS": "1",
},
"extra_args": [
"--width=1024",
"--height=1024",
"--num-frames=1",
"--num-inference-steps=35",
],
},
"cosmos3-nano-t2v": {
"path": "nvidia/Cosmos3-Nano",
"prompt": "A blue box slides across a clean warehouse floor.",
"env": {
"SGLANG_DISABLE_COSMOS3_GUARDRAILS": "1",
},
"extra_args": [
"--width=832",
"--height=480",
"--num-frames=9",
"--num-inference-steps=4",
],
},
"ideogram4-fp8": {
"path": "ideogram-ai/ideogram-4-fp8",
"prompt": "A clean product poster for a new open-source inference engine",
"extra_args": [
"--width=1024",
"--height=1024",
],
},
"ernie-image-turbo": {
"path": "baidu/ERNIE-Image-Turbo",
"prompt": "A futuristic cyberpunk city at night, neon lights reflecting on wet streets",
"extra_args": [
"--width=1024",
"--height=1024",
],
},
"glm-image": {
"path": "zai-org/GLM-Image",
"prompt": "A futuristic cyberpunk city at night, neon lights reflecting on wet streets",
"extra_args": [
"--width=1024",
"--height=1024",
],
},
"sana-1.5-1.6b": {
"path": "Efficient-Large-Model/SANA1.5_1.6B_1024px_diffusers",
"prompt": "A futuristic cyberpunk city at night, neon lights reflecting on wet streets",
"extra_args": [
"--width=1024",
"--height=1024",
],
},
"fastwan22-ti2v-5b": {
"path": "FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers",
"prompt": "The cat starts walking slowly towards the camera.",
"image_path": str(ASSET_DIR / "cat.png"),
"extra_args": [
"--width=1280",
"--height=720",
"--num-frames=81",
],
},
"ltx23-hq-two-stage": {
"path": "Lightricks/LTX-2.3",
"prompt": "A beautiful sunset over the ocean",
"env": {
"PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True",
"SGLANG_LTX2_SNAPSHOT_RELEASE_EMPTY_CACHE": "true",
},
"extra_args": [
"--pipeline-class-name=LTX2TwoStageHQPipeline",
"--ltx2-two-stage-device-mode=snapshot",
"--width=1920",
"--height=1088",
"--num-frames=121",
],
},
# Skill-only extra preset
"ltx23-one-stage": {
"path": "Lightricks/LTX-2.3",
"prompt": "A beautiful sunset over the ocean",
@@ -233,7 +368,7 @@ MODELS = {
"--num-gpus=2",
],
},
# 12. Skill-only extra preset
# Skill-only extra preset
"ltx23-two-stage": {
"path": "Lightricks/LTX-2.3",
"prompt": "A beautiful sunset over the ocean",
@@ -250,7 +385,7 @@ MODELS = {
"--num-gpus=2",
],
},
# 13. Skill-only extra preset
# Skill-only extra preset
"ltx23-two-stage-cfg-parallel": {
"path": "Lightricks/LTX-2.3",
"prompt": "A beautiful sunset over the ocean",
@@ -268,7 +403,7 @@ MODELS = {
"--cfg-parallel-size=2",
],
},
# 14. Skill-only extra preset
# Skill-only extra preset
"hunyuanvideo": {
"path": "hunyuanvideo-community/HunyuanVideo",
"prompt": "A cat and a dog baking a cake together in a kitchen. The cat is carefully measuring flour, while the dog is stirring the batter with a wooden spoon. The kitchen is cozy, with sunlight streaming through the window.",
@@ -281,7 +416,7 @@ MODELS = {
"--num-inference-steps=30",
],
},
# 15. Skill-only extra preset
# Skill-only extra preset
# Requires: <repo>/inputs/diffusion_benchmark/figs/mova_single_person.jpg
"mova-720p": {
"path": "OpenMOSS-Team/MOVA-720p",
@@ -297,7 +432,7 @@ MODELS = {
"--num-inference-steps=2",
],
},
# 16. Skill-only extra preset
# Skill-only extra preset
"helios": {
"path": "BestWishYsh/Helios-Base",
"prompt": "A curious raccoon",
@@ -315,14 +450,13 @@ MODELS = {
"false",
],
},
# 16. Skill-only extra preset
# Skill-only extra preset
# Requires: <repo>/inputs/diffusion_benchmark/figs/cat.png
"joyai-edit": {
"path": "jdopensource/JoyAI-Image-Edit-Diffusers",
"prompt": "Make the cat wear a red hat",
"image_path": str(ASSET_DIR / "cat.png"),
"extra_args": [
"--backend=sglang",
"--width=1024",
"--height=1024",
"--num-inference-steps=40",
@@ -336,14 +470,13 @@ MODELS = {
"--ulysses-degree=1",
],
},
# 17. Skill-only extra preset
# Skill-only extra preset
# Requires: <repo>/inputs/diffusion_benchmark/figs/cat.png
"firered-edit-1.0": {
"path": "FireRedTeam/FireRed-Image-Edit-1.0",
"prompt": "Make the cat wear a red hat",
"image_path": str(ASSET_DIR / "cat.png"),
"extra_args": [
"--backend=sglang",
"--width=1024",
"--height=1024",
"--num-inference-steps=40",
@@ -357,14 +490,13 @@ MODELS = {
"--ulysses-degree=1",
],
},
# 18. Skill-only extra preset
# Skill-only extra preset
# Requires: <repo>/inputs/diffusion_benchmark/figs/cat.png
"firered-edit-1.1": {
"path": "FireRedTeam/FireRed-Image-Edit-1.1",
"prompt": "Make the cat wear a red hat",
"image_path": str(ASSET_DIR / "cat.png"),
"extra_args": [
"--backend=sglang",
"--width=1024",
"--height=1024",
"--num-inference-steps=40",
@@ -378,7 +510,7 @@ MODELS = {
"--ulysses-degree=1",
],
},
# 19. Skill-only extra preset
# Skill-only extra preset
# Requires: <repo>/inputs/diffusion_benchmark/figs/cat.png
"hunyuan3d-shape": {
"path": "tencent/Hunyuan3D-2",
@@ -388,7 +520,6 @@ MODELS = {
"paint_enable": False,
},
"extra_args": [
"--backend=sglang",
"--num-inference-steps=50",
"--guidance-scale=5.0",
"--dit-layerwise-offload",
@@ -544,17 +675,17 @@ def print_model_catalog():
"""Print preset order, model path, and whether each preset maps to nightly."""
print()
print("=" * CATALOG_TABLE_WIDTH)
print("MODEL PRESETS — Nightly-aligned first, skill-only extras after")
print("MODEL PRESETS — Nightly-aligned, then current-source and skill-only extras")
print("=" * CATALOG_TABLE_WIDTH)
print(f"{'Preset':<24} {'Nightly':<28} {'Model Path':<46} {'GPUs':>4}")
print(f"{'Preset':<32} {'Nightly':<30} {'Model Path':<66} {'GPUs':>4}")
print("-" * CATALOG_TABLE_WIDTH)
for model_key, cfg in MODELS.items():
print(
f"{model_key:<24} {model_nightly_case_id(model_key):<28} {cfg['path']:<46} {required_gpus_for_model(model_key):>4}"
f"{model_key:<32} {model_nightly_case_id(model_key):<30} {cfg['path']:<66} {required_gpus_for_model(model_key):>4}"
)
print("-" * CATALOG_TABLE_WIDTH)
print(
"Nightly column shows the comparison_configs.json case id; '-' means skill-only."
"Nightly column shows the comparison_configs.json case id; '-' means no nightly mapping."
)
@@ -575,6 +706,7 @@ def build_sglang_cmd(
cmd = [
"sglang",
"generate",
"--backend=sglang",
f"--model-path={cfg['path']}",
f"--prompt={cfg['prompt']}",
]
@@ -631,6 +763,9 @@ def run_benchmark_once(
env = os.environ.copy()
env.setdefault("FLASHINFER_DISABLE_VERSION_CHECK", "1")
cfg = MODELS[model_key]
for key, value in cfg.get("env", {}).items():
env.setdefault(key, str(value))
if env.get("HF_TOKEN") and not env.get("HUGGINGFACE_HUB_TOKEN"):
env["HUGGINGFACE_HUB_TOKEN"] = env["HF_TOKEN"]
@@ -54,6 +54,7 @@ This repo now contains:
- diffusion-side NVFP4 loading from ModelOpt exports
- FLUX.2 packed-QKV detection that distinguishes packed NVFP4 checkpoints from standard diffusers exports
- automatic protection against incompatible FP8 CPU offload while keeping layerwise DiT offload available
- separate online diffusion quantization paths such as `--quantization fp8` / `mxfp4`; keep those out of this ModelOpt PTQ/export workflow unless the user explicitly asks for runtime quantization
- FP8 transformer build:
[`python/sglang/multimodal_gen/tools/build_modelopt_fp8_transformer.py`](../../../tools/build_modelopt_fp8_transformer.py)
- NVFP4 mixed transformer build:
@@ -67,6 +68,11 @@ Validated documentation and CI coverage currently center on these ModelOpt diffu
- NVFP4: FLUX.1-dev, FLUX.2-dev, Wan2.2
Treat a new family, a new precision, or a new checkpoint layout as unsupported until it has a documented matrix row and a matching validation story.
Current B200 CI also contains an Ideogram4 NVFP4 native load case
(`ideogram4_nvfp4_t2i` via `Comfy-Org/Ideogram-4`). Treat that as source
evidence for an existing NVFP4 path, but do not expand the ModelOpt support
matrix to Ideogram4 unless `docs/diffusion/quantization.md` is updated with the
exact checkpoint, loader path, quality check, and benchmark scope.
Before writing CLI examples, re-read the active branch's `docs/diffusion/quantization.md`: FLUX.2 NVFP4 is an official `black-forest-labs/*` repo rather than a `lmsys/*` converted repo, and its preferred flag depends on the current documented loader flow. Use `--transformer-path` for a component override directory with `config.json`; use `--transformer-weights-path` when the repo or path should be probed as raw weights.
B200 CI coverage can include loose BF16-vs-quantized quality checks. Inspect the active branch's `run_suite.py` before assuming they are part of the suite; mainline and feature branches may differ. Those checks are intended to catch blank, corrupted, or obviously divergent images, not exact image parity.
@@ -78,9 +84,9 @@ you are explicitly testing a historical branch.
## Related PR Watchlist
As of 2026-05-04, these related SGLang PRs are relevant to ModelOpt diffusion
support. Treat unmerged items as future support or migration work until the
docs/CI matrix is updated.
These related SGLang PRs are useful as ModelOpt diffusion support history.
Re-check the PR state and the active source tree before treating any item as
current behavior, and keep the docs/CI matrix as the support boundary.
- #23155 added Qwen Image ModelOpt FP8 support.
- #23199 adds HunyuanVideo ModelOpt FP8 support.
@@ -252,21 +252,27 @@ Use these as first commands to benchmark, not as universal winners.
| Model family | First performance shape | Starting flags | Notes |
|---|---|---|---|
| FLUX.1 / FLUX.2 image | 1024x1024, runtime-default steps/guidance, 1 GPU | `--enable-torch-compile --warmup --dit-layerwise-offload false` | `black-forest-labs/FLUX.*` repos are gated; for FP8/NVFP4 use validated `--transformer-path` or `--transformer-weights-path` flows from the quant skill. |
| FLUX.2 Klein / Klein Base | 1024x1024, runtime-default steps/guidance, 1 GPU | `--enable-torch-compile --warmup --dit-layerwise-offload false` | Current registry has `black-forest-labs/FLUX.2-klein-4B`, `FLUX.2-klein-9B`, and base variants. Klein is step-distilled; Klein Base is not. |
| Qwen-Image / Qwen-Image-Edit | 1024x1024, runtime-default steps/guidance, 1 GPU | `--enable-torch-compile --warmup`; optionally native `SGLANG_CACHE_DIT_ENABLED=true` | Cache-DiT is lossy. For edit tasks, keep reference image, seed, and output size fixed. |
| Z-Image-Turbo | 1024x1024, runtime-default steps/guidance, 1 GPU | `--enable-torch-compile --warmup` | Mainline has Z-Image tanh/gate norm fusions; PR #21912 tracks FP8 plus CUDA Graph work. |
| Z-Image / Z-Image-Turbo | 1024x1024, runtime-default steps/guidance, 1 GPU | `--enable-torch-compile --warmup` | Keep base Z-Image separate from Turbo: base uses 50-step CFG defaults, Turbo uses 9-step zero-CFG defaults. Mainline has Z-Image tanh/gate norm fusions. |
| Wan2.2 A14B T2V/I2V | 1280x720, 81 frames | Nightly: `--num-gpus 4 --enable-cfg-parallel --ulysses-degree 2 --text-encoder-cpu-offload --pin-cpu-memory` | For lowest latency, also benchmark pure Ulysses on the same GPUs. |
| Wan2.2 TI2V 5B | 1280x720, 81 frames, 1 GPU | `--enable-torch-compile --warmup` | Keep the input image and motion prompt fixed when comparing sparse attention or Cache-DiT. |
| LTX-2 / LTX-2.3 | 768x512, 121 frames, runtime-default steps/guidance, 2 GPUs | `--pipeline-class-name LTX2TwoStagePipeline --enable-torch-compile --warmup`; LTX-2 uses `--enable-cfg-parallel`, LTX-2.3 TI2V uses `--cfg-parallel-size 2` | Use the benchmark/profile skill presets for exact nightly alignment. PRs #22441, #24025, and #23736 track additional LTX2 perf/parallel work. |
| Wan2.1 / FastWan / TurboWan variants | 480p or 720p video, family defaults | `--enable-torch-compile --warmup`; add `--ulysses-degree` / CFG parallel only after measuring | Current registry includes Wan2.1, FastWan2.1, FastWan2.2 TI2V, TurboWan2.1, TurboWan2.2 I2V, and Wan2.1-Fun InP. Use the compatibility matrix and benchmark presets before choosing topology. |
| Cosmos3 Nano / Super | T2I: 1024x1024 with `--num-frames 1`; T2V/I2V: 480p/720p video | `SGLANG_DISABLE_COSMOS3_GUARDRAILS=1` for benchmark isolation; `--enable-torch-compile --warmup` | One checkpoint serves T2I/T2V/I2V. Mode is request-driven: `num_frames == 1` means T2I, `--image-path` means I2V. |
| Ideogram 4 FP8/NVFP4 | 1024x1024, native preset defaults | `--enable-torch-compile --warmup` | Do not set `--num-inference-steps` or `--guidance-scale` directly unless you also update the Ideogram preset; sampling params derive them from `preset`. |
| ERNIE-Image / GLM-Image / SANA / SD3 | 1024-class image, family defaults | `--enable-torch-compile --warmup`; disable offload only after checking VRAM | Treat these as current native image families. Start with benchmark/profile presets for ERNIE, GLM, and SANA; use registry/config defaults for SD3 unless you add a new preset. |
| LTX-2 / LTX-2.3 | 768x512 or HQ 1920x1088, 121 frames | `--pipeline-class-name LTX2TwoStagePipeline --enable-torch-compile --warmup`; HQ uses `LTX2TwoStageHQPipeline` and `--ltx2-two-stage-device-mode snapshot` by default | Use benchmark/profile presets for nightly alignment, one-stage, high-resolution stress, and HQ. Device mode choices are `original`, `snapshot`, and `resident`; `resident` is fastest but uses more VRAM. |
| HunyuanVideo | 848x480 or 720p class video | `--text-encoder-cpu-offload --pin-cpu-memory --enable-torch-compile --warmup` | Check VAE decode separately. GroupNorm+SiLU is default-eligible in mainline when wrapper guards pass; use `bench_group_norm_silu.py` when VAE residual blocks are hot. |
| JoyAI-Image-Edit | 1024-class TI2I, 40 steps, guidance 4.0 | `--backend=sglang --num-gpus 2 --enable-cfg-parallel --ulysses-degree 1 --enable-torch-compile --warmup --dit-layerwise-offload false --dit-cpu-offload false` | Newly supported image-edit path. Keep the input image, prompt, seed, and output size fixed; 2-GPU CFG parallel is the validated H100 starting point. |
| FireRed-Image-Edit 1.0 / 1.1 | 1024x1024 image edit, 40 steps, guidance 4.0 | `--backend=sglang --num-gpus 2 --enable-cfg-parallel --ulysses-degree 1 --enable-torch-compile --warmup --dit-layerwise-offload false --dit-cpu-offload false` | Uses the native `QwenImageEditPlusPipeline` path. 2-GPU CFG parallel is the validated H100 starting point; benchmark 1.0 and 1.1 separately because checkpoint differences can change denoise latency. |
| Hunyuan3D-2 shape | Shape generation, 50 steps, guidance 5.0 | `--backend=sglang --enable-torch-compile --warmup --dit-layerwise-offload false --dit-cpu-offload false` | Focus on `Hunyuan3DShapeDenoisingStage`; keep mesh export/paint timings separate from denoise. |
| MOVA / Helios | Use the benchmark/profile presets first | `--enable-torch-compile --warmup`; pin offload flags explicitly | PR #20530 tracks MOVA fused RMSNorm+RoPE; PR #24059 tracks Helios fused norm modulation. |
| MOVA / Helios / LingBot World | Use the benchmark/profile presets or server test cases first | `--enable-torch-compile --warmup`; pin offload and topology flags explicitly | These video/realtime families have model-specific stages and condition handling. Keep prompt/image/action inputs fixed and prefer perf dumps over wall time alone. |
## Open PR Watchlist
## Historical PR Watchlist
As of 2026-05-02, these performance PRs were open. Treat them as direction and
prior art until merged:
Treat these performance PRs as direction and prior art only. Re-check the PR
state and the active source tree before relying on any path, flag, or claim
about whether the work has merged:
- Fusion/kernel: #24025 LTX2 QK norm, #24059 Helios norm modulation, #24117 Z-Image packed QKV, #19488 Wan elementwise cross-block fusion, #19249 Z-Image gate/norm fusion, #20429 Qwen-Image layernorm/modulation, #20530 MOVA RMSNorm+RoPE.
- VAE/decode: #22531 LTX2 parallel VAE, #20927 batched tiled VAE decode.
@@ -279,6 +285,6 @@ prior art until merged:
- **Offload tuning**: after the first request, the runtime logs peak GPU memory and which components could stay resident. Use this to decide which `--*-cpu-offload` flags to disable.
- **Backend selection**: `--backend sglang` (default, auto-detected) enables native optimizations (fused kernels, SP, native Cache-DiT env knobs, etc.). `--backend diffusers` falls back to Diffusers pipelines and is the path that accepts `--cache-dit-config` plus diffusers attention backend names.
- **Wan2.2-I2V sizing**: explicit `--width/--height` on `Wan2.2-I2V-A14B` control the target area while preserving the condition-image aspect ratio.
- **Mainline diffusion fast paths**: before proposing a new kernel or overlap scheme, check `sglang-diffusion-benchmark-profile/existing-fast-paths.md`. It covers GroupNorm+SiLU, Z-Image residual-form modulation, fused diffusion `QK norm + RoPE`, packed QKV/NVFP4 expectations, and existing multi-GPU overlap families such as Ulysses / USP and turbo-layer async all-to-all.
- **Mainline diffusion fast paths**: before proposing a new kernel or overlap scheme, check `sglang-diffusion-benchmark-profile/existing-fast-paths.md`. It covers GroupNorm+SiLU, Z-Image residual-form modulation, fused diffusion `QK norm + RoPE`, LTX2 split RoPE, varlen USP pack/scatter, packed QKV/NVFP4 expectations, and existing multi-GPU overlap families such as Ulysses / USP and turbo-layer async all-to-all.
- **NVFP4 trace interpretation**: on FLUX.2 NVFP4 and Nunchaku-style checkpoints, packed QKV is expected. SGLang intentionally uses fused projection modules such as `to_qkv` / `to_added_qkv` instead of separate `to_q` / `to_k` / `to_v`, so a split-QKV trace usually means the quantized path did not engage rather than a brand new fusion opportunity.
- **Hotspot workflow split**: use `sglang-diffusion-benchmark-profile` to prove and classify a slowdown with perf dumps plus `torch.profiler`; hand concrete kernel work to `sglang-diffusion-ako4all-kernel` or another specialized optimization skill instead of expanding the benchmark skill.
- **Hotspot workflow split**: use `sglang-diffusion-benchmark-profile` to prove and classify a slowdown with perf dumps plus `torch.profiler`; hand concrete kernel work off with the perf/profile evidence attached instead of expanding the benchmark skill.