diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-add-model/SKILL.md b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-add-model/SKILL.md index cb325ce25..e213bd763 100644 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-add-model/SKILL.md +++ b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-add-model/SKILL.md @@ -64,6 +64,8 @@ See existing Modular examples: `QwenImagePipeline` (uses `add_standard_t2i_stage | Scheduler implementations | `python/sglang/multimodal_gen/runtime/models/schedulers/` | | Model/VAE/DiT configs | `python/sglang/multimodal_gen/configs/models/dits/`, `vaes/`, `encoders/` | | Central registry | `python/sglang/multimodal_gen/registry.py` | +| Model component registry | `python/sglang/multimodal_gen/runtime/models/registry.py` | +| Current support list | `docs/diffusion/compatibility_matrix.md` | --- @@ -92,7 +94,7 @@ Once you have the reference code, study it thoroughly: **Before creating any new files, check whether an existing pipeline or stage can be reused or extended.** Only create new pipelines/stages when the existing ones would require extensive modifications or when no similar implementation exists. Specifically: -1. **Compare the new model's architecture against existing pipelines** (Flux, Wan, Qwen-Image, GLM-Image, HunyuanVideo, LTX, etc.). If the new model shares most of its structure with an existing one (e.g., same text encoders, similar latent format, compatible denoising loop), prefer: +1. **Compare the new model's architecture against existing pipelines** before creating files. Current native families include LTX-2/2.3, HunyuanVideo/FastHunyuan, Wan/FastWan/TurboWan/LingBot World, MOVA, FLUX/FLUX.2/Klein, Z-Image, Qwen-Image/edit/layered, GLM-Image, SD3, Hunyuan3D, Helios, Cosmos3, SANA, FireRed, ERNIE-Image, JoyAI, and Ideogram4. If the new model shares most of its structure with an existing one (e.g., same text encoders, similar latent format, compatible denoising loop), prefer: - Adding a new config variant to the existing pipeline rather than creating a new pipeline class - Reusing the existing `BeforeDenoisingStage` with minor parameter differences - Using `add_standard_t2i_stages()` / `add_standard_ti2i_stages()` / `add_standard_ti2v_stages()` if the model fits standard patterns @@ -524,7 +526,9 @@ register_configs( `register_configs()` does not take a `model_family` argument. It registers the sampling and pipeline config classes, then resolves models by exact -`hf_model_paths` or optional detector predicates. +`hf_model_paths` or optional detector predicates. Prefer exact `hf_model_paths` +for public checkpoints used in docs or tests; use detector predicates only for +families where local mirrors, renamed repos, or generated paths are common. The `EntryClass` in your pipeline file is automatically discovered by the registry's `_discover_and_register_pipelines()` function -- no additional registration needed for the pipeline class itself. @@ -552,6 +556,10 @@ After implementation, **you must verify that the generated output is not noise** |-------|----------|---------------------|----------------| | GLM-Image | `runtime/pipelines/glm_image.py` | `stages/model_specific_stages/glm_image.py` | `configs/pipeline_configs/glm_image.py` | | Qwen-Image-Layered | `runtime/pipelines/qwen_image.py` (`QwenImageLayeredPipeline`) | `stages/model_specific_stages/qwen_image_layered.py` | `configs/pipeline_configs/qwen_image.py` (`QwenImageLayeredPipelineConfig`) | +| Cosmos3 | `runtime/pipelines/cosmos3_pipeline.py` | `stages/model_specific_stages/cosmos3.py` | `configs/pipeline_configs/cosmos3.py` | +| ErnieImage | `runtime/pipelines/ernie_image.py` | `runtime/pipelines/ernie_image.py` | `configs/pipeline_configs/ernie_image.py` | +| Hunyuan3D | `runtime/pipelines/hunyuan3d_pipeline.py` | `stages/model_specific_stages/hunyuan3d.py` | `configs/pipeline_configs/hunyuan3d.py` | +| LingBot World realtime | `runtime/pipelines/lingbot_world_causal_dmd_pipeline.py` | `stages/model_specific_stages/lingbot_world/` | `configs/pipeline_configs/lingbot_world.py` | ### Modular Style (when standard stages fit well) @@ -560,6 +568,10 @@ After implementation, **you must verify that the generated output is not noise** | Qwen-Image (T2I) | `runtime/pipelines/qwen_image.py` | Uses `add_standard_t2i_stages()` — standard text encoding + latent prep fits this model | | Qwen-Image-Edit | `runtime/pipelines/qwen_image.py` | Uses `add_standard_ti2i_stages()` — standard image-to-image flow | | Flux | `runtime/pipelines/flux.py` | Uses `add_standard_t2i_stages()` with custom `prepare_mu` | +| FLUX.2 / FLUX.2 Klein | `runtime/pipelines/flux_2.py`, `flux_2_klein.py` | Reuses FLUX.2 stages; Klein differences live in config and sampling params | +| Z-Image | `runtime/pipelines/zimage_pipeline.py` | Uses standard image pipeline stages plus Z-Image-specific config/model code | +| Ideogram4 | `runtime/pipelines/ideogram.py` | Uses dedicated text encoding and denoising stages while keeping standard latent prep | +| SANA | `runtime/pipelines/sana.py` | Spatial image pipeline; reuse the spatial image config pattern | | Wan | `runtime/pipelines/wan_pipeline.py` | Uses `add_standard_ti2v_stages()` | --- diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-ako4all-kernel/SKILL.md b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-ako4all-kernel/SKILL.md deleted file mode 100644 index 423ea408e..000000000 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-ako4all-kernel/SKILL.md +++ /dev/null @@ -1,145 +0,0 @@ ---- -name: sglang-diffusion-ako4all-kernel -description: Use when optimizing an existing SGLang diffusion kernel with AKO4ALL, including AKO4ALL repo hygiene, custom microbench setup, ncu-guided iteration, and end-to-end denoise validation. Also use when a sibling AKO4ALL repo must be cloned or refreshed before starting kernel tuning work. ---- - -# SGLang Diffusion AKO4ALL Kernel - -Use this skill to run the full AKO4ALL-based optimization loop for an existing SGLang diffusion kernel. -It is the default implementation path once the benchmark/profile skill has already shown that a hotspot is real and not covered by an existing fast path. This workflow bootstraps a custom AKO harness, benchmarks and profiles the kernel, iterates with `ncu`, ports the best version back to `sglang`, then validates with targeted tests and model-level denoise runs. - -This skill assumes a sibling repo layout like: - -```text -/ -├── sglang/ -└── AKO4ALL/ -``` - -If `AKO4ALL/` is missing under the current base directory, clone it first. - -## Use This Skill When - -- tuning an existing diffusion Triton, CUDA JIT, CuTeDSL, or runtime-integrated kernel in `sglang` -- `sglang-diffusion-benchmark-profile` has already ruled out an existing in-repo fast path or overlap family -- creating a custom AKO4ALL harness for a real diffusion kernel instead of using the default benchmark tasks -- validating that a kernel-level win transfers to Qwen, FLUX, Wan, Hunyuan, MOVA, or other diffusion denoise latency -- preparing PR artifacts such as microbench tables, `ncu` before/after data, and proof image outputs - -Do not start here when the bottleneck has not been proven yet. -First use [../sglang-diffusion-benchmark-profile/SKILL.md](../sglang-diffusion-benchmark-profile/SKILL.md) to: -- measure the real denoise regression -- collect the perf dump baseline -- capture one representative `torch.profiler` trace -- rule out existing mainline fast paths -- prove the run stayed on the native SGLang diffusion backend, not a diffusers fallback - -Before opening AKO, also read -[../sglang-diffusion-benchmark-profile/existing-fast-paths.md](../sglang-diffusion-benchmark-profile/existing-fast-paths.md). -It records current mainline fusions plus the open PR watchlist for diffusion -kernel, VAE, attention, cache, and scheduling work. If an open PR already covers -the same shape family, use it as prior art or decide whether to rebase/extend it -instead of starting a duplicate kernel. - -If a future specialized optimization skill matches the kernel family better than AKO4ALL, hand off there instead. The diagnosis contract stays the same. - -## Mandatory AKO4ALL Preflight - -Before any AKO work: - -1. Run `scripts/ensure_ako4all_clean.sh [base-dir]`. -2. If `/AKO4ALL` does not exist, the script clones it. -3. Do not continue unless `AKO4ALL` is: - - on the upstream default branch, usually `main` - - fully clean with no tracked or untracked local changes - - exactly synced to `upstream/` -4. If the script reports local commits, divergence, or a dirty worktree, stop and clean or re-clone the repo before continuing. - -The script creates an `upstream` remote automatically when missing. -By default it uses the existing `origin` URL, or `AKO4ALL_URL` if you need to override the clone source. - -## Workflow - -### 1. Scope the Kernel - -- Identify the exact kernel entry point and runtime call sites in `sglang`. -- Record the target shapes, dtypes, model families, and whether the kernel is on a hot path. -- Reuse existing unit tests and benchmark entry points when they already exist. -- Record whether the hotspot overlaps an open PR from `existing-fast-paths.md`; - if it does, note the PR number in the AKO context and final PR artifacts. - -### 2. Bootstrap the AKO Harness - -Inside the clean `AKO4ALL` repo: - -- read `TASK.md` and `HINTS.md` -- create a custom harness instead of relying on the stock benchmark tasks -- mirror the real SGLang kernel into: - - `input/reference.py` - - `input/.py` - - `solution/.py` - - `bench/bench_.py` -- keep a short context note in `context/` when the kernel has model-specific shape assumptions or perf conclusions - -The custom benchmark should: - -- cover representative diffusion shapes -- check correctness against the reference kernel -- report aggregate runtime plus per-shape results when useful - -### 3. Establish the Baseline - -- run the AKO custom microbench before changing the kernel -- capture one representative `ncu` baseline on the hottest meaningful shape -- note whether the bottleneck looks like registers, occupancy, instruction count, launch config, or memory latency - -### 4. Iterate in AKO4ALL - -- change one idea at a time -- rerun the microbench after every change -- update `ITERATIONS.md` with hypothesis, result, and next step -- prefer simple, explainable wins over clever rewrites that do not transfer - -After 3 consecutive no-improvement or regression iterations: - -- rerun `ncu` -- re-read `ITERATIONS.md` -- change direction instead of continuing blind sweeps - -### 5. Port the Best Version Back to SGLang - -- apply the best candidate to the real `sglang` kernel file -- run import or syntax checks and targeted tests first -- keep the AKO `solution/` version aligned with the main-tree version you actually want to keep - -### 6. Validate on Real Models - -- use the benchmark/profile skill for denoise perf dumps and before/after comparison -- prefer exact local snapshot validation when testing local edits on a GPU box -- run targeted kernel tests first -- run model-level denoise benchmarks with perf dumps -- compare baseline vs optimized runs with `compare_perf.py` -- if the PR needs proof that generation still works, save one real model output image - -### 7. Prepare PR Artifacts - -At minimum, keep: - -- one microbench table -- one denoise-stage table -- one end-to-end table -- one `ncu` before/after pair on the most representative kernel shape -- one generated image when the kernel affects production inference - -See [references/ako-loop.md](references/ako-loop.md) for the checklist and common stop rules. - -## Operating Rules - -- Treat AKO4ALL repo hygiene as a gate, not a suggestion. -- Prefer exact local snapshot validation over hand-wavy “remote tree is close enough”. -- Do not start or justify kernel work from traces collected after - `Falling back to diffusers backend`, `Using diffusers backend`, or - `Loaded diffusers pipeline`; fix backend selection and rerun the - benchmark/profile workflow first. -- Keep model-level validation honest: if microbench improves but denoise does not, do not keep the AKO-only variant in the main code path. -- When writing conclusions, explain the win in terms of measurable causes such as lower registers per thread, higher occupancy, fewer executed instructions, or better scheduler eligibility. diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-ako4all-kernel/references/ako-loop.md b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-ako4all-kernel/references/ako-loop.md deleted file mode 100644 index 557af9dba..000000000 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-ako4all-kernel/references/ako-loop.md +++ /dev/null @@ -1,54 +0,0 @@ -# AKO Loop Checklist - -Use this checklist after `scripts/ensure_ako4all_clean.sh` succeeds. - -## Minimum Repo Layout - -Inside `AKO4ALL/`, prefer these files for a diffusion kernel task: - -- `input/reference.py` -- `input/.py` -- `solution/.py` -- `bench/bench_.py` -- `context/_notes.md` - -## Baseline Checklist - -- Reproduce the current SGLang kernel exactly in AKO first. -- Run the custom microbench before making edits. -- Record one representative `ncu` report on a real hot shape. -- Note the baseline bottleneck in plain language. - -## Iteration Discipline - -- One optimization idea per iteration. -- Re-benchmark after every code change. -- Log the result in `ITERATIONS.md`. -- Keep the best candidate easy to identify. - -Stop a direction early when: - -- 3 consecutive iterations do not beat the best runtime -- correctness gets fragile -- AKO-only gains stop transferring to real denoise runs - -## Real Validation Gate - -Before calling a kernel "done", validate all of: - -- syntax or import checks -- targeted unit test or regression test -- kernel or op-level benchmark -- model-level denoise benchmark with perf dumps -- one generated image if the PR needs production proof - -## PR Artifact Checklist - -Prepare these artifacts: - -- microbench table -- denoise-stage table -- end-to-end table -- one `ncu` before or after pair -- one short explanation of why the kernel got faster -- one generated output image when applicable diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-ako4all-kernel/scripts/ensure_ako4all_clean.sh b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-ako4all-kernel/scripts/ensure_ako4all_clean.sh deleted file mode 100755 index 055ad56ec..000000000 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-ako4all-kernel/scripts/ensure_ako4all_clean.sh +++ /dev/null @@ -1,84 +0,0 @@ -#!/usr/bin/env bash -set -euo pipefail - -BASE_DIR="${1:-$PWD}" -BASE_DIR="$(cd "$BASE_DIR" && pwd)" -AKO_DIR="${BASE_DIR}/AKO4ALL" -CANONICAL_UPSTREAM_URL="https://github.com/TongmingLAIC/AKO4ALL.git" -UPSTREAM_URL="${AKO4ALL_UPSTREAM_URL:-$CANONICAL_UPSTREAM_URL}" -CLONE_URL="${AKO4ALL_URL:-$UPSTREAM_URL}" - -say() { - printf '[ako4all] %s\n' "$*" -} - -fail() { - printf '[ako4all] ERROR: %s\n' "$*" >&2 - exit 1 -} - -if [[ ! -d "$AKO_DIR/.git" ]]; then - say "AKO4ALL not found under ${BASE_DIR}; cloning ${CLONE_URL}" - git clone "$CLONE_URL" "$AKO_DIR" -fi - -cd "$AKO_DIR" - -if ! git remote get-url origin >/dev/null 2>&1; then - fail "AKO4ALL exists but has no origin remote." -fi - -if ! git remote get-url upstream >/dev/null 2>&1; then - say "Adding missing upstream remote -> ${UPSTREAM_URL}" - git remote add upstream "$UPSTREAM_URL" -fi - -git fetch upstream --prune -git remote set-head upstream -a >/dev/null 2>&1 || true - -default_branch="${AKO4ALL_BRANCH:-}" -if [[ -z "$default_branch" ]]; then - if upstream_head="$(git symbolic-ref --quiet --short refs/remotes/upstream/HEAD 2>/dev/null)"; then - default_branch="${upstream_head#upstream/}" - else - default_branch="main" - fi -fi - -if [[ -n "$(git status --porcelain)" ]]; then - fail "AKO4ALL worktree is dirty. Clean all local changes before using this skill." -fi - -if git show-ref --verify --quiet "refs/heads/${default_branch}"; then - git switch "$default_branch" >/dev/null -else - git switch -c "$default_branch" --track "upstream/${default_branch}" >/dev/null -fi - -git fetch upstream --prune - -local_head="$(git rev-parse HEAD)" -upstream_head="$(git rev-parse "upstream/${default_branch}")" - -if [[ "$local_head" != "$upstream_head" ]]; then - if git merge-base --is-ancestor "$local_head" "$upstream_head"; then - say "Fast-forwarding ${default_branch} to upstream/${default_branch}" - git merge --ff-only "upstream/${default_branch}" >/dev/null - else - fail "Local ${default_branch} diverges from upstream/${default_branch}. Reset or re-clone AKO4ALL before continuing." - fi -fi - -if [[ -n "$(git status --porcelain)" ]]; then - fail "AKO4ALL became dirty after sync; stop and inspect the repo." -fi - -final_head="$(git rev-parse HEAD)" -expected_head="$(git rev-parse "upstream/${default_branch}")" -if [[ "$final_head" != "$expected_head" ]]; then - fail "AKO4ALL is not exactly at upstream/${default_branch}." -fi - -say "Ready: ${AKO_DIR}" -say "Branch: ${default_branch}" -say "Commit: ${final_head}" diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/SKILL.md b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/SKILL.md index 7a848d6f7..aba54feb9 100644 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/SKILL.md +++ b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/SKILL.md @@ -12,7 +12,7 @@ This skill is diagnosis-first. It owns: - perf dump collection and before/after comparison - `torch.profiler` trace capture and quick hotspot ranking - mapping hot kernels back to known fast paths and fusion families -- handing confirmed kernel work to a specialized optimization skill such as [../sglang-diffusion-ako4all-kernel/SKILL.md](../sglang-diffusion-ako4all-kernel/SKILL.md) +- packaging confirmed kernel work with enough evidence for the appropriate kernel, Nsight, or framework-specific optimization workflow This skill does not own low-level kernel authoring or standalone Nsight workflows. @@ -42,7 +42,7 @@ If any benchmark, perf-dump, or `torch.profiler` command prints one of those sig ## Main Reference -- [benchmark-and-profile.md](benchmark-and-profile.md) — canonical denoise benchmark, perf dump, and `torch.profiler` workflow; uses checked-in nightly-aligned presets plus skill-only stress recipes such as `LTX-2.3` one-stage/two-stage, HunyuanVideo, MOVA, Helios, JoyAI/FireRed image edit, and Hunyuan3D shape +- [benchmark-and-profile.md](benchmark-and-profile.md) — canonical denoise benchmark, perf dump, and `torch.profiler` workflow; uses checked-in nightly-aligned presets plus current-source extras such as FLUX.2 Klein, Cosmos3, Ideogram4, ERNIE/GLM/SANA image models, FastWan2.2, `LTX-2.3` one-stage/two-stage/HQ, HunyuanVideo, MOVA, Helios, JoyAI/FireRed image edit, and Hunyuan3D shape - [existing-fast-paths.md](existing-fast-paths.md) — map bottlenecks to existing fused kernels, packed QKV paths, fused `QK norm + RoPE`, distributed overlap patterns, and open optimization PRs before proposing new code - [scripts/diffusion_skill_env.py](scripts/diffusion_skill_env.py) — preflight helper: repo root discovery via `sglang.__file__`, write-access probe, benchmark/profile output directories, idle GPU selection - [scripts/bench_diffusion_denoise.py](scripts/bench_diffusion_denoise.py) — end-to-end denoise benchmark preset runner via `sglang generate`; supports `--no-torch-compile`, validates nightly preset drift with `--validate-nightly-alignment`, and saves perf dumps by label for `compare_perf.py` @@ -53,8 +53,11 @@ Before calling a diffusion hotspot "new", first classify it with `existing-fast- Always rule out these existing families first: - HunyuanVideo VAE GroupNorm+SiLU +- LTX upsampler GroupNorm+SiLU - Z-Image residual-form modulation - fused diffusion `QK norm + RoPE` +- LTX2 split RoPE +- varlen USP attention pack/scatter - NVFP4 / Nunchaku packed QKV - Nunchaku fused GELU MLP - Ulysses / USP attention overlap diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/benchmark-and-profile.md b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/benchmark-and-profile.md index a744c0721..0132ab656 100644 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/benchmark-and-profile.md +++ b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/benchmark-and-profile.md @@ -21,10 +21,10 @@ This guide intentionally stops at: - hotspot ranking - mapping hotspots to known fast paths -If the hotspot survives this checklist, hand the work to -`sglang-diffusion-ako4all-kernel` or another specialized kernel-optimization -skill. Do not grow this skill back into a general Nsight or kernel-authoring -guide. +If the hotspot survives this checklist, package the perf dump, profiler trace, +exact command, and shape/topology notes for the appropriate kernel, Nsight, or +framework-specific optimization workflow. Do not grow this skill back into a +general Nsight or kernel-authoring guide. ## Prerequisites @@ -163,7 +163,8 @@ PYTHONPATH=python python3 "$BENCH_PY" \ --output-dir "${BENCH_DIR}" ``` -Run the full preset sweep: +Run the full preset sweep only when you have enough GPU time for both the +nightly-aligned cases and the source-tracked extras: ```bash PYTHONPATH=python python3 "$BENCH_PY" \ @@ -172,7 +173,21 @@ PYTHONPATH=python python3 "$BENCH_PY" \ --output-dir "${BENCH_DIR}" ``` -Nightly-aligned presets come first; skill-only presets stay available after them. +Nightly-aligned presets come first, followed by current-source extras from the +registry / GPU test cases, then broader skill-only stress presets. + +Use the preset categories this way: +- **Nightly-aligned**: exact mirrors of + `scripts/ci/utils/diffusion/comparison_configs.json`; use these when the goal + is apples-to-apples comparison with CI / nightly coverage. +- **Current-source extras**: models or request shapes with explicit support + evidence in the current registry, GPU cases, compatibility matrix, pipeline + files, or unit tests, but without a nightly comparison case yet. +- **Skill-only stress / coverage presets**: extra profiling scenarios kept by + this skill to stress a topology, high-resolution path, multi-GPU mode, or + model-specific stage. These may be older than the latest registry additions, + so re-check the active source tree before treating them as support-matrix + commitments. | Preset | Model | Nightly | Notes | | --- | --- | --- | --- | @@ -186,6 +201,19 @@ Nightly-aligned presets come first; skill-only presets stay available after them | `ltx2` | `Lightricks/LTX-2` | Yes: `ltx2_twostage_t2v` | `LTX2TwoStagePipeline`, 2 GPUs, CFG parallel, 768x512, 121 frames, seed 42 | | `ltx23-ti2v-two-stage` | `Lightricks/LTX-2.3` | Yes: `ltx2.3_twostage_ti2v_2gpus` | Nightly cat image, motion prompt, `LTX2TwoStagePipeline`, 2 GPUs, `--cfg-parallel-size 2`, 768x512, 121 frames, seed 42 | | `wan-i2v` | `Wan-AI/Wan2.2-I2V-A14B-Diffusers` | Yes: `wan22_i2v_a14b_720p` | Nightly cat image and motion prompt, 1280x720, 81 frames, 4 GPUs, CFG parallel, Ulysses degree 2, text encoder CPU offload and pinned CPU memory | +| `qwen-image` | `Qwen/Qwen-Image` | No | Current-source extra covering the base Qwen-Image native path, separate from the nightly `Qwen-Image-2512` case | +| `qwen-edit-2509` | `Qwen/Qwen-Image-Edit-2509` | No | Current-source extra for the pre-2511 edit-plus path; uses the cat image, 1024x1024 | +| `zimage-base` | `Tongyi-MAI/Z-Image` | No | Current-source extra for non-turbo Z-Image; keep it separate from `zimage` / `Z-Image-Turbo` | +| `flux2-klein` | `black-forest-labs/FLUX.2-klein-4B` | No | Current-source extra for the distilled FLUX.2 Klein path; gated repo, 1024x1024, DiT layerwise offload disabled | +| `flux2-klein-base` | `black-forest-labs/FLUX.2-klein-base-4B` | No | Current-source extra for the undistilled FLUX.2 Klein Base path; gated repo, 1024x1024, DiT layerwise offload disabled | +| `cosmos3-nano-t2i` | `nvidia/Cosmos3-Nano` | No | Current-source extra for the single-frame Cosmos3 image path; sets `SGLANG_DISABLE_COSMOS3_GUARDRAILS=1` in the helper environment | +| `cosmos3-nano-t2v` | `nvidia/Cosmos3-Nano` | No | Current-source extra for a short Cosmos3 video path; sets `SGLANG_DISABLE_COSMOS3_GUARDRAILS=1` in the helper environment | +| `ideogram4-fp8` | `ideogram-ai/ideogram-4-fp8` | No | Current-source extra matching the native Ideogram 4 FP8 pipeline; do not override steps/guidance directly because the sampling preset owns them | +| `ernie-image-turbo` | `baidu/ERNIE-Image-Turbo` | No | Current-source extra for ERNIE-Image Turbo | +| `glm-image` | `zai-org/GLM-Image` | No | Current-source extra for GLM-Image | +| `sana-1.5-1.6b` | `Efficient-Large-Model/SANA1.5_1.6B_1024px_diffusers` | No | Current-source extra for a SANA native image path | +| `fastwan22-ti2v-5b` | `FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers` | No | Current-source extra matching the FastWan2.2 TI2V registered path | +| `ltx23-hq-two-stage` | `Lightricks/LTX-2.3` | No | Current-source extra for `LTX2TwoStageHQPipeline` with `snapshot` device mode; high-resolution and VRAM-heavy | | `ltx23-one-stage` | `Lightricks/LTX-2.3` | No | Skill-only extra preset for the native `LTX-2.3` one-stage baseline; 2 GPUs, 768x512, 121 frames, fps 24, 30 steps, guidance 3.0, seed 1234 | | `ltx23-two-stage` | `Lightricks/LTX-2.3` | No | Skill-only high-resolution stress preset for the native `LTX-2.3` two-stage path; uses `LTX2TwoStagePipeline`, 2 GPUs, 1536x1024, 121 frames, fps 24, 30 steps, guidance 3.0, seed 1234 | | `ltx23-two-stage-cfg-parallel` | `Lightricks/LTX-2.3` | No | Skill-only high-resolution CFG-parallel stress preset matching `ltx23-two-stage` plus `--cfg-parallel-size 2` | @@ -461,6 +489,8 @@ the known mainline families. | `fused_inplace_qknorm_rope` missing, but separate qk norm plus rope show up | Check whether the fused diffusion `QK norm + RoPE` path should have engaged | | `to_q -> to_k -> to_v` on NVFP4 or Nunchaku FLUX-family checkpoints | Treat as a packed-QKV fast-path miss or checkpoint-format mismatch | | `fused_norm_tanh_mul_add*` missing on Z-Image | Treat as a missing mainline modulation path, not a new fusion request | +| LTX-2 split RoPE appears as a long PyTorch elementwise chain | Check the `apply_ltx2_split_rotary_emb` Triton path and its shape guards | +| masked attention spends time packing/unpacking Q/K/V | Check whether fused varlen USP pack/scatter should have engaged | | `all_to_all`, ring attention, or async A2A dominate | Classify against Ulysses, USP, or turbo-layer overlap first | | split `fc1 -> gelu -> quant -> fc2.lora_down` on Nunchaku FLUX | Treat as a missing fused GELU MLP path | | attention kernels dominate | Confirm backend, topology, and shape guards before proposing a new kernel | @@ -475,7 +505,7 @@ Only after the hotspot survives the fast-path checklist: 1. save a baseline perf dump 2. save a representative `torch.profiler` trace 3. note the exact model, shape, dtype, and GPU topology -4. hand the work to `sglang-diffusion-ako4all-kernel` or another future specialized optimization skill +4. hand the work to the appropriate kernel, Nsight, or framework-specific optimization workflow This skill intentionally stops here. It tells you whether you are looking at: - a missing existing optimization @@ -490,4 +520,4 @@ This skill intentionally stops here. It tells you whether you are looking at: - [ ] one representative `torch.profiler` trace saved - [ ] hotspot classified against `existing-fast-paths.md` - [ ] reference image or video checked for correctness -- [ ] any remaining kernel work handed to a specialized optimization skill +- [ ] any remaining kernel work handed off with perf/profile evidence attached diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/existing-fast-paths.md b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/existing-fast-paths.md index 24399eee3..932a49e07 100644 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/existing-fast-paths.md +++ b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/existing-fast-paths.md @@ -2,11 +2,13 @@ Use this guide when mapping a diffusion bottleneck to an existing fused path or distributed overlap pattern in `sglang.multimodal_gen`. Prefer reuse and -configuration first before handing the problem to a specialized kernel-optimization skill. +configuration first before handing the problem to a kernel, Nsight, or +framework-specific optimization workflow. **Key Files** - `python/sglang/multimodal_gen/runtime/layers/layernorm.py` - `python/sglang/multimodal_gen/runtime/layers/elementwise.py` +- `python/sglang/multimodal_gen/runtime/layers/fused_scale_shift_gate.py` - `python/sglang/multimodal_gen/runtime/layers/rotary_embedding/utils.py` - `python/sglang/jit_kernel/diffusion/triton/scale_shift.py` - `python/sglang/jit_kernel/diffusion/group_norm_silu.py` @@ -14,8 +16,14 @@ configuration first before handing the problem to a specialized kernel-optimizat - `python/sglang/jit_kernel/diffusion/triton/norm.py` - `python/sglang/jit_kernel/diffusion/triton/rmsnorm_onepass.py` - `python/sglang/jit_kernel/diffusion/triton/rotary.py` +- `python/sglang/jit_kernel/diffusion/triton/ltx2_rotary.py` +- `python/sglang/jit_kernel/diffusion/triton/varlen_pack_pad.py` - `python/sglang/jit_kernel/diffusion/cutedsl/scale_residual_norm_scale_shift.py` +- `python/sglang/jit_kernel/tests/diffusion/test_qwen_image_modulation.py` - `python/sglang/jit_kernel/tests/diffusion/test_group_norm_silu.py` +- `python/sglang/jit_kernel/tests/diffusion/test_varlen_pack_pad.py` +- `python/sglang/jit_kernel/tests/diffusion/test_varlen_uspattn_equivalence.py` +- `python/sglang/jit_kernel/benchmark/diffusion/bench_qwen_image_modulation.py` - `python/sglang/jit_kernel/benchmark/diffusion/bench_group_norm_silu.py` - `python/sglang/jit_kernel/norm.py` - `python/sglang/multimodal_gen/runtime/platforms/cuda.py` @@ -24,10 +32,10 @@ configuration first before handing the problem to a specialized kernel-optimizat **Core Fusion Patterns** -1. Scale/Shift elementwise fusion (AdaLN modulation) -- Kernels: `fuse_scale_shift_kernel`, `fuse_scale_shift_gate_select01_kernel` -- Locations: `elementwise.py`, `layernorm.py`, `qwen_image.py`, `triton/scale_shift.py` -- Use cases: `x * (1 + scale) + shift` and `a * (k + b) + c` +1. Scale/Shift elementwise and gate fusion (AdaLN modulation) +- Kernels: `fuse_scale_shift_kernel`, `fuse_layernorm_scale_shift_gate_select01_kernel`, `fuse_residual_layernorm_scale_shift_gate_select01_kernel` +- Locations: `elementwise.py`, `layernorm.py`, `fused_scale_shift_gate.py`, `qwen_image.py`, `triton/scale_shift.py` +- Use cases: `x * (1 + scale) + shift`, `a * (k + b) + c`, and Qwen-style `(layernorm/residual layernorm) + scale/shift + gate select`. - Constraints: `x` must be CUDA and contiguous. `scale/shift` support 0D/1D/2D/3D/4D broadcast. 4D `[B, F, 1, C]` requires `L % F == 0`. - NPU fallback: `scale_shift.py` swaps to `npu_fallback` native path. - Validation: `python/sglang/jit_kernel/tests/diffusion/test_qwen_image_modulation.py`. @@ -72,11 +80,18 @@ configuration first before handing the problem to a specialized kernel-optimizat - NPU fallback: `npu_fallback.apply_rotary_embedding_native`. - Validation: `python/sglang/jit_kernel/tests/test_rope.py`. -7. HunyuanVideo VAE GroupNorm + SiLU fusion +7. LTX2 split RoPE fusion +- Kernel: `apply_ltx2_split_rotary_emb` +- Locations: `triton/ltx2_rotary.py`, `runtime/models/dits/ltx_2.py` +- Use case: LTX-2 split rotary embedding over `[B, S, num_heads * head_dim]` with separate `cos` and `sin` tensors. +- Constraints: `cos` and `sin` shapes must match `[B, H, S, head_dim / 2]`, and `inner_dim == H * head_dim`. +- Workflow rule: if LTX-2 traces show a large split-RoPE PyTorch chain, check whether the LTX2-specific Triton path was disabled by shape or dtype before proposing a new RoPE kernel. + +8. HunyuanVideo / LTX upsampler GroupNorm + SiLU fusion - Kernel: `triton_group_norm_silu` -- Locations: `diffusion/group_norm_silu.py`, `triton/group_norm_silu.py`, `runtime/models/vaes/hunyuanvae.py` +- Locations: `diffusion/group_norm_silu.py`, `triton/group_norm_silu.py`, `runtime/models/vaes/hunyuanvae.py`, `runtime/models/upsampler/latent_upsampler.py` - Use case: `activation(group_norm(x))` when the activation is non-inplace `nn.SiLU` and the GroupNorm is affine. -- Enablement: mainline uses `apply_group_norm_silu(...)` in HunyuanVideo VAE paths by default; there is no env toggle. The wrapper dispatches to Triton only when guards pass. +- Enablement: mainline uses `apply_group_norm_silu(...)` in HunyuanVideo VAE paths and LTX latent upsampler paths by default; there is no env toggle. The wrapper dispatches to Triton only when guards pass. - Constraints: CUDA inference path only; no grad, `x.requires_grad == False`, `nn.GroupNorm`, `nn.SiLU(inplace=False)`, affine norm with weight and bias. Unsupported cases fall back to native `activation(norm(x))`. - Validation: `python/sglang/jit_kernel/tests/diffusion/test_group_norm_silu.py`. - Microbench: `python/sglang/jit_kernel/benchmark/diffusion/bench_group_norm_silu.py`. @@ -99,6 +114,12 @@ configuration first before handing the problem to a specialized kernel-optimizat - Location: `rotary_embedding/utils.py` - Behavior: `flashinfer.rope.apply_rope_with_cos_sin_cache_inplace` when available, otherwise Triton RoPE fallback. +4. Varlen USP attention pack/scatter +- Locations: `runtime/layers/attention/layer.py`, `triton/varlen_pack_pad.py` +- Behavior: masked `USPAttention.forward` can gather dense Q/K/V into packed `[total_valid, H, D]` rows with `fused_pack_qkv`, run varlen attention, then scatter back with `fused_scatter_to_padded`. +- Validation: `python/sglang/jit_kernel/tests/diffusion/test_varlen_pack_pad.py` and `test_varlen_uspattn_equivalence.py`. +- Workflow rule: if a masked attention trace spends time in Python/advanced indexing pack or scatter, first check whether this fused varlen path should have engaged. + **QK Norm Optimization** - Entry point: `apply_qk_norm` in `layernorm.py`. @@ -125,7 +146,7 @@ configuration first before handing the problem to a specialized kernel-optimizat - Supported head dims: `64, 128, 256`. - Behavior: `apply_qk_norm_rope` prefers the fused JIT kernel when all guards pass; otherwise it falls back to `apply_qk_norm(...)` plus `apply_flashinfer_rope_qk_inplace(...)`. - Validation: `python/sglang/jit_kernel/tests/diffusion/test_qknorm_rope.py`. -- Watchlist: PR #24025 adds LTX2-specific QK norm fusion work. Until merged, treat LTX2 traces that miss the generic fused path as an enablement/shape-guard issue first, not proof that the PR path exists locally. +- Workflow rule: treat LTX2 traces that miss the generic fused path as an enablement/shape-guard issue first, and check the separate LTX2 split-RoPE path before proposing new attention-prep kernels. **Nunchaku Fused GELU MLP** @@ -144,11 +165,13 @@ configuration first before handing the problem to a specialized kernel-optimizat **Common Entry Points in Diffusion Models** - AdaLN modulation: `LayerNormScaleShift`, `RMSNormScaleShift`, `ScaleResidual*` in `layernorm.py`. -- Qwen-Image gating: `fuse_scale_shift_gate_select01_kernel` in `qwen_image.py`. +- Qwen-Image gating: `fuse_layernorm_scale_shift_gate_select01_kernel` and `fuse_residual_layernorm_scale_shift_gate_select01_kernel` through `fused_scale_shift_gate.py` and `qwen_image.py`. - Z-Image residual-form modulation: `fused_norm_tanh_mul_add` and `fused_norm_tanh_mul_add_norm_scale` in `zimage.py`. -- HunyuanVideo VAE GroupNorm+SiLU: `apply_group_norm_silu` in `hunyuanvae.py`; default-eligible when wrapper guards pass. +- HunyuanVideo VAE and LTX upsampler GroupNorm+SiLU: `apply_group_norm_silu` in `hunyuanvae.py` and `latent_upsampler.py`; default-eligible when wrapper guards pass. - QK norm: `apply_qk_norm` used in `flux.py`, `flux_2.py`, `qwen_image.py`, `zimage.py`, `wanvideo.py`, `ltx_2.py`, `hunyuanvideo.py`. - QK norm + RoPE: `apply_qk_norm_rope` in `layernorm.py`; use this path when the model wants fused attention prep instead of separate QK norm and RoPE calls. +- LTX2 split RoPE: `apply_ltx2_split_rotary_emb` in `ltx_2.py`. +- Varlen USP attention: `fused_pack_qkv` and `fused_scatter_to_padded` in `attention/layer.py`. - Nunchaku fused GELU MLP: `_fused_gelu_mlp` in `flux.py` for quantized FLUX-family checkpoints. - NVFP4 / packed QKV attention: `to_qkv`, `to_added_qkv`, and `to_qkv_mlp_proj` in FLUX-family quantized paths. - RoPE: `_apply_rotary_emb` prefers Triton; Q/K RoPE prefers FlashInfer when present. @@ -161,11 +184,11 @@ configuration first before handing the problem to a specialized kernel-optimizat - Dual-stream diffusion models: `use_dual_stream = True` in models such as `hunyuan3d.py` is an existing overlap family. - Workflow rule: if a hotspot is communication-heavy, rule out these in-repo overlap families before proposing a brand new overlap design. -**Open PR Watchlist** +**Historical PR Watchlist** -As of 2026-05-02, these SGLang PRs were still open. Use them as upstream -direction and prior art, not as current-main behavior. Re-check the PR state -before relying on any file path or flag. +These SGLang PRs are useful as upstream direction and prior art, not as +current-main behavior. Re-check the PR state and the active source tree before +relying on any file path, flag, or claim about whether the work has merged. - Norm, modulation, and packed projection fusions: - #24025 LTX2 QK norm fusion. @@ -211,4 +234,4 @@ before relying on any file path or flag. - Keep CuTe compile cache keys aligned to `(dtype, ndim, D)`. - Avoid implicit broadcasts that force hidden `contiguous()` copies. - Preserve NPU and ROCm fallback paths. -- If none of the families above match, package the evidence from the benchmark/profile skill and hand the kernel work to a specialized optimization skill such as `sglang-diffusion-ako4all-kernel`. +- If none of the families above match, package the evidence from the benchmark/profile skill and hand the kernel work to the appropriate kernel, Nsight, or framework-specific optimization workflow. diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py index bbf99cdb3..b4003ff23 100755 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py +++ b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py @@ -12,7 +12,7 @@ Usage: # Tag the run for later compare_perf.py usage python3 python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py --model flux --label tuned - # All 20 preset models + # All preset models python3 python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py --all # Show preset order, model path, and nightly mapping @@ -56,13 +56,18 @@ ASSET_DIR = ensure_dir(get_assets_dir(REPO_ROOT)) NIGHTLY_CONFIG_PATH = ( REPO_ROOT / "scripts" / "ci" / "utils" / "diffusion" / "comparison_configs.json" ) -GATED_MODELS = {"flux", "flux2"} +GATED_MODELS = { + "flux", + "flux2", + "flux2-klein", + "flux2-klein-base", +} DIFFUSERS_FALLBACK_SIGNALS = ( "falling back to diffusers backend", "using diffusers backend", "loaded diffusers pipeline", ) -CATALOG_TABLE_WIDTH = 105 +CATALOG_TABLE_WIDTH = 140 RESULTS_TABLE_WIDTH = 105 NIGHTLY_PRESET_ORDER = ( "flux", @@ -80,7 +85,7 @@ NIGHTLY_PRESET_ORDER = ( # --------------------------------------------------------------------------- # Model configs — kept in exact sync with benchmark-and-profile.md # Nightly-aligned presets mirror scripts/ci/utils/diffusion/comparison_configs.json -# first, followed by skill-only extras. +# first, followed by current-source extras and skill-only stress / coverage presets. # Each entry produces the same `sglang generate` command as shown in that doc. # --------------------------------------------------------------------------- MODELS = { @@ -217,7 +222,137 @@ MODELS = { "--pin-cpu-memory", ], }, - # 11. Skill-only extra preset + # Source-tracked extras from current registry / GPU test coverage. + "qwen-image": { + "path": "Qwen/Qwen-Image", + "prompt": "A futuristic cyberpunk city at night, neon lights reflecting on wet streets", + "extra_args": [ + "--width=1024", + "--height=1024", + ], + }, + # Requires: /inputs/diffusion_benchmark/figs/cat.png + "qwen-edit-2509": { + "path": "Qwen/Qwen-Image-Edit-2509", + "prompt": "Make the cat wear a red hat", + "image_path": str(ASSET_DIR / "cat.png"), + "extra_args": [ + "--width=1024", + "--height=1024", + ], + }, + "zimage-base": { + "path": "Tongyi-MAI/Z-Image", + "prompt": "A futuristic cyberpunk city at night, neon lights reflecting on wet streets", + "extra_args": [ + "--width=1024", + "--height=1024", + ], + }, + "flux2-klein": { + "path": "black-forest-labs/FLUX.2-klein-4B", + "prompt": "A futuristic cyberpunk city at night, neon lights reflecting on wet streets", + "extra_args": [ + "--width=1024", + "--height=1024", + "--dit-layerwise-offload", + "false", + ], + }, + "flux2-klein-base": { + "path": "black-forest-labs/FLUX.2-klein-base-4B", + "prompt": "A futuristic cyberpunk city at night, neon lights reflecting on wet streets", + "extra_args": [ + "--width=1024", + "--height=1024", + "--dit-layerwise-offload", + "false", + ], + }, + "cosmos3-nano-t2i": { + "path": "nvidia/Cosmos3-Nano", + "prompt": "A red cube on a white table, product photo.", + "env": { + "SGLANG_DISABLE_COSMOS3_GUARDRAILS": "1", + }, + "extra_args": [ + "--width=1024", + "--height=1024", + "--num-frames=1", + "--num-inference-steps=35", + ], + }, + "cosmos3-nano-t2v": { + "path": "nvidia/Cosmos3-Nano", + "prompt": "A blue box slides across a clean warehouse floor.", + "env": { + "SGLANG_DISABLE_COSMOS3_GUARDRAILS": "1", + }, + "extra_args": [ + "--width=832", + "--height=480", + "--num-frames=9", + "--num-inference-steps=4", + ], + }, + "ideogram4-fp8": { + "path": "ideogram-ai/ideogram-4-fp8", + "prompt": "A clean product poster for a new open-source inference engine", + "extra_args": [ + "--width=1024", + "--height=1024", + ], + }, + "ernie-image-turbo": { + "path": "baidu/ERNIE-Image-Turbo", + "prompt": "A futuristic cyberpunk city at night, neon lights reflecting on wet streets", + "extra_args": [ + "--width=1024", + "--height=1024", + ], + }, + "glm-image": { + "path": "zai-org/GLM-Image", + "prompt": "A futuristic cyberpunk city at night, neon lights reflecting on wet streets", + "extra_args": [ + "--width=1024", + "--height=1024", + ], + }, + "sana-1.5-1.6b": { + "path": "Efficient-Large-Model/SANA1.5_1.6B_1024px_diffusers", + "prompt": "A futuristic cyberpunk city at night, neon lights reflecting on wet streets", + "extra_args": [ + "--width=1024", + "--height=1024", + ], + }, + "fastwan22-ti2v-5b": { + "path": "FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers", + "prompt": "The cat starts walking slowly towards the camera.", + "image_path": str(ASSET_DIR / "cat.png"), + "extra_args": [ + "--width=1280", + "--height=720", + "--num-frames=81", + ], + }, + "ltx23-hq-two-stage": { + "path": "Lightricks/LTX-2.3", + "prompt": "A beautiful sunset over the ocean", + "env": { + "PYTORCH_CUDA_ALLOC_CONF": "expandable_segments:True", + "SGLANG_LTX2_SNAPSHOT_RELEASE_EMPTY_CACHE": "true", + }, + "extra_args": [ + "--pipeline-class-name=LTX2TwoStageHQPipeline", + "--ltx2-two-stage-device-mode=snapshot", + "--width=1920", + "--height=1088", + "--num-frames=121", + ], + }, + # Skill-only extra preset "ltx23-one-stage": { "path": "Lightricks/LTX-2.3", "prompt": "A beautiful sunset over the ocean", @@ -233,7 +368,7 @@ MODELS = { "--num-gpus=2", ], }, - # 12. Skill-only extra preset + # Skill-only extra preset "ltx23-two-stage": { "path": "Lightricks/LTX-2.3", "prompt": "A beautiful sunset over the ocean", @@ -250,7 +385,7 @@ MODELS = { "--num-gpus=2", ], }, - # 13. Skill-only extra preset + # Skill-only extra preset "ltx23-two-stage-cfg-parallel": { "path": "Lightricks/LTX-2.3", "prompt": "A beautiful sunset over the ocean", @@ -268,7 +403,7 @@ MODELS = { "--cfg-parallel-size=2", ], }, - # 14. Skill-only extra preset + # Skill-only extra preset "hunyuanvideo": { "path": "hunyuanvideo-community/HunyuanVideo", "prompt": "A cat and a dog baking a cake together in a kitchen. The cat is carefully measuring flour, while the dog is stirring the batter with a wooden spoon. The kitchen is cozy, with sunlight streaming through the window.", @@ -281,7 +416,7 @@ MODELS = { "--num-inference-steps=30", ], }, - # 15. Skill-only extra preset + # Skill-only extra preset # Requires: /inputs/diffusion_benchmark/figs/mova_single_person.jpg "mova-720p": { "path": "OpenMOSS-Team/MOVA-720p", @@ -297,7 +432,7 @@ MODELS = { "--num-inference-steps=2", ], }, - # 16. Skill-only extra preset + # Skill-only extra preset "helios": { "path": "BestWishYsh/Helios-Base", "prompt": "A curious raccoon", @@ -315,14 +450,13 @@ MODELS = { "false", ], }, - # 16. Skill-only extra preset + # Skill-only extra preset # Requires: /inputs/diffusion_benchmark/figs/cat.png "joyai-edit": { "path": "jdopensource/JoyAI-Image-Edit-Diffusers", "prompt": "Make the cat wear a red hat", "image_path": str(ASSET_DIR / "cat.png"), "extra_args": [ - "--backend=sglang", "--width=1024", "--height=1024", "--num-inference-steps=40", @@ -336,14 +470,13 @@ MODELS = { "--ulysses-degree=1", ], }, - # 17. Skill-only extra preset + # Skill-only extra preset # Requires: /inputs/diffusion_benchmark/figs/cat.png "firered-edit-1.0": { "path": "FireRedTeam/FireRed-Image-Edit-1.0", "prompt": "Make the cat wear a red hat", "image_path": str(ASSET_DIR / "cat.png"), "extra_args": [ - "--backend=sglang", "--width=1024", "--height=1024", "--num-inference-steps=40", @@ -357,14 +490,13 @@ MODELS = { "--ulysses-degree=1", ], }, - # 18. Skill-only extra preset + # Skill-only extra preset # Requires: /inputs/diffusion_benchmark/figs/cat.png "firered-edit-1.1": { "path": "FireRedTeam/FireRed-Image-Edit-1.1", "prompt": "Make the cat wear a red hat", "image_path": str(ASSET_DIR / "cat.png"), "extra_args": [ - "--backend=sglang", "--width=1024", "--height=1024", "--num-inference-steps=40", @@ -378,7 +510,7 @@ MODELS = { "--ulysses-degree=1", ], }, - # 19. Skill-only extra preset + # Skill-only extra preset # Requires: /inputs/diffusion_benchmark/figs/cat.png "hunyuan3d-shape": { "path": "tencent/Hunyuan3D-2", @@ -388,7 +520,6 @@ MODELS = { "paint_enable": False, }, "extra_args": [ - "--backend=sglang", "--num-inference-steps=50", "--guidance-scale=5.0", "--dit-layerwise-offload", @@ -544,17 +675,17 @@ def print_model_catalog(): """Print preset order, model path, and whether each preset maps to nightly.""" print() print("=" * CATALOG_TABLE_WIDTH) - print("MODEL PRESETS — Nightly-aligned first, skill-only extras after") + print("MODEL PRESETS — Nightly-aligned, then current-source and skill-only extras") print("=" * CATALOG_TABLE_WIDTH) - print(f"{'Preset':<24} {'Nightly':<28} {'Model Path':<46} {'GPUs':>4}") + print(f"{'Preset':<32} {'Nightly':<30} {'Model Path':<66} {'GPUs':>4}") print("-" * CATALOG_TABLE_WIDTH) for model_key, cfg in MODELS.items(): print( - f"{model_key:<24} {model_nightly_case_id(model_key):<28} {cfg['path']:<46} {required_gpus_for_model(model_key):>4}" + f"{model_key:<32} {model_nightly_case_id(model_key):<30} {cfg['path']:<66} {required_gpus_for_model(model_key):>4}" ) print("-" * CATALOG_TABLE_WIDTH) print( - "Nightly column shows the comparison_configs.json case id; '-' means skill-only." + "Nightly column shows the comparison_configs.json case id; '-' means no nightly mapping." ) @@ -575,6 +706,7 @@ def build_sglang_cmd( cmd = [ "sglang", "generate", + "--backend=sglang", f"--model-path={cfg['path']}", f"--prompt={cfg['prompt']}", ] @@ -631,6 +763,9 @@ def run_benchmark_once( env = os.environ.copy() env.setdefault("FLASHINFER_DISABLE_VERSION_CHECK", "1") + cfg = MODELS[model_key] + for key, value in cfg.get("env", {}).items(): + env.setdefault(key, str(value)) if env.get("HF_TOKEN") and not env.get("HUGGINGFACE_HUB_TOKEN"): env["HUGGINGFACE_HUB_TOKEN"] = env["HF_TOKEN"] diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-modelopt-quant/SKILL.md b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-modelopt-quant/SKILL.md index 12c10511d..1ac5a5d40 100644 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-modelopt-quant/SKILL.md +++ b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-modelopt-quant/SKILL.md @@ -54,6 +54,7 @@ This repo now contains: - diffusion-side NVFP4 loading from ModelOpt exports - FLUX.2 packed-QKV detection that distinguishes packed NVFP4 checkpoints from standard diffusers exports - automatic protection against incompatible FP8 CPU offload while keeping layerwise DiT offload available +- separate online diffusion quantization paths such as `--quantization fp8` / `mxfp4`; keep those out of this ModelOpt PTQ/export workflow unless the user explicitly asks for runtime quantization - FP8 transformer build: [`python/sglang/multimodal_gen/tools/build_modelopt_fp8_transformer.py`](../../../tools/build_modelopt_fp8_transformer.py) - NVFP4 mixed transformer build: @@ -67,6 +68,11 @@ Validated documentation and CI coverage currently center on these ModelOpt diffu - NVFP4: FLUX.1-dev, FLUX.2-dev, Wan2.2 Treat a new family, a new precision, or a new checkpoint layout as unsupported until it has a documented matrix row and a matching validation story. +Current B200 CI also contains an Ideogram4 NVFP4 native load case +(`ideogram4_nvfp4_t2i` via `Comfy-Org/Ideogram-4`). Treat that as source +evidence for an existing NVFP4 path, but do not expand the ModelOpt support +matrix to Ideogram4 unless `docs/diffusion/quantization.md` is updated with the +exact checkpoint, loader path, quality check, and benchmark scope. Before writing CLI examples, re-read the active branch's `docs/diffusion/quantization.md`: FLUX.2 NVFP4 is an official `black-forest-labs/*` repo rather than a `lmsys/*` converted repo, and its preferred flag depends on the current documented loader flow. Use `--transformer-path` for a component override directory with `config.json`; use `--transformer-weights-path` when the repo or path should be probed as raw weights. B200 CI coverage can include loose BF16-vs-quantized quality checks. Inspect the active branch's `run_suite.py` before assuming they are part of the suite; mainline and feature branches may differ. Those checks are intended to catch blank, corrupted, or obviously divergent images, not exact image parity. @@ -78,9 +84,9 @@ you are explicitly testing a historical branch. ## Related PR Watchlist -As of 2026-05-04, these related SGLang PRs are relevant to ModelOpt diffusion -support. Treat unmerged items as future support or migration work until the -docs/CI matrix is updated. +These related SGLang PRs are useful as ModelOpt diffusion support history. +Re-check the PR state and the active source tree before treating any item as +current behavior, and keep the docs/CI matrix as the support boundary. - #23155 added Qwen Image ModelOpt FP8 support. - #23199 adds HunyuanVideo ModelOpt FP8 support. diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-performance/SKILL.md b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-performance/SKILL.md index 1368d4ce5..914899b96 100644 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-performance/SKILL.md +++ b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-performance/SKILL.md @@ -252,21 +252,27 @@ Use these as first commands to benchmark, not as universal winners. | Model family | First performance shape | Starting flags | Notes | |---|---|---|---| | FLUX.1 / FLUX.2 image | 1024x1024, runtime-default steps/guidance, 1 GPU | `--enable-torch-compile --warmup --dit-layerwise-offload false` | `black-forest-labs/FLUX.*` repos are gated; for FP8/NVFP4 use validated `--transformer-path` or `--transformer-weights-path` flows from the quant skill. | +| FLUX.2 Klein / Klein Base | 1024x1024, runtime-default steps/guidance, 1 GPU | `--enable-torch-compile --warmup --dit-layerwise-offload false` | Current registry has `black-forest-labs/FLUX.2-klein-4B`, `FLUX.2-klein-9B`, and base variants. Klein is step-distilled; Klein Base is not. | | Qwen-Image / Qwen-Image-Edit | 1024x1024, runtime-default steps/guidance, 1 GPU | `--enable-torch-compile --warmup`; optionally native `SGLANG_CACHE_DIT_ENABLED=true` | Cache-DiT is lossy. For edit tasks, keep reference image, seed, and output size fixed. | -| Z-Image-Turbo | 1024x1024, runtime-default steps/guidance, 1 GPU | `--enable-torch-compile --warmup` | Mainline has Z-Image tanh/gate norm fusions; PR #21912 tracks FP8 plus CUDA Graph work. | +| Z-Image / Z-Image-Turbo | 1024x1024, runtime-default steps/guidance, 1 GPU | `--enable-torch-compile --warmup` | Keep base Z-Image separate from Turbo: base uses 50-step CFG defaults, Turbo uses 9-step zero-CFG defaults. Mainline has Z-Image tanh/gate norm fusions. | | Wan2.2 A14B T2V/I2V | 1280x720, 81 frames | Nightly: `--num-gpus 4 --enable-cfg-parallel --ulysses-degree 2 --text-encoder-cpu-offload --pin-cpu-memory` | For lowest latency, also benchmark pure Ulysses on the same GPUs. | | Wan2.2 TI2V 5B | 1280x720, 81 frames, 1 GPU | `--enable-torch-compile --warmup` | Keep the input image and motion prompt fixed when comparing sparse attention or Cache-DiT. | -| LTX-2 / LTX-2.3 | 768x512, 121 frames, runtime-default steps/guidance, 2 GPUs | `--pipeline-class-name LTX2TwoStagePipeline --enable-torch-compile --warmup`; LTX-2 uses `--enable-cfg-parallel`, LTX-2.3 TI2V uses `--cfg-parallel-size 2` | Use the benchmark/profile skill presets for exact nightly alignment. PRs #22441, #24025, and #23736 track additional LTX2 perf/parallel work. | +| Wan2.1 / FastWan / TurboWan variants | 480p or 720p video, family defaults | `--enable-torch-compile --warmup`; add `--ulysses-degree` / CFG parallel only after measuring | Current registry includes Wan2.1, FastWan2.1, FastWan2.2 TI2V, TurboWan2.1, TurboWan2.2 I2V, and Wan2.1-Fun InP. Use the compatibility matrix and benchmark presets before choosing topology. | +| Cosmos3 Nano / Super | T2I: 1024x1024 with `--num-frames 1`; T2V/I2V: 480p/720p video | `SGLANG_DISABLE_COSMOS3_GUARDRAILS=1` for benchmark isolation; `--enable-torch-compile --warmup` | One checkpoint serves T2I/T2V/I2V. Mode is request-driven: `num_frames == 1` means T2I, `--image-path` means I2V. | +| Ideogram 4 FP8/NVFP4 | 1024x1024, native preset defaults | `--enable-torch-compile --warmup` | Do not set `--num-inference-steps` or `--guidance-scale` directly unless you also update the Ideogram preset; sampling params derive them from `preset`. | +| ERNIE-Image / GLM-Image / SANA / SD3 | 1024-class image, family defaults | `--enable-torch-compile --warmup`; disable offload only after checking VRAM | Treat these as current native image families. Start with benchmark/profile presets for ERNIE, GLM, and SANA; use registry/config defaults for SD3 unless you add a new preset. | +| LTX-2 / LTX-2.3 | 768x512 or HQ 1920x1088, 121 frames | `--pipeline-class-name LTX2TwoStagePipeline --enable-torch-compile --warmup`; HQ uses `LTX2TwoStageHQPipeline` and `--ltx2-two-stage-device-mode snapshot` by default | Use benchmark/profile presets for nightly alignment, one-stage, high-resolution stress, and HQ. Device mode choices are `original`, `snapshot`, and `resident`; `resident` is fastest but uses more VRAM. | | HunyuanVideo | 848x480 or 720p class video | `--text-encoder-cpu-offload --pin-cpu-memory --enable-torch-compile --warmup` | Check VAE decode separately. GroupNorm+SiLU is default-eligible in mainline when wrapper guards pass; use `bench_group_norm_silu.py` when VAE residual blocks are hot. | | JoyAI-Image-Edit | 1024-class TI2I, 40 steps, guidance 4.0 | `--backend=sglang --num-gpus 2 --enable-cfg-parallel --ulysses-degree 1 --enable-torch-compile --warmup --dit-layerwise-offload false --dit-cpu-offload false` | Newly supported image-edit path. Keep the input image, prompt, seed, and output size fixed; 2-GPU CFG parallel is the validated H100 starting point. | | FireRed-Image-Edit 1.0 / 1.1 | 1024x1024 image edit, 40 steps, guidance 4.0 | `--backend=sglang --num-gpus 2 --enable-cfg-parallel --ulysses-degree 1 --enable-torch-compile --warmup --dit-layerwise-offload false --dit-cpu-offload false` | Uses the native `QwenImageEditPlusPipeline` path. 2-GPU CFG parallel is the validated H100 starting point; benchmark 1.0 and 1.1 separately because checkpoint differences can change denoise latency. | | Hunyuan3D-2 shape | Shape generation, 50 steps, guidance 5.0 | `--backend=sglang --enable-torch-compile --warmup --dit-layerwise-offload false --dit-cpu-offload false` | Focus on `Hunyuan3DShapeDenoisingStage`; keep mesh export/paint timings separate from denoise. | -| MOVA / Helios | Use the benchmark/profile presets first | `--enable-torch-compile --warmup`; pin offload flags explicitly | PR #20530 tracks MOVA fused RMSNorm+RoPE; PR #24059 tracks Helios fused norm modulation. | +| MOVA / Helios / LingBot World | Use the benchmark/profile presets or server test cases first | `--enable-torch-compile --warmup`; pin offload and topology flags explicitly | These video/realtime families have model-specific stages and condition handling. Keep prompt/image/action inputs fixed and prefer perf dumps over wall time alone. | -## Open PR Watchlist +## Historical PR Watchlist -As of 2026-05-02, these performance PRs were open. Treat them as direction and -prior art until merged: +Treat these performance PRs as direction and prior art only. Re-check the PR +state and the active source tree before relying on any path, flag, or claim +about whether the work has merged: - Fusion/kernel: #24025 LTX2 QK norm, #24059 Helios norm modulation, #24117 Z-Image packed QKV, #19488 Wan elementwise cross-block fusion, #19249 Z-Image gate/norm fusion, #20429 Qwen-Image layernorm/modulation, #20530 MOVA RMSNorm+RoPE. - VAE/decode: #22531 LTX2 parallel VAE, #20927 batched tiled VAE decode. @@ -279,6 +285,6 @@ prior art until merged: - **Offload tuning**: after the first request, the runtime logs peak GPU memory and which components could stay resident. Use this to decide which `--*-cpu-offload` flags to disable. - **Backend selection**: `--backend sglang` (default, auto-detected) enables native optimizations (fused kernels, SP, native Cache-DiT env knobs, etc.). `--backend diffusers` falls back to Diffusers pipelines and is the path that accepts `--cache-dit-config` plus diffusers attention backend names. - **Wan2.2-I2V sizing**: explicit `--width/--height` on `Wan2.2-I2V-A14B` control the target area while preserving the condition-image aspect ratio. -- **Mainline diffusion fast paths**: before proposing a new kernel or overlap scheme, check `sglang-diffusion-benchmark-profile/existing-fast-paths.md`. It covers GroupNorm+SiLU, Z-Image residual-form modulation, fused diffusion `QK norm + RoPE`, packed QKV/NVFP4 expectations, and existing multi-GPU overlap families such as Ulysses / USP and turbo-layer async all-to-all. +- **Mainline diffusion fast paths**: before proposing a new kernel or overlap scheme, check `sglang-diffusion-benchmark-profile/existing-fast-paths.md`. It covers GroupNorm+SiLU, Z-Image residual-form modulation, fused diffusion `QK norm + RoPE`, LTX2 split RoPE, varlen USP pack/scatter, packed QKV/NVFP4 expectations, and existing multi-GPU overlap families such as Ulysses / USP and turbo-layer async all-to-all. - **NVFP4 trace interpretation**: on FLUX.2 NVFP4 and Nunchaku-style checkpoints, packed QKV is expected. SGLang intentionally uses fused projection modules such as `to_qkv` / `to_added_qkv` instead of separate `to_q` / `to_k` / `to_v`, so a split-QKV trace usually means the quantized path did not engage rather than a brand new fusion opportunity. -- **Hotspot workflow split**: use `sglang-diffusion-benchmark-profile` to prove and classify a slowdown with perf dumps plus `torch.profiler`; hand concrete kernel work to `sglang-diffusion-ako4all-kernel` or another specialized optimization skill instead of expanding the benchmark skill. +- **Hotspot workflow split**: use `sglang-diffusion-benchmark-profile` to prove and classify a slowdown with perf dumps plus `torch.profiler`; hand concrete kernel work off with the perf/profile evidence attached instead of expanding the benchmark skill.