diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-add-model/SKILL.md b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-add-model/SKILL.md index 6a06878cf..3e7e0075f 100644 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-add-model/SKILL.md +++ b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-add-model/SKILL.md @@ -586,55 +586,8 @@ Before submitting, verify: ## After Implementation: Tests and Performance Data -### Component Accuracy When Adding a New Testcase Config - -If you add a new entry to `python/sglang/multimodal_gen/test/server/testcase_configs.py`, you must treat component accuracy as part of the model-adding workflow. Do not assume the new testcase will automatically fit the existing component-accuracy harness. - -The component-accuracy harness compares SGLang components against Diffusers/HF reference components. This is stricter than pipeline-level inference. New testcase configs commonly fail here for one of three reasons: - -1. **The model family needs explicit hook wiring** in `python/sglang/multimodal_gen/test/server/accuracy_hooks.py`. - - Add hook logic only when the harness cannot call the raw component correctly without it. - - Valid examples: - - required forward arguments are missing from the synthetic input bundle - - a known runtime execution context must be matched for the component to run at all, such as transformer autocast - - the reference and SGLang expose the same component contract, but the harness needs family-specific input preparation to reach it - - Invalid examples: - - changing the compared output mode just to make shapes or values line up - - adding a harness-side behavior override that changes the component contract instead of matching it - -2. **The component is already covered by another testcase with the same source component and topology**. - - In that case, do not add redundant component-accuracy coverage. - - Add a skip entry in `python/sglang/multimodal_gen/test/server/accuracy_config.py` with a concrete reason such as: - - `Representative VAE accuracy is already covered by ... for the same source component and topology` - - This is the preferred path for variant-only cases such as LoRA, cache-dit, upscaling, or other testcases that reuse the same underlying component weights and topology. - -3. **The HF/Diffusers reference component cannot be loaded or compared faithfully in the harness**. - - Add a skip entry in `python/sglang/multimodal_gen/test/server/accuracy_config.py` with the exact technical failure. - - Good reasons include: - - missing or unsupported HF component layout - - incomplete or partially initialized HF checkpoint - - unsupported raw component contract for trustworthy comparison - - proven divergence after matched weight transfer and matching output shape - - Keep the skip reason concrete and technical. Do not write vague reasons like "component accuracy flaky" or "needs investigation." - -When adding a new testcase config, make this decision explicitly: -- if the model family needs minimal harness wiring, add the smallest possible change in `accuracy_hooks.py` -- if the testcase is only a variant of an already covered source component and topology, add a skip in `accuracy_config.py` -- if the HF/Diffusers reference component cannot be compared faithfully, add a skip in `accuracy_config.py` - -Do not add a new testcase config and wait for CI to discover missing component-accuracy wiring. Do not use `accuracy_hooks.py` to change the compared component contract just to make the test pass. - -Once the model is working and output quality is verified, **ask the user** whether they would like to: - -1. **Add tests** — Create unit tests and/or integration tests for the new model. Tests should cover: - - Pipeline construction and stage wiring - - Single-GPU inference producing non-noise output - - Multi-GPU inference (TP/SP) if supported - - See the `write-sglang-test` skill for test conventions and placement guidelines - -2. **Generate performance data** — Run benchmarks and collect perf metrics: - - Single-GPU latency and throughput (look for `Pixel data generated successfully in xxxx seconds` in console output; use the `warmup excluded` line for accurate timing) - - Multi-GPU scaling (TP/SP) throughput comparison - - Use `python/sglang/multimodal_gen/benchmarks/bench_serving.py` for serving benchmarks - -Do not skip this step — always ask the user before proceeding, as test and benchmark requirements vary per model. +After the model produces non-noise output, read +[references/testing-and-accuracy.md](references/testing-and-accuracy.md) before +adding GPU cases, component-accuracy skips/hooks, suite entries, or benchmark +claims. That reference tracks the current `gpu_cases.py` / `testcase_configs.py` +/ `run_suite.py` split and the component-accuracy decision rules. diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-add-model/references/testing-and-accuracy.md b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-add-model/references/testing-and-accuracy.md new file mode 100644 index 000000000..846d3b82e --- /dev/null +++ b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-add-model/references/testing-and-accuracy.md @@ -0,0 +1,74 @@ +# Testing And Accuracy + +Use this reference after a new diffusion model or pipeline variant can already +produce a non-noise image or video. + +## Test Placement + +- Add concrete GPU integration cases in `python/sglang/multimodal_gen/test/server/gpu_cases.py`. +- Keep reusable dataclasses, constants, thresholds, and testcase factory helpers in `python/sglang/multimodal_gen/test/server/testcase_configs.py`. +- Let `python/sglang/multimodal_gen/test/run_suite.py` own suite selection, runtime-based partitioning, and standalone test files. Do not hard-code CI shard lists elsewhere. +- If a new standalone test file is added to a suite, update `STANDALONE_FILE_EST_TIMES` after the first measured CI/runtime value is known. + +Useful local entrypoints from repo root: + +```bash +PYTHONPATH=python python3 python/sglang/multimodal_gen/test/run_suite.py --suite unit +PYTHONPATH=python python3 python/sglang/multimodal_gen/test/run_suite.py --suite component-accuracy-1-gpu -k +PYTHONPATH=python python3 python/sglang/multimodal_gen/test/run_suite.py --suite 1-gpu --total-partitions 1 --partition-id 0 -k +``` + +## Component Accuracy When Adding A GPU Case + +If you add a new entry to `ONE_GPU_CASES`, `TWO_GPU_CASES`, or a B200-specific +case group in `gpu_cases.py`, treat component accuracy as part of the +model-adding workflow. Do not assume the new testcase will automatically fit the +existing component-accuracy harness. + +The component-accuracy harness compares SGLang components against Diffusers/HF +reference components. This is stricter than pipeline-level inference. New GPU +cases commonly fail here for one of three reasons: + +1. The model family needs explicit hook wiring in `python/sglang/multimodal_gen/test/server/accuracy_hooks.py`. + - Add hook logic only when the harness cannot call the raw component correctly without it. + - Valid reasons include missing required forward arguments, required autocast/runtime context, or family-specific input preparation for the same component contract. + - Do not change the compared output mode or add harness-side behavior that changes the component contract just to make the test pass. + +2. The component is already covered by another testcase with the same source component and topology. + - Do not add redundant component-accuracy coverage. + - Add a skip entry in `python/sglang/multimodal_gen/test/server/accuracy_config.py` with a concrete reason such as `Representative VAE accuracy is already covered by ... for the same source component and topology`. + - This is the preferred path for variant-only cases such as LoRA, Cache-DiT, upscaling, or other cases that reuse the same underlying component weights and topology. + +3. The HF/Diffusers reference component cannot be loaded or compared faithfully in the harness. + - Add a skip entry in `accuracy_config.py` with the exact technical failure. + - Good reasons include missing/unsupported HF component layout, incomplete checkpoints, unsupported raw component contract, or proven divergence after matched weight transfer and matching output shape. + - Keep the skip reason concrete and technical. Do not write vague reasons like "component accuracy flaky" or "needs investigation." + +When adding a new GPU case, make this decision explicitly: + +- if the family needs minimal harness wiring, add the smallest possible change in `accuracy_hooks.py` +- if the case is only a variant of an already covered source component and topology, add a skip in `accuracy_config.py` +- if the HF/Diffusers reference component cannot be compared faithfully, add a skip in `accuracy_config.py` + +Do not add a new GPU case and wait for CI to discover missing component-accuracy +wiring. + +## Follow-up Scope + +Once the model is working and output quality is verified, cover the follow-up +scope the user requested. If the user did not specify test or benchmark depth, +propose the smallest useful validation set before launching long GPU runs. + +Tests should cover: + +- pipeline construction and stage wiring +- single-GPU inference producing non-noise output +- multi-GPU inference if TP/SP is supported +- relevant unit tests for new math, parsing, scheduling, or loader behavior + +For performance data: + +- use the `warmup excluded` latency line for command-line generation +- keep prompt, seed, shape, step count, model path, backend, and GPU topology fixed +- use `sglang-diffusion-benchmark-profile` for denoise perf dumps and profiler traces +- use `python/sglang/multimodal_gen/benchmarks/bench_serving.py` for serving benchmarks diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-ako4all-kernel/SKILL.md b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-ako4all-kernel/SKILL.md index 94c3a884b..e35497b33 100644 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-ako4all-kernel/SKILL.md +++ b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-ako4all-kernel/SKILL.md @@ -31,7 +31,7 @@ First use [../sglang-diffusion-benchmark-profile/SKILL.md](../sglang-diffusion-b - measure the real denoise regression - collect the perf dump baseline - capture one representative `torch.profiler` trace -- rule out existing merged fast paths +- rule out existing mainline fast paths If a future specialized optimization skill matches the kernel family better than AKO4ALL, hand off there instead. The diagnosis contract stays the same. diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/SKILL.md b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/SKILL.md index 29b12dab0..28fd91a1e 100644 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/SKILL.md +++ b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/SKILL.md @@ -25,19 +25,34 @@ Before running any benchmark, profiler, or kernel-validation command: - export `FLASHINFER_DISABLE_VERSION_CHECK=1` - choose idle GPU(s) before starting perf work +## Native Backend Gate + +All diffusion benchmark and profiling results owned by this skill must come from the native SGLang diffusion backend. + +Treat any of the following as a hard stop condition: +- `Falling back to diffusers backend` +- `Using diffusers backend` +- `Loaded diffusers pipeline` + +If any benchmark, perf-dump, or `torch.profiler` command prints one of those signals: +- stop the workflow immediately +- do not keep the generated numbers or traces as SGLang benchmark evidence +- do not continue to hotspot classification or kernel work +- first fix model resolution, pipeline selection, overlay/materialization, or other backend-selection issues so the model runs on the native SGLang diffusion path + ## Main Reference - [benchmark-and-profile.md](benchmark-and-profile.md) — canonical denoise benchmark, perf dump, and `torch.profiler` workflow; uses the checked-in nightly-aligned presets, plus `LTX-2`, `LTX-2.3` one-stage, and `LTX-2.3` two-stage benchmark recipes - [existing-fast-paths.md](existing-fast-paths.md) — map bottlenecks to existing fused kernels, packed QKV paths, fused `QK norm + RoPE`, and distributed overlap patterns before proposing new code - [scripts/diffusion_skill_env.py](scripts/diffusion_skill_env.py) — preflight helper: repo root discovery via `sglang.__file__`, write-access probe, benchmark/profile output directories, idle GPU selection -- [scripts/bench_diffusion_denoise.py](scripts/bench_diffusion_denoise.py) — end-to-end denoise benchmark preset runner via `sglang generate`; use `--list-models` to inspect preset order, then save perf dumps by label and compare them with `compare_perf.py` +- [scripts/bench_diffusion_denoise.py](scripts/bench_diffusion_denoise.py) — end-to-end denoise benchmark preset runner via `sglang generate`; pins `--backend=sglang`, supports `--no-torch-compile`, and saves perf dumps by label for `compare_perf.py` ## Opportunity Discovery Rule Before calling a diffusion hotspot "new", first classify it with `existing-fast-paths.md`. Always rule out these existing families first: -- merged Z-Image residual-form modulation +- Z-Image residual-form modulation - fused diffusion `QK norm + RoPE` - NVFP4 / Nunchaku packed QKV - Nunchaku fused GELU MLP @@ -45,3 +60,15 @@ Always rule out these existing families first: - turbo-layer async all-to-all overlap - `torch.compile` compute / communication reorder - dual-stream diffusion execution + +If the user explicitly requires `torch.compile` to stay off, do not use the +default benchmark preset invocation unchanged. Either pass the checked-in +benchmark helper its no-compile switch or run the equivalent manual command +without `--enable-torch-compile`. + +For FLUX-family manual profiling runs with a quantized transformer override: +- use `sglang generate` directly +- pass the override as `--transformer-path ` +- prefer `--prompt-path ` when also fixing `--output-file-name` +- if the base model is already cached locally and the machine has unreliable HF access, use the local cached `--model-path` plus `HF_HUB_OFFLINE=1` +- remember that `--profile` changes latency substantially; use the non-profile perf dump for the real before/after benchmark claim diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/benchmark-and-profile.md b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/benchmark-and-profile.md index 4ea949fb7..9d0da0810 100644 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/benchmark-and-profile.md +++ b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/benchmark-and-profile.md @@ -30,6 +30,7 @@ guide. ```bash ENV_PY=python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/diffusion_skill_env.py +BENCH_PY=python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py ROOT=$(python3 "$ENV_PY" print-root) cd "$ROOT" python3 "$ENV_PY" check-write-access >/dev/null @@ -54,6 +55,26 @@ check "torch+CUDA" python3 -c "import torch; assert torch.cuda.is_available()" check "torch.profiler" python3 -c "import torch.profiler" ``` +## Native Backend Gate + +Every benchmark and profile result in this guide must come from the native SGLang diffusion backend. + +If the command log contains any of: +- `Falling back to diffusers backend` +- `Using diffusers backend` +- `Loaded diffusers pipeline` + +then stop immediately: +- do not record the perf dump or trace as valid benchmark evidence +- do not compare it against other runs +- do not continue to hotspot ranking or kernel optimization +- first fix backend selection so the model stays on the native SGLang diffusion path + +The checked-in benchmark helper pins `--backend=sglang` so native presets fail +fast instead of silently falling back through `--backend=auto`. Do the same for +manual native profiling commands unless you are intentionally collecting a +diffusers baseline. + Environment notes: - all commands below assume you are inside the configured diffusion container shell - export `HF_TOKEN` before any gated Hugging Face model run @@ -72,9 +93,7 @@ wget -O "${ASSET_DIR}/mova_single_person.jpg" \ ## Benchmark Presets -Treat -`python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py` -as the source of truth for preset order. +Treat `"$BENCH_PY"` as the source of truth for preset order. Nightly diffusion comparison is server/API based (`sglang serve` plus requests). This skill stays on `sglang generate` for local benchmarking and profiling, but @@ -85,26 +104,32 @@ count, and any explicitly overridden sampling or parallelism flags. List the current preset order: ```bash -PYTHONPATH=python python3 \ - python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py \ - --list-models +PYTHONPATH=python python3 "$BENCH_PY" --list-models ``` Run one preset and save a perf dump: ```bash -PYTHONPATH=python python3 \ - python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py \ +PYTHONPATH=python python3 "$BENCH_PY" \ --model ltx2 \ --label baseline \ --output-dir "${BENCH_DIR}" ``` +Keep `torch.compile` off when the task requires it: + +```bash +PYTHONPATH=python python3 "$BENCH_PY" \ + --model flux \ + --label baseline \ + --output-dir "${BENCH_DIR}" \ + --no-torch-compile +``` + Run the `LTX-2.3` one-stage skill preset: ```bash -PYTHONPATH=python python3 \ - python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py \ +PYTHONPATH=python python3 "$BENCH_PY" \ --model ltx23-one-stage \ --label baseline \ --output-dir "${BENCH_DIR}" @@ -113,8 +138,7 @@ PYTHONPATH=python python3 \ Run the `LTX-2.3` two-stage skill preset: ```bash -PYTHONPATH=python python3 \ - python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py \ +PYTHONPATH=python python3 "$BENCH_PY" \ --model ltx23-two-stage \ --label baseline \ --output-dir "${BENCH_DIR}" @@ -123,8 +147,7 @@ PYTHONPATH=python python3 \ Run the full preset sweep: ```bash -PYTHONPATH=python python3 \ - python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py \ +PYTHONPATH=python python3 "$BENCH_PY" \ --all \ --label prXXXX \ --output-dir "${BENCH_DIR}" @@ -169,7 +192,6 @@ sglang generate \ --save-output --enable-torch-compile --warmup ``` -After [PR #20707](https://github.com/sgl-project/sglang/pull/20707), `LTX2TwoStagePipeline` is a native path. The spatial upsampler and distilled LoRA are auto-resolved from the same model snapshot unless you override them. @@ -205,8 +227,8 @@ sglang generate \ --save-output --enable-torch-compile --warmup ``` -This matches the new `ltx23-two-stage` skill preset and is a good benchmark target for -the recently merged `LTX-2.3` two-stage path. +This matches the `ltx23-two-stage` skill preset and is a good benchmark target +for the native `LTX-2.3` two-stage path. ### Manual command example: Wan2.2-I2V-A14B 720P @@ -224,7 +246,6 @@ sglang generate \ --warmup --enable-torch-compile ``` -After [PR #21390](https://github.com/sgl-project/sglang/pull/21390), `Wan2.2-I2V-A14B` uses the 720p max-area config by default, and explicit `--width/--height` overrides control the target area while preserving the reference-image aspect ratio. @@ -251,20 +272,22 @@ Always keep: - peak GPU memory - exact command line, model shape, dtype, and GPU topology +Never keep a perf dump produced after a diffusers-backend fallback. + ## `torch.profiler` Workflow ### 1. Establish the baseline ```bash -PYTHONPATH=python python3 \ - python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py \ +PYTHONPATH=python python3 "$BENCH_PY" \ --model flux \ --label baseline \ --output-dir "${BENCH_DIR}" ``` Keep model shape, seed, and GPU topology fixed for every comparison. Save one -reference image or video before changing code. +reference image or video before changing code. If the active task requires +`torch.compile` off, add `--no-torch-compile` here too. ### 2. Capture a representative trace @@ -327,18 +350,18 @@ attention, norm, modulation, MLP, or communication boundaries and re-run. ### 4. Classify the hotspot with `existing-fast-paths.md` Do not jump from a hot kernel straight into new code. First classify it against -the known merged families. +the known mainline families. | What the trace shows | First interpretation | | --- | --- | | `fused_inplace_qknorm_rope` missing, but separate qk norm plus rope show up | Check whether the fused diffusion `QK norm + RoPE` path should have engaged | | `to_q -> to_k -> to_v` on NVFP4 or Nunchaku FLUX-family checkpoints | Treat as a packed-QKV fast-path miss or checkpoint-format mismatch | -| `fused_norm_tanh_mul_add*` missing on Z-Image | Treat as a missing merged modulation path, not a new fusion request | +| `fused_norm_tanh_mul_add*` missing on Z-Image | Treat as a missing mainline modulation path, not a new fusion request | | `all_to_all`, ring attention, or async A2A dominate | Classify against Ulysses, USP, or turbo-layer overlap first | | split `fc1 -> gelu -> quant -> fc2.lora_down` on Nunchaku FLUX | Treat as a missing fused GELU MLP path | | attention kernels dominate | Confirm backend, topology, and shape guards before proposing a new kernel | -If the hot path is already covered by a merged optimization family, fix the +If the hot path is already covered by a mainline optimization family, fix the enablement, shape guard, backend choice, or checkpoint mapping first. ### 5. Hand off only real kernel work diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/existing-fast-paths.md b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/existing-fast-paths.md index a0e0c4c60..2bbb0c88c 100644 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/existing-fast-paths.md +++ b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/existing-fast-paths.md @@ -26,6 +26,7 @@ configuration first before handing the problem to a specialized kernel-optimizat - Use cases: `x * (1 + scale) + shift` and `a * (k + b) + c` - Constraints: `x` must be CUDA and contiguous. `scale/shift` support 0D/1D/2D/3D/4D broadcast. 4D `[B, F, 1, C]` requires `L % F == 0`. - NPU fallback: `scale_shift.py` swaps to `npu_fallback` native path. +- Validation: `python/sglang/jit_kernel/tests/diffusion/test_qwen_image_modulation.py`. 2. Norm + Scale/Shift fusion (CuTe DSL) - Kernels: `fused_norm_scale_shift`, `fused_scale_residual_norm_scale_shift` @@ -44,13 +45,14 @@ configuration first before handing the problem to a specialized kernel-optimizat - `y, y2 = tanh(gate) * norm(x) + shift`, then `y2 = norm(y) * (1 + scale)` - Constraints: same CuTe DSL envelope as the norm+scale/shift family in practice: contiguous last dim, fp16/bf16/fp32, and `D % 256 == 0`, `D <= 8192`. - Validation: `python/sglang/jit_kernel/tests/diffusion/test_norm_tanh_mul_add_norm_scale.py` -- Behavior: this is already a merged fast path, so if Z-Image traces show the unfused chain, treat it as a missing or regressed existing optimization before proposing a new kernel. +- Behavior: this is already a mainline fast path, so if Z-Image traces show the unfused chain, treat it as a missing or regressed existing optimization before proposing a new kernel. 4. Triton LayerNorm/RMSNorm fusion - Kernels: `rms_norm_fn`, `layer_norm_fn`, `norm_infer` - Locations: `triton/norm.py`, `layernorm.py` - Use cases: fp32 RMSNorm with residual/dropout/rowscale/x1 branches, and inference-friendly `norm_infer`. - Constraints: last dim must be contiguous, and `N * element_size < 64KB`. +- Validation: `python/sglang/jit_kernel/tests/test_rmsnorm.py`. 5. Triton one-pass RMSNorm (small hidden size fast path) - Kernel: `triton_one_pass_rms_norm` @@ -64,6 +66,7 @@ configuration first before handing the problem to a specialized kernel-optimizat - Use case: GPT-J style RoPE when not Neox. - Constraints: `head_size` must be even. - NPU fallback: `npu_fallback.apply_rotary_embedding_native`. +- Validation: `python/sglang/jit_kernel/tests/test_rope.py`. **Faster CUDA Kernel Usage Points** @@ -93,6 +96,7 @@ configuration first before handing the problem to a specialized kernel-optimizat - `can_use_fused_inplace_qknorm(head_dim, dtype)` returns true. - Supported head dims: `64, 128, 256, 512, 1024`. - Behavior: Fused path operates on `q` and `k` in place after reshaping to `[B, -1, head_dim]`. If preconditions fail, fall back to per-tensor RMSNorm. +- Validation: `python/sglang/jit_kernel/tests/test_qknorm.py` and `python/sglang/jit_kernel/tests/test_qknorm_across_heads.py`. **QK Norm + RoPE Optimization** @@ -107,6 +111,7 @@ configuration first before handing the problem to a specialized kernel-optimizat - `can_use_fused_inplace_qknorm_rope(head_dim, rope_dim, is_neox, dtype)` returns true. - Supported head dims: `64, 128, 256`. - Behavior: `apply_qk_norm_rope` prefers the fused JIT kernel when all guards pass; otherwise it falls back to `apply_qk_norm(...)` plus `apply_flashinfer_rope_qk_inplace(...)`. +- Validation: `python/sglang/jit_kernel/tests/diffusion/test_qknorm_rope.py`. **Nunchaku Fused GELU MLP** diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py index 224462896..3ffce18e0 100755 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py +++ b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py @@ -53,6 +53,11 @@ from diffusion_skill_env import ( REPO_ROOT = get_repo_root() ASSET_DIR = ensure_dir(get_assets_dir(REPO_ROOT)) GATED_MODELS = {"flux", "flux2"} +DIFFUSERS_FALLBACK_SIGNALS = ( + "falling back to diffusers backend", + "using diffusers backend", + "loaded diffusers pipeline", +) # --------------------------------------------------------------------------- # Model configs — kept in exact sync with benchmark-and-profile.md @@ -326,6 +331,7 @@ def build_sglang_cmd( "generate", f"--model-path={cfg['path']}", f"--prompt={cfg['prompt']}", + "--backend=sglang", "--log-level=info", ] @@ -358,6 +364,7 @@ def run_benchmark_once( label: str, output_dir: Path, warmup: bool = True, + torch_compile: bool = True, ) -> dict: """Run a single benchmark pass and return results dict.""" perf_path = output_dir / f"{model_key}_{label}.json" @@ -366,6 +373,7 @@ def run_benchmark_once( model_key, perf_dump_path=str(perf_path), warmup=warmup, + torch_compile=torch_compile, ) env = os.environ.copy() @@ -397,11 +405,32 @@ def run_benchmark_once( print() t0 = time.time() - result = subprocess.run(cmd, env=env, text=True) + process = subprocess.Popen( + cmd, + env=env, + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.STDOUT, + bufsize=1, + ) + fallback_detected = False + assert process.stdout is not None + for line in process.stdout: + print(line, end="") + if any(signal in line.lower() for signal in DIFFUSERS_FALLBACK_SIGNALS): + fallback_detected = True + returncode = process.wait() elapsed = time.time() - t0 - if result.returncode != 0: - print(f" ERROR: exit code {result.returncode}") + if fallback_detected: + print( + " ERROR: model fell back to the diffusers backend. " + "Fix native SGLang diffusion backend selection before collecting perf data." + ) + return {"model": model_key, "label": label, "error": True, "elapsed_s": elapsed} + + if returncode != 0: + print(f" ERROR: exit code {returncode}") return {"model": model_key, "label": label, "error": True, "elapsed_s": elapsed} metrics = {"model": model_key, "label": label, "elapsed_s": elapsed, "error": False} @@ -528,6 +557,11 @@ def main(): help="Directory for perf dump JSON files", ) parser.add_argument("--no-warmup", action="store_true", help="Skip warmup") + parser.add_argument( + "--no-torch-compile", + action="store_true", + help="Keep torch.compile disabled for eager-mode comparisons.", + ) args = parser.parse_args() @@ -538,12 +572,21 @@ def main(): output_dir = Path(args.output_dir) output_dir.mkdir(parents=True, exist_ok=True) warmup = not args.no_warmup + torch_compile = not args.no_torch_compile models_to_run = list(MODELS.keys()) if args.all else [args.model or "flux"] results = [] for model_key in models_to_run: - results.append(run_benchmark_once(model_key, args.label, output_dir, warmup)) + results.append( + run_benchmark_once( + model_key, + args.label, + output_dir, + warmup=warmup, + torch_compile=torch_compile, + ) + ) if results: print_results_table(results) diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-modelopt-quant/SKILL.md b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-modelopt-quant/SKILL.md index 4e81ef4eb..227bea369 100644 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-modelopt-quant/SKILL.md +++ b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-modelopt-quant/SKILL.md @@ -23,6 +23,7 @@ This skill owns the ModelOpt-to-SGLang bridge. It is not a generic kernel-tuning - Benchmark only when BF16 and quantized commands are identical except for the checkpoint override being tested. - For diffusion FP8, keep `dit_cpu_offload=false`. `dit_layerwise_offload=true` is valid on the fixed path when you want lower DiT residency. - For multi-transformer pipelines, use per-component overrides when different components need different checkpoints. +- For B200 NVFP4 validation, keep backend-sensitive environment variables explicit. Wan2.2 NVFP4 is commonly validated with `SGLANG_DIFFUSION_FLASHINFER_FP4_GEMM_BACKEND=cudnn`; benchmark the default CUTLASS path separately if that is what you are evaluating. - When a branch is missing the validated helper tools, refresh `python/sglang/multimodal_gen/tools/build_modelopt_fp8_transformer.py`, `python/sglang/multimodal_gen/tools/build_modelopt_nvfp4_transformer.py`, and `python/sglang/multimodal_gen/tools/compare_diffusion_trajectory_similarity.py` instead of inventing one-off scripts elsewhere. - After validating a new ModelOpt quant path, update the ModelOpt support matrix in `docs/diffusion/quantization.md` before closing the task. @@ -55,9 +56,21 @@ This repo now contains: - automatic protection against incompatible FP8 CPU offload while keeping layerwise DiT offload available - FP8 transformer build: [`python/sglang/multimodal_gen/tools/build_modelopt_fp8_transformer.py`](../../../tools/build_modelopt_fp8_transformer.py) +- NVFP4 mixed transformer build: + [`python/sglang/multimodal_gen/tools/build_modelopt_nvfp4_transformer.py`](../../../tools/build_modelopt_nvfp4_transformer.py) - trajectory similarity validation: [`python/sglang/multimodal_gen/tools/compare_diffusion_trajectory_similarity.py`](../../../tools/compare_diffusion_trajectory_similarity.py) +Validated documentation and CI coverage currently center on six ModelOpt diffusion transformer override families: + +- FP8: FLUX.1-dev, FLUX.2-dev, Wan2.2 +- NVFP4: FLUX.1-dev, FLUX.2-dev, Wan2.2 + +Treat a new family, a new precision, or a new checkpoint layout as unsupported until it has a documented matrix row and a matching validation story. +Before writing CLI examples, re-read the active branch's `docs/diffusion/quantization.md`: FLUX.2 NVFP4 is an official `black-forest-labs/*` repo rather than a `BBuf/*` converted repo, and its preferred flag depends on the current documented loader flow. Use `--transformer-path` for a component override directory with `config.json`; use `--transformer-weights-path` when the repo or path should be probed as raw weights. + +B200 CI coverage can include loose BF16-vs-quantized quality smoke checks. Inspect the active branch's `run_suite.py` before assuming they are part of the suite; mainline and feature branches may differ. Those checks are intended to catch blank, corrupted, or obviously divergent images, not exact image parity. + ## Documentation Maintenance - Keep the validated ModelOpt support matrix in `docs/diffusion/quantization.md`. @@ -79,6 +92,7 @@ NVFP4: - the official diffusers export often already contains packed FP4 weights, scale tensors, and enough safetensors metadata for SGLang to rebuild the quant config - in that case SGLang mainly needs to detect the checkpoint family and rearrange tensors into the runtime layout - this is why NVFP4 often does not need an extra offline conversion pass like FP8 does +- backend choice matters on B200; record whether the run used the default CUTLASS path or a cuDNN-backed FlashInfer FP4 GEMM path Important caveat: @@ -310,4 +324,8 @@ When documenting results: | `runtime/loader/transformer_load_utils.py` | guards incompatible FP8 offload modes | | `runtime/models/dits/flux_2.py` | packed-QKV handling for the packed FLUX.2 NVFP4 family | | `tools/build_modelopt_fp8_transformer.py` | Build an SGLang-loadable FP8 transformer from a ModelOpt export | +| `tools/build_modelopt_nvfp4_transformer.py` | Build mixed BF16+NVFP4 transformer directories when a family needs preserved BF16 layers | | `tools/compare_diffusion_trajectory_similarity.py` | reduced deterministic BF16-vs-quantized validation | +| `docs/diffusion/quantization.md` | public ModelOpt support matrix and CLI examples | +| `test/server/testcase_configs.py` | reusable ModelOpt testcase constants, thresholds, and helpers | +| `test/server/gpu_cases.py` | concrete GPU and B200 ModelOpt CI case lists | diff --git a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-performance/SKILL.md b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-performance/SKILL.md index eadc83899..a66232ada 100644 --- a/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-performance/SKILL.md +++ b/python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-performance/SKILL.md @@ -13,6 +13,15 @@ Before running any `sglang generate` command below inside the diffusion containe - export `FLASHINFER_DISABLE_VERSION_CHECK=1` - `cd` to the repo root resolved from `sglang.__file__` +## Native Backend Gate + +Performance numbers are useful only when the intended backend actually ran. + +- Treat any log containing `Falling back to diffusers backend`, `Using diffusers backend`, or `Loaded diffusers pipeline` as invalid for native SGLang performance tuning. +- Use `--backend diffusers` only for an explicit diffusers baseline. For native recipes, leave the default backend or pin `--backend sglang`. +- If a fallback happened, fix pipeline registration/model-path/config issues first, then rerun. Do not compare perf dumps collected from a fallback run. +- When the runtime auto-selects parallel settings because the user omitted them, keep the result as an auto-tuned baseline. For reproducible tuning, pin `--num-gpus`, `--ulysses-degree`, `--ring-degree`, and `--enable-cfg-parallel` explicitly. + Reference: [SGLang-Diffusion Advanced Optimizations Blog](https://lmsys.org/blog/2026-02-16-sglang-diffusion-advanced-optimizations/) --- @@ -28,8 +37,8 @@ These options are intended to preserve output quality. In practice, some paths ( | **Warmup Resolutions** | `--warmup-resolutions 256x256 720x720` | Pre-compiles and warms up specific resolutions at server startup (instead of lazily on first request). | Faster first request per resolution | Each resolution adds to startup time. Serving mode only; useful when you know your target resolutions in advance. | | **Multi-GPU (SP)** | `--num-gpus N --ulysses-degree N` | Sequence parallelism across GPUs. Shards sequence tokens (not frames) to minimize padding. | Near-linear scaling with N GPUs | Requires NCCL; inter-GPU bandwidth matters. `ulysses_degree * ring_degree = sp_degree`. For Wan2.2 video, start by benchmarking pure Ulysses before assuming a mixed Ulysses/Ring layout is fastest. | | **CFG Parallel** | `--enable-cfg-parallel` | Runs conditional and unconditional CFG branches in parallel across GPUs. For CFG models on multi-GPU, benchmark this against pure Ulysses on your topology instead of assuming one always wins. | Often faster than pure SP for CFG models | Requires `num_gpus >= 2`. Halves the Ulysses group size (e.g. 8 GPU → two 4-GPU groups). Only for models that use CFG. Nightly coverage configs may intentionally use smaller Ulysses groups to keep ring behavior exercised; that does not automatically make them the lowest-latency choice. | -| **Layerwise Offload** | `--dit-layerwise-offload` | Async layer-by-layer H2D prefetch with compute overlap. Only ~2 DiT layers reside on GPU at a time, dramatically reducing VRAM. For some video models the copy stream can be almost fully hidden behind compute ([PR #15511](https://github.com/sgl-project/sglang/pull/15511)). | Saves VRAM (40 GB → ~11 GB for Wan A14B); can be near-zero speed cost on the right workload | Enabled by default for Wan/MOVA video models. Incompatible with Cache-DiT. For **image models** or highly parallelized setups (many GPUs, small per-GPU compute), the copy stream may not be fully hidden and can cause slowdown. | -| **Offload Prefetch Size** | `--dit-offload-prefetch-size F` | Fine-grained control over layerwise offload: how many layers to prefetch ahead. `0.0` = 1 layer (min VRAM), `0.1` = 10% of layers, `≥1` = absolute layer count. | Tune for cases where default offload has copy stream interference (e.g. image models). 0.05–0.1 is a good starting point. | Values ≥ 0.5 approach no-offload VRAM with worse performance. See [PR #17693](https://github.com/sgl-project/sglang/pull/17693) for benchmarks on image models. | +| **Layerwise Offload** | `--dit-layerwise-offload` | Async layer-by-layer H2D prefetch with compute overlap. Only ~2 DiT layers reside on GPU at a time, dramatically reducing VRAM. For some video models the copy stream can be almost fully hidden behind compute. | Saves VRAM (40 GB → ~11 GB for Wan A14B); can be near-zero speed cost on the right workload | Enabled by default for Wan/MOVA video models. Incompatible with Cache-DiT. For **image models** or highly parallelized setups (many GPUs, small per-GPU compute), the copy stream may not be fully hidden and can cause slowdown. | +| **Offload Prefetch Size** | `--dit-offload-prefetch-size F` | Fine-grained control over layerwise offload: how many layers to prefetch ahead. `0.0` = 1 layer (min VRAM), `0.1` = 10% of layers, `≥1` = absolute layer count. | Tune for cases where default offload has copy stream interference (e.g. image models). 0.05–0.1 is a good starting point. | Values ≥ 0.5 approach no-offload VRAM with worse performance. Use lower values when copy overlap is weak; disable offload when memory allows and latency dominates. | | **FSDP Inference** | `--use-fsdp-inference` | Uses PyTorch FSDP to shard model weights across GPUs with prefetch. Low latency, low VRAM. | Reduces per-GPU VRAM | Mutually exclusive with `--dit-layerwise-offload`. More overhead than SP on high-bandwidth interconnects. | | **CPU Offload (components)** | `--text-encoder-cpu-offload`, `--image-encoder-cpu-offload`, `--vae-cpu-offload`, `--dit-cpu-offload` | Offloads specific pipeline components to CPU when not in use. | Reduces peak VRAM | Adds H2D transfer latency when the component is needed. Auto-enabled for low-VRAM GPUs (<30 GB). **Tip:** after the first request completes, the console prints a peak VRAM analysis with suggestions on which offload flags can be safely disabled — look for the `"Components that could stay resident"` log line. | | **Pin CPU Memory** | `--pin-cpu-memory` | Uses pinned (page-locked) memory for CPU offload transfers. | Faster H2D transfers | Slightly higher host memory usage. Enabled by default; disable only as workaround for CUDA errors. | @@ -47,7 +56,7 @@ These options **trade output quality** for speed or VRAM savings. Results will d | **Approximate Attention** | `--attention-backend sage_attn` / `sage_attn_3` / `sliding_tile_attn` / `video_sparse_attn` / `sparse_video_gen_2_attn` / `vmoba_attn` / `sla_attn` / `sage_sla_attn` | Replaces exact attention with approximate or sparse variants. `sage_attn`: INT8/FP8 quantized Q·K; `sliding_tile_attn`: spatial-temporal tile skipping; others: model-specific sparse patterns. | ~1.5–2x on attention (varies by backend) | Quality degradation varies by backend and model. `sage_attn` is the most general; sparse backends (`sliding_tile_attn`, `video_sparse_attn`, etc.) are video-model-specific and may require config files (e.g. `--mask-strategy-file-path` for STA). Requires corresponding packages installed. | | **Cache-DiT** | `SGLANG_CACHE_DIT_ENABLED=true` + `--cache-dit-config ` | Caches intermediate residuals across denoising steps and skips redundant computations via a Selective Computation Mask (SCM). | ~1.5–2x on supported models | Quality depends on SCM config. Incompatible with `--dit-layerwise-offload`. Requires correct per-model config YAML. | | **Quantized Models (Nunchaku / SVDQuant)** | `--enable-svdquant --transformer-weights-path ` + optional `--quantization-precision int4\|nvfp4`, `--quantization-rank 32` | W4A4-style quantization via [Nunchaku](https://nunchaku.tech). Reduces DiT weight memory by ~4x. Precision/rank can be auto-inferred from weight filename or set explicitly. | ~1.5–2x compute speedup | Lossy quantization; quality depends on rank and precision. Requires pre-quantized weights. Ampere (SM8x) or SM12x only (no Hopper SM90). Higher rank = better quality but more memory. | -| **Pre-quantized Weights** | `--transformer-weights-path ` | Load any pre-quantized transformer weights (FP8, INT8, etc.) from a single `.safetensors` file, a directory, or a HuggingFace repo ID. | ~1.3–1.5x compute (dtype dependent) | Requires a validated quantized transformer override, such as one produced by `tools/build_modelopt_fp8_transformer.py` for ModelOpt FP8. Quality slightly worse than BF16; varies by quantization format. | +| **Pre-quantized Transformer Override** | `--transformer-path ` / `--transformer-weights-path ` | Load a quantized transformer component or raw transformer weights. For converted ModelOpt FP8/NVFP4 directories, prefer `--transformer-path`; use `--transformer-weights-path` for weight-only artifacts the model loader expects. | ~1.3–1.5x compute (dtype dependent) | Requires a validated quantized transformer override, such as one produced by the ModelOpt helper tools. Quality is usually slightly worse than BF16 and depends on the format, fallback layers, and calibration scope. | | **Component Precision Override** | `--dit-precision fp16`, `--vae-precision fp16\|bf16` | On-the-fly dtype conversion for individual components. E.g. convert a BF16 model to FP16 at load time, or run VAE in BF16 instead of FP32. | Reduces memory; FP16 can be faster on some GPUs | May affect numerical stability. VAE is FP32 by default for accuracy; lowering it is lossy. DiT defaults to BF16. | | **Fewer Inference Steps** | `--num-inference-steps N` (sampling param) | Reduces the number of denoising steps. Fewer steps = faster. | Linear speedup | Quality degrades with too few steps. Model-dependent optimal range. | @@ -86,7 +95,7 @@ sglang generate --model-path Lightricks/LTX-2 \ --enable-torch-compile --warmup --save-output ``` -Note: this generate recipe is aligned with the nightly comparison case `ltx2_twostage_t2v`. After [PR #20707](https://github.com/sgl-project/sglang/pull/20707), `LTX2TwoStagePipeline` is a native path and auto-resolves the spatial upsampler plus distilled LoRA from the same model snapshot unless you override them. +Note: this generate recipe is aligned with the nightly comparison case `ltx2_twostage_t2v`. `LTX2TwoStagePipeline` is a native path and auto-resolves the spatial upsampler plus distilled LoRA from the same model snapshot unless you override them. ### Native baseline, 2 GPUs: LTX-2.3 one-stage @@ -117,7 +126,7 @@ sglang generate --model-path Lightricks/LTX-2.3 \ --enable-torch-compile --warmup --save-output ``` -Note: this is the recommended benchmark command for the new `LTX-2.3` two-stage path. It uses the native `LTX2TwoStagePipeline` and matches the `ltx23-two-stage` benchmark preset in `sglang-diffusion-benchmark-profile`. +Note: this is the recommended benchmark command for the `LTX-2.3` two-stage path. It uses the native `LTX2TwoStagePipeline` and matches the `ltx23-two-stage` benchmark preset in `sglang-diffusion-benchmark-profile`. ### Maximum speed, image model, single GPU, lossless @@ -159,7 +168,7 @@ SGLANG_CACHE_DIT_ENABLED=true sglang generate --model-path \ - **Perf dump**: use `--perf-dump-path result.json` to save structured metrics, then compare with `python python/sglang/multimodal_gen/benchmarks/compare_perf.py baseline.json result.json`. - **Offload tuning**: after the first request, the runtime logs peak GPU memory and which components could stay resident. Use this to decide which `--*-cpu-offload` flags to disable. - **Backend selection**: `--backend sglang` (default, auto-detected) enables all native optimizations (fused kernels, SP, etc.). `--backend diffusers` falls back to vanilla Diffusers pipelines but supports `--cache-dit-config` and diffusers attention backends. -- **Wan2.2-I2V sizing**: after [PR #21390](https://github.com/sgl-project/sglang/pull/21390), explicit `--width/--height` on `Wan2.2-I2V-A14B` control the target area while preserving the condition-image aspect ratio. -- **Merged diffusion fast paths**: before proposing a new kernel or overlap scheme, check `sglang-diffusion-benchmark-profile/existing-fast-paths.md`. It now covers merged Z-Image residual-form modulation, fused diffusion `QK norm + RoPE`, and existing multi-GPU overlap families such as Ulysses / USP and turbo-layer async all-to-all. +- **Wan2.2-I2V sizing**: explicit `--width/--height` on `Wan2.2-I2V-A14B` control the target area while preserving the condition-image aspect ratio. +- **Mainline diffusion fast paths**: before proposing a new kernel or overlap scheme, check `sglang-diffusion-benchmark-profile/existing-fast-paths.md`. It covers Z-Image residual-form modulation, fused diffusion `QK norm + RoPE`, and existing multi-GPU overlap families such as Ulysses / USP and turbo-layer async all-to-all. - **NVFP4 trace interpretation**: on FLUX.2 NVFP4 and Nunchaku-style checkpoints, packed QKV is expected. SGLang intentionally uses fused projection modules such as `to_qkv` / `to_added_qkv` instead of separate `to_q` / `to_k` / `to_v`, so a split-QKV trace usually means the quantized path did not engage rather than a brand new fusion opportunity. - **Hotspot workflow split**: use `sglang-diffusion-benchmark-profile` to prove and classify a slowdown with perf dumps plus `torch.profiler`; hand concrete kernel work to `sglang-diffusion-ako4all-kernel` or another specialized optimization skill instead of expanding the benchmark skill.