[codex] Align diffusion skills with nightly Nvidia benchmarks (#25842)

This commit is contained in:
Xiaoyu Zhang
2026-05-20 12:18:05 +08:00
committed by GitHub
parent 579fed2090
commit af22390af7
4 changed files with 192 additions and 55 deletions
@@ -45,7 +45,7 @@ If any benchmark, perf-dump, or `torch.profiler` command prints one of those sig
- [benchmark-and-profile.md](benchmark-and-profile.md) — canonical denoise benchmark, perf dump, and `torch.profiler` workflow; uses checked-in nightly-aligned presets plus skill-only stress recipes such as `LTX-2.3` one-stage/two-stage, HunyuanVideo, MOVA, Helios, JoyAI/FireRed image edit, and Hunyuan3D shape - [benchmark-and-profile.md](benchmark-and-profile.md) — canonical denoise benchmark, perf dump, and `torch.profiler` workflow; uses checked-in nightly-aligned presets plus skill-only stress recipes such as `LTX-2.3` one-stage/two-stage, HunyuanVideo, MOVA, Helios, JoyAI/FireRed image edit, and Hunyuan3D shape
- [existing-fast-paths.md](existing-fast-paths.md) — map bottlenecks to existing fused kernels, packed QKV paths, fused `QK norm + RoPE`, distributed overlap patterns, and open optimization PRs before proposing new code - [existing-fast-paths.md](existing-fast-paths.md) — map bottlenecks to existing fused kernels, packed QKV paths, fused `QK norm + RoPE`, distributed overlap patterns, and open optimization PRs before proposing new code
- [scripts/diffusion_skill_env.py](scripts/diffusion_skill_env.py) — preflight helper: repo root discovery via `sglang.__file__`, write-access probe, benchmark/profile output directories, idle GPU selection - [scripts/diffusion_skill_env.py](scripts/diffusion_skill_env.py) — preflight helper: repo root discovery via `sglang.__file__`, write-access probe, benchmark/profile output directories, idle GPU selection
- [scripts/bench_diffusion_denoise.py](scripts/bench_diffusion_denoise.py) — end-to-end denoise benchmark preset runner via `sglang generate`; pins `--backend=sglang`, supports `--no-torch-compile`, and saves perf dumps by label for `compare_perf.py` - [scripts/bench_diffusion_denoise.py](scripts/bench_diffusion_denoise.py) — end-to-end denoise benchmark preset runner via `sglang generate`; supports `--no-torch-compile`, validates nightly preset drift with `--validate-nightly-alignment`, and saves perf dumps by label for `compare_perf.py`
## Opportunity Discovery Rule ## Opportunity Discovery Rule
@@ -99,8 +99,10 @@ Nightly diffusion comparison is server/API based (`sglang serve` plus requests).
This skill stays on `sglang generate` for local benchmarking and profiling, but This skill stays on `sglang generate` for local benchmarking and profiling, but
the nightly-aligned presets in `bench_diffusion_denoise.py` mirror the nightly-aligned presets in `bench_diffusion_denoise.py` mirror
`scripts/ci/utils/diffusion/comparison_configs.json` on model, task, prompt, `scripts/ci/utils/diffusion/comparison_configs.json` on model, task, prompt,
reference image, size, frames, seed, GPU count, serve args, and the request reference image, size, frames, seed, GPU count, and SGLang serve args. If
defaults used by `run_comparison.py` when a case omits steps or guidance. `comparison_configs.json` omits sampling params such as steps or guidance, the
nightly-aligned `sglang generate` preset omits them too and relies on the same
runtime defaults.
When in doubt, re-check that JSON before trusting this reference. When in doubt, re-check that JSON before trusting this reference.
List the current preset order: List the current preset order:
@@ -109,6 +111,12 @@ List the current preset order:
PYTHONPATH=python python3 "$BENCH_PY" --list-models PYTHONPATH=python python3 "$BENCH_PY" --list-models
``` ```
Check that the nightly presets still match the Nvidia nightly comparison config:
```bash
PYTHONPATH=python python3 "$BENCH_PY" --validate-nightly-alignment
```
Run one preset and save a perf dump: Run one preset and save a perf dump:
```bash ```bash
@@ -168,18 +176,19 @@ Nightly-aligned presets come first; skill-only presets stay available after them
| Preset | Model | Nightly | Notes | | Preset | Model | Nightly | Notes |
| --- | --- | --- | --- | | --- | --- | --- | --- |
| `flux` | `black-forest-labs/FLUX.1-dev` | Yes: `flux1_dev_t2i_1024` | Aligned to nightly prompt plus `--dit-layerwise-offload false` | | `flux` | `black-forest-labs/FLUX.1-dev` | Yes: `flux1_dev_t2i_1024` | Prompt, 1024x1024, seed 42, `--dit-layerwise-offload false`; no explicit steps/guidance override |
| `flux2` | `black-forest-labs/FLUX.2-dev` | Yes: `flux2_dev_t2i_1024` | Aligned to nightly prompt, 50 steps, guidance 4.0 | | `flux2` | `black-forest-labs/FLUX.2-dev` | Yes: `flux2_dev_t2i_1024` | Prompt, 1024x1024, seed 42, `--dit-layerwise-offload false`; no explicit steps/guidance override |
| `qwen` | `Qwen/Qwen-Image-2512` | Yes: `qwen_image_2512_t2i_1024` | Aligned to nightly prompt and steps | | `qwen` | `Qwen/Qwen-Image-2512` | Yes: `qwen_image_2512_t2i_1024` | Prompt, 1024x1024, seed 42; no explicit steps/guidance override |
| `qwen-edit` | `Qwen/Qwen-Image-Edit-2511` | Yes: `qwen_image_edit_2511` | Uses the nightly cat image and edit prompt | | `qwen-edit` | `Qwen/Qwen-Image-Edit-2511` | Yes: `qwen_image_edit_2511` | Uses the nightly cat image and edit prompt |
| `zimage` | `Tongyi-MAI/Z-Image-Turbo` | Yes: `zimage_turbo_t2i_1024` | Aligned to nightly prompt and guidance 4.0 | | `zimage` | `Tongyi-MAI/Z-Image-Turbo` | Yes: `zimage_turbo_t2i_1024` | Prompt, 1024x1024, seed 42; no explicit steps/guidance override |
| `wan-t2v` | `Wan-AI/Wan2.2-T2V-A14B-Diffusers` | Yes: `wan22_t2v_a14b_720p` | Aligned to nightly CFG-parallel 4-GPU launch | | `wan-t2v` | `Wan-AI/Wan2.2-T2V-A14B-Diffusers` | Yes: `wan22_t2v_a14b_720p` | 1280x720, 81 frames, 4 GPUs, CFG parallel, Ulysses degree 2, text encoder CPU offload and pinned CPU memory |
| `wan-ti2v` | `Wan-AI/Wan2.2-TI2V-5B-Diffusers` | Yes: `wan22_ti2v_5b_720p` | Uses the nightly cat image and motion prompt | | `wan-ti2v` | `Wan-AI/Wan2.2-TI2V-5B-Diffusers` | Yes: `wan22_ti2v_5b_720p` | Nightly cat image and motion prompt, 1280x720, 81 frames, seed 42 |
| `ltx2` | `Lightricks/LTX-2` | Yes: `ltx2_twostage_t2v` | Uses `LTX2TwoStagePipeline`; 2 GPUs, CFG parallel, 768x512, 121 frames, seed 42 | | `ltx2` | `Lightricks/LTX-2` | Yes: `ltx2_twostage_t2v` | `LTX2TwoStagePipeline`, 2 GPUs, CFG parallel, 768x512, 121 frames, seed 42 |
| `ltx23-ti2v-two-stage` | `Lightricks/LTX-2.3` | Yes: `ltx2.3_twostage_ti2v_2gpus` | Uses the nightly cat image, motion prompt, `LTX2TwoStagePipeline`, 2 GPUs, 768x512, 121 frames, seed 42 | | `ltx23-ti2v-two-stage` | `Lightricks/LTX-2.3` | Yes: `ltx2.3_twostage_ti2v_2gpus` | Nightly cat image, motion prompt, `LTX2TwoStagePipeline`, 2 GPUs, `--cfg-parallel-size 2`, 768x512, 121 frames, seed 42 |
| `wan-i2v` | `Wan-AI/Wan2.2-I2V-A14B-Diffusers` | Yes: `wan22_i2v_a14b_720p` | Aligned to nightly CFG-parallel 4-GPU launch | | `wan-i2v` | `Wan-AI/Wan2.2-I2V-A14B-Diffusers` | Yes: `wan22_i2v_a14b_720p` | Nightly cat image and motion prompt, 1280x720, 81 frames, 4 GPUs, CFG parallel, Ulysses degree 2, text encoder CPU offload and pinned CPU memory |
| `ltx23-one-stage` | `Lightricks/LTX-2.3` | No | Skill-only extra preset for the native `LTX-2.3` one-stage baseline; 2 GPUs, 768x512, 121 frames, fps 24, 30 steps, guidance 3.0, seed 1234 | | `ltx23-one-stage` | `Lightricks/LTX-2.3` | No | Skill-only extra preset for the native `LTX-2.3` one-stage baseline; 2 GPUs, 768x512, 121 frames, fps 24, 30 steps, guidance 3.0, seed 1234 |
| `ltx23-two-stage` | `Lightricks/LTX-2.3` | No | Skill-only high-resolution stress preset for the native `LTX-2.3` two-stage path; uses `LTX2TwoStagePipeline`, 2 GPUs, 1536x1024, 121 frames, fps 24, 30 steps, guidance 3.0, seed 1234 | | `ltx23-two-stage` | `Lightricks/LTX-2.3` | No | Skill-only high-resolution stress preset for the native `LTX-2.3` two-stage path; uses `LTX2TwoStagePipeline`, 2 GPUs, 1536x1024, 121 frames, fps 24, 30 steps, guidance 3.0, seed 1234 |
| `ltx23-two-stage-cfg-parallel` | `Lightricks/LTX-2.3` | No | Skill-only high-resolution CFG-parallel stress preset matching `ltx23-two-stage` plus `--cfg-parallel-size 2` |
| `hunyuanvideo` | `hunyuanvideo-community/HunyuanVideo` | No | Skill-only extra preset | | `hunyuanvideo` | `hunyuanvideo-community/HunyuanVideo` | No | Skill-only extra preset |
| `mova-720p` | `OpenMOSS-Team/MOVA-720p` | No | Skill-only extra preset | | `mova-720p` | `OpenMOSS-Team/MOVA-720p` | No | Skill-only extra preset |
| `helios` | `BestWishYsh/Helios-Base` | No | Skill-only extra preset | | `helios` | `BestWishYsh/Helios-Base` | No | Skill-only extra preset |
@@ -203,7 +212,6 @@ sglang generate \
--prompt="A cat and a dog baking a cake together in a kitchen." \ --prompt="A cat and a dog baking a cake together in a kitchen." \
--width=768 --height=512 \ --width=768 --height=512 \
--num-frames=121 \ --num-frames=121 \
--num-inference-steps=50 --guidance-scale=4.0 \
--seed=42 --num-gpus=2 --enable-cfg-parallel \ --seed=42 --num-gpus=2 --enable-cfg-parallel \
--save-output --enable-torch-compile --warmup --save-output --enable-torch-compile --warmup
``` ```
@@ -221,8 +229,7 @@ sglang generate \
--image-path="${ASSET_DIR}/cat.png" \ --image-path="${ASSET_DIR}/cat.png" \
--width=768 --height=512 \ --width=768 --height=512 \
--num-frames=121 \ --num-frames=121 \
--num-inference-steps=50 --guidance-scale=4.0 \ --seed=42 --num-gpus=2 --cfg-parallel-size=2 \
--seed=42 --num-gpus=2 \
--save-output --enable-torch-compile --warmup --save-output --enable-torch-compile --warmup
``` ```
@@ -331,8 +338,8 @@ sglang generate \
--model-path=Wan-AI/Wan2.2-I2V-A14B-Diffusers \ --model-path=Wan-AI/Wan2.2-I2V-A14B-Diffusers \
--prompt="The cat starts walking slowly towards the camera." \ --prompt="The cat starts walking slowly towards the camera." \
--image-path="${ASSET_DIR}/cat.png" \ --image-path="${ASSET_DIR}/cat.png" \
--720p --num-inference-steps=2 --num-frames=81 \ --width=1280 --height=720 --num-frames=81 \
--guidance-scale=5.0 --seed=42 --save-output \ --seed=42 --save-output \
--num-gpus=4 --enable-cfg-parallel --ulysses-degree=2 \ --num-gpus=4 --enable-cfg-parallel --ulysses-degree=2 \
--text-encoder-cpu-offload --pin-cpu-memory \ --text-encoder-cpu-offload --pin-cpu-memory \
--warmup --enable-torch-compile --warmup --enable-torch-compile
@@ -12,7 +12,7 @@ Usage:
# Tag the run for later compare_perf.py usage # Tag the run for later compare_perf.py usage
python3 python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py --model flux --label tuned python3 python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py --model flux --label tuned
# All 19 preset models # All 20 preset models
python3 python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py --all python3 python/sglang/multimodal_gen/.claude/skills/sglang-diffusion-benchmark-profile/scripts/bench_diffusion_denoise.py --all
# Show preset order, model path, and nightly mapping # Show preset order, model path, and nightly mapping
@@ -32,6 +32,7 @@ Input images required for image-guided models:
import argparse import argparse
import json import json
import os import os
import shlex
import subprocess import subprocess
import sys import sys
import time import time
@@ -52,6 +53,9 @@ from diffusion_skill_env import ( # noqa: E402
REPO_ROOT = get_repo_root() REPO_ROOT = get_repo_root()
ASSET_DIR = ensure_dir(get_assets_dir(REPO_ROOT)) ASSET_DIR = ensure_dir(get_assets_dir(REPO_ROOT))
NIGHTLY_CONFIG_PATH = (
REPO_ROOT / "scripts" / "ci" / "utils" / "diffusion" / "comparison_configs.json"
)
GATED_MODELS = {"flux", "flux2"} GATED_MODELS = {"flux", "flux2"}
DIFFUSERS_FALLBACK_SIGNALS = ( DIFFUSERS_FALLBACK_SIGNALS = (
"falling back to diffusers backend", "falling back to diffusers backend",
@@ -60,6 +64,18 @@ DIFFUSERS_FALLBACK_SIGNALS = (
) )
CATALOG_TABLE_WIDTH = 105 CATALOG_TABLE_WIDTH = 105
RESULTS_TABLE_WIDTH = 105 RESULTS_TABLE_WIDTH = 105
NIGHTLY_PRESET_ORDER = (
"flux",
"flux2",
"qwen",
"qwen-edit",
"zimage",
"wan-t2v",
"wan-ti2v",
"ltx2",
"ltx23-ti2v-two-stage",
"wan-i2v",
)
# --------------------------------------------------------------------------- # ---------------------------------------------------------------------------
# Model configs — kept in exact sync with benchmark-and-profile.md # Model configs — kept in exact sync with benchmark-and-profile.md
@@ -76,8 +92,6 @@ MODELS = {
"extra_args": [ "extra_args": [
"--width=1024", "--width=1024",
"--height=1024", "--height=1024",
"--num-inference-steps=50",
"--guidance-scale=4.0",
"--dit-layerwise-offload", "--dit-layerwise-offload",
"false", "false",
], ],
@@ -90,8 +104,6 @@ MODELS = {
"extra_args": [ "extra_args": [
"--width=1024", "--width=1024",
"--height=1024", "--height=1024",
"--num-inference-steps=50",
"--guidance-scale=4.0",
"--dit-layerwise-offload", "--dit-layerwise-offload",
"false", "false",
], ],
@@ -104,8 +116,6 @@ MODELS = {
"extra_args": [ "extra_args": [
"--width=1024", "--width=1024",
"--height=1024", "--height=1024",
"--num-inference-steps=50",
"--guidance-scale=4.0",
], ],
}, },
# 4. Nightly: qwen_image_edit_2511 # 4. Nightly: qwen_image_edit_2511
@@ -118,8 +128,6 @@ MODELS = {
"extra_args": [ "extra_args": [
"--width=1024", "--width=1024",
"--height=1024", "--height=1024",
"--num-inference-steps=50",
"--guidance-scale=4.0",
], ],
}, },
# 5. Nightly: zimage_turbo_t2i_1024 # 5. Nightly: zimage_turbo_t2i_1024
@@ -130,8 +138,6 @@ MODELS = {
"extra_args": [ "extra_args": [
"--width=1024", "--width=1024",
"--height=1024", "--height=1024",
"--num-inference-steps=9",
"--guidance-scale=4.0",
], ],
}, },
# 6. Nightly: wan22_t2v_a14b_720p # 6. Nightly: wan22_t2v_a14b_720p
@@ -140,10 +146,9 @@ MODELS = {
"path": "Wan-AI/Wan2.2-T2V-A14B-Diffusers", "path": "Wan-AI/Wan2.2-T2V-A14B-Diffusers",
"prompt": "A cat and a dog baking a cake together in a kitchen.", "prompt": "A cat and a dog baking a cake together in a kitchen.",
"extra_args": [ "extra_args": [
"--720p", "--width=1280",
"--num-inference-steps=2", "--height=720",
"--num-frames=81", "--num-frames=81",
"--guidance-scale=5.0",
"--num-gpus=4", "--num-gpus=4",
"--enable-cfg-parallel", "--enable-cfg-parallel",
"--ulysses-degree=2", "--ulysses-degree=2",
@@ -159,10 +164,9 @@ MODELS = {
"prompt": "The cat starts walking slowly towards the camera.", "prompt": "The cat starts walking slowly towards the camera.",
"image_path": str(ASSET_DIR / "cat.png"), "image_path": str(ASSET_DIR / "cat.png"),
"extra_args": [ "extra_args": [
"--720p", "--width=1280",
"--height=720",
"--num-frames=81", "--num-frames=81",
"--num-inference-steps=50",
"--guidance-scale=5.0",
], ],
}, },
# 8. Nightly: ltx2_twostage_t2v # 8. Nightly: ltx2_twostage_t2v
@@ -175,8 +179,6 @@ MODELS = {
"--width=768", "--width=768",
"--height=512", "--height=512",
"--num-frames=121", "--num-frames=121",
"--num-inference-steps=50",
"--guidance-scale=4.0",
"--num-gpus=2", "--num-gpus=2",
"--enable-cfg-parallel", "--enable-cfg-parallel",
], ],
@@ -193,9 +195,8 @@ MODELS = {
"--width=768", "--width=768",
"--height=512", "--height=512",
"--num-frames=121", "--num-frames=121",
"--num-inference-steps=50",
"--guidance-scale=4.0",
"--num-gpus=2", "--num-gpus=2",
"--cfg-parallel-size=2",
], ],
}, },
# 10. Nightly: wan22_i2v_a14b_720p # 10. Nightly: wan22_i2v_a14b_720p
@@ -206,10 +207,9 @@ MODELS = {
"prompt": "The cat starts walking slowly towards the camera.", "prompt": "The cat starts walking slowly towards the camera.",
"image_path": str(ASSET_DIR / "cat.png"), "image_path": str(ASSET_DIR / "cat.png"),
"extra_args": [ "extra_args": [
"--720p", "--width=1280",
"--num-inference-steps=2", "--height=720",
"--num-frames=81", "--num-frames=81",
"--guidance-scale=5.0",
"--num-gpus=4", "--num-gpus=4",
"--enable-cfg-parallel", "--enable-cfg-parallel",
"--ulysses-degree=2", "--ulysses-degree=2",
@@ -322,6 +322,7 @@ MODELS = {
"prompt": "Make the cat wear a red hat", "prompt": "Make the cat wear a red hat",
"image_path": str(ASSET_DIR / "cat.png"), "image_path": str(ASSET_DIR / "cat.png"),
"extra_args": [ "extra_args": [
"--backend=sglang",
"--width=1024", "--width=1024",
"--height=1024", "--height=1024",
"--num-inference-steps=40", "--num-inference-steps=40",
@@ -342,6 +343,7 @@ MODELS = {
"prompt": "Make the cat wear a red hat", "prompt": "Make the cat wear a red hat",
"image_path": str(ASSET_DIR / "cat.png"), "image_path": str(ASSET_DIR / "cat.png"),
"extra_args": [ "extra_args": [
"--backend=sglang",
"--width=1024", "--width=1024",
"--height=1024", "--height=1024",
"--num-inference-steps=40", "--num-inference-steps=40",
@@ -362,6 +364,7 @@ MODELS = {
"prompt": "Make the cat wear a red hat", "prompt": "Make the cat wear a red hat",
"image_path": str(ASSET_DIR / "cat.png"), "image_path": str(ASSET_DIR / "cat.png"),
"extra_args": [ "extra_args": [
"--backend=sglang",
"--width=1024", "--width=1024",
"--height=1024", "--height=1024",
"--num-inference-steps=40", "--num-inference-steps=40",
@@ -385,6 +388,7 @@ MODELS = {
"paint_enable": False, "paint_enable": False,
}, },
"extra_args": [ "extra_args": [
"--backend=sglang",
"--num-inference-steps=50", "--num-inference-steps=50",
"--guidance-scale=5.0", "--guidance-scale=5.0",
"--dit-layerwise-offload", "--dit-layerwise-offload",
@@ -406,6 +410,7 @@ def required_gpus_for_model(model_key: str) -> int:
"ltx23-ti2v-two-stage", "ltx23-ti2v-two-stage",
"ltx23-one-stage", "ltx23-one-stage",
"ltx23-two-stage", "ltx23-two-stage",
"ltx23-two-stage-cfg-parallel",
"joyai-edit", "joyai-edit",
"firered-edit-1.0", "firered-edit-1.0",
"firered-edit-1.1", "firered-edit-1.1",
@@ -418,6 +423,123 @@ def model_nightly_case_id(model_key: str) -> str:
return MODELS[model_key].get("nightly_case_id", "-") return MODELS[model_key].get("nightly_case_id", "-")
def _parse_cli_args(args: list[str]) -> dict[str, object]:
parsed: dict[str, object] = {}
i = 0
while i < len(args):
arg = args[i]
if not isinstance(arg, str) or not arg.startswith("--"):
i += 1
continue
if "=" in arg:
key, value = arg[2:].split("=", 1)
parsed[key] = value
elif i + 1 < len(args) and not str(args[i + 1]).startswith("--"):
parsed[arg[2:]] = str(args[i + 1])
i += 1
else:
parsed[arg[2:]] = True
i += 1
return parsed
def _normalize_cli_value(value: object) -> str:
if isinstance(value, bool):
return "true" if value else "false"
return str(value)
def _expected_nightly_cli_args(case: dict) -> dict[str, str]:
expected = {
"width": str(case["width"]),
"height": str(case["height"]),
}
for key, flag in (
("num_frames", "num-frames"),
("fps", "fps"),
("num_inference_steps", "num-inference-steps"),
("guidance_scale", "guidance-scale"),
):
if key in case:
expected[flag] = str(case[key])
if case.get("num_gpus", 1) > 1:
expected["num-gpus"] = str(case["num_gpus"])
serve_args = shlex.split(case["frameworks"]["sglang"].get("serve_args", ""))
parsed_serve_args = _parse_cli_args(serve_args)
for flag, value in parsed_serve_args.items():
if flag in {"enable-torch-compile", "warmup"}:
continue
expected[flag] = _normalize_cli_value(value)
return expected
def validate_nightly_alignment() -> int:
"""Validate nightly presets against diffusion comparison_configs.json."""
if not NIGHTLY_CONFIG_PATH.exists():
print(f"Missing nightly config: {NIGHTLY_CONFIG_PATH}")
return 1
with open(NIGHTLY_CONFIG_PATH) as f:
config = json.load(f)
cases = {case["id"]: case for case in config["cases"]}
errors: list[str] = []
preset_case_ids = [
MODELS[model_key].get("nightly_case_id") for model_key in NIGHTLY_PRESET_ORDER
]
if preset_case_ids != list(cases):
errors.append(
"Nightly preset order differs from comparison_configs.json: "
f"skill={preset_case_ids}, ci={list(cases)}"
)
for model_key in NIGHTLY_PRESET_ORDER:
preset = MODELS[model_key]
case_id = preset["nightly_case_id"]
case = cases.get(case_id)
if case is None:
errors.append(f"{model_key}: missing CI case {case_id}")
continue
if preset["path"] != case["model"]:
errors.append(f"{model_key}: model path differs")
if preset["prompt"] != case["prompt"]:
errors.append(f"{model_key}: prompt differs")
if bool(preset.get("image_path")) != bool(case.get("reference_image")):
errors.append(f"{model_key}: reference image presence differs")
if preset.get("seed", 42) != case.get("seed"):
errors.append(f"{model_key}: seed differs")
actual_args = {
key: _normalize_cli_value(value)
for key, value in _parse_cli_args(preset["extra_args"]).items()
}
expected_args = _expected_nightly_cli_args(case)
if actual_args != expected_args:
errors.append(
f"{model_key}: CLI args differ\n"
f" skill={actual_args}\n"
f" ci={expected_args}"
)
if errors:
print("Nightly alignment check failed:")
for error in errors:
print(f"- {error}")
return 1
print(
"Nightly alignment check passed: presets match "
"scripts/ci/utils/diffusion/comparison_configs.json."
)
return 0
def print_model_catalog(): def print_model_catalog():
"""Print preset order, model path, and whether each preset maps to nightly.""" """Print preset order, model path, and whether each preset maps to nightly."""
print() print()
@@ -455,8 +577,6 @@ def build_sglang_cmd(
"generate", "generate",
f"--model-path={cfg['path']}", f"--model-path={cfg['path']}",
f"--prompt={cfg['prompt']}", f"--prompt={cfg['prompt']}",
"--backend=sglang",
"--log-level=info",
] ]
effective_seed = cfg.get("seed", seed) effective_seed = cfg.get("seed", seed)
@@ -671,12 +791,19 @@ def main():
choices=list(MODELS.keys()), choices=list(MODELS.keys()),
help="Model to benchmark (default: flux)", help="Model to benchmark (default: flux)",
) )
parser.add_argument("--all", action="store_true", help="Benchmark all 19 models") parser.add_argument(
"--all", action="store_true", help=f"Benchmark all {len(MODELS)} models"
)
parser.add_argument( parser.add_argument(
"--list-models", "--list-models",
action="store_true", action="store_true",
help="List preset order, nightly mapping, and exit", help="List preset order, nightly mapping, and exit",
) )
parser.add_argument(
"--validate-nightly-alignment",
action="store_true",
help="Validate nightly presets against scripts/ci/utils/diffusion/comparison_configs.json and exit.",
)
parser.add_argument( parser.add_argument(
"--label", "--label",
type=str, type=str,
@@ -702,6 +829,9 @@ def main():
print_model_catalog() print_model_catalog()
return return
if args.validate_nightly_alignment:
raise SystemExit(validate_nightly_alignment())
output_dir = Path(args.output_dir) output_dir = Path(args.output_dir)
output_dir.mkdir(parents=True, exist_ok=True) output_dir.mkdir(parents=True, exist_ok=True)
warmup = not args.no_warmup warmup = not args.no_warmup
@@ -89,12 +89,12 @@ sglang generate --model-path Lightricks/LTX-2 \
--pipeline-class-name LTX2TwoStagePipeline \ --pipeline-class-name LTX2TwoStagePipeline \
--prompt "A cat and a dog baking a cake together in a kitchen." \ --prompt "A cat and a dog baking a cake together in a kitchen." \
--width 768 --height 512 \ --width 768 --height 512 \
--num-frames 121 --num-inference-steps 50 --guidance-scale 4.0 \ --num-frames 121 \
--seed 42 --num-gpus 2 --enable-cfg-parallel \ --seed 42 --num-gpus 2 --enable-cfg-parallel \
--enable-torch-compile --warmup --save-output --enable-torch-compile --warmup --save-output
``` ```
Note: this generate recipe is aligned with the nightly comparison case `ltx2_twostage_t2v`. `LTX2TwoStagePipeline` is a native path and auto-resolves the spatial upsampler plus distilled LoRA from the same model snapshot unless you override them. Note: this generate recipe is aligned with the nightly comparison case `ltx2_twostage_t2v`. The nightly config omits explicit steps and guidance, so this command omits them too and uses runtime defaults. `LTX2TwoStagePipeline` is a native path and auto-resolves the spatial upsampler plus distilled LoRA from the same model snapshot unless you override them.
### Nightly-aligned model, 2 GPUs: LTX-2.3 TI2V two-stage ### Nightly-aligned model, 2 GPUs: LTX-2.3 TI2V two-stage
@@ -104,12 +104,12 @@ sglang generate --model-path Lightricks/LTX-2.3 \
--prompt "The cat starts walking slowly towards the camera." \ --prompt "The cat starts walking slowly towards the camera." \
--image-path "${ASSET_DIR}/cat.png" \ --image-path "${ASSET_DIR}/cat.png" \
--width 768 --height 512 \ --width 768 --height 512 \
--num-frames 121 --num-inference-steps 50 --guidance-scale 4.0 \ --num-frames 121 \
--seed 42 --num-gpus 2 \ --seed 42 --num-gpus 2 --cfg-parallel-size 2 \
--enable-torch-compile --warmup --save-output --enable-torch-compile --warmup --save-output
``` ```
Note: this matches the nightly comparison case `ltx2.3_twostage_ti2v_2gpus`. Download `${ASSET_DIR}/cat.png` with the benchmark/profile skill before running it. Note: this matches the nightly comparison case `ltx2.3_twostage_ti2v_2gpus`. The nightly config omits explicit steps and guidance, so this command omits them too and uses runtime defaults. Download `${ASSET_DIR}/cat.png` with the benchmark/profile skill before running it.
### Native baseline, 2 GPUs: LTX-2.3 one-stage ### Native baseline, 2 GPUs: LTX-2.3 one-stage
@@ -251,12 +251,12 @@ Use these as first commands to benchmark, not as universal winners.
| Model family | First performance shape | Starting flags | Notes | | Model family | First performance shape | Starting flags | Notes |
|---|---|---|---| |---|---|---|---|
| FLUX.1 / FLUX.2 image | 1024x1024, 50 steps, 1 GPU | `--enable-torch-compile --warmup --dit-layerwise-offload false` | `black-forest-labs/FLUX.*` repos are gated; for FP8/NVFP4 use validated `--transformer-path` or `--transformer-weights-path` flows from the quant skill. | | FLUX.1 / FLUX.2 image | 1024x1024, runtime-default steps/guidance, 1 GPU | `--enable-torch-compile --warmup --dit-layerwise-offload false` | `black-forest-labs/FLUX.*` repos are gated; for FP8/NVFP4 use validated `--transformer-path` or `--transformer-weights-path` flows from the quant skill. |
| Qwen-Image / Qwen-Image-Edit | 1024x1024, 50 steps, 1 GPU | `--enable-torch-compile --warmup`; optionally native `SGLANG_CACHE_DIT_ENABLED=true` | Cache-DiT is lossy. For edit tasks, keep reference image, seed, and output size fixed. | | Qwen-Image / Qwen-Image-Edit | 1024x1024, runtime-default steps/guidance, 1 GPU | `--enable-torch-compile --warmup`; optionally native `SGLANG_CACHE_DIT_ENABLED=true` | Cache-DiT is lossy. For edit tasks, keep reference image, seed, and output size fixed. |
| Z-Image-Turbo | 1024x1024, 9 steps, guidance 4.0 | `--enable-torch-compile --warmup` | Mainline has Z-Image tanh/gate norm fusions; PR #21912 tracks FP8 plus CUDA Graph work. | | Z-Image-Turbo | 1024x1024, runtime-default steps/guidance, 1 GPU | `--enable-torch-compile --warmup` | Mainline has Z-Image tanh/gate norm fusions; PR #21912 tracks FP8 plus CUDA Graph work. |
| Wan2.2 A14B T2V/I2V | 720p, 81 frames | Nightly: `--num-gpus 4 --enable-cfg-parallel --ulysses-degree 2 --text-encoder-cpu-offload --pin-cpu-memory` | For lowest latency, also benchmark pure Ulysses on the same GPUs. | | Wan2.2 A14B T2V/I2V | 1280x720, 81 frames | Nightly: `--num-gpus 4 --enable-cfg-parallel --ulysses-degree 2 --text-encoder-cpu-offload --pin-cpu-memory` | For lowest latency, also benchmark pure Ulysses on the same GPUs. |
| Wan2.2 TI2V 5B | 720p, 81 frames, 1 GPU | `--enable-torch-compile --warmup` | Keep the input image and motion prompt fixed when comparing sparse attention or Cache-DiT. | | Wan2.2 TI2V 5B | 1280x720, 81 frames, 1 GPU | `--enable-torch-compile --warmup` | Keep the input image and motion prompt fixed when comparing sparse attention or Cache-DiT. |
| LTX-2 / LTX-2.3 | 768x512, 121 frames, 2 GPUs | `--pipeline-class-name LTX2TwoStagePipeline --enable-torch-compile --warmup`; LTX-2 nightly also uses `--enable-cfg-parallel` | Use the benchmark/profile skill presets for exact nightly alignment. PRs #22441, #24025, and #23736 track additional LTX2 perf/parallel work. | | LTX-2 / LTX-2.3 | 768x512, 121 frames, runtime-default steps/guidance, 2 GPUs | `--pipeline-class-name LTX2TwoStagePipeline --enable-torch-compile --warmup`; LTX-2 uses `--enable-cfg-parallel`, LTX-2.3 TI2V uses `--cfg-parallel-size 2` | Use the benchmark/profile skill presets for exact nightly alignment. PRs #22441, #24025, and #23736 track additional LTX2 perf/parallel work. |
| HunyuanVideo | 848x480 or 720p class video | `--text-encoder-cpu-offload --pin-cpu-memory --enable-torch-compile --warmup` | Check VAE decode separately. GroupNorm+SiLU is default-eligible in mainline when wrapper guards pass; use `bench_group_norm_silu.py` when VAE residual blocks are hot. | | HunyuanVideo | 848x480 or 720p class video | `--text-encoder-cpu-offload --pin-cpu-memory --enable-torch-compile --warmup` | Check VAE decode separately. GroupNorm+SiLU is default-eligible in mainline when wrapper guards pass; use `bench_group_norm_silu.py` when VAE residual blocks are hot. |
| JoyAI-Image-Edit | 1024-class TI2I, 40 steps, guidance 4.0 | `--backend=sglang --num-gpus 2 --enable-cfg-parallel --ulysses-degree 1 --enable-torch-compile --warmup --dit-layerwise-offload false --dit-cpu-offload false` | Newly supported image-edit path. Keep the input image, prompt, seed, and output size fixed; 2-GPU CFG parallel is the validated H100 starting point. | | JoyAI-Image-Edit | 1024-class TI2I, 40 steps, guidance 4.0 | `--backend=sglang --num-gpus 2 --enable-cfg-parallel --ulysses-degree 1 --enable-torch-compile --warmup --dit-layerwise-offload false --dit-cpu-offload false` | Newly supported image-edit path. Keep the input image, prompt, seed, and output size fixed; 2-GPU CFG parallel is the validated H100 starting point. |
| FireRed-Image-Edit 1.0 / 1.1 | 1024x1024 image edit, 40 steps, guidance 4.0 | `--backend=sglang --num-gpus 2 --enable-cfg-parallel --ulysses-degree 1 --enable-torch-compile --warmup --dit-layerwise-offload false --dit-cpu-offload false` | Uses the native `QwenImageEditPlusPipeline` path. 2-GPU CFG parallel is the validated H100 starting point; benchmark 1.0 and 1.1 separately because checkpoint differences can change denoise latency. | | FireRed-Image-Edit 1.0 / 1.1 | 1024x1024 image edit, 40 steps, guidance 4.0 | `--backend=sglang --num-gpus 2 --enable-cfg-parallel --ulysses-degree 1 --enable-torch-compile --warmup --dit-layerwise-offload false --dit-cpu-offload false` | Uses the native `QwenImageEditPlusPipeline` path. 2-GPU CFG parallel is the validated H100 starting point; benchmark 1.0 and 1.1 separately because checkpoint differences can change denoise latency. |