From 6880a4795533640f41ebb3db9e4ae0af5a371a1f Mon Sep 17 00:00:00 2001 From: Mick Date: Sun, 20 Sep 2026 20:39:26 +0800 Subject: [PATCH] [diffusion] docs: simplify Qwen-Image 2.1 cookbook (#40455) Co-authored-by: Mick Qian --- .../diffusion/Qwen-Image/Qwen-Image-2.1.mdx | 623 +++--------------- .../sglang-diffusion/compatibility_matrix.mdx | 2 +- .../sglang-diffusion/dynamic_batching.mdx | 2 +- docs/docs/sglang-diffusion/quantization.mdx | 6 +- .../snippets/configs/Qwen/qwen-image-2.1.jsx | 16 +- 5 files changed, 103 insertions(+), 546 deletions(-) diff --git a/docs/cookbook/diffusion/Qwen-Image/Qwen-Image-2.1.mdx b/docs/cookbook/diffusion/Qwen-Image/Qwen-Image-2.1.mdx index 18a324199..33d1be645 100644 --- a/docs/cookbook/diffusion/Qwen-Image/Qwen-Image-2.1.mdx +++ b/docs/cookbook/diffusion/Qwen-Image/Qwen-Image-2.1.mdx @@ -27,8 +27,8 @@ steps, and output count. Set reference PNG paths under **Variables**; edits upload files from the machine running cURL, so they need not exist on the server. Hardware selection applies the recommended placement for that GPU. H200, -B200, and RTX PRO 6000 96GB keep weights resident; RTX 5090 and RTX 4090 use -offload to fit the full pipeline. +B200, and RTX PRO 6000 96GB keep weights resident; RTX 5090 and RTX 4090 +offload selected components to fit the full pipeline. Untested topologies and feature combinations remain selectable and are labeled **Unverified**. Invalid topology combinations disable Copy. This integration currently uses the Python/source command; no published Docker image is verified. @@ -47,338 +47,75 @@ for i, item in enumerate(json.loads(Path("response.json").read_text())["data"]): PY ``` -### Current native recipes +### Recommended hardware settings -The picker defaults use native BF16/FP32 precision, exact attention, eager -execution, and full-image VAE decoding. The table below uses checkpoint -`840b4adb1e2c21c7d77967203188b55b678c535f` and source `1f89064b419`, -measured on 2026-09-20. These are the best configurations among the tested -candidates for this workload, not a claim of a global optimum. +The picker defaults to native BF16/FP32 precision, exact attention, eager +execution, and full-image VAE decoding. -| GPU | Single-output recipe | Generation median | Edit median | Request-phase peak | +| GPU | Placement / attention | Generation | Edit | Peak VRAM | | --- | --- | --- | --- | --- | | H200 141GB | Resident / FlashAttention | 4.48 s | 5.29 s | 38.4 GiB | | B200 192GB | Resident / FlashAttention | 2.46 s | 3.02 s | 38.5 GiB | | RTX PRO 6000 96GB | Resident / Torch SDPA | 8.03 s | 9.63 s | 38.4 GiB | -| RTX 4090 24GB | DiT offload, 8 resident layers, encoder CPU offload / FlashAttention | 23.25 s | 24.63 s | 23.1 GiB | +| RTX 4090 24GB | DiT and VAE resident, encoder layerwise offload / FlashAttention | 18.68 s | 21.68 s | 22.7 GiB | -Each shape uses two 1024px/40-step warmups, then five generation and five edit -requests; the RTX 4090 8-layer candidate uses three measured requests per mode. -All use seed 42, CFG 1, CPU noise, and one RGBA PNG per request. HTTP time -includes encoding and PNG serialization, excluding startup. Memory is the peak -sampled every 0.2 seconds during these requests, excluding startup. The H200, B200, and -RTX PRO 6000 batch-matrix servers use a ceiling of four images and a 20 ms batching -window; the single-output 4090 candidate has dynamic batching off. -PyTorch is 2.13.0+cu130, Transformers 5.12.1, and Diffusers 0.37.0; native -conditioning explicitly preserves the reference's Transformers 4.57.3 semantics. +Measured on 2026-09-20 at 1024×1024, 40 steps, CFG 1, and one RGBA PNG per +request. Times are median HTTP latency after warmup, including PNG serialization +and excluding startup; VRAM is the sampled request-phase peak. Prompts and +software versions affect both latency and memory use. -FlashAttention beats SDPA on H200 (4.48 vs 4.91 s generation; 5.29 vs 6.23 s -editing) and B200 (2.46 vs 2.91 s; 3.02 vs 4.03 s). SDPA uses three measured -requests after two warmups. On RTX 4090, retaining eight DiT layers reduces -generation from 24.50 to 23.25 s and editing from 25.19 to 24.63 s. Its generation, -editing, transparent generation, and transparent-input editing pixels match -the fully streamed recipe exactly in this comparison. The tradeoff is a larger -request-phase peak: 19.9 to 23.1 GiB. - -On RTX PRO 6000, resident weights are faster than DiT offload: 8.03 vs 10.24 s -for generation and 9.63 vs 10.59 s for editing. Offload lowers the request-phase -peak from 38.4 to 25.3 GiB; its timings use two warmups and three measurements. - -The picker applies the 8-layer flag only to a native, eager, single-output -RTX 4090 FlashAttention recipe with request batching off. Multiple outputs or -request batching select the fully streamed recipe instead; use the updated -**Server** command when switching. Multi-reference and other untested shapes -remain marked Unverified. - -The RTX 5090 recipe retains its earlier validation below; it has not been rerun -with this checkpoint. Both RTX 5090 and RTX PRO 6000 map FA selection to SDPA, so those -labels do not represent two different attention kernels. - -A 1024px H200 BCG server captured its warmup graph, but these prompts and image -prefixes missed that signature and ran eagerly. Text buckets do not pad this -model's condition KV layout. Keep eager execution as the default; declaring a -resolution alone does not establish graph replay or a speedup. +RTX 5090 uses DiT layerwise offload and Torch SDPA; its recipe has not been +retested with the updated checkpoint. Both RTX 5090 and RTX PRO 6000 use SDPA +when FlashAttention is selected in this runtime. CPU offload requires host RAM. ### Batching -Keep **Request batching → Off** for interactive use. For a concurrent -text-to-image workload on one RTX 4090, the tested two-image server command is: +Keep **Request batching → Off** and **Outputs → 1** for interactive use. +Batching increases individual request latency and does not guarantee higher +throughput. Measure your workload before enabling it. -```bash Command -sglang serve \ - --model-path /models/qwen-image-2.1 --model-id Qwen-Image-2.1 \ - --num-gpus 1 --ulysses-degree 1 --encoder-parallel auto \ - --performance-mode manual \ - --dit-layerwise-offload true --text-encoder-cpu-offload true \ - --attention-backend fa \ - --batching-max-size 2 --batching-delay-ms 20 \ - --host 0.0.0.0 --port 30010 -``` +Cross-request batching merges compatible text-to-image requests. Image edits +run separately; **Outputs** controls multiple images within one request. +On RTX 4090, selecting multiple outputs or request batching switches to DiT +layerwise offload for memory headroom. Restart with the updated **Server** command. -Send requests concurrently to use dynamic batching. The merge limit counts -output images, including each request's `n`; a client that waits for one -response before sending the next does not supply concurrent work. Image-edit -requests are not merged across requests. Set **Outputs** (the HTTP `n` field) -to produce multiple images within one generation or edit request. - -| GPU | Generation, 1 / 2 / 4 outputs | Editing, 1 / 2 / 4 outputs | Recommendation | -| --- | --- | --- | --- | -| H200 | 4.48 / 9.13 / 18.32 s | 5.29 / 10.60 / 21.13 s | Off; no throughput gain | -| B200 | 2.46 / 5.17 / 9.91 s | 3.02 / 5.98 / 12.04 s | Off; no material throughput gain | -| RTX PRO 6000 | 8.03 / 16.19 / 32.55 s | 9.63 / 19.08 / 38.11 s | Off; no material throughput gain | -| RTX 4090 | 23.25 / 38.33 / untested | 24.63 / 44.01 / untested | Off for latency; consider 2 for throughput | - -The H200/B200/RTX PRO 6000 and two-output 4090 measurements use five requests after two -warmups per shape. The one-output 4090 numbers use the faster 8-layer recipe -above with three measurements; batches stream all DiT layers. Against that -single-output recipe, two outputs improve 4090 generation throughput by about -**21%** and editing throughput by about **12%**, while increasing request latency. -Two concurrent 4090 generation requests complete in **39.75 s** wall time -(median of three rounds with different prompts and seeds). Concurrent 2/4-request -batches did not improve resident H200/B200/RTX PRO 6000 throughput. All measurements use -the checkpoint, software, 1024px/40-step settings, and HTTP timing scope above. -Transparent generation and transparent-input editing also passed for these -batch sizes, with alpha values from 0 to 255. - -Each sample retains its own condition-prefix KV cache, prompt, seed, and output -position. DiT target projections and MLPs run as a batch, and layerwise offload -transfers each block once per batch. This amortizes transfers on consumer GPUs. -Resident weights do not incur those transfers, so batching is not automatically -faster there. Full-checkpoint TP2 and Ulysses2 generation, editing, alpha, and -dynamic two-request batching also passed on two B200s; those are functional -checks, not latency recommendations for every multi-GPU topology. - -Batching retains native BF16/FP32 precision but changes GEMM shapes and -floating-point rounding. Batched images are **not guaranteed to match singleton -pixels**, even with the same seeds. Use one output and leave request batching -off when reproducing a singleton image exactly. Quantization, SageAttention, -and approximate denoising caches remain separate options. Larger batches and -other prompts need their own memory and throughput measurements. See +Batching preserves native precision but can change floating-point rounding and +output pixels, even with the same seed. See [Inference batching](/docs/sglang-diffusion/dynamic_batching) for admission rules -and batch metrics. - -### Earlier platform measurements - -The following four-platform comparison and the fusion measurements below precede -the training-template and VAE normalization corrections in `c2a31b2693c`; -their output comparisons should not be treated as baselines for that revision. -The separate RTX PRO 6000 measurement uses the corrected implementation. - -| GPU | Recommended placement / attention | Generation median | Single edit | Peak device memory | -| --- | --- | --- | --- | --- | -| H200 141GB | Resident / FlashAttention | Functional verification only | Passed | Not measured in this comparison | -| B200 192GB | Resident / FlashAttention | 3.44 s | 3.84 s | 40.1 GiB | -| RTX 5090 32GB | DiT layerwise offload / SDPA | 14.30 s | 16.84 s | 26.9 GiB | -| RTX 4090 24GB | DiT layerwise + encoder CPU offload / FlashAttention | 24.60 s | 25.59 s | 21.4 GiB | - -The recommendations compare exact attention backends and memory placement on -one GPU per platform. Each run warms up with one 512px, 4-step request, then -measures three 1024px, 40-step generations, one single-image edit, and one -transparent generation. All use seed 42, CFG 1, eager execution, full-image VAE -decoding, and PNG output. Generation latency is the median of three sequential -HTTP requests; editing is one request. Times include encoding and PNG response -serialization, but exclude server startup. Device memory is the highest sampled -`nvidia-smi` usage across loading and requests, sampled every 0.5 seconds. - -Measured on 2026-09-16 with source revision `128ae46cc`, PyTorch 2.13.0+cu130, -Transformers 5.12.1, and Diffusers 0.37.0. SGLang's native encoder uses the -Transformers 4.57.3 numerical semantics described below. The RTX 5090 runs used -a 50 GiB process-group memory limit on a roughly 60 GiB host; this is a tested -budget, not a minimum host-memory requirement. - -B200 FlashAttention was faster than SDPA in this comparison (3.44 vs 3.70 s). -On RTX 5090, both commands used Torch SDPA: this runtime falls back to SDPA -when `--attention-backend fa` is selected on SM120. The measured 14.30 s -(explicit SDPA) and 14.39 s (FA selection with SDPA fallback) therefore do not -compare different backends. The picker defaults to SDPA and rejects Ring with -either selection on RTX 5090. Keeping eight DiT layers resident -did not improve the RTX 5090 generation median, so that flag is omitted. -On RTX 4090, DiT offload alone passed generation but ran out of memory during -editing. The recommended command also sets `--text-encoder-cpu-offload true`; -this complete recipe passed generation, editing, and transparent PNG output. - -These are measurements of this small workload, not universal latency or image -quality guarantees. Different prompts, reference sizes, batching, and software -versions can change memory use and latency. Multi-reference recipes retain their -separate H200 verification scope. See [Batching](#batching) for the current -multi-output and concurrent-request matrix. - -### Earlier RTX PRO 6000 Blackwell 96GB measurements - -The recommended single-GPU command keeps all weights resident and selects Torch -SDPA. This is the 96GB Blackwell Server Edition (SM120). This runtime also maps -`--attention-backend fa` to SDPA on this GPU; Ring therefore requires another -supported backend and is rejected with either selection in the picker. - -Source revision `1eab5de5990` was measured on 2026-09-18: - -| Placement | Generation median | Edit median | Peak device memory | -| --- | --- | --- | --- | -| Resident (recommended) | 8.23 s | 9.85 s | 40.1 GiB | -| DiT layerwise offload | 10.28 s | 10.66 s | 26.1 GiB | - -Both runs used PyTorch 2.13.0+cu130, Transformers 5.12.1, Diffusers 0.37.0, -native precision, eager execution, and full-image VAE decoding. -After two 1024px/40-step warmups, each measured five generations and three edits -at that same resolution and step count, with seed 42, CFG 1, and CPU noise -generation. HTTP latency includes PNG serialization and excludes server startup; -device memory was sampled every 0.5 seconds across startup and requests. - -Transparent generation and two repeated edits of the same transparent input passed -with both placements, retaining alpha values from 0 to 255. Repeated requests -and corresponding outputs across placements produced identical RGBA pixels for -this workload. Quantized checkpoints and multi-GPU recipes on RTX PRO 6000 remain -unverified. - -### Lossless RoPE fusion - -The native DiT fuses the float conversion, complex rotary multiplication, and -output cast on supported CUDA tensors. Its first eager call checks exact -agreement with the original PyTorch operation; a mismatch disables the fusion. -No additional command flag is needed. - -A separate comparison on 2026-09-17 used native revision `6b190085c48` as the -baseline and `63ed20bbedb` with the fusion. Both used the software versions -listed above, full-image VAE decode, eager execution, and the recommended -placement and attention backend for each GPU: - -- B200: generation **3.42 → 3.27 s** (4.5% lower latency), editing - **4.03 → 3.89 s** (3.4% lower). -- RTX 5090: generation **14.49 → 14.20 s** (2.0% lower), editing - **16.97 → 16.68 s** (1.7% lower). - -Each GPU ran four fresh servers in optimized/baseline/baseline/optimized order. -Each startup used two full-size warmups followed by five generations and three -edits. The medians pool 10 generations and six edits per variant, all at -1024px, 40 steps, seed 42, CFG 1, CPU noise generation, and one RGBA PNG per -request. The workload generated a red teapot and edited the same reference -image to blue. HTTP times include PNG serialization and exclude startup. -All corresponding output pixels were identical between revisions on each GPU. -These measurements cover this fixed workload; other prompts and configurations -can have different gains. - -### Lossless MLP and residual fusion - -The native DiT also uses the shared BF16 SiLU-multiply and gated-residual -kernels, preserving the eager operations' intermediate rounding. SiLU-multiply -checks its first eager call and falls back on mismatch. These optimizations -are automatic on supported CUDA inputs. - -A second B200 comparison on 2026-09-17 used `f874eae18be` (already including -the RoPE fusion) versus `a3d14531474`. With resident weights, FlashAttention, -and the same four-startup protocol and workload above, generation decreased -from **3.272 to 3.134 s** (4.23%) and editing from **3.886 to 3.762 s** (3.18%). -All corresponding RGBA pixels were identical across the 10 generation and six -editing samples per variant. These are additional gains over the RoPE baseline; -this comparison does not establish the gain on other GPUs. - -### Lossless Q/K normalization - -Q/K RMSNorm fuses the input conversion and square, then the normalization, -output cast, and weight multiply. It retains the original FP32 mean reduction -with the same tensor shape, preserving the eager reduction order and -cast-before-weight rounding. The native DiT verifies its first eager call and -uses the original implementation if the outputs differ. No flag is needed. - -A B200 comparison on 2026-09-17 used revision `4e5459e0eda` (including the -RoPE, MLP, and residual fusions) versus `d9e1e5dac96`. With resident weights, -FlashAttention, and the four-startup protocol above, generation decreased from -**3.114 to 2.828 s** (9.17%) and editing from **3.742 to 3.450 s** (7.81%). -Each variant has 10 generation and six editing measurements at 1024px, -40 steps, seed 42, and CFG 1. Every corresponding RGBA pixel was identical. -These gains apply to this fixed B200 workload; other GPUs were not measured -in this comparison. - -### Lossless LayerNorm modulation - -The DiT fuses affine-free LayerNorm and `* (1 + scale)` while retaining the -eager Welford reduction and BF16 rounding order. Scale-only modulation skips -the shift addition, including its effect on signed zeros. The first eager call -checks the fused result against the native path and falls back on a mismatch. - -A B200 comparison on 2026-09-17 used `5bddbfca9b1` (including the preceding -fusions) versus `162181ff0ec`. With resident weights, FlashAttention, eager -execution, and the same four-startup protocol, generation decreased from -**2.831 to 2.748 s** (2.92%) and editing from **3.436 to 3.358 s** (2.26%). -Each variant has 10 generation and six editing measurements at 1024px, -40 steps, seed 42, and CFG 1. Every corresponding RGBA pixel was identical. -This comparison measures this B200 workload only. +and metrics. ## 2. Model capabilities -Qwen-Image 2.1 supports text-to-image generation and image-conditioned editing -through one pipeline. Qwen3-VL encodes the instruction and reference images; -a single-stream transformer inserts each reference image's latents into its -corresponding position in that sequence. Block-causal attention keeps each -image internally bidirectional while respecting the order of text and images. +Qwen-Image 2.1 supports text-to-image generation, single- and multi-image editing, +and RGBA output. Use one checkpoint for all modes. -For successive edits, send the previous output as the next request's reference -image. Requests do not retain dialogue history. Conditional KV is reused across -denoising steps within one request and released afterward; cross-request caching -and incremental dialogue-history caching are not implemented. - -Choose this pipeline for checkpoints declaring `QwenImage21Pipeline`, -`QwenImage21Transformer2DModel`, and `AutoencoderKLQwenImage21`. The older -Qwen-Image and Qwen-Image-Edit checkpoints use different components and latent -packing. They cannot share this model's VAE or transformer weights. Text and -condition-image activations use timestep zero, allowing their attention keys -and values to be reused for the remaining denoising steps. +For multi-round editing, send the previous output as the next reference image. +The server does not retain conversation state. Condition-prefix KV caches are +reused within one request; cross-request and dialogue-history caching are not +implemented. ## 3. Checkpoint layout The checkpoint directory must contain `model_index.json` and the `processor`, -`text_encoder`, `transformer`, `vae`, and `scheduler` subdirectories. The -processor must include the Qwen3-VL tokenizer assets. SGLang loads all three -neural components natively. A separate tokenizer directory is not required. +`text_encoder`, `transformer`, `vae`, and `scheduler` subdirectories. The processor +includes the Qwen3-VL tokenizer assets; no separate tokenizer directory is needed. +Use `--model-id Qwen-Image-2.1` when your local checkpoint directory has another +name. Older Qwen-Image and Qwen-Image-Edit transformer/VAE weights are incompatible. -The checkpoint's VAE uses RGBA input and output with 64-channel latents. PNG -reference images retain their alpha channel; RGB inputs receive an opaque -alpha channel. Save generated images as PNG to preserve transparency. - -Text conditioning uses the last decoder layer's output before the final -normalization, matching the reference implementation with Transformers -4.57.3. Vision position interpolation also follows its BF16 rounding order. -SGLang selects these native semantics explicitly, so keep the -repository's installed dependencies instead of downgrading the entire runtime. -The updated [Diffusers reference](https://github.com/huggingface/diffusers/pull/14804) -also selects pre-normalization hidden states explicitly on newer Transformers. - -Editing uses the training markers ``, ``, and so on. The vision -encoder sees alpha composited over white, while the VAE receives the original -RGBA pixels. Empty prompts become a space. The VAE normalizes features in -FP32 before casting back to the activation dtype and compresses spatial -dimensions by a factor of 16. - -Use `--model-id Qwen-Image-2.1` when the checkpoint directory has a different -name. The model ID is a routing identifier; it does not grant access to model -weights. Keep checkpoint access credentials in your environment. - -### Two-GPU end-to-end test - -The `qwen_image21_t2i_tp2` case is temporarily disabled until the checkpoint is -accessible to fork PR CI. Its configuration and pinned reference image are -retained for re-enabling the test. - -The case uses TP 2 with sequence -parallelism disabled, 1024 × 1024 PNG output, 40 steps, CFG 1, and seed 42. -It sends two consecutive requests and checks the model API and image consistency. -This case does not enforce a latency baseline or run a component accuracy check. +Keep SGLang's installed dependencies. Its native encoder preserves the +reference's Transformers 4.57.3 conditioning semantics without requiring a +runtime-wide downgrade. ### Transparent PNG output -Choose **Transparent / alpha** under Request to generate an isolated subject -or preserve a transparent reference during editing. The picker adds the -transparency instruction to the prompt and sets `output_format: "png"`. -`background: "transparent"` alone only selects an output format; it does not -remove the background or change model conditioning. JPEG cannot retain alpha. +Choose **Transparent / alpha** under Request and describe an isolated subject +on a transparent background in the prompt. The picker adds this instruction +and selects PNG. `background: "transparent"` alone does not change conditioning +or remove the background; JPEG cannot retain alpha. -The model predicts continuous alpha values, including partly transparent edges. -No thresholding or background-removal postprocessing is applied. Transparent -generation and transparent-input editing were compared against the reference -at 1024 × 1024 and 40 steps; that check does not guarantee perfect cutouts for -every prompt. The updated checkpoint also passed transparent generation and -transparent-input editing on H200, B200, RTX PRO 6000, and RTX 4090. See -[Batching](#batching) for the tested output counts and two-B200 topologies. +PNG references retain their alpha channel during editing; RGB references +receive an opaque alpha channel. The model predicts continuous alpha values, +including partly transparent edges, without thresholding or background removal. ## 4. Offline requests @@ -415,247 +152,67 @@ noise seeds and independent prefix caches. ## 5. Runtime features -The API requires a text prompt; precomputed embeddings alone do not provide -the image-token positions needed by this pipeline. +The default is 40 Euler flow-matching steps with CFG disabled. For CFG, provide +`--negative-prompt` and `--guidance-scale` greater than one. The API requires a +text prompt; precomputed embeddings alone are insufficient. -The default is 40 Euler flow-matching steps with CFG disabled. To use CFG, -provide `--negative-prompt` and a `--guidance-scale` greater than one. CFG uses -the ordinary linear combination without the older Qwen-Image norm correction. -Positive and negative prompts have separate request-owned prefix caches. +- **Parallelism:** TP, Ulysses, Ring, CFG parallelism, and encoder folding are + available in the picker. The target token count, `(height / 16) × (width / 16)`, + must be divisible by the SP degree. Ring requires FlashAttention or SageAttention. +- **Memory:** use the hardware's recommended placement. **All components + layerwise** also streams encoder and VAE blocks, trading transfers for lower + device memory. +- **VAE:** full-image decoding is the default. Tiling can change pixels near + boundaries. With two or more GPUs, **Spatial shard** distributes full-image + decoding without enabling tiling; floating-point rounding can still differ. -TP uses native parallel projections. Ulysses and Ring shard target-image -attention while keeping the condition prefix replicated. The target token -count, `(height / 16) × (width / 16)`, must be divisible by the SP degree. Encoder -folding shards Qwen3-VL's language projections using the native encoder TP group. -Full-checkpoint editing passed with TP2 × Ulysses2 and TP2 × Ring2 + FlashAttention -on four B200 GPUs. These CLI checks do not mark every HTTP topology as verified. +See the [compatibility inventory](/docs/sglang-diffusion/compatibility_matrix) +for configuration support and the +[performance guide](/docs/sglang-diffusion/performance-optimization) for shared +runtime options. -VAE tiling is disabled by default for both encoding and decoding. Enable -`--vae-tiling true` for tiled encoding and decoding; `--vae-sp true` also distributes tiles -across the configured GPUs. These paths use the standard VAE runtime; tiled -decode can differ from full image decode near tile boundaries. +### Quantization -For full-image spatial parallel decode, select **Spatial shard** or pass -`--vae-config.parallel-decode-mode spatial_shard` with at least two GPUs. -This mode splits feature-map height, exchanges convolution halos, and gathers -the full map for VAE attention. It does not require `--vae-tiling` or `--vae-sp`. -Two-B200 checks cover TP2, CFG parallelism, and all-component layerwise offload. -FP64 component comparisons match full decode; BF16 full-checkpoint output can -differ through floating-point rounding. +Native precision is the default. Quantization changes image and alpha values; +check quality on your own prompts and reference images. Set compatible component +paths under **Variables** when choosing an exported format. Adding quantization +metadata to native weights does not convert them. -Select **All components layerwise** or pass `--layerwise-offload-components all` -to stream repeated blocks in the DiT, Qwen3-VL language and vision encoders, and -VAE encoder/decoder. Full-checkpoint 512px editing passed on one B200 and on -two B200s with TP2 plus spatial VAE decode. This setting reduces device memory -at the cost of host-device transfers; it is not the measured default for the -consumer-GPU recipes above. - -Revision `f1f3366c7c` fixes CPU/GPU initialization rounding in the vision -encoder's rotary frequencies after device transfer. On one B200, native -1024px/40-step generation, editing, and transparent output with all-component -layerwise offload matched resident RGBA pixels exactly. Repeated editing after -a transparent-generation request also matched. Resident output was unchanged -from revision `6ee35b52fb`. These checks use FlashAttention, seed 42, and CFG 1. - -Revision `81c8c550fa` also preserves the loader's FP8 weights and FP32 rotary -buffers when moving the whole encoder between CPU and GPU. With that fix, -`--text-encoder-cpu-offload true` matched resident generation, editing, and -transparent RGBA pixels for both native precision and the combined serialized -FP8 export in the same B200 workload, including repeated editing. - -The pipeline also supports the shared -[disaggregated runtime](/docs/sglang-diffusion/disaggregation). The encoder role -loads both Qwen3-VL and the VAE to prepare reference-image conditioning; nested -condition tensors and complex RoPE tensors transfer with the request. Separate -encoder, denoiser, and decoder processes matched monolithic RGBA output for -512px/4-step generation, editing, different prompt lengths, and CFG on B200. -That check used same-host Mooncake TCP; multi-host RDMA remains unverified. - -Online FP8 is available independently for the DiT and encoder through -`--component-quantizations.transformer fp8` and -`--component-quantizations.text_encoder fp8`. Each component and the combination -passed 1024px/40-step HTTP generation and editing on a resident B200. FP8 changes -the output: in one generation/edit pair, DiT-only FP8 gave RGBA PSNR -37.56/41.07 dB against native precision; quantizing both gave 32.66/40.99 dB. -These samples do not establish general image or alpha quality. Native precision -remains the default. +For online FP8, use `--component-quantizations.transformer fp8`, +`--component-quantizations.text_encoder fp8`, or both. ### Serialized FP8 components -Select a **Serialized FP8** precision option in the picker and set the component -directories under **Variables**. The tested format is E4M3FN weights with one -FP32 `weight_scale` per linear and dynamic activation quantization. Each -component directory contains its own architecture `config.json`, weight shards, -and index; merge this top-level quantization configuration into its `config.json`: - -```json -{ - "quantization_config": { - "quant_method": "fp8", - "activation_scheme": "dynamic" - } -} -``` - -Load compatible exported components through the shared loader: - -```bash Command -sglang serve \ - --model-path /models/qwen-image-2.1 \ - --model-id Qwen-Image-2.1 \ - --component-paths.transformer /models/qwen-image-2.1-fp8/transformer \ - --component-paths.text_encoder /models/qwen-image-2.1-fp8/text_encoder \ - --num-gpus 1 --performance-mode speed --attention-backend fa \ - --host 0.0.0.0 --port 30010 -``` - -Use either override independently, or both as shown. Omit online quantization -flags: the component metadata selects serialized loading. Adding metadata to -BF16 weights does not convert them. The validated export quantizes 224 DiT -attention/MLP matrices and 252 Qwen3-VL language matrices; the vision encoder, -embeddings, output head, other DiT projections, and VAE retain native precision. -All 476 loaded matrices and scales matched their serialized values. - -At revision `5a117c9f3f`, DiT-only, encoder-only, and combined exports passed -1024px/40-step generation, editing, and transparent PNG requests on B200 with -FlashAttention, seed 42, and CFG 1. The combined export also passed TP2 with -encoder folding and single-GPU `--layerwise-offload-components all`. -At that revision, offload matched resident generation and transparent output -exactly, but editing differed at 49.50 dB RGBA PSNR. Revision `f1f3366c7c` fixes -the vision rotary initialization difference: a new 1024px/40-step comparison -matched resident generation, editing, and transparent RGBA pixels exactly -with all-component layerwise offload. Resident outputs were unchanged. TP2 -still changes numerical results. - -| Serialized FP8 scope | Generation RGBA PSNR vs native | Edit RGBA PSNR vs native | -| --- | --- | --- | -| DiT | 38.35 dB | 40.94 dB | -| Encoder | 34.46 dB | 49.19 dB | -| Both | 34.93 dB | 41.25 dB | - -For the combined export, the transparent cat's alpha channel measured 32.03 dB -PSNR and 0.81 mean absolute error on the 0–255 scale against native precision; -individual boundary pixels can differ substantially. Online FP8 for both -components also produced a real transparent PNG in this check. These are -single-example comparisons, not a quality guarantee. Offline tensorwise scales -differ from B200 online FP8's channelwise scales. +Select a **Serialized FP8** option and set the exported component directories. +Each directory needs its architecture `config.json`, weights, and quantization +metadata. Use `--component-paths.transformer` and/or +`--component-paths.text_encoder`; omit online quantization flags. +See the [quantization guide](/docs/sglang-diffusion/quantization) for formats. ### GGUF components -Select **GGUF DiT**, **GGUF encoder**, or **GGUF DiT + encoder** under Server -precision, then set the corresponding `.gguf` files under **Variables**. -The picker uses `--component-weights-paths.transformer` and -`--component-weights-paths.text_encoder`, retaining each component's architecture -config from the base checkpoint. Each file must contain the entire component -with native checkpoint tensor names. No online quantization flag is needed; -the loader reads the quantization type from each GGUF tensor. - -The tested Q4_0 export quantizes the same 224 DiT and 252 language-encoder -matrices listed above. Other tensors retain native precision, including the -vision tower, embeddings, output head, and VAE. Its DiT and encoder files are -3.91 and 7.03 GiB respectively. All 476 loaded packed matrices matched the -exported bytes; sampled CUDA dequantization matched the GGUF CPU reference -after conversion to BF16. - -At revision `7e0d4e9185`, DiT-only, encoder-only, and combined Q4_0 exports -passed 1024px/40-step HTTP generation, editing, and transparent PNG output on -B200 with FlashAttention, seed 42, and CFG 1. These are private validation -exports, not published download targets. Use a compatible export of weights -you are authorized to access. - -The combined export also passed TP2 with encoder folding. On one GPU, -all-component layerwise offload and whole-encoder CPU offload each matched -resident generation, editing, and transparent RGBA pixels exactly. TP2 changed -numerical results. Quantization itself is lossy: - -| Q4_0 scope | Generation RGBA PSNR vs native | Edit RGBA PSNR vs native | -| --- | --- | --- | -| DiT | 24.99 dB | 33.66 dB | -| Encoder | 28.97 dB | 43.26 dB | -| Both | 23.86 dB | 33.46 dB | - -The combined export's transparent cat retained alpha values from 0 to 255, -with 66.8% of pixels at alpha 5 or below. Against native precision, its alpha -PSNR was 21.20 dB and mean absolute error was 3.29/255; individual boundary -pixels differed by up to 255. These single-example comparisons do not establish -general image or cutout quality. Keep native precision when exact output is -required. - -GGUF reduces weight storage; it is not a promise of lower latency. The runtime -dequantizes packed linears before BF16 matrix multiplication. Other GGUF tensor -types, exports, and hardware need separate validation. -See the shared [GGUF guide](/docs/sglang-diffusion/quantization#gguf) -for loader and parallelism constraints. +Select a **GGUF** option and set the `.gguf` files. The picker uses +`--component-weights-paths.transformer` and/or +`--component-weights-paths.text_encoder`, retaining architecture configs from the +base checkpoint. Each file must contain the entire component with native tensor +names. GGUF reduces weight storage but does not guarantee lower latency. +See the [GGUF guide](/docs/sglang-diffusion/quantization#gguf). ### NVFP4 components -Select **NVFP4 DiT**, **NVFP4 encoder**, or **NVFP4 DiT + encoder** in the -picker, then set the component directories under **Variables**. These options -require Blackwell; H200 and RTX 4090 cannot run this native FP4 path. B200 has -completed the checks below. RTX PRO 6000 and RTX 5090 remain unverified for this -model's NVFP4 exports; their FlashInfer backend defaults to `auto`, because -TensorRT-LLM FP4 GEMM does not support SM120. Keep that default on these GPUs. - -Each exported directory contains its architecture config, weight shards, and -index. The config declares `quant_method: modelopt`, `quant_algo: NVFP4`, and -block size 16, with exclusions for native-precision layers. Use -`--component-paths.transformer` and/or `--component-paths.text_encoder` to load -the exported directories. Omit online quantization flags; metadata alone does -not convert native weights into an NVFP4 checkpoint. - -The private validation export quantizes the same 224 DiT and 252 language -matrices as the FP8 example. Vision, embeddings, the output head, other DiT -projections, and VAE retain native precision. Weight quantization uses ModelOpt -0.46.1 with max calibration; static activation scales come from six separate -1024px/40-step requests, including two edits and one transparent generation. -This small calibration set does not establish general quality. It does not -use SVDQuant or AWQ. All 476 loaded packed weights, block scales, and global -scales matched the export after the runtime's layout transforms. - -At revision `57b625d3e3`, each component and both together passed 1024px/40-step -HTTP generation, editing, and transparent PNG output on B200 with -FlashAttention, seed 42, CFG 1, and FlashInfer TensorRT-LLM FP4 GEMM. The combined -export also passed TP2 with encoder folding. Single-GPU all-component layerwise -offload and whole-encoder CPU offload each matched the combined resident RGBA -pixels exactly. TP2 changed numerical results. - -| NVFP4 scope | Generation RGBA PSNR vs native | Edit RGBA PSNR vs native | -| --- | --- | --- | -| DiT | 24.97 dB | 31.56 dB | -| Encoder | 26.48 dB | 36.63 dB | -| Both | 19.36 dB | 29.96 dB | - -The combined export's transparent cat retained alpha from 0 to 255, with -67.8% of pixels at alpha 5 or below. Against native precision, alpha PSNR was -23.81 dB and mean absolute error was 2.22/255; some boundary pixels differed -by 255. These are single-example comparisons of private exports, not download -targets or quality guarantees. Native precision remains the default. See the -shared [NVFP4 guide](/docs/sglang-diffusion/quantization#modelopt-nvfp4) for loader -details. +NVFP4 requires Blackwell and compatible ModelOpt exports. Select the component +directories using `--component-paths.transformer` and/or +`--component-paths.text_encoder`. Keep the FlashInfer backend at `auto` on +RTX 5090 and RTX PRO 6000: TensorRT-LLM FP4 GEMM does not support SM120. +These GPUs remain unverified for this model's NVFP4 exports. See the +[NVFP4 guide](/docs/sglang-diffusion/quantization#modelopt-nvfp4). ### LoRA and execution options -LoRA uses the shared `--lora-path` and `--lora-merge-mode dynamic|merge` options -and runtime adapter APIs. Diffusers keys prefixed with `transformer.` map to -the native DiT. A synthetic adapter covering attention and MLP projections -passed dynamic loading, merging, and removal on one B200 and TP2 with encoder -folding. Both removal paths restored the base image exactly. This verifies -adapter application and lifecycle, not the quality of a trained LoRA. +Use `--lora-path` and `--lora-merge-mode dynamic|merge` or the runtime adapter APIs. +Diffusers adapter keys prefixed with `transformer.` map to the native DiT. -Cache-DiT hooks operate on target-image transformer blocks; two-output generation, -editing, transparent generation, and transparent-input editing passed with it -enabled on B200 at 1024px/40 steps. This is a functional check of an approximate -cache, not a lossless recipe. Breakable CUDA -Graph execution fills each request's prefix caches eagerly, then replays -matching warmup graphs with those cache tensors as inputs. Warmup and request -condition-prefix lengths must match, in addition to the output resolution; -unseen shapes run eagerly. Text buckets alone cannot pad condition KV without -changing attention semantics. FlashAttention, Sage -attention and Torch SDPA are wired through the native attention layers; -causal text runs use exact masked SDPA. Sage and Cache-DiT can change numerical -results and require application-specific quality checks. - -See the [compatibility inventory](/docs/sglang-diffusion/compatibility_matrix) -for tested configurations and remaining validation boundaries. These checks -are functional and numerical comparisons. The platform measurements above cover -their stated HTTP workload; broader image quality is not evaluated. +Keep eager execution as the default. Breakable CUDA Graph replay requires +matching resolution and condition-prefix length; unseen shapes run eagerly. +Text buckets alone do not guarantee replay. SageAttention and Cache-DiT can +change numerical results and require quality checks for your workload. diff --git a/docs/docs/sglang-diffusion/compatibility_matrix.mdx b/docs/docs/sglang-diffusion/compatibility_matrix.mdx index 1aa611353..daca9f373 100644 --- a/docs/docs/sglang-diffusion/compatibility_matrix.mdx +++ b/docs/docs/sglang-diffusion/compatibility_matrix.mdx @@ -41,7 +41,7 @@ tests, including batched targets and independent variable-length prefixes. Compatible text-to-image requests support opt-in dynamic batching; image-edit requests remain separate, while multiple outputs within one request are supported. See the cookbook's [batching guidance](/cookbook/diffusion/Qwen-Image/Qwen-Image-2.1#batching) -for measured throughput and floating-point reproducibility limits. +for deployment settings and floating-point reproducibility limits. The deployment picker marks only its exact tested HTTP combinations as verified, including H200, B200, RTX PRO 6000 96GB, RTX 5090, and RTX 4090. CLI-only combinations remain Unverified in the picker. diff --git a/docs/docs/sglang-diffusion/dynamic_batching.mdx b/docs/docs/sglang-diffusion/dynamic_batching.mdx index 1085b3282..14c845e51 100644 --- a/docs/docs/sglang-diffusion/dynamic_batching.mdx +++ b/docs/docs/sglang-diffusion/dynamic_batching.mdx @@ -102,7 +102,7 @@ An initial implementation of dynamic batching for T2I and T2V models can be foun Qwen Image 2.1 supports merging compatible text-to-image requests. Image edits are not merged across requests; `n > 1` still produces multiple outputs within one edit request. See its [cookbook](/cookbook/diffusion/Qwen-Image/Qwen-Image-2.1#batching) -for platform measurements and when batching is useful. +for deployment settings and batching tradeoffs. ### Video diff --git a/docs/docs/sglang-diffusion/quantization.mdx b/docs/docs/sglang-diffusion/quantization.mdx index 8742ffe45..90873d9c8 100644 --- a/docs/docs/sglang-diffusion/quantization.mdx +++ b/docs/docs/sglang-diffusion/quantization.mdx @@ -780,7 +780,7 @@ component directories through `--component-paths.transformer` and generation, editing, transparent RGBA, offload, and TP2 checks. These checks do not cover arbitrary exports or RTX 5090. See its [cookbook](/cookbook/diffusion/Qwen-Image/Qwen-Image-2.1#nvfp4-components) for -calibration scope and image/alpha error measurements. +component-loading instructions and hardware requirements. ### Notes @@ -936,8 +936,8 @@ The MiniMax-H3 measurements do not validate other quantization types, the Qwen-Image 2.1's Q4_0 exports use the base component configs and native tensor names. Its [cookbook](/cookbook/diffusion/Qwen-Image/Qwen-Image-2.1#gguf-components) -documents the tested matrix selection, offload modes, TP2, and image/alpha -differences against native precision. These private exports establish loading +describes how to load compatible component files and choose precision settings. +These private exports establish loading and execution compatibility, not general output quality or compatibility with other GGUF exports. diff --git a/docs/src/snippets/configs/Qwen/qwen-image-2.1.jsx b/docs/src/snippets/configs/Qwen/qwen-image-2.1.jsx index 4c34f7ade..69ad357ac 100644 --- a/docs/src/snippets/configs/Qwen/qwen-image-2.1.jsx +++ b/docs/src/snippets/configs/Qwen/qwen-image-2.1.jsx @@ -39,7 +39,7 @@ const config = { id: "placement", title: "Placement", scope: "serve", - description: "Hardware selection applies its recommended placement. Stream DiT layers when the full pipeline exceeds device memory.", + description: "Hardware selection applies its recommended placement. Offload selected components when the full pipeline exceeds device memory.", learnMore: "#5-runtime-features", default: "resident", options: [ @@ -53,14 +53,14 @@ const config = { }, { id: "offload", label: "CPU offload", - flags: (s) => ["--performance-mode manual", "--dit-layerwise-offload true", ...(s.hw === "rtx4090" ? ["--text-encoder-cpu-offload true"] : []), - ...(s.hw === "rtx4090" && Number(s.gpus_per_node) === 1 && effectiveAttention(s) === "fa" && s.precision === "native" && s.execution === "eager" - && ["text", "edit"].includes(s.mode) && Number(s.outputs) === 1 && (!s.batching || s.batching === "off") - ? ["--dit-layerwise-resident-layers 8"] : [])], + flags: (s) => s.hw === "rtx4090" && Number(s.gpus_per_node) === 1 && effectiveAttention(s) === "fa" && s.precision === "native" && s.execution === "eager" + && ["text", "edit"].includes(s.mode) && Number(s.outputs) === 1 && (!s.batching || s.batching === "off") + ? ["--performance-mode manual", "--component-residency dit=resident text_encoder=layerwise-offload vae=resident", `--warmup-resolutions ${s.resolution || "1024"}x${s.resolution || "1024"}`] + : ["--performance-mode manual", "--dit-layerwise-offload true", ...(s.hw === "rtx4090" ? ["--text-encoder-cpu-offload true"] : [])], recommendedWhen: (s) => ["rtx5090", "rtx4090"].includes(s.hw), soft: (s) => !["rtxpro6000", "rtx5090", "rtx4090"].includes(s.hw) || Number(s.gpus_per_node) !== 1, softReason: "This offload topology has not completed an HTTP verification run.", - description: "Streams DiT layers. RTX 4090 also offloads the encoder. Its native single-output FlashAttention recipe keeps 8 DiT layers resident; batch recipes stream every layer for memory headroom. Requires sufficient host RAM.", + description: "RTX 4090 native single-output FlashAttention keeps the DiT and VAE resident and streams encoder layers. Other offload recipes stream DiT layers; RTX 4090 also offloads the encoder. Requires sufficient host RAM.", }, { id: "all_offload", label: "All components layerwise", @@ -228,7 +228,7 @@ const config = { { id: "2", label: "Up to 2 images", flags: ["--batching-max-size 2", "--batching-delay-ms 20"], - description: "Wait up to 20 ms to merge compatible queued requests. The tested RTX 4090 offload recipe benefits under concurrent load; each response takes longer.", + description: "Wait up to 20 ms to merge compatible queued requests. Benchmark throughput and response latency on your workload.", }, { id: "4", label: "Up to 4 images", @@ -339,7 +339,7 @@ const config = { if (topology.ring_degree > 1) flags.push(`--ring-degree ${topology.ring_degree}`); flags.push("--host {{HOST_IP}}", "--port {{PORT}}"); const warnings = []; - if (s.hw === "rtx4090" && (Number(s.outputs) > 1 || (s.batching && s.batching !== "off"))) warnings.push("This recipe streams all DiT layers. Use the updated Server command if switching from the single-output recipe with 8 resident layers."); + if (s.hw === "rtx4090" && (Number(s.outputs) > 1 || (s.batching && s.batching !== "off"))) warnings.push("This recipe streams DiT layers for batch memory headroom. Restart with the updated Server command when changing output count or request batching."); if (s.batching && s.batching !== "off" && s.mode !== "text") warnings.push("Cross-request batching applies to text-to-image requests. Image edits run separately; use Outputs for multiple images in one edit request."); if (!serveVerified && !errors.length) warnings.push("This server combination has not completed an exact HTTP verification run."); if (!requestVerified && !errors.length) warnings.push("This request shape is outside the verified HTTP matrix.");