From 42b5af8c62e37b744e351e85867ada4a383f17d3 Mon Sep 17 00:00:00 2001 From: Haocheng Xi <87399272+haochengxi@users.noreply.github.com> Date: Sun, 13 Sep 2026 20:50:40 -0700 Subject: [PATCH] [diffusion] docs: refresh the VDN-H3 on b200 numbers in cookbook (#39244) Co-authored-by: haochengxi Co-authored-by: Claude Opus 5 (1M context) Co-authored-by: Kevin Mi --- .../cookbook/diffusion/MiniMax/MiniMax-H3.mdx | 209 ++++++++---------- .../runtime/utils/model_overlay.py | 2 +- 2 files changed, 93 insertions(+), 118 deletions(-) diff --git a/docs/cookbook/diffusion/MiniMax/MiniMax-H3.mdx b/docs/cookbook/diffusion/MiniMax/MiniMax-H3.mdx index c49f66eb5..13c523c0b 100644 --- a/docs/cookbook/diffusion/MiniMax/MiniMax-H3.mdx +++ b/docs/cookbook/diffusion/MiniMax/MiniMax-H3.mdx @@ -802,32 +802,31 @@ latency. ## 7. VDN-H3: hybrid attention, 8-step distill +[Video DeltaNet](https://openvdn.github.io/) (VDN) is a hybrid attention +architecture that approaches the speed of linear attention while maintaining +near-lossless quality. It splits the video-video attention into two +complementary branches: a sliding-window softmax branch for local frame pairs +(five consecutive latent frames form a chunk, and each chunk attends to +itself, the previous chunk, and the following chunk; the first and last +frames, text and audio stay dense) and a bidirectional linear branch, Video +Delta Attention, for long-range context. [OpenVDN/vdn-minimax-h3](https://huggingface.co/OpenVDN/vdn-minimax-h3) -(Video DeltaNet MiniMax-H3, [openvdn.github.io](https://openvdn.github.io/)) -replaces every DiT block's dense self-attention with a hybrid of two branches: -an exact, gated softmax over a chunk-aligned frame window (chunk 5, radius 1, -the first and last frames dense as anchors, text and audio dense both ways) -and a frame-wise linear-attention branch (the Video Delta rule, scanned -forward and backward over frames) that covers exactly the rest. The released -`stage-dmd-step-250` checkpoint is an 8-NFE DMD2 distill and adds a 4.3 GB -linear branch plus two small LoRA adapters on top of the untouched MiniMax-H3 -backbone. The weights are the FL2VA partition, so the deployment serves `t2va` -and `fl2va` (first / last keyframes); the keyframe rows are attended densely, -like text and audio. `ref2va` was not trained and is rejected. The weights -inherit the MiniMax-H3 Community License, including its territorial exclusions. +applies it to MiniMax-H3: the released `stage-dmd-step-250` checkpoint adds a +4.3 GB linear branch and OpenVDN's own 8-step DMD2 LoRA on the unchanged H3 +backbone. It is the FL2VA partition, so it serves `t2va` and `fl2va` (first, +last, or both keyframes); `ref2va` is rejected. The weights inherit the +MiniMax-H3 Community License. -Pass the repository directly to `--model-path`. A registered model overlay -materializes the base-H3 layout once: both LoRA adapters are prefused into the -transformer weights (a real 62 GB write), the linear branch is attached as an -extra shard, the Qwen3-VL conditioner is linked from `MiniMaxAI/MiniMax-H3`, -and the video VAE is re-serialized as for FastH3. Point -`SGLANG_DIFFUSION_CACHE_ROOT` at a volume with at least 90 GB free before the -first launch. +Pass the repository directly to `--model-path`. On first launch a registered +overlay prefuses both adapters into the transformer weights (a 62 GB write) +and attaches the linear branch; the Qwen3-VL conditioner and the VAEs are +hard-linked from `MiniMaxAI/MiniMax-H3`. Keep `SGLANG_DIFFUSION_CACHE_ROOT` on +the same filesystem as the Hugging Face cache with at least 90 GB free. -```bash 4×B200 hybrid window attention +```bash 8×B200 sglang serve \ --model-path OpenVDN/vdn-minimax-h3 \ - --num-gpus 4 \ + --num-gpus 8 \ --attention-backend hybrid_window_attn_h3 \ --performance-mode speed \ --warmup-num-frames 345 \ @@ -835,45 +834,33 @@ sglang serve \ --port 30010 ``` -`--warmup-num-frames` / `--warmup-resolutions` make the startup warmup run -at the clip length and canvas you will serve (here the 14.375 s paper -workload). Without them H3 warms up on a 5-second clip, and the first -forward of the first longer request pays allocator growth and kernel setup -in every DiT block (2 to 3 s on this workload, for base H3 as well). +`--num-gpus 4`, `2` or `1` gives the smaller Ulysses recipes in the benchmark +table; `sglang generate` takes the same flags. `--warmup-num-frames` / +`--warmup-resolutions` warm up at the clip length and canvas you will serve; +without them the first long request pays 2 to 3 s of allocator growth and +kernel setup. -Requests use the same asynchronous video endpoint as the base model, with -`task: "t2va"`, `conditions: []` (or `task: "fl2va"` with the keyframe -conditions of the [FL2VA request](#4-generate-video-and-audio)), and a target such as -`{"short_edge": 768, "aspect_ratio": "16:9", "duration_seconds": 14.375}`. Each -keyframe adds about 2,000 global rows (the conditioner's image tokens plus the -latent keyframe), dense both ways in the window softmax, so an `fl2va` forward -costs more than a `t2va` one. The -request default is `num_inference_steps: 9`: nine points on the standard -shift-12/shift-3 sigma grid, i.e. the eight distilled DiT evaluations (VDN -counts NFEs, SGLang counts grid points). Any other step count is rejected. The -paper workload is 1344×768 at 24 fps for 14.375 s: 345 frames, already 17n+5 -aligned, 102 latent frames and about 104.5k packed rows. +Requests use the same video endpoint as the base model with `task: "t2va"` +and `conditions: []`, or `task: "fl2va"` with the +[keyframe conditions](#4-generate-video-and-audio), and a target such as +`{"short_edge": 768, "aspect_ratio": "16:9", "duration_seconds": 14.375}` (the +345-frame paper workload). The default `num_inference_steps: 9` is the only +accepted value: nine sigma grid points, eight DiT forwards. Each keyframe adds +about 2,000 dense rows, so an `fl2va` forward costs a little more than `t2va`. -`hybrid_window_attn_h3` is required for the transformer: a dense backend on -these weights would silently skip the linear branch and the softmax gates and -produce the wrong model, so it is rejected. The window softmax runs as a union -of dense FlashAttention varlen calls, exact to bf16 rounding. The -linear branch runs on fused Triton kernels (temporal conv + SiLU + L2 norm, -statistics prologue, gated RMSNorm epilogue). On Blackwell (SM100 / SM103, -and SM120 such as the RTX PRO 6000) the transformer defaults to online MXFP8 -(`--quantization fp8` selects it too, `--quantization bf16` opts out); the -gates, beta, alpha and the conv stay bf16. Before SM100 `--quantization fp8` -is the per-channel fp8 path of base H3 (SM89 and SM90); SM80 has no fp8 -tensor cores, so it runs the bf16 DiT (62 GB, budget for -`--layerwise-offload-components` or `--dit-layerwise-offload`). Ampere and -Ada run the window on FA3's Sm80 mainloop, which is FA2-class throughput; they -are enabled but not benchmarked. Ulysses sequence parallelism is supported; -`--model-variant`, `quality: "high"`, `--ring-degree` greater than 1, -`torch.compile`, and breakable CUDA graph execution are rejected. See -[Attention Backends](/docs/sglang-diffusion/attention_backends) for the -backend options. +`hybrid_window_attn_h3` is required: a dense backend would silently skip the +linear branch and the gates and produce the wrong model. On Blackwell +(SM100 / SM103 / SM120) the transformer defaults to online MXFP8; +`--quantization bf16` opts out. Before SM100, `--quantization fp8` is base +H3's per-channel fp8 path, and SM80 runs the bf16 DiT (62 GB; budget for +`--layerwise-offload-components`). Ampere and Ada are enabled but not +benchmarked. Ulysses sequence parallelism is supported; `--model-variant`, +`quality: "high"`, `--ring-degree` greater than 1, `torch.compile`, and +breakable CUDA graph execution are rejected. The backend's +`--attention-backend-config` keys are listed in +[Attention Backends](/docs/sglang-diffusion/attention_backends). -Measured latencies for the 4× B200 recipe are in +Measured latencies for 1, 2, 4 and 8 B200s are in [VDN-H3 on B200](#vdn-h3-on-b200); single-card and PCIe multi-card numbers for the RTX PRO 6000 are in [VDN-H3 on RTX PRO 6000](#vdn-h3-on-rtx-pro-6000). @@ -1531,74 +1518,62 @@ peak memory. ### VDN-H3 on B200 -A 4× B200 (SM100, 183 GB) host served [VDN-H3](#7-vdn-h3-hybrid-attention-8-step-distill) +An 8× B200 (SM100, 183 GB) host served [VDN-H3](#7-vdn-h3-hybrid-attention-8-step-distill) on the paper workload: 1344×768 at 24 fps with audio for 14.375 s (345 frames, 102 latent frames, about 104k packed rows), `task: "t2va"`, -`num_inference_steps: 9` (8 DiT forwards), seed 1000, `hybrid_window_attn_h3` -with the decomposed window kernel, eager, and the warmup run at the served -clip shape (`--warmup-num-frames 345 --warmup-resolutions 1344x768`), so -every forward of the served request is steady state; without those flags the -first forward pays about 3 s of allocator growth and kernel setup. "Steady +`num_inference_steps: 9` (8 DiT forwards), seed 1000, `hybrid_window_attn_h3`, +eager, and the warmup run at the served clip shape +(`--warmup-num-frames 345 --warmup-resolutions 1344x768`), so every forward +of the served request is steady state; without those flags the first forward +pays 2 to 3 s of allocator growth and kernel setup. "Steady s/NFE" is the mean of forwards 2 to 8. The OpenVDN reference rows ran the released inference stack (`8nfe_tuned_fp8.yaml`, `infer_ulysses.py`) on the same host -and the same clip length: +and the same clip length, with its Ulysses layout (`parallel.softmax_ranks`) +swept at every GPU count rather than assumed. Every row below comes from one +campaign, one arm at a time on an otherwise idle node; a repeated arm moves by +about 3% across nodes, so treat smaller differences as noise: | Config | Steady s/NFE | Denoise (8 forwards) | Decode | Peak/GPU | | --- | ---: | ---: | ---: | ---: | -| OpenVDN published, FP8, 8× B200 Ulysses 5+3 ([openvdn.github.io](https://openvdn.github.io/)) | 1.40 | 11.2 s | – | – | -| OpenVDN reference, BF16 (`8nfe_tuned.yaml`), 1× B200 | 7.91 | 63.3 s | – | – | -| OpenVDN reference, FP8, 1× B200 | 6.47 | 51.7 s | – | – | -| OpenVDN reference, FP8, 4× B200 Ulysses (standard / 3+1 branch-parallel) | 2.68 / 2.61 | 21.4 / 20.9 s | – | – | -| SGLang, BF16, 1× B200, DiT layerwise offload (auto policy) | 7.86 | 64.1 s | 9.5 s | 147,608 MB | -| SGLang, FP8, 1× B200, DiT layerwise offload (auto policy) | 7.49 | 61.7 s | 9.4 s | 111,616 MB | -| SGLang, FP8, 1× B200, DiT resident (`--layerwise-offload-components text_encoder`) | 7.49 | 61.1 s | 9.3 s | 63,922 MB | -| SGLang, `--quantization bf16`, 4× B200 Ulysses4, `--performance-mode speed`, served-shape warmup | **2.53** | **20.3 s** | 3.5 s | 97,894 MB | -| SGLang, FP8, 4× B200 Ulysses4, `--performance-mode speed`, served-shape warmup | **2.34** | **18.7 s** | 3.4 s | 63,022 MB | -| SGLang default (online `mxfp8`), 8× B200 Ulysses8, `--performance-mode speed`, served-shape warmup | **0.93** | **7.45 s** | 1.5 s | 77,704 MB | -| SGLang, per-channel fp8 weight scales (the online `fp8` path of other models), 8× B200 Ulysses8 | 1.05 | 8.4 s | 1.5 s | 76,828 MB | +| OpenVDN published, FP8, 8× B200 Ulysses 5+3 ([openvdn.github.io](https://openvdn.github.io/)) | 1.40 | 11.2 s | — | — | +| OpenVDN reference, BF16 (`8nfe_tuned.yaml`), 1× B200 | 7.66 | 61.2 s | — | — | +| OpenVDN reference, FP8, 1× B200 | 6.25 | 50.0 s | — | — | +| OpenVDN reference, FP8, 2× B200 Ulysses (standard / 1+1 branch-parallel) | 4.87 / 5.04 | 39.0 / 40.3 s | — | — | +| OpenVDN reference, FP8, 4× B200 Ulysses (3+1 branch-parallel) | 2.59 | 20.7 s | — | — | +| OpenVDN reference, FP8, 8× B200 Ulysses (5+3 / 4+4 / standard) | 1.40 / 1.47 / 1.52 | 11.2 / 11.7 / 12.2 s | — | — | +| SGLang, `--quantization bf16`, 1× B200 | 7.45 | 58.8 s | 8.8 s | 156,298 MB | +| SGLang, `--quantization fp8` (→ online `mxfp8` on SM100+), 1× B200 | 6.03 | 47.6 s | 8.7 s | 128,418 MB | +| SGLang default (online `mxfp8`), 2× B200 Ulysses2, `--performance-mode speed`, served-shape warmup | **3.30** | **25.9 s** | 4.7 s | 106,738 MB | +| SGLang, `--quantization bf16`, 4× B200 Ulysses4, `--performance-mode speed`, served-shape warmup | 2.06 | 16.2 s | 2.3 s | 108,434 MB | +| SGLang default (online `mxfp8`), 4× B200 Ulysses4, `--performance-mode speed`, served-shape warmup | **1.67** | **13.1 s** | 2.4 s | 85,244 MB | +| SGLang default (online `mxfp8`), 8× B200 Ulysses8, `--performance-mode speed`, served-shape warmup | **0.88** | **6.9 s** | 1.5 s | 79,972 MB | +| SGLang, per-channel fp8 weight scales (the online `fp8` path of other models; on SM100+ `--quantization fp8` resolves to `mxfp8`, so this arm selects the per-channel path explicitly), 8× B200 Ulysses8 | 0.98 | 7.7 s | 1.5 s | 79,504 MB | -On 8× B200 the SGLang Ulysses8 path runs the paper workload at 0.93 s/NFE -(7.5 GPU-seconds per NFE) against the published 1.40 s/NFE of OpenVDN's 5+3 +On 8× B200 the SGLang Ulysses8 path runs the paper workload at 0.88 s/NFE +(7.0 GPU-seconds per NFE) against the published 1.40 s/NFE of OpenVDN's 5+3 branch-parallel layout; the whole request (text encoding, 8 forwards, joint -decode) completes in 9.3 s after warmup. At this GPU count the step is -launch- and copy-bound in the linear branch rather than FLOP-bound: the scans -fold the 102 frames into per-chunk composites plus one chain over the chunks -with both directions per launch (the chunked window only reads states at -chunk boundaries), q/k/v and the per-head scalars travel as four async -field-major all-to-alls that land contiguous (no relayout copies for the -window K/V gathers or the branch's conv), only the 128-wide output-gate -hidden crosses the fabric (the gate's `up` runs on the head shard), the frame -mean is a reshape-sum instead of an atomic index_add, and the transformer -runs online `mxfp8` by default (e4m3 with one E8M0 scale per 32 elements into -cuBLASLt's block-scaled GEMM; the activation quant is fused into the adaLN -modulation and SwiGLU producers, so no standalone quant pass runs). The -prompt's text state joins -the frames' Cholesky batch as a virtual frame. OpenVDN's branch-parallel layout was measured slower here (1.39 -s/NFE at 5+3, 1.49 at 6+2): with 11 to 12 heads per softmax rank the window -attention and its gathers grow faster than the linear ranks shrink; splitting -each rank's heads into two pipelined all-to-all groups also lost (1.10 s/NFE) -to smaller attention kernels and doubled branch launches. The profile of one -rank is 31% window FlashAttention, 24% fp8 GEMM, 12% NCCL, the rest small -kernels. -Against the published 8× B200 headline (1.40 s/NFE, 11.2 -GPU-seconds per NFE), the SGLang 4× B200 FP8 run spends 9.8 GPU-seconds per -NFE, with half the GPUs and half the all-to-all fan-out. On the same host and GPU count the -SGLang path is 10-13% faster per NFE than the released stack (2.34 vs 2.61 / -2.68 s), and the 8-forward denoise is 18.7 s against their 20.9 / 21.4 s. In BF16 the single-GPU path matches the released tuned BF16 stack -(7.86 vs 7.91 s/NFE); the single-GPU FP8 gap (7.49 vs 6.47 s) is the online -FP8 GEMM path, not the hybrid attention. +decode) completes in about 8.7 s after warmup. At this GPU count the step is +launch- and copy-bound in the linear branch rather than FLOP-bound, and the +transformer runs online `mxfp8` by default (e4m3 with one E8M0 scale per 32 +elements into cuBLASLt's block-scaled GEMM, with the activation quant fused +into the adaLN and SwiGLU producers). -On one B200 the auto memory policy streams the DiT layers from host memory -(the 62 GB DiT, the 66 GB conditioner and the VAEs do not all fit resident -with the 120 GB headroom the policy keeps); keeping the FP8 DiT resident with -`--layerwise-offload-components text_encoder` changes nothing measurable, so -the single-GPU rows are compute-bound. The BF16 DiT does not fit resident at -this clip length on one 183 GB card (the 104k-row activations alone take -about 100 GB). Per DiT block the hybrid attention costs about 117 ms at this -shape (window FlashAttention 46 ms, projections 24 ms, the linear branch and -its fused kernels the rest). A static-tile block-sparse Triton kernel was -measured at 226 ms against 169 ms for the decomposed FlashAttention path and -is not shipped. +Against the published 8× B200 headline (1.40 s/NFE, 11.2 +GPU-seconds per NFE), the SGLang 4× B200 run spends 6.7 GPU-seconds per +NFE, with half the GPUs and half the all-to-all fan-out. At equal GPU count +SGLang is 1.48× faster at two cards (3.30 vs 4.87 s/NFE), 1.55× at four +(1.67 vs 2.59) and 1.60× at eight (0.88 vs 1.40), while on one card the two +stacks are within 4% of each other (6.03 vs 6.25 s/NFE quantized, 7.45 vs 7.66 +in BF16). The advantage is therefore parallel scaling rather than per-GPU +kernel work: SGLang holds 86-91% parallel efficiency from 2 to 8 cards against +the reference stack's 56-64%, and sweeping the reference stack's own layout +does not close it (its best split is 8% better than standard Ulysses at eight +cards, 1.40 against 1.52 s/NFE). + +On one B200 both single-GPU rows run with the DiT resident under the auto +memory policy (128 GB peak quantized, 156 GB in BF16, of the card's 183 GB), +so no `--layerwise-offload-components` flag is needed at this clip length and +the single-GPU rows are compute-bound. ### VDN-H3 on RTX PRO 6000 @@ -1627,8 +1602,8 @@ measured 15.7 to 17.0 s/NFE on a shared host. | 8× RTX PRO 6000 Ulysses8 (default `mxfp8`) | 5.43 | 43.1 s | 3.5 s | 77,704 MB | | 8× RTX PRO 6000 Ulysses8, per-channel `fp8` | 5.47 | 43.4 s | 3.5 s | 77,236 MB | -For comparison the same code on B200 measures 6.23 / 3.34 / 1.73 / 0.89 s/NFE -at 1 / 2 / 4 / 8 cards (NVLink, 87% parallel efficiency at 8). +For comparison the same code on B200 measures 6.03 / 3.30 / 1.67 / 0.88 s/NFE +at 1 / 2 / 4 / 8 cards (NVLink, 86% parallel efficiency at 8). ### H200 topology comparison diff --git a/python/sglang/multimodal_gen/runtime/utils/model_overlay.py b/python/sglang/multimodal_gen/runtime/utils/model_overlay.py index 3daf4cc0d..6d957a3df 100644 --- a/python/sglang/multimodal_gen/runtime/utils/model_overlay.py +++ b/python/sglang/multimodal_gen/runtime/utils/model_overlay.py @@ -48,7 +48,7 @@ BUILTIN_MODEL_OVERLAY_REGISTRY: dict[str, dict[str, Any]] = { }, "OpenVDN/vdn-minimax-h3": { "overlay_repo_id": "kevin-mi/VDN-H3-overlay", - "overlay_revision": "0ad315a05b914c4003af4d26152d288c2506a609", + "overlay_revision": "7de18275dddfe59da36a234e222bcdd274963bc3", }, }