[diffusion] docs: sync snapshot and minimax-h3 subblock features (#38784)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
This commit is contained in:
@@ -1187,6 +1187,59 @@ benchmark latency as well as visual and audio quality on the target workload.
|
|||||||
|
|
||||||
</Tab>
|
</Tab>
|
||||||
|
|
||||||
|
<Tab title="SubBlock sparse attention">
|
||||||
|
|
||||||
|
SubBlock sparse attention applies training-free block sparsity to MiniMax-H3's
|
||||||
|
long, non-causal DiT self-attention. The BF16 path supports CUDA SM90, SM100,
|
||||||
|
and SM120. Other compute capabilities, including SM103/B300, are rejected.
|
||||||
|
Ulysses sequence parallelism is supported; Ring parallelism is not.
|
||||||
|
|
||||||
|
The global backend selects the H3 DiT lazily, so keep the Qwen text encoder on
|
||||||
|
its compatible dense backend with a component override:
|
||||||
|
|
||||||
|
```bash Command
|
||||||
|
sglang serve \
|
||||||
|
--model-path MiniMaxAI/MiniMax-H3 \
|
||||||
|
--model-variant fl2va \
|
||||||
|
--num-gpus 8 \
|
||||||
|
--ulysses-degree 8 \
|
||||||
|
--attention-backend subblock_sparse_attn \
|
||||||
|
--component-attention-backends text_encoder=fa \
|
||||||
|
--attention-backend-config '{"sparsity": 0.75, "skip_first_steps": 10}' \
|
||||||
|
--port 30010
|
||||||
|
```
|
||||||
|
|
||||||
|
Use `text_encoder=torch_sdpa` on SM120. The default configuration drops at
|
||||||
|
most 75% of key blocks after the first 10 denoise forwards; short sequences,
|
||||||
|
cross-attention, and unsupported shapes continue on the dense path. Tune
|
||||||
|
`sparsity` and `skip_first_steps` together with visual and audio quality checks.
|
||||||
|
|
||||||
|
On H100/H200 (SM90), `compute_mode=sage_fp8` switches the sparse kernel to
|
||||||
|
online INT8 Q/K and FP8 P/V compute. Install the optional kernel and add the
|
||||||
|
compute mode:
|
||||||
|
|
||||||
|
```bash Command
|
||||||
|
pip install git+https://github.com/thu-ml/SpargeAttn.git --no-build-isolation
|
||||||
|
|
||||||
|
sglang serve \
|
||||||
|
--model-path MiniMaxAI/MiniMax-H3 \
|
||||||
|
--model-variant fl2va \
|
||||||
|
--num-gpus 8 \
|
||||||
|
--ulysses-degree 8 \
|
||||||
|
--attention-backend subblock_sparse_attn \
|
||||||
|
--component-attention-backends text_encoder=fa \
|
||||||
|
--attention-backend-config '{"compute_mode": "sage_fp8", "sparsity": 0.75}' \
|
||||||
|
--port 30010
|
||||||
|
```
|
||||||
|
|
||||||
|
<Warning>
|
||||||
|
Both SubBlock routing and `sage_fp8` are approximate. Use dense BF16 attention
|
||||||
|
for consistency ground truth, and validate the selected sparsity on the target
|
||||||
|
resolution, duration, task, and checkpoint.
|
||||||
|
</Warning>
|
||||||
|
|
||||||
|
</Tab>
|
||||||
|
|
||||||
<Tab title="Encoder scheduling">
|
<Tab title="Encoder scheduling">
|
||||||
|
|
||||||
The picker explicitly writes `--encoder-parallel auto` in every single-node
|
The picker explicitly writes `--encoder-parallel auto` in every single-node
|
||||||
|
|||||||
@@ -79,6 +79,11 @@ For SGLang-native pipelines, the CLI accepts the lowercase names of `AttentionBa
|
|||||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)", whiteSpace: "nowrap"}}>`VIDEO_SPARSE_ATTN`</td>
|
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)", whiteSpace: "nowrap"}}>`VIDEO_SPARSE_ATTN`</td>
|
||||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>Requires <code>vsa</code>. Configure <code>sparsity</code> via <code>--attention-backend-config</code>.</td>
|
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>Requires <code>vsa</code>. Configure <code>sparsity</code> via <code>--attention-backend-config</code>.</td>
|
||||||
</tr>
|
</tr>
|
||||||
|
<tr>
|
||||||
|
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>`subblock_sparse_attn`</td>
|
||||||
|
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)", whiteSpace: "nowrap"}}>`SUBBLOCK_SPARSE_ATTN`</td>
|
||||||
|
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>Training-free SubBlock sparsity for MiniMax-H3 on CUDA SM90, SM100, and SM120. The default compute mode is BF16; <code>compute_mode=sage_fp8</code> is an approximate SM90-only path that requires <code>SpargeAttn</code>. See the <a href="/cookbook/diffusion/MiniMax/MiniMax-H3#subblock-sparse-attention">MiniMax-H3 recipe</a>.</td>
|
||||||
|
</tr>
|
||||||
<tr>
|
<tr>
|
||||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>`video_sparse_attn_h3`</td>
|
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>`video_sparse_attn_h3`</td>
|
||||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)", whiteSpace: "nowrap"}}>`VIDEO_SPARSE_ATTN_H3`</td>
|
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)", whiteSpace: "nowrap"}}>`VIDEO_SPARSE_ATTN_H3`</td>
|
||||||
|
|||||||
@@ -84,6 +84,8 @@ loaded component:
|
|||||||
|
|
||||||
- `resident` keeps the complete component on the accelerator.
|
- `resident` keeps the complete component on the accelerator.
|
||||||
- `component-offload` stores the complete component on CPU between uses.
|
- `component-offload` stores the complete component on CPU between uses.
|
||||||
|
- `snapshot-offload` keeps a CPU weight snapshot while the complete component
|
||||||
|
runs on the accelerator, avoiding a weight copy back to CPU after each use.
|
||||||
- `layerwise-offload` streams the component's declared layers from CPU.
|
- `layerwise-offload` streams the component's declared layers from CPU.
|
||||||
|
|
||||||
`COMPONENT` can be an exact `model_index.json` key or one of `all`, `dit`,
|
`COMPONENT` can be an exact `model_index.json` key or one of `all`, `dit`,
|
||||||
|
|||||||
Reference in New Issue
Block a user