[diffusion] doc: document verified GB300 and derived GB200 H3 recipes (#38296)
This commit is contained in:
@@ -28,7 +28,8 @@ reload the model.
|
|||||||
The Docker form installs the platform-specific diffusion extra from the source
|
The Docker form installs the platform-specific diffusion extra from the source
|
||||||
bundled in the image. For conditioned requests, set **Host media directory**
|
bundled in the image. For conditioned requests, set **Host media directory**
|
||||||
under **Variables**; the builder mounts it read-only at `/data/minimax-h3`.
|
under **Variables**; the builder mounts it read-only at `/data/minimax-h3`.
|
||||||
AMD currently offers the Python form, while NVIDIA also offers Docker.
|
AMD and Grace Blackwell currently offer the Python form; the other NVIDIA
|
||||||
|
recipes also offer Docker.
|
||||||
|
|
||||||
To use ModelScope through the same normal `sglang serve` path, prefix the copied
|
To use ModelScope through the same normal `sglang serve` path, prefix the copied
|
||||||
command with `SGLANG_USE_MODELSCOPE=true` and replace the model path with
|
command with `SGLANG_USE_MODELSCOPE=true` and replace the model path with
|
||||||
@@ -72,6 +73,29 @@ The builder accepts legal custom GPU counts and topologies, marking them
|
|||||||
Static H3 head or partition violations disable Copy before they reach
|
Static H3 head or partition violations disable Copy before they reach
|
||||||
`sglang serve`.
|
`sglang serve`.
|
||||||
|
|
||||||
|
### Grace Blackwell
|
||||||
|
|
||||||
|
Select **GB300** or **GB200** in the builder. The default is one Linux ARM64
|
||||||
|
host with four GPUs, Ulysses4, and resident weights. This follows the
|
||||||
|
[four-GPU compute-tray layout](https://docs.nvidia.com/dgx/dgxgb200-user-guide/hardware.html);
|
||||||
|
an NVL72 rack is not a single 72-GPU process host.
|
||||||
|
|
||||||
|
For two four-GPU hosts, set **Nodes = 2** and **GPUs / node = 4**. The builder
|
||||||
|
emits Ulysses4 inside each host, Ring2 across hosts, and replicated encoders.
|
||||||
|
Set the same reachable **Head address** on both hosts and use node ranks 0 and
|
||||||
|
1. Run the generated command once per host, with the same SGLang revision and
|
||||||
|
dependencies. Let `sglang serve` resolve and download the selected weights.
|
||||||
|
|
||||||
|
GB200 is a **derived, unverified recipe**, not a measured GB300-equivalent
|
||||||
|
performance claim. Keep the native precision and platform-default attention
|
||||||
|
for the baseline; the B200/B300 quantization measurements and H200 Cache-DiT
|
||||||
|
quality audit do not establish coverage on Grace Blackwell.
|
||||||
|
|
||||||
|
GB300 end-to-end coverage currently includes single-host FL2VA text-to-video
|
||||||
|
with audio, native precision, eager execution, and `quality="lossless"`.
|
||||||
|
Ref2VA and multi-host commands remain **Unverified**. See the
|
||||||
|
[GB300 measurements](#gb300-single-host) for the tested workload.
|
||||||
|
|
||||||
### Checkpoint and adapter formats
|
### Checkpoint and adapter formats
|
||||||
|
|
||||||
Start with the command emitted by the builder. Every row below is an overlay on
|
Start with the command emitted by the builder. Every row below is an overlay on
|
||||||
@@ -1181,6 +1205,8 @@ the configurations with collected measurements:
|
|||||||
| --- | --- | --- |
|
| --- | --- | --- |
|
||||||
| B300 | 8× Ulysses8 resident | 8× FSDP + Ulysses8; the 8-GPU sweep is not a minimum-GPU claim. |
|
| B300 | 8× Ulysses8 resident | 8× FSDP + Ulysses8; the 8-GPU sweep is not a minimum-GPU claim. |
|
||||||
| B200 | 8× Ulysses8 resident | 4× FSDP + Ulysses4 |
|
| B200 | 8× Ulysses8 resident | 4× FSDP + Ulysses4 |
|
||||||
|
| GB300 | 4× Ulysses4 resident (FL2VA T2VA) | 2 nodes × 4× Ulysses4×Ring2 is unverified. |
|
||||||
|
| GB200 | 4× Ulysses4 resident, derived/unverified | No performance measurements. |
|
||||||
| H200 | 4× Ulysses4 resident | 4× FSDP + Ulysses4; 4× TP2 + Ulysses2; 2 nodes × 8× Ulysses8×Ring2 cross-node |
|
| H200 | 4× Ulysses4 resident | 4× FSDP + Ulysses4; 4× TP2 + Ulysses2; 2 nodes × 8× Ulysses8×Ring2 cross-node |
|
||||||
| H100 | 4× TP2 + Ulysses2 resident | 4× TP4 + Ulysses1; 4× FSDP + Ulysses4 |
|
| H100 | 4× TP2 + Ulysses2 resident | 4× TP4 + Ulysses1; 4× FSDP + Ulysses4 |
|
||||||
| Ascend NPU | 8 NPUs, TP2 + SP4, Laser Attention | 4 NPUs, TP2 + SP2, Laser Attention |
|
| Ascend NPU | 8 NPUs, TP2 + SP4, Laser Attention | 4 NPUs, TP2 + SP2, Laser Attention |
|
||||||
@@ -1188,6 +1214,33 @@ the configurations with collected measurements:
|
|||||||
| RTX 5090 | 2× TP2 + layerwise offload | — |
|
| RTX 5090 | 2× TP2 + layerwise offload | — |
|
||||||
| RTX 4090 24 GB | 1× layerwise offload + `kitchen_int8` | Approximate attention backends are opt-in |
|
| RTX 4090 24 GB | 1× layerwise offload + `kitchen_int8` | Approximate attention backends are opt-in |
|
||||||
|
|
||||||
|
### GB300 single-host
|
||||||
|
|
||||||
|
Four requests on **1 host × 4 NVIDIA GB300**, using Linux ARM64, SGLang main
|
||||||
|
[`15d2cbcc90fc`](https://github.com/sgl-project/sglang/commit/15d2cbcc90fc66b4d08745994cc707a45b067ffc),
|
||||||
|
PyTorch 2.13.0+cu130, and NCCL 2.29.7. The source checkout was installed with
|
||||||
|
`python3 -m pip install -e "python[diffusion]" --upgrade` inside
|
||||||
|
`lmsysorg/sglang:latest`; this is a source-validation environment, not a claim
|
||||||
|
that the image's bundled release contains that revision.
|
||||||
|
|
||||||
|
The command matches the GB300 default in the builder: Ulysses4, resident
|
||||||
|
weights, automatic encoder folding, native BF16/FP32, platform-default
|
||||||
|
Dynamic cuDNN/FA attention, and eager execution. The model was downloaded by
|
||||||
|
`sglang serve` from `MiniMaxAI/MiniMax-H3`.
|
||||||
|
|
||||||
|
The workload uses the builder's cat-band prompt, FL2VA T2VA, a 5-second request
|
||||||
|
at a 768-pixel short edge, 50 inference steps, `quality="lossless"`, one output,
|
||||||
|
and seed 1101. Every MP4 contained video and audio and passed a full decode check.
|
||||||
|
|
||||||
|
| First full request after server warmup | Following three requests | Warm median |
|
||||||
|
| ---: | --- | ---: |
|
||||||
|
| 35.16 s | 33.11 / 33.10 / 33.10 s | **33.10 s** |
|
||||||
|
|
||||||
|
These are client-observed completion times, including up to one second of
|
||||||
|
status-polling delay. Download, model loading, and server warmup are excluded.
|
||||||
|
This is a deployment smoke/latency measurement, not a cross-framework quality
|
||||||
|
audit, a minimum-GPU claim, or an extrapolation to GB200/NVL72 scaling.
|
||||||
|
|
||||||
### Ascend NPU topology comparison
|
### Ascend NPU topology comparison
|
||||||
|
|
||||||
Both topologies used Laser Attention and the explicit Cache-DiT configuration
|
Both topologies used Laser Attention and the explicit Cache-DiT configuration
|
||||||
|
|||||||
@@ -300,6 +300,33 @@ for (const path of walk(CONFIGS)) {
|
|||||||
checkH3("B200 1x8", { hw: "b200", nodes: 1, gpus_per_node: 8, placement: "resident" }, { tp_size: 1, ulysses_degree: 8, ring_degree: 1 });
|
checkH3("B200 1x8", { hw: "b200", nodes: 1, gpus_per_node: 8, placement: "resident" }, { tp_size: 1, ulysses_degree: 8, ring_degree: 1 });
|
||||||
checkH3("H100 1x4", { hw: "h100", nodes: 1, gpus_per_node: 4, placement: "resident" }, { tp_size: 2, ulysses_degree: 2, ring_degree: 1 });
|
checkH3("H100 1x4", { hw: "h100", nodes: 1, gpus_per_node: 4, placement: "resident" }, { tp_size: 2, ulysses_degree: 2, ring_degree: 1 });
|
||||||
checkH3("H200 2x8", { hw: "h200", nodes: 2, gpus_per_node: 8, placement: "resident" }, { tp_size: 1, ulysses_degree: 8, ring_degree: 2 });
|
checkH3("H200 2x8", { hw: "h200", nodes: 2, gpus_per_node: 8, placement: "resident" }, { tp_size: 1, ulysses_degree: 8, ring_degree: 2 });
|
||||||
|
for (const hw of ["gb200", "gb300"]) {
|
||||||
|
checkH3(`${hw} 1x4`, { hw, nodes: 1, gpus_per_node: 4, placement: "resident" }, { tp_size: 1, ulysses_degree: 4, ring_degree: 1 }, hw === "gb300");
|
||||||
|
checkH3(`${hw} 2x4`, { hw, nodes: 2, gpus_per_node: 4, placement: "resident" }, { tp_size: 1, ulysses_degree: 4, ring_degree: 2 }, false);
|
||||||
|
const selection = selectionOf({ hw, nodes: 2, gpus_per_node: 4 });
|
||||||
|
const encoder = config.overlayDims.find((dim) => dim.id === "encoder").options.find((option) => option.id === "auto");
|
||||||
|
if (!encoder.flags(selection).includes("--encoder-parallel replicate")) {
|
||||||
|
fail(where, `${hw} cross-node auto encoder must replicate`);
|
||||||
|
}
|
||||||
|
if (config.runModes(selection).includes("docker")) {
|
||||||
|
fail(where, `${hw} must not advertise an unvalidated Docker command`);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
for (const extra of [
|
||||||
|
{ hw: "gb200" },
|
||||||
|
{ hw: "gb300", mode: "i2va" },
|
||||||
|
{ hw: "gb300", weights: "ref2va", mode: "v2v" },
|
||||||
|
{ hw: "gb300", quality: "extra-high" },
|
||||||
|
{ hw: "gb300", outputs: "2" },
|
||||||
|
{ hw: "gb300", precision: "fp8" },
|
||||||
|
{ hw: "gb300", attention: "sage" },
|
||||||
|
]) {
|
||||||
|
const selection = selectionOf({ nodes: 1, gpus_per_node: 4, placement: "resident", ...extra });
|
||||||
|
const resolved = validateResolved(selection, "H3 Grace Blackwell coverage");
|
||||||
|
if (resolved?.builder.verification?.request !== "unverified") {
|
||||||
|
fail(where, `H3 Grace Blackwell request is outside the measured scope: ${JSON.stringify(extra)}`);
|
||||||
|
}
|
||||||
|
}
|
||||||
for (const hw of ["mi300x", "mi355x"]) {
|
for (const hw of ["mi300x", "mi355x"]) {
|
||||||
for (const count of [1, 2, 4, 8]) {
|
for (const count of [1, 2, 4, 8]) {
|
||||||
checkH3(`${hw} 1x${count}`, { hw, nodes: 1, gpus_per_node: count, placement: "resident" }, { tp_size: 1, ulysses_degree: count, ring_degree: 1 });
|
checkH3(`${hw} 1x${count}`, { hw, nodes: 1, gpus_per_node: count, placement: "resident" }, { tp_size: 1, ulysses_degree: count, ring_degree: 1 });
|
||||||
|
|||||||
@@ -147,6 +147,8 @@ return {
|
|||||||
supportedHardware: [
|
supportedHardware: [
|
||||||
"b200",
|
"b200",
|
||||||
"b300",
|
"b300",
|
||||||
|
"gb300",
|
||||||
|
"gb200",
|
||||||
"h200",
|
"h200",
|
||||||
"h100",
|
"h100",
|
||||||
"mi300x",
|
"mi300x",
|
||||||
@@ -550,6 +552,8 @@ return {
|
|||||||
{ id: "b200-fsdp-4", hw: "b200", nodes: 1, gpus_per_node: 4, placement: "fsdp", tp_size: 1, ulysses_degree: 4, ring_degree: 1, encoder: "auto" },
|
{ id: "b200-fsdp-4", hw: "b200", nodes: 1, gpus_per_node: 4, placement: "fsdp", tp_size: 1, ulysses_degree: 4, ring_degree: 1, encoder: "auto" },
|
||||||
{ id: "b300-resident-8", hw: "b300", nodes: 1, gpus_per_node: 8, placement: "resident", tp_size: 1, ulysses_degree: 8, ring_degree: 1, encoder: "auto", default: true },
|
{ id: "b300-resident-8", hw: "b300", nodes: 1, gpus_per_node: 8, placement: "resident", tp_size: 1, ulysses_degree: 8, ring_degree: 1, encoder: "auto", default: true },
|
||||||
{ id: "b300-fsdp-8", hw: "b300", nodes: 1, gpus_per_node: 8, placement: "fsdp", tp_size: 1, ulysses_degree: 8, ring_degree: 1, encoder: "auto" },
|
{ id: "b300-fsdp-8", hw: "b300", nodes: 1, gpus_per_node: 8, placement: "fsdp", tp_size: 1, ulysses_degree: 8, ring_degree: 1, encoder: "auto" },
|
||||||
|
{ id: "gb300-resident-4", hw: "gb300", nodes: 1, gpus_per_node: 4, placement: "resident", tp_size: 1, ulysses_degree: 4, ring_degree: 1, encoder: "auto", default: true },
|
||||||
|
{ id: "gb200-resident-4", hw: "gb200", nodes: 1, gpus_per_node: 4, placement: "resident", tp_size: 1, ulysses_degree: 4, ring_degree: 1, encoder: "auto", default: true, unverified: true },
|
||||||
{ id: "h200-resident-4", hw: "h200", nodes: 1, gpus_per_node: 4, placement: "resident", tp_size: 1, ulysses_degree: 4, ring_degree: 1, encoder: "auto", default: true },
|
{ id: "h200-resident-4", hw: "h200", nodes: 1, gpus_per_node: 4, placement: "resident", tp_size: 1, ulysses_degree: 4, ring_degree: 1, encoder: "auto", default: true },
|
||||||
{ id: "h200-fsdp-4", hw: "h200", nodes: 1, gpus_per_node: 4, placement: "fsdp", tp_size: 1, ulysses_degree: 4, ring_degree: 1, encoder: "auto" },
|
{ id: "h200-fsdp-4", hw: "h200", nodes: 1, gpus_per_node: 4, placement: "fsdp", tp_size: 1, ulysses_degree: 4, ring_degree: 1, encoder: "auto" },
|
||||||
{ id: "h200-cross-node-16", hw: "h200", nodes: 2, gpus_per_node: 8, placement: "resident", tp_size: 1, ulysses_degree: 8, ring_degree: 2, encoder: "replicate" },
|
{ id: "h200-cross-node-16", hw: "h200", nodes: 2, gpus_per_node: 8, placement: "resident", tp_size: 1, ulysses_degree: 8, ring_degree: 2, encoder: "replicate" },
|
||||||
@@ -673,9 +677,12 @@ return {
|
|||||||
|| (s.precision === "fp8" && ["b200", "b300"].includes(s.hw));
|
|| (s.precision === "fp8" && ["b200", "b300"].includes(s.hw));
|
||||||
const executionVerified = s.execution === "eager"
|
const executionVerified = s.execution === "eager"
|
||||||
|| (s.execution === "bcg" && ["b200", "h200"].includes(s.hw) && s.weights === "ref2va");
|
|| (s.execution === "bcg" && ["b200", "h200"].includes(s.hw) && s.weights === "ref2va");
|
||||||
|
const checkpointVerified = s.hw !== "gb300" || s.weights === "fl2va";
|
||||||
const serveVerified = topologyVerified && encoderVerified && attentionVerified
|
const serveVerified = topologyVerified && encoderVerified && attentionVerified
|
||||||
&& precisionVerified && executionVerified;
|
&& precisionVerified && executionVerified && checkpointVerified;
|
||||||
const requestVerified = topologyVerified && (["lossless", "extra-high"].includes(s.quality)
|
const requestCovered = s.hw !== "gb300" || (serveVerified && s.weights === "fl2va"
|
||||||
|
&& s.mode === "t2va" && s.quality === "lossless" && Number(s.outputs) === 1);
|
||||||
|
const requestVerified = topologyVerified && requestCovered && (["lossless", "extra-high"].includes(s.quality)
|
||||||
|| (s.quality === "high" && highAudited && s.execution === "eager"));
|
|| (s.quality === "high" && highAudited && s.execution === "eager"));
|
||||||
|
|
||||||
const topologyParts = [];
|
const topologyParts = [];
|
||||||
@@ -712,7 +719,7 @@ return {
|
|||||||
let automaticAttention = "FlashAttention (auto)";
|
let automaticAttention = "FlashAttention (auto)";
|
||||||
if (["mi300x", "mi355x"].includes(s.hw)) {
|
if (["mi300x", "mi355x"].includes(s.hw)) {
|
||||||
automaticAttention = "AITER (auto)";
|
automaticAttention = "AITER (auto)";
|
||||||
} else if (topology.ring_degree === 1 && ["b200", "b300"].includes(s.hw)) {
|
} else if (topology.ring_degree === 1 && ["b200", "b300", "gb200", "gb300"].includes(s.hw)) {
|
||||||
automaticAttention = "Dynamic cuDNN / FA (auto)";
|
automaticAttention = "Dynamic cuDNN / FA (auto)";
|
||||||
} else if (topology.ring_degree === 1 && ["rtx5090", "rtx4090"].includes(s.hw)) {
|
} else if (topology.ring_degree === 1 && ["rtx5090", "rtx4090"].includes(s.hw)) {
|
||||||
automaticAttention = "Torch SDPA (auto)";
|
automaticAttention = "Torch SDPA (auto)";
|
||||||
@@ -964,9 +971,9 @@ return {
|
|||||||
? `bash -lc 'python -m pip install -e "/sgl-workspace/sglang/python[diffusion_hip]" && exec sglang serve "$@"' --`
|
? `bash -lc 'python -m pip install -e "/sgl-workspace/sglang/python[diffusion_hip]" && exec sglang serve "$@"' --`
|
||||||
: `bash -lc 'python -m pip install -e "/sgl-workspace/sglang/python[diffusion]" && exec sglang serve "$@"' --`,
|
: `bash -lc 'python -m pip install -e "/sgl-workspace/sglang/python[diffusion]" && exec sglang serve "$@"' --`,
|
||||||
|
|
||||||
// Publish AMD Docker only after an H3-capable ROCm image has been validated.
|
// Publish Docker only after the platform's H3 image/command has been validated.
|
||||||
runModes: (s) =>
|
runModes: (s) =>
|
||||||
["mi300x", "mi355x"].includes(s.hw)
|
["mi300x", "mi355x", "gb200", "gb300"].includes(s.hw)
|
||||||
? ["python"]
|
? ["python"]
|
||||||
: ["python", "docker"],
|
: ["python", "docker"],
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user