[docs] Re-measure the Qwen3.8-27B RTX 5090, RTX PRO 6000 and DGX Spark grids on 1cf2b8c (#35825)

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Jimmy Shong
2026-08-22 13:03:18 +08:00
committed by GitHub
co-authored by Claude Fable 5
parent 96bfd2476c
commit 4cb5aebfe0
2 changed files with 118 additions and 149 deletions
@@ -19,14 +19,10 @@ For all methods and hardware platforms, see the [official SGLang installation gu
```bash Command
pip install --upgrade pip
pip install uv
uv pip install --prerelease=allow sglang
# For the DFLASH2 cells only — DFlash2 selector support is newer than the
# latest release, so build from the commit those cells were validated on
# instead of the line above:
# git clone https://github.com/sgl-project/sglang.git && cd sglang
# git checkout 1cf2b8c54d81802abc15dcf23a29b9cc687bc01e # PR #35496
# uv pip install --prerelease=allow -e "python[all]"
git clone https://github.com/sgl-project/sglang.git && cd sglang
git checkout 1cf2b8c54d81802abc15dcf23a29b9cc687bc01e
uv pip install --prerelease=allow -e "python[all]"
```
Then run the **Python** output of the command panel below in that environment.
@@ -36,13 +32,7 @@ Then run the **Python** output of the command panel below in that environment.
<Tab title="Docker">
```bash Command
docker pull lmsysorg/sglang:qwen38-27b
# For the DFLASH2 cells only — that tag predates DFlash2 selector support.
# Build the image from the commit those cells were validated on instead:
# git clone https://github.com/sgl-project/sglang.git && cd sglang
# git checkout 1cf2b8c54d81802abc15dcf23a29b9cc687bc01e # PR #35496
# docker build -t sglang:dflash2 -f docker/Dockerfile .
docker pull lmsysorg/sglang:dev-qwen38-27b-dflash2
```
For how to launch the image, see [Install → Method 3: Using Docker](../../../docs/get-started/install#method-3-using-docker). Substitute the inner `sglang serve ...` with what the command generator below produces.
@@ -51,19 +41,6 @@ For how to launch the image, see [Install → Method 3: Using Docker](../../../d
</Tabs>
<Warning>
**DFLASH2 needs a build from `1cf2b8c` (PR
[#35496](https://github.com/sgl-project/sglang/pull/35496)) or newer.** DFlash2
landed in [#35371](https://github.com/sgl-project/sglang/pull/35371) and its
quantized-`lm_head` path — what the NVFP4 cells need — in #35496, both newer
than the pinned `lmsysorg/sglang:qwen38-27b` tag and than the latest PyPI
release. A build without #35496 fails on the NVFP4 cells at boot with `requires
a dense FP16/BF16/FP32 target lm_head`. The DFLASH2 pins on this page were
measured on `1cf2b8c` exactly, which is why both install paths above check that
commit out rather than tracking a moving branch. Every other recipe on this
page — no speculation, MTP, DSpark — runs on the pinned tag as written.
</Warning>
</Accordion>
Pick your card + checkpoint precision to generate the launch command. The model runs single-GPU on every supported card — H200, RTX PRO 6000, RTX 5090 and DGX Spark — and ships one operating point.
@@ -82,34 +59,15 @@ import { Qwen38MambaRatioCalculator } from "/src/snippets/_qwen38_mamba_ratio_ca
<Deployment config={config} />
<Warning>
**Picking DFLASH2 requires SGLang built from main.** DFlash2 landed in
[#35371](https://github.com/sgl-project/sglang/pull/35371), and DFlash2 +
NVFP4 — the quantized `lm_head` path — in
[#35496](https://github.com/sgl-project/sglang/pull/35496). Both are newer
than the `lmsysorg/sglang:qwen38-27b` tag pinned in
[Install SGLang](#install) and than the latest PyPI release; on NVFP4 a build
without #35496 fails at boot with `requires a dense FP16/BF16/FP32 target
lm_head`. Every other Speculative Decoding pick — None, EAGLE, DSPARK — runs
on the pinned tag as written.
</Warning>
<Note>
The RTX 5090 and RTX PRO 6000 cells above — including every Speculative
Decoding / Serving Strategy / SSM dtype combination — were validated at
ISL 8192 / OSL 1024, concurrency 1; for DFLASH2, to that full standard on
NVFP4 and to boot-and-serve on the RTX PRO 6000 BF16/FP8 cells. The DGX Spark
cells cover the same combination set minus DFLASH2, but to a weaker standard:
each was confirmed to **boot and serve** at ISL 8192 / OSL 1024,
concurrency 1, with no throughput or acceptance-length numbers taken. The
H200 and GB300 recipes carry their original validation, which covers the
default overlay picks (plus MTP on GB300); non-default overlay picks there
are valid but unmeasured.
DFLASH2 has not been exercised on H200, DGX Spark or GB300: selecting it
there switches the cell's badge to **Final Verification In Progress** until
that round lands. The other picks on those platforms keep the validation
described above.
cells cover the full combination set DFLASH2 included — re-measured end to
end on `1cf2b8c`, but to a weaker standard: each of the 48 was confirmed to
**boot and serve** at ISL 8192 / OSL 1024, concurrency 1, with no throughput
or acceptance-length numbers taken.
</Note>
### Mamba ratio calculator
@@ -228,17 +186,26 @@ checkpoint's calibration scales automatically.
requires a FlashInfer build whose prefill `plan` accepts `uniform_q_len`
(newer than 0.6.15.post1); otherwise run spec with `--attention-backend triton`.
On DGX Spark the 128GB is unified memory shared with the host CPU, so all
three checkpoints fit, and its cells reuse the RTX PRO 6000 recipe verbatim
rather than a separate operating point. **Validated on SM121 / aarch64**: all
36 configurations (3 checkpoints x Speculative Decoding x Serving Strategy x
Mamba SSM Dtype) booted and served on GB10 under `lmsysorg/sglang:qwen38-27b`
at ISL 8192 / OSL 1024, concurrency 1. That is boot-and-serve coverage only —
no throughput or acceptance-length numbers — and it includes the FlashInfer `plan` /
`uniform_q_len` path above, which raised no arity error on that image. Two
host quirks when reproducing on GB10: docker GPU access is CDI-only
(`--device nvidia.com/gpu=all`, as no `nvidia` runtime is registered), and
`nvidia-smi` reports `Not Supported` for memory because it is unified with the
CPU — gate a relaunch on `MemAvailable` in `/proc/meminfo` instead.
three checkpoints fit, and its cells reuse the RTX PRO 6000 recipe at
`--mem-fraction-static 0.80` rather than a separate operating point. The one
lower pin is the unified pool pricing the host's memory too: 0.85 of 128GB
leaves ~8GB for the OS — exactly DGX OS earlyoom's SIGTERM threshold — and
the first long prefill or boot-time graph capture dips under it and gets the
scheduler killed with `exit code -15` and no traceback (`journalctl -u
earlyoom` shows the kill). At 0.85, 15 of the 48 cells were killed that way,
and which cells is margin noise; at 0.80 every cell served on every attempt.
**Validated on SM121 / aarch64**: all 48 configurations (3 checkpoints x
Speculative Decoding x Serving Strategy x Mamba SSM Dtype, DFLASH2 included)
booted and served on GB10 at `1cf2b8c` at ISL 8192 / OSL 1024, concurrency 1.
That is boot-and-serve coverage only — no throughput or acceptance-length
numbers — and it includes the FlashInfer `plan` / `uniform_q_len` path above,
which raised no arity error on that build. Three host quirks when reproducing
on GB10: docker GPU access is CDI-only (`--device nvidia.com/gpu=all`, as no
`nvidia` runtime is registered); `nvidia-smi` reports `Not Supported` for
memory because it is unified with the CPU — gate a relaunch on `MemAvailable`
in `/proc/meminfo` instead; and the BF16 checkpoint takes ~6.5 minutes just
to load its 18 shards from NVMe, so budget ~10 minutes to READY before
calling a boot hung.
- **H200 (SM90)**: BF16 and FP8 only — the card has no FP4 tensor cores, so the
NVFP4 checkpoint's MLP would fall back to the Marlin W4A16 weight-only path
and its cell is greyed out. The H200 recipes use 32768-token prefill chunks
@@ -274,21 +241,23 @@ checkpoint's calibration scales automatically.
including quantized heads, so it runs on the NVFP4 checkpoint (whose head
is NVFP4-packed; the BF16 and FP8 checkpoints keep a dense head).
Validation: NVFP4 measured end-to-end on RTX PRO 6000 and RTX 5090; the
RTX PRO 6000 BF16/FP8 cells boot and serve; on H200, DGX Spark and GB300
those cells carry the **Final Verification In Progress**
badge. The
RTX PRO 6000 recipe needs no changes. On the 32GB RTX 5090 the pins differ per
state dtype, and the panel applies them automatically. bfloat16 serves at
`--mem-fraction-static 0.88` on the balanced ratio (0.90, DSpark's pin, OOMs
on the first request). float32 reaches only the High-Throughput tier, at
`--mem-fraction-static 0.895` with `--mamba-full-memory-ratio 10` overriding
the balanced value — these cells pin `--max-running-requests 1`, so the
balanced ratio provisions KV for concurrency the recipe never uses and
starves the state pool of the slots fp32 needs. The Low-Latency tier is
greyed out for float32: it needs five fp32 slots plus a full request's KV,
and no mem-fraction holds both — buying the fifth slot cuts KV to 7,752
tokens against the 9,216 one 8192/1024 request needs, while every
mem-fraction with a large enough pool dies in prefill graph capture.
RTX PRO 6000 BF16/FP8 cells boot and serve; all 12 DGX Spark DFLASH2 cells
boot and serve on `1cf2b8c` with the selector folded into the draft CUDA
graph; on H200 and GB300 those cells carry the
**Final Verification In Progress** badge. The
RTX PRO 6000 recipe needs no changes. On the 32GB RTX 5090 every pin is
re-measured against that commit, and the panel applies them automatically:
DFlash2 at `--mem-fraction-static 0.91` with `--chunked-prefill-size 1024` —
the only cell on this page needing a smaller prefill chunk, because at 0.91
the pools fit but a 2048-token chunk's activations do not — DSpark at 0.88,
EAGLE at 0.93 (bfloat16) and 0.94 (float32), and no-speculation at 0.90.
float32 is greyed out for both draft-model picks on that card: an fp32 state
slot costs 154 MB against bfloat16's 78 MB, so below ~0.92 the state pool
never reaches the tier's slot count, while at or above it prefill graph
capture or the first request runs out of memory. That was verified across
0.860.96 at both chunk sizes and with balanced-ratio overrides up to 20.
bfloat16 is also the faster choice: DFlash2 posts 4.92 ms median TPOT at an
accept length of 4.29, the best result on this card.
- **Hardware fit**: FP8 weights ~28.5GB (not serviceable beyond bs≤2 on
32GB cards); NVFP4 weights ~16.5GB (recommended for RTX 5090-class GPUs).
- `--mamba-radix-cache-strategy extra_buffer_lazy` lowers the state cost per
+73 -73
View File
@@ -83,12 +83,14 @@ export const config = {
...(["rtx5090", "rtx6000", "dgx-spark"].includes(sel.hw)
? ["--enable-linear-replayssm-spec"]
: []),
// Measured on the 5090: bf16 state serves at 0.92, fp32 needs
// 0.94 (an fp32 slot is 146.81 MiB vs bf16's 74.81).
// Measured on the 5090 at commit 1cf2b8c: fp32 serves at 0.94,
// bf16 at 0.93. bf16 moved UP from 0.92 with the dense-lm_head
// checkpoint -- the heavier weights need a larger static budget
// before the state pool fits.
...(sel.hw === "rtx5090"
? [sel.ssmDtype === "float32"
? "--mem-fraction-static 0.94"
: "--mem-fraction-static 0.92"]
: "--mem-fraction-static 0.93"]
: []),
],
},
@@ -107,13 +109,13 @@ export const config = {
"--speculative-algorithm DSPARK",
"--speculative-draft-model-path RadixArk/Qwen3.8-27B-DSpark",
"--speculative-draft-attention-backend flashinfer",
// Measured on the 5090: bf16 state serves at 0.90, fp32 needs
// 0.92 — the opposite correction to EAGLE's (see above).
...(sel.hw === "rtx5090"
? [sel.ssmDtype === "float32"
? "--mem-fraction-static 0.92"
: "--mem-fraction-static 0.90"]
: []),
// Measured on the 5090 at commit 1cf2b8c: bf16 serves at 0.88,
// below the 0.90 this recipe carried when it was measured on an
// older build, because a draft model plus the automatic prefill
// CUDA-graph capture no longer fit there. fp32 is greyed out by the
// SSM dtype row. EAGLE and no-speculation are unaffected: replayssm
// keeps EAGLE's state pool tiny and no-spec loads no draft weights.
...(sel.hw === "rtx5090" ? ["--mem-fraction-static 0.88"] : []),
],
},
{
@@ -128,31 +130,21 @@ export const config = {
disabled: (sel) => sel.hw === "rtx5090" && sel.quant !== "nvfp4",
disableReason:
"On the 32GB RTX 5090 the DFlash2 draft model only fits on top of the NVFP4 weights",
// fp32 is the one case that needs the balanced ratio overridden, so
// that family is stripped too and re-emitted below.
stripPrefixes: (sel) =>
sel.hw === "rtx5090"
? sel.ssmDtype === "float32"
? ["--mem-fraction-static", "--mamba-full-memory-ratio"]
: ["--mem-fraction-static"]
: [],
sel.hw === "rtx5090" ? ["--mem-fraction-static"] : [],
flags: (sel) => [
"--speculative-algorithm DFLASH",
"--speculative-draft-model-path incoai/Qwen3.8-27B-DFlash2",
"--speculative-num-draft-tokens 8",
// Measured on the 5090 at commit 1cf2b8c, the build the Install
// accordion pins for this pick. bf16 serves at 0.88 on the balanced
// ratio (0.90, DSPARK's pin, OOMs on the first request). fp32 needs
// 0.895 AND the balanced ratio overridden to 10: these cells pin
// --max-running-requests 1, so the balanced value provisions KV for
// concurrency this recipe never uses, starving the state pool of the
// slots fp32 needs. Only High-Throughput reaches fp32 (S=4); the SSM
// dtype row greys fp32 out for Low-Latency (S=5).
// accordion pins. This is the only cell on the page that also needs
// a prefill chunk smaller than the engine default: at 0.91 the pools
// fit but a 2048-token chunk's activations do not. The pair together
// is the fastest recipe on this card (4.92ms median TPOT, 4.29
// accept length). fp32 is greyed out by the SSM dtype row.
...(sel.hw === "rtx5090"
? sel.ssmDtype === "float32"
? ["--mem-fraction-static 0.895",
"--mamba-full-memory-ratio 10"]
: ["--mem-fraction-static 0.88"]
? ["--mem-fraction-static 0.91",
"--chunked-prefill-size 1024"]
: []),
],
},
@@ -198,21 +190,22 @@ export const config = {
// serves, and is the faster cell there anyway.
{
id: "float32", label: "float32",
// Only the Low-Latency tier is out of reach: it needs S=5 fp32 slots
// (735MB) plus >=9216 KV tokens for one request, and no mem-fraction
// holds both -- at 0.8975/r14 the pool buys the 5th slot but KV falls
// to 7752 tokens and generation stops after one token, while every
// mem-fraction with a big enough pool (>=0.90) dies in graph capture.
// High-Throughput needs one slot fewer and does fit; see the DFLASH2
// option's pins.
// On the 32GB RTX 5090 a draft model plus an fp32 state pool no longer
// fit together: the checkpoint's dense lm_head adds ~3.2GB of weights,
// which pushes the pools up into the mem-fraction range where prefill
// CUDA-graph capture no longer fits. Measured across 0.86-0.96 at both
// prefill chunk sizes, plus balanced-ratio overrides up to 20: below
// ~0.92 the state pool never reaches the tier's slot count, and at or
// above it capture or the first request OOMs. An fp32 slot is 154MB
// against bfloat16's 78MB, which is why only fp32 is caught. EAGLE and
// no-speculation are unaffected -- replayssm keeps EAGLE's pool tiny
// and no-spec loads no draft weights at all.
disabled: (sel) =>
sel.hw === "rtx5090" &&
sel.spec === "dflash" &&
sel.tier === "low-latency",
(sel.spec === "dflash" || sel.spec === "dspark"),
disableReason:
"On the 32GB RTX 5090 the Low-Latency tier cannot hold five fp32 state " +
"slots and a full request's KV at once — use bfloat16, or the " +
"High-Throughput tier which fits fp32",
"On the 32GB RTX 5090 an fp32 GDN state pool and a speculative draft model " +
"do not fit together — use bfloat16",
flags: ["--mamba-ssm-dtype float32"],
},
{
@@ -274,11 +267,14 @@ export const config = {
dockerImages: {
h200: "lmsysorg/sglang:qwen38-27b",
rtx6000: "lmsysorg/sglang:qwen38-27b",
rtx5090: "lmsysorg/sglang:qwen38-27b",
// Both SM120 cards are validated on this image (built from 1cf2b8c, the
// commit every pin on those cards was measured against).
rtx6000: "lmsysorg/sglang:dev-qwen38-27b-dflash2",
rtx5090: "lmsysorg/sglang:dev-qwen38-27b-dflash2",
// Multi-arch: this tag ships both linux/amd64 and linux/arm64, so it pulls
// natively on DGX Spark (GB10 is aarch64).
"dgx-spark": "lmsysorg/sglang:qwen38-27b",
// Multi-arch (linux/amd64 + linux/arm64), so GB10 pulls it natively.
"dgx-spark": "lmsysorg/sglang:dev-qwen38-27b-dflash2",
gb300: "lmsysorg/sglang:dev",
},
@@ -558,32 +554,41 @@ export const config = {
},
// DGX Spark (GB10, SM121): single node, 128GB coherent unified memory
// shared with the CPU — every checkpoint fits, so all three quants get a
// cell. These cells reuse the RTX PRO 6000 recipe verbatim rather than a
// separate SM121 operating point: both cards are SM12x Blackwell, and
// GB10's 128GB unified pool is larger than the 6000's 96GB, so a recipe
// that fits the smaller card has headroom here.
// cell. These cells reuse the RTX PRO 6000 recipe at one lower
// mem-fraction rather than a separate SM121 operating point: both cards
// are SM12x Blackwell, and GB10's 128GB unified pool is larger than the
// 6000's 96GB, so a recipe that fits the smaller card has headroom here.
//
// Validated on GB10 (SM121 / aarch64): all 36 configurations booted and
// served at ISL 8192 / OSL 1024, concurrency 1. Boot-and-serve only -- no
// throughput or acceptance-length numbers were taken, so this is a weaker
// standard than the SM120 pair's validation, and the Deploy-panel Note says
// so.
// Why 0.80 and not the 0.85 every other SM12x cell pins: the pool is
// unified, so mem-fraction prices the HOST's memory too. 0.85 of 128GB
// leaves ~8GB for the OS — exactly DGX OS earlyoom's SIGTERM threshold —
// and the first long prefill or boot-time graph capture dips under it and
// gets the scheduler killed (exit code -15, no traceback; check
// `journalctl -u earlyoom`). Re-measured on 1cf2b8c (2026-08-21): at 0.85,
// 15 of 48 cells were SIGTERMed, and which 15 is margin noise, biased
// toward the big-state configs (bfloat16 SSM, DSPARK/DFLASH2 ratios); at
// 0.80 every cell served on every attempt.
//
// Validated on GB10 (SM121 / aarch64) at 1cf2b8c: all 48 configurations —
// DFLASH2 included — booted and served at ISL 8192 / OSL 1024,
// concurrency 1. Boot-and-serve only -- no throughput or acceptance-length
// numbers were taken, so this is a weaker standard than the SM120 pair's
// validation, and the Deploy-panel Note says so. NVFP4 was exercised with
// the BF16-LMHead export (same as the SM120 re-measurement); the
// packed-head export also served its DFLASH2 cells on this platform in the
// 12-cell DFLASH2 pass.
{
match: { hw: "dgx-spark", variant: "default", quant: "nvfp4", nodes: "single" },
// All 12 overlay combinations served on GB10. DSPARK here also
// exercises the 4-bit `lm_head` this checkpoint quantizes, with no shape
// error.
// All 16 overlay combinations served on GB10 at 1cf2b8c, DFLASH2
// included — its selector folded into the draft CUDA graph in all four
// of its cells here.
verified: true,
// DFLASH2 has not been exercised on this platform; every other overlay
// pick keeps this cell's original validation.
verificationStatus: (sel) =>
sel.spec === "dflash" ? "in-progress" : "verified",
env: [],
flags: [
"--trust-remote-code",
"--model-path {{MODEL_NAME}}",
"--kv-cache-dtype fp8_e4m3",
"--mem-fraction-static 0.85",
"--mem-fraction-static 0.80",
"--attention-backend flashinfer",
"--chunked-prefill-size 2048",
"--reasoning-parser qwen3",
@@ -594,18 +599,16 @@ export const config = {
},
{
match: { hw: "dgx-spark", variant: "default", quant: "fp8", nodes: "single" },
// All 12 overlay combinations served on GB10.
// All 16 overlay combinations served on GB10 at 1cf2b8c, DFLASH2
// included. This checkpoint held the sweep's most earlyoom-prone cells
// at 0.85 (every bfloat16-SSM pick was killed); all clean at 0.80.
verified: true,
// DFLASH2 has not been exercised on this platform; every other overlay
// pick keeps this cell's original validation.
verificationStatus: (sel) =>
sel.spec === "dflash" ? "in-progress" : "verified",
env: [],
flags: [
"--trust-remote-code",
"--model-path {{MODEL_NAME}}",
"--kv-cache-dtype fp8_e4m3",
"--mem-fraction-static 0.85",
"--mem-fraction-static 0.80",
"--attention-backend flashinfer",
"--chunked-prefill-size 2048",
"--reasoning-parser qwen3",
@@ -616,19 +619,16 @@ export const config = {
},
{
match: { hw: "dgx-spark", variant: "default", quant: "bf16", nodes: "single" },
// All 12 overlay combinations served on GB10. Heaviest checkpoint, so
// it holds the sweep's tightest cell: DSPARK + float32 + extra_buffer.
// All 16 overlay combinations served on GB10 at 1cf2b8c, DFLASH2
// included. Heaviest checkpoint (52GB, ~6.5 min to load its 18 shards
// from NVMe — budget ~10 min to READY before calling a boot hung).
verified: true,
// DFLASH2 has not been exercised on this platform; every other overlay
// pick keeps this cell's original validation.
verificationStatus: (sel) =>
sel.spec === "dflash" ? "in-progress" : "verified",
env: [],
flags: [
"--trust-remote-code",
"--model-path {{MODEL_NAME}}",
"--kv-cache-dtype fp8_e4m3",
"--mem-fraction-static 0.85",
"--mem-fraction-static 0.80",
"--attention-backend flashinfer",
"--chunked-prefill-size 2048",
"--reasoning-parser qwen3",