[docs] Add the NVIDIA NVFP4 export to the Qwen3.8-27B cookbook (#38611)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Jiminator <jimmysh341@gmail.com>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
This commit is contained in:
zijiexia
2026-09-11 11:58:30 -07:00
committed by GitHub
co-authored by Claude Opus 5 Jiminator Jimmy Shong
parent d7c284b894
commit df6424967a
3 changed files with 283 additions and 96 deletions
@@ -19,9 +19,7 @@ For all methods and hardware platforms, see the [official SGLang installation gu
pip install --upgrade pip pip install --upgrade pip
pip install uv pip install uv
git clone https://github.com/sgl-project/sglang.git && cd sglang uv pip install --prerelease=allow sglang
git checkout 1cf2b8c54d81802abc15dcf23a29b9cc687bc01e
uv pip install --prerelease=allow -e "python[all]"
``` ```
Then run the **Python** output of the command panel below in that environment. Then run the **Python** output of the command panel below in that environment.
@@ -31,7 +29,7 @@ Then run the **Python** output of the command panel below in that environment.
<Tab title="Docker"> <Tab title="Docker">
```bash Command ```bash Command
docker pull lmsysorg/sglang:dev-qwen38-27b-dflash2 docker pull lmsysorg/sglang:latest
``` ```
For how to launch the image, see [Install → Method 3: Using Docker](../../../docs/get-started/install#method-3-using-docker). Substitute the inner `sglang serve ...` with what the command generator below produces. For how to launch the image, see [Install → Method 3: Using Docker](../../../docs/get-started/install#method-3-using-docker). Substitute the inner `sglang serve ...` with what the command generator below produces.
@@ -59,14 +57,12 @@ import { Qwen38MambaRatioCalculator } from "/src/snippets/_qwen38_mamba_ratio_ca
<Deployment config={config} /> <Deployment config={config} />
<Note> <Note>
The RTX 5090 and RTX PRO 6000 cells above — including every Speculative Every cell above — RTX 5090, RTX PRO 6000 and DGX Spark, across all five
Decoding / Serving Strategy / SSM dtype combination — were validated at checkpoints and every Speculative Decoding / Serving Strategy / SSM dtype
ISL 8192 / OSL 1024, concurrency 1; for DFLASH2, to that full standard on combination — is measured on **v0.5.19**. That is 202 cells, each one served
NVFP4 and to boot-and-serve on the RTX PRO 6000 BF16/FP8 cells. The DGX Spark and scored on the full 1319-question GSM8K (93.18-95.15%). The serving
cells cover the full combination set — DFLASH2 included — re-measured end to envelope behind the pins is ISL 8192 / OSL 1024 at concurrency 1; throughput
end on `1cf2b8c`, but to a weaker standard: each of the 48 was confirmed to and acceptance-length numbers were not re-taken in that sweep.
**boot and serve** at ISL 8192 / OSL 1024, concurrency 1, with no throughput
or acceptance-length numbers taken.
</Note> </Note>
### Mamba ratio calculator ### Mamba ratio calculator
@@ -176,18 +172,49 @@ context from earlier messages.
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Same body, `lm_head` left dense in BF16</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Same body, `lm_head` left dense in BF16</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="https://huggingface.co/RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead">RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead</a></td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="https://huggingface.co/RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead">RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead</a></td>
</tr> </tr>
<tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3.8-27B-NVFP4 (NVIDIA)</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>NVIDIA's ModelOpt export of the same W4A4 body, `lm_head` packed to FP4</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}><a href="https://huggingface.co/nvidia/Qwen3.8-27B-NVFP4">nvidia/Qwen3.8-27B-NVFP4</a></td>
</tr>
</tbody> </tbody>
</table> </table>
The two NVFP4 exports differ only in the `lm_head`: one packs it to FP4, the The two RadixArk NVFP4 exports differ only in the `lm_head`: one packs it to
other leaves it dense in BF16. The dense head is ~1.7 GB larger on disk and FP4, the other leaves it dense in BF16. The dense head is ~1.7 GB larger on disk
~3.2 GB larger at runtime, so it is the harder of the two to fit — every and ~3.2 GB larger at runtime, so it is the harder of the two to fit — every
recipe on this page was measured against it, and the FP4-head cells reuse recipe on this page was measured against it, and the FP4-head cells reuse
those pins unchanged. those pins unchanged.
Both NVFP4 checkpoints declare `kv_cache_quant_algo: FP8`; SGLang's default NVIDIA's own export is that same W4A4 body with that same FP4 head: identical
`--kv-cache-dtype auto` honors it, so the KV pool runs in `fp8_e4m3` with the quantized-layer map (FP8 attention and GDN projections, NVFP4 MLPs), identical
checkpoint's calibration scales automatically. tensor set, identical 21.9 GB on disk. On GB300, RTX PRO 6000 and DGX Spark its
cells reuse the FP4-head pins unchanged, and both SM12x grids have been
re-measured against this export on v0.5.19: all 16 overlay combinations per
card serve and score 94.01-95.00% (RTX PRO 6000) and 94.16-95.07% (DGX Spark)
on the full 1319-question GSM8K.
The RTX 5090 is measured too — all 15 overlay combinations it offers serve and
score 93.93-94.92% — and every winning launch command there is identical to the
FP4-head export's, which is the strongest form of the claim above. What the
32GB card does need is the draft-model rows pinning their own pools: those
recipes pin `--max-running-requests 1`, but nothing caps the pools to match, so
the KV pool sizes itself for 127,332 tokens against the 9,216 one
8192-in/1024-out request needs, and the engine's default split then leaves the
GDN state pool far short of the slots it needs once the draft model's weights
are counted against `--mem-fraction-static`. The DSPARK row therefore pins
`--max-total-tokens` and a measured `--mamba-full-memory-ratio`, as do DFLASH2
and MTP on the dense-lm_head export at float32 state. Those pins override the
calculator's live value for the selections that carry them. The no-speculation
row needs none of it and runs at the pins shown.
The two RadixArk checkpoints declare `kv_cache_quant_algo: FP8`, so SGLang's
default `--kv-cache-dtype auto` already puts their KV pool in `fp8_e4m3`. The
NVIDIA export ships no `kv_cache_scheme`, so `auto` would leave its pool in
BF16 instead. Every recipe on this page pins `--kv-cache-dtype fp8_e4m3`
explicitly, so all three run the same `fp8_e4m3` pool regardless; the
difference only shows up if you switch the Playground's **KV Cache Precision**
row back to Auto.
## 2. Configuration Tips ## 2. Configuration Tips
@@ -204,25 +231,26 @@ checkpoint's calibration scales automatically.
scheduler killed with `exit code -15` and no traceback (`journalctl -u scheduler killed with `exit code -15` and no traceback (`journalctl -u
earlyoom` shows the kill). At 0.85, 15 of the 48 cells were killed that way, earlyoom` shows the kill). At 0.85, 15 of the 48 cells were killed that way,
and which cells is margin noise; at 0.80 every cell served on every attempt. and which cells is margin noise; at 0.80 every cell served on every attempt.
**Validated on SM121 / aarch64**: all 48 configurations (3 checkpoints x **Validated on SM121 / aarch64**: all 80 configurations (5 checkpoints x
Speculative Decoding x Serving Strategy x Mamba SSM Dtype, DFLASH2 included) Speculative Decoding x Serving Strategy x Mamba SSM Dtype, DFLASH2 included)
booted and served on GB10 at `1cf2b8c` at ISL 8192 / OSL 1024, concurrency 1. served on GB10 on `v0.5.19` at ISL 8192 / OSL 1024, concurrency 1, and each
That is boot-and-serve coverage only — no throughput or acceptance-length scored the full 1319-question GSM8K (93.18-95.15%); the float32 and bfloat16
numbers — and it includes the FlashInfer `plan` / `uniform_q_len` path above, halves ran on two separate GB10 boxes. No throughput or acceptance-length
which raised no arity error on that build. Three host quirks when reproducing numbers were re-taken. The sweep exercises the FlashInfer `plan` /
`uniform_q_len` path above, which raised no arity error on that build. Three host quirks when reproducing
on GB10: docker GPU access is CDI-only (`--device nvidia.com/gpu=all`, as no on GB10: docker GPU access is CDI-only (`--device nvidia.com/gpu=all`, as no
`nvidia` runtime is registered); `nvidia-smi` reports `Not Supported` for `nvidia` runtime is registered); `nvidia-smi` reports `Not Supported` for
memory because it is unified with the CPU — gate a relaunch on `MemAvailable` memory because it is unified with the CPU — gate a relaunch on `MemAvailable`
in `/proc/meminfo` instead; and the BF16 checkpoint takes ~6.5 minutes just in `/proc/meminfo` instead; and the BF16 checkpoint takes ~6.5 minutes just
to load its 18 shards from NVMe, so budget ~10 minutes to READY before to load its 18 shards from NVMe, so budget ~10 minutes to READY before
calling a boot hung. calling a boot hung.
- **H200 (SM90)**: BF16 and FP8 only — the card has no FP4 tensor cores, so the - **H200 (SM90)**: BF16 and FP8 only — the card has no FP4 tensor cores, so an
NVFP4 checkpoint's MLP would fall back to the Marlin W4A16 weight-only path NVFP4 checkpoint's MLP would fall back to the Marlin W4A16 weight-only path,
and its cell is greyed out. The H200 recipes use 32768-token prefill chunks and all three NVFP4 cells are greyed out. The H200 recipes use 32768-token
(SM90 prefill is fast enough that a big chunk barely stalls decode, unlike prefill chunks (SM90 prefill is fast enough that a big chunk barely stalls
the SM120 guidance below), and the FlashInfer GDN prefill backend engages by decode, unlike the SM120 guidance below), and the FlashInfer GDN prefill
default under them. `--attention-backend fa3` is a valid alternative, backend engages by default under them. `--attention-backend fa3` is a valid
measured slightly faster at bs=1. alternative, measured slightly faster at bs=1.
- **MTP**: `--speculative-algorithm EAGLE --speculative-num-steps 3 - **MTP**: `--speculative-algorithm EAGLE --speculative-num-steps 3
--speculative-eagle-topk 1 --speculative-num-draft-tokens 4` uses the --speculative-eagle-topk 1 --speculative-num-draft-tokens 4` uses the
in-checkpoint MTP head. (This recipe was originally documented with `NEXTN`, in-checkpoint MTP head. (This recipe was originally documented with `NEXTN`,
@@ -263,23 +291,24 @@ checkpoint's calibration scales automatically.
--mamba-radix-cache-strategy extra_buffer`, and disabled RadixCache for both --mamba-radix-cache-strategy extra_buffer`, and disabled RadixCache for both
baseline and DFlash2 to exclude cache warm-up and prefix reuse. The DFlash2 baseline and DFlash2 to exclude cache warm-up and prefix reuse. The DFlash2
run added the three flags shown above. run added the three flags shown above.
Accuracy used zero-shot GSM8K with greedy sampling, `max_new_tokens=2048`, That comparison's accuracy used zero-shot GSM8K with greedy sampling,
128 examples, and concurrency levels 1, 2, 4, 8, and 16. `max_new_tokens=2048`, 128 examples, and concurrency levels 1, 2, 4, 8
Validation: NVFP4 measured end-to-end on RTX PRO 6000 and RTX 5090; the and 16 — a different protocol from this page's own sweep below.
RTX PRO 6000 BF16/FP8 cells boot and serve; all 12 DGX Spark DFLASH2 cells Validation: every SM12x cell on this page is measured end to end on
boot and serve on `1cf2b8c` with the selector folded into the draft CUDA v0.5.19 — 202 cells over the five checkpoints, four speculative options, two
graph; on H200 and GB300 those cells carry the serving tiers and two GDN state dtypes, full 1319-question GSM8K on each,
**Final Verification In Progress** badge. The 93.18-95.15%. The RTX PRO 6000 and DGX Spark recipes need no changes. On the
RTX PRO 6000 recipe needs no changes. On the 32GB RTX 5090 every pin is 32GB RTX 5090 the panel applies the measured pins automatically: DFlash2 at
re-measured against that commit, and the panel applies them automatically: `--mem-fraction-static 0.91` with `--chunked-prefill-size 1024` — at 0.91 the
DFlash2 at `--mem-fraction-static 0.91` with `--chunked-prefill-size 1024` — pools fit but a 2048-token chunk's activations do not — DSpark at 0.88
the only cell on this page needing a smaller prefill chunk, because at 0.91 (bfloat16), 0.91 (float32) and 0.92 on the dense-lm_head export, all three
the pools fit but a 2048-token chunk's activations do not — DSpark at 0.88, with their pools pinned and the last two also cutting the prefill chunk to
EAGLE at 0.93 (bfloat16) and 0.94 (float32), and no-speculation at 0.90. 1024 and 512, EAGLE at 0.93 (bfloat16) and 0.94 (float32), and
no-speculation at 0.90.
Whether float32 is available with a draft model depends on the `lm_head`: on Whether float32 is available with a draft model depends on the `lm_head`: on
the BF16-head export it is greyed out for both DSpark and DFlash2, since the the BF16-head export it is greyed out for both DSpark and DFlash2, since the
dense head's ~3.2 GB leave no fp32 state pool that also clears prefill graph dense head's ~3.2 GB leave no fp32 state pool that also clears prefill graph
capture. The FP4-head export frees that headroom back — DSpark serves at 0.89 capture. The FP4-head export frees that headroom back — DSpark serves at 0.91
and DFlash2 High-Throughput at 0.895 with `--mamba-full-memory-ratio 10` and DFlash2 High-Throughput at 0.895 with `--mamba-full-memory-ratio 10`
overriding the balanced value — and only DFlash2 Low-Latency stays out of overriding the balanced value — and only DFlash2 Low-Latency stays out of
reach, where five fp32 slots and a full request's KV never coexist. bfloat16 reach, where five fp32 slots and a full request's KV never coexist. bfloat16
@@ -89,17 +89,19 @@ export const Qwen38MambaRatioCalculator = () => {
// reported as out of range rather than silently mis-computed. // reported as out of range rather than silently mis-computed.
const tp = Number(flagArg("--tp")) || Number(flagArg("--tp-size")) || 1; const tp = Number(flagArg("--tp")) || Number(flagArg("--tp-size")) || 1;
// The NVFP4 checkpoint declares kv_cache_quant_algo: FP8, so the default // The two RadixArk NVFP4 exports declare kv_cache_quant_algo: FP8, so the
// --kv-cache-dtype auto lands on fp8_e4m3 there with no flag present; the // default --kv-cache-dtype auto lands on fp8_e4m3 there with no flag
// BF16 / FP8 checkpoints keep a bf16 KV pool under the same default. An // present. The BF16 / FP8 checkpoints keep a bf16 KV pool under the same
// explicit flag always wins over the checkpoint's declaration. // default, and so does NVIDIA's NVFP4 export — it is the same W4A4 body,
// but it ships no kv_cache_scheme, so this cannot key off the `nvfp4`
// prefix. An explicit flag always wins over the checkpoint's declaration.
const kvFlag = flagArg("--kv-cache-dtype"); const kvFlag = flagArg("--kv-cache-dtype");
const kvDtype = const kvDtype =
kvFlag === "fp8_e4m3" kvFlag === "fp8_e4m3"
? "fp8_e4m3" ? "fp8_e4m3"
: kvFlag === "bfloat16" || kvFlag === "bf16" : kvFlag === "bfloat16" || kvFlag === "bf16"
? "bfloat16" ? "bfloat16"
: String(quant).startsWith("nvfp4") : quant === "nvfp4-bf16-head" || quant === "nvfp4-fp4-head"
? "fp8_e4m3" ? "fp8_e4m3"
: "bfloat16"; : "bfloat16";
+203 -47
View File
@@ -28,10 +28,11 @@ export const config = {
], ],
// Every cell pins `--kv-cache-dtype fp8_e4m3` at the maintainers' direction // Every cell pins `--kv-cache-dtype fp8_e4m3` at the maintainers' direction
// (sign-off recorded in the PR description). NVFP4: a no-op made visible // (sign-off recorded in the PR description). The two RadixArk NVFP4 exports:
// (the checkpoint's `kv_cache_quant_algo: FP8` already resolved `auto` to // a no-op made visible (their `kv_cache_quant_algo: FP8` already resolved
// fp8_e4m3). BF16/FP8: a real quality/capacity trade — halves // `auto` to fp8_e4m3). BF16/FP8 and the NVIDIA NVFP4 export: a real
// kv_bytes_per_token but those checkpoints carry no fp8 KV calibration. // quality/capacity trade — halves kv_bytes_per_token, and those checkpoints
// declare no KV scheme at all, so `auto` would leave the pool in bf16.
// //
// Speculative decoding and GDN state precision are orthogonal knobs, so // Speculative decoding and GDN state precision are orthogonal knobs, so
// they are overlay rows, not match dims (3 x 2 would turn 12 cells into // they are overlay rows, not match dims (3 x 2 would turn 12 cells into
@@ -51,6 +52,15 @@ export const config = {
// BF16-head recipes verbatim. // BF16-head recipes verbatim.
{ id: "nvfp4-bf16-head", label: "NVFP4-BF16-Head" }, { id: "nvfp4-bf16-head", label: "NVFP4-BF16-Head" },
{ id: "nvfp4-fp4-head", label: "NVFP4-FP4-Head" }, { id: "nvfp4-fp4-head", label: "NVFP4-FP4-Head" },
// NVIDIA's own ModelOpt export of the same W4A4 body: identical
// quantized-layer map (401 layers, FP8 attention projections + NVFP4
// MLPs), identical tensor set, identical 21.9GB on disk, and the same
// FP4-packed lm_head as RadixArk/Qwen3.8-27B-NVFP4 — so every cell
// reuses that checkpoint's recipe verbatim. The one difference is that
// it declares no `kv_cache_scheme`, so `--kv-cache-dtype auto` resolves
// to bf16 here rather than fp8_e4m3; the cells pin fp8_e4m3 explicitly,
// which makes the launch command and the KV pool identical either way.
{ id: "nvfp4-nvidia", label: "NVFP4-NVIDIA" },
] }, ] },
{ id: "nodes", title: "Nodes", options: [ { id: "nodes", title: "Nodes", options: [
{ id: "single", label: "Single Node" }, { id: "single", label: "Single Node" },
@@ -76,7 +86,10 @@ export const config = {
// DSpark starves runtime activations and wants it DOWN), so each // DSpark starves runtime activations and wants it DOWN), so each
// option strips the cell's value and re-pins its own. // option strips the cell's value and re-pins its own.
stripPrefixes: (sel) => stripPrefixes: (sel) =>
sel.hw === "rtx5090" ? ["--mem-fraction-static"] : [], sel.hw === "rtx5090"
? ["--mem-fraction-static", "--mamba-full-memory-ratio",
"--max-total-tokens"]
: [],
flags: (sel) => [ flags: (sel) => [
"--speculative-algorithm EAGLE", "--speculative-algorithm EAGLE",
"--speculative-num-steps 3", "--speculative-num-steps 3",
@@ -89,7 +102,7 @@ export const config = {
...(["rtx5090", "rtx6000", "dgx-spark"].includes(sel.hw) ...(["rtx5090", "rtx6000", "dgx-spark"].includes(sel.hw)
? ["--enable-linear-replayssm-spec"] ? ["--enable-linear-replayssm-spec"]
: []), : []),
// Measured on the 5090 at commit 1cf2b8c: fp32 serves at 0.94, // Measured on the 5090 on v0.5.19: fp32 serves at 0.94,
// bf16 at 0.93. bf16 moved UP from 0.92 with the dense-lm_head // bf16 at 0.93. bf16 moved UP from 0.92 with the dense-lm_head
// checkpoint -- the heavier weights need a larger static budget // checkpoint -- the heavier weights need a larger static budget
// before the state pool fits. // before the state pool fits.
@@ -98,6 +111,21 @@ export const config = {
? "--mem-fraction-static 0.94" ? "--mem-fraction-static 0.94"
: "--mem-fraction-static 0.93"] : "--mem-fraction-static 0.93"]
: []), : []),
// The dense-lm_head export is the one case where replayssm's tiny
// state pool still is not enough: its head costs ~3.2GB more at
// runtime, and at fp32 the default split leaves the pool short of
// its slots. Measured on v0.5.19 - the published pins alone, and
// the KV cap alone, both fail to boot here. The pins below are the
// measured pair; pinning the ratio here overrides the calculator's
// live value for this selection.
...(sel.hw === "rtx5090" &&
sel.quant === "nvfp4-bf16-head" &&
sel.ssmDtype === "float32"
? ["--max-total-tokens 16384",
sel.tier === "low-latency"
? "--mamba-full-memory-ratio 2.67"
: "--mamba-full-memory-ratio 2.14"]
: []),
], ],
}, },
{ {
@@ -110,26 +138,48 @@ export const config = {
disableReason: disableReason:
"On the 32GB RTX 5090 the DSpark draft model only fits on top of the NVFP4 weights", "On the 32GB RTX 5090 the DSpark draft model only fits on top of the NVFP4 weights",
stripPrefixes: (sel) => stripPrefixes: (sel) =>
sel.hw === "rtx5090" ? ["--mem-fraction-static"] : [], sel.hw === "rtx5090"
? ["--mem-fraction-static", "--mamba-full-memory-ratio",
"--chunked-prefill-size", "--max-total-tokens"]
: [],
flags: (sel) => [ flags: (sel) => [
"--speculative-algorithm DSPARK", "--speculative-algorithm DSPARK",
"--speculative-draft-model-path RadixArk/Qwen3.8-27B-DSpark", "--speculative-draft-model-path RadixArk/Qwen3.8-27B-DSpark",
"--speculative-draft-attention-backend flashinfer", "--speculative-draft-attention-backend flashinfer",
// Measured on the 5090 at commit 1cf2b8c: bf16 serves at 0.88, // On the 32GB 5090 this row owns the pools outright. Two things
// below the 0.90 this recipe carried when it was measured on an // fail if it does not: the KV pool sizes itself for concurrency
// older build, because a draft model plus the automatic prefill // --max-running-requests 1 forbids (127,332 tokens against the
// CUDA-graph capture no longer fit there. fp32 is greyed out by the // 9,216 one 8192-in/1024-out request needs), and the engine's
// SSM dtype row. EAGLE and no-speculation are unaffected: replayssm // default split leaves the GDN state pool a fraction of what it
// keeps EAGLE's state pool tiny and no-spec loads no draft weights. // needs once the draft model's 3.64GB is counted against
// Measured on the 5090 at the commit the Install accordion pins: // --mem-fraction-static, so boot dies with
// bf16 serves at 0.88, and on the FP4-head export fp32 serves at // `max_mamba_cache_size=0 ... max_num_reqs=0`.
// 0.89 on the balanced ratio (pool 25,911 / K=6 low-latency, //
// 29,490 / K=5 high-throughput). fp32 on the BF16-head export is // Every value below is measured on v0.5.19 at ISL 8192 / OSL 1024,
// greyed out by the SSM dtype row. // concurrency 1. Pinning the ratio here overrides the calculator's
// live value for these selections. The two FP4-head exports share
// one set of pins; the dense-lm_head export needs its own because
// its head costs ~3.2GB more at runtime.
...(sel.hw === "rtx5090" ...(sel.hw === "rtx5090"
? [sel.ssmDtype === "float32" ? [
? "--mem-fraction-static 0.89" "--max-total-tokens 16384",
: "--mem-fraction-static 0.88"] ...(sel.quant === "nvfp4-bf16-head"
? ["--mem-fraction-static 0.92",
"--chunked-prefill-size 512",
sel.tier === "low-latency"
? "--mamba-full-memory-ratio 3.38"
: "--mamba-full-memory-ratio 2.71"]
: sel.ssmDtype === "float32"
? ["--mem-fraction-static 0.91",
"--chunked-prefill-size 1024",
sel.tier === "low-latency"
? "--mamba-full-memory-ratio 6.94"
: "--mamba-full-memory-ratio 5.56"]
: ["--mem-fraction-static 0.88",
sel.tier === "low-latency"
? "--mamba-full-memory-ratio 3.38"
: "--mamba-full-memory-ratio 3.12"]),
]
: []), : []),
], ],
}, },
@@ -157,12 +207,12 @@ export const config = {
"--speculative-algorithm DFLASH", "--speculative-algorithm DFLASH",
"--speculative-draft-model-path incoai/Qwen3.8-27B-DFlash2", "--speculative-draft-model-path incoai/Qwen3.8-27B-DFlash2",
"--speculative-num-draft-tokens 8", "--speculative-num-draft-tokens 8",
// Measured on the 5090 at commit 1cf2b8c, the build the Install // Measured on the 5090 on v0.5.19. This cell needs a prefill chunk
// accordion pins. This is the only cell on the page that also needs // smaller than the engine default (DSPARK is the other row that
// a prefill chunk smaller than the engine default: at 0.91 the pools // does): at 0.91 the pools fit but a 2048-token chunk's activations
// fit but a 2048-token chunk's activations do not. The pair together // do not. The pair together is the fastest recipe on this card
// is the fastest recipe on this card (4.92ms median TPOT, 4.29 // (4.92ms median TPOT, 4.29 accept length). fp32 is greyed out by
// accept length). fp32 is greyed out by the SSM dtype row. // the SSM dtype row.
...(sel.hw === "rtx5090" ...(sel.hw === "rtx5090"
? sel.ssmDtype === "float32" ? sel.ssmDtype === "float32"
// FP4-head export, High-Throughput only (the SSM dtype row // FP4-head export, High-Throughput only (the SSM dtype row
@@ -175,6 +225,15 @@ export const config = {
: ["--mem-fraction-static 0.91", : ["--mem-fraction-static 0.91",
"--chunked-prefill-size 1024"] "--chunked-prefill-size 1024"]
: []), : []),
// The dense-lm_head export carries ~3.2GB more weight at runtime,
// which is the difference between serving at the pins above and
// needing the pools pinned outright. Measured on v0.5.19.
...(sel.hw === "rtx5090" && sel.quant === "nvfp4-bf16-head"
? ["--max-total-tokens 16384",
sel.tier === "low-latency"
? "--mamba-full-memory-ratio 3.38"
: "--mamba-full-memory-ratio 3.12"]
: []),
], ],
}, },
], ],
@@ -235,7 +294,8 @@ export const config = {
// clears prefill CUDA-graph capture, for either draft model. // clears prefill CUDA-graph capture, for either draft model.
// Measured across 0.86-0.96 at both chunk sizes, plus balanced- // Measured across 0.86-0.96 at both chunk sizes, plus balanced-
// ratio overrides to 20. // ratio overrides to 20.
// FP4 head — the packed head frees that headroom back: DSpark // FP4 head — either FP4-head export (RadixArk or NVIDIA; same
// packed head, same footprint) frees that headroom back: DSpark
// serves at 0.89 on the balanced ratio and DFlash2 High-Throughput // serves at 0.89 on the balanced ratio and DFlash2 High-Throughput
// at 0.895 with the ratio overridden to 10. Only DFlash2 // at 0.895 with the ratio overridden to 10. Only DFlash2
// Low-Latency stays out of reach: S=5 fp32 slots plus a full // Low-Latency stays out of reach: S=5 fp32 slots plus a full
@@ -268,6 +328,7 @@ export const config = {
"default|fp8": "Qwen/Qwen3.8-27B-FP8", "default|fp8": "Qwen/Qwen3.8-27B-FP8",
"default|nvfp4-bf16-head": "RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead", "default|nvfp4-bf16-head": "RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead",
"default|nvfp4-fp4-head": "RadixArk/Qwen3.8-27B-NVFP4", "default|nvfp4-fp4-head": "RadixArk/Qwen3.8-27B-NVFP4",
"default|nvfp4-nvidia": "nvidia/Qwen3.8-27B-NVFP4",
}, },
placeholders: { placeholders: {
@@ -311,16 +372,11 @@ export const config = {
], ],
dockerImages: { dockerImages: {
h200: "lmsysorg/sglang:qwen38-27b", h200: "lmsysorg/sglang:latest",
// Both SM120 cards are validated on this image (built from 1cf2b8c, the rtx6000: "lmsysorg/sglang:latest",
// commit every pin on those cards was measured against). rtx5090: "lmsysorg/sglang:latest",
rtx6000: "lmsysorg/sglang:dev-qwen38-27b-dflash2", "dgx-spark": "lmsysorg/sglang:latest",
rtx5090: "lmsysorg/sglang:dev-qwen38-27b-dflash2", gb300: "lmsysorg/sglang:latest",
// Multi-arch: this tag ships both linux/amd64 and linux/arm64, so it pulls
// natively on DGX Spark (GB10 is aarch64).
// Multi-arch (linux/amd64 + linux/arm64), so GB10 pulls it natively.
"dgx-spark": "lmsysorg/sglang:dev-qwen38-27b-dflash2",
gb300: "lmsysorg/sglang:dev",
}, },
github: { github: {
@@ -462,9 +518,10 @@ export const config = {
// is the H200-validated setting: SM90 prefill is fast enough that a big // is the H200-validated setting: SM90 prefill is fast enough that a big
// chunk stalls decode far less than on SM120, and the SM90 FlashInfer GDN // chunk stalls decode far less than on SM120, and the SM90 FlashInfer GDN
// prefill default engages under it (fp32 state pool, chunk <= 32768). // prefill default engages under it (fp32 state pool, chunk <= 32768).
// No NVFP4 cell on this card: SM90 has no FP4 tensor cores, so the W4A4 // No NVFP4 cell on this card, for any of the three exports: SM90 has no
// checkpoint's MLP would fall back to the Marlin W4A16 weight-only path — // FP4 tensor cores, so a W4A4 checkpoint's MLP would fall back to the
// runnable, but not a recipe this page ships. // Marlin W4A16 weight-only path — runnable, but not a recipe this page
// ships.
match: { hw: "h200", variant: "default", quant: "fp8", nodes: "single" }, match: { hw: "h200", variant: "default", quant: "fp8", nodes: "single" },
verified: true, verified: true,
// DFLASH2 has not been exercised on this platform; every other overlay // DFLASH2 has not been exercised on this platform; every other overlay
@@ -549,6 +606,28 @@ export const config = {
"--port {{PORT}}", "--port {{PORT}}",
], ],
}, },
{
// NVIDIA's ModelOpt export of the same W4A4 body and FP4 lm_head as the
// RadixArk FP4-head checkpoint above, so it reuses that recipe verbatim.
// Re-measured against this export on v0.5.19: all 16 overlay
// combinations (spec x tier x state dtype) serve at these pins and score
// 94.01-95.00% on the full 1319-question GSM8K.
match: { hw: "rtx6000", variant: "default", quant: "nvfp4-nvidia", nodes: "single" },
verified: true,
env: [],
flags: [
"--trust-remote-code",
"--model-path {{MODEL_NAME}}",
"--kv-cache-dtype fp8_e4m3",
"--mem-fraction-static 0.85",
"--attention-backend flashinfer",
"--chunked-prefill-size 2048",
"--reasoning-parser qwen3",
"--tool-call-parser qwen3_coder",
"--host {{HOST_IP}}",
"--port {{PORT}}",
],
},
{ {
// FP8 blockwise, ~28.5GB of weights — comfortable on 96GB. // FP8 blockwise, ~28.5GB of weights — comfortable on 96GB.
match: { hw: "rtx6000", variant: "default", quant: "fp8", nodes: "single" }, match: { hw: "rtx6000", variant: "default", quant: "fp8", nodes: "single" },
@@ -653,6 +732,40 @@ export const config = {
"--port {{PORT}}", "--port {{PORT}}",
], ],
}, },
{
// NVIDIA's ModelOpt export: same body, same FP4 lm_head, same 21.9GB of
// weights as the RadixArk FP4-head checkpoint, so the 32GB fit and every
// mem-fraction pin the overlay rows apply carry over unchanged.
match: { hw: "rtx5090", variant: "default", quant: "nvfp4-nvidia", nodes: "single" },
// Measured on v0.5.19 against this export: all 15 offered overlay
// combinations serve and score 93.93-94.92% on the full 1319-question
// GSM8K. Every winning launch command is identical to the FP4-head
// export's, which is what "reuses that recipe verbatim" above is claiming.
verified: true,
// Rendered with the cell so nobody ships the bs=1 pins into a
// multi-user deployment unaware.
warn:
"This recipe serves ONE request at a time: --max-running-requests 1 " +
"and --cuda-graph-max-bs-decode 1 pin it to the validated single-stream " +
"envelope. To handle more concurrent requests, raise both flags " +
"together and re-derive --mamba-full-memory-ratio (and mem-fraction) " +
"with the [Mamba ratio calculator](#mamba-ratio-calculator) — on this " +
"32GB card the GDN state pool, not KV, is what runs out first.",
env: [],
flags: [
"--trust-remote-code",
"--model-path {{MODEL_NAME}}",
"--kv-cache-dtype fp8_e4m3",
"--mem-fraction-static 0.9",
"--attention-backend flashinfer",
"--max-running-requests 1",
"--cuda-graph-max-bs-decode 1",
"--reasoning-parser qwen3",
"--tool-call-parser qwen3_coder",
"--host {{HOST_IP}}",
"--port {{PORT}}",
],
},
// DGX Spark (GB10, SM121): single node, 128GB coherent unified memory // DGX Spark (GB10, SM121): single node, 128GB coherent unified memory
// shared with the CPU — every checkpoint fits, so all three quants get a // shared with the CPU — every checkpoint fits, so all three quants get a
// cell. These cells reuse the RTX PRO 6000 recipe at one lower // cell. These cells reuse the RTX PRO 6000 recipe at one lower
@@ -665,12 +778,12 @@ export const config = {
// leaves ~8GB for the OS — exactly DGX OS earlyoom's SIGTERM threshold — // leaves ~8GB for the OS — exactly DGX OS earlyoom's SIGTERM threshold —
// and the first long prefill or boot-time graph capture dips under it and // and the first long prefill or boot-time graph capture dips under it and
// gets the scheduler killed (exit code -15, no traceback; check // gets the scheduler killed (exit code -15, no traceback; check
// `journalctl -u earlyoom`). Re-measured on 1cf2b8c (2026-08-21): at 0.85, // `journalctl -u earlyoom`). Re-measured on v0.5.19: at 0.85,
// 15 of 48 cells were SIGTERMed, and which 15 is margin noise, biased // 15 of 48 cells were SIGTERMed, and which 15 is margin noise, biased
// toward the big-state configs (bfloat16 SSM, DSPARK/DFLASH2 ratios); at // toward the big-state configs (bfloat16 SSM, DSPARK/DFLASH2 ratios); at
// 0.80 every cell served on every attempt. // 0.80 every cell served on every attempt.
// //
// Validated on GB10 (SM121 / aarch64) at 1cf2b8c: all 48 configurations — // Validated on GB10 (SM121 / aarch64) on v0.5.19: all 80 configurations —
// DFLASH2 included — booted and served at ISL 8192 / OSL 1024, // DFLASH2 included — booted and served at ISL 8192 / OSL 1024,
// concurrency 1. Boot-and-serve only -- no throughput or acceptance-length // concurrency 1. Boot-and-serve only -- no throughput or acceptance-length
// numbers were taken, so this is a weaker standard than the SM120 pair's // numbers were taken, so this is a weaker standard than the SM120 pair's
@@ -680,7 +793,7 @@ export const config = {
// 12-cell DFLASH2 pass. // 12-cell DFLASH2 pass.
{ {
match: { hw: "dgx-spark", variant: "default", quant: "nvfp4-bf16-head", nodes: "single" }, match: { hw: "dgx-spark", variant: "default", quant: "nvfp4-bf16-head", nodes: "single" },
// All 16 overlay combinations served on GB10 at 1cf2b8c, DFLASH2 // All 16 overlay combinations served on GB10 on v0.5.19, DFLASH2
// included — its selector folded into the draft CUDA graph in all four // included — its selector folded into the draft CUDA graph in all four
// of its cells here. // of its cells here.
verified: true, verified: true,
@@ -702,7 +815,7 @@ export const config = {
// Same recipe as the BF16-head cell above: the FP4 head is smaller, // Same recipe as the BF16-head cell above: the FP4 head is smaller,
// so anything that fits the bf16 head fits here with room to spare. // so anything that fits the bf16 head fits here with room to spare.
match: { hw: "dgx-spark", variant: "default", quant: "nvfp4-fp4-head", nodes: "single" }, match: { hw: "dgx-spark", variant: "default", quant: "nvfp4-fp4-head", nodes: "single" },
// All 16 overlay combinations served on GB10 at 1cf2b8c, DFLASH2 // All 16 overlay combinations served on GB10 on v0.5.19, DFLASH2
// included — its selector folded into the draft CUDA graph in all four // included — its selector folded into the draft CUDA graph in all four
// of its cells here. // of its cells here.
verified: true, verified: true,
@@ -720,9 +833,32 @@ export const config = {
"--port {{PORT}}", "--port {{PORT}}",
], ],
}, },
{
// NVIDIA's ModelOpt export of the same W4A4 body as the RadixArk FP4-head
// checkpoint, on that cell's recipe.
match: { hw: "dgx-spark", variant: "default", quant: "nvfp4-nvidia", nodes: "single" },
// Re-measured against this export on v0.5.19: all 16 overlay
// combinations serve at these pins and score 94.16-95.07% on the full
// 1319-question GSM8K (float32 and bfloat16 halves run on two separate
// GB10 boxes).
verified: true,
env: [],
flags: [
"--trust-remote-code",
"--model-path {{MODEL_NAME}}",
"--kv-cache-dtype fp8_e4m3",
"--mem-fraction-static 0.80",
"--attention-backend flashinfer",
"--chunked-prefill-size 2048",
"--reasoning-parser qwen3",
"--tool-call-parser qwen3_coder",
"--host {{HOST_IP}}",
"--port {{PORT}}",
],
},
{ {
match: { hw: "dgx-spark", variant: "default", quant: "fp8", nodes: "single" }, match: { hw: "dgx-spark", variant: "default", quant: "fp8", nodes: "single" },
// All 16 overlay combinations served on GB10 at 1cf2b8c, DFLASH2 // All 16 overlay combinations served on GB10 on v0.5.19, DFLASH2
// included. This checkpoint held the sweep's most earlyoom-prone cells // included. This checkpoint held the sweep's most earlyoom-prone cells
// at 0.85 (every bfloat16-SSM pick was killed); all clean at 0.80. // at 0.85 (every bfloat16-SSM pick was killed); all clean at 0.80.
verified: true, verified: true,
@@ -742,7 +878,7 @@ export const config = {
}, },
{ {
match: { hw: "dgx-spark", variant: "default", quant: "bf16", nodes: "single" }, match: { hw: "dgx-spark", variant: "default", quant: "bf16", nodes: "single" },
// All 16 overlay combinations served on GB10 at 1cf2b8c, DFLASH2 // All 16 overlay combinations served on GB10 on v0.5.19, DFLASH2
// included. Heaviest checkpoint (52GB, ~6.5 min to load its 18 shards // included. Heaviest checkpoint (52GB, ~6.5 min to load its 18 shards
// from NVMe — budget ~10 min to READY before calling a boot hung). // from NVMe — budget ~10 min to READY before calling a boot hung).
verified: true, verified: true,
@@ -806,6 +942,26 @@ export const config = {
"--port {{PORT}}", "--port {{PORT}}",
], ],
}, },
{
match: { hw: "gb300", variant: "default", quant: "nvfp4-nvidia", nodes: "single" },
verified: true,
// DFLASH2 has not been exercised on this platform; every other overlay
// pick keeps this cell's original validation.
verificationStatus: (sel) =>
sel.spec === "dflash" ? "in-progress" : "verified",
env: [],
flags: [
"--trust-remote-code",
"--model-path {{MODEL_NAME}}",
"--kv-cache-dtype fp8_e4m3",
"--mem-fraction-static 0.85",
"--chunked-prefill-size 2048",
"--reasoning-parser qwen3",
"--tool-call-parser qwen3_coder",
"--host {{HOST_IP}}",
"--port {{PORT}}",
],
},
{ {
match: { hw: "gb300", variant: "default", quant: "fp8", nodes: "single" }, match: { hw: "gb300", variant: "default", quant: "fp8", nodes: "single" },
verified: true, verified: true,