From df6424967ab850cc53bf0321904a11a0cce099b2 Mon Sep 17 00:00:00 2001 From: zijiexia <37504505+zijiexia@users.noreply.github.com> Date: Fri, 11 Sep 2026 11:58:30 -0700 Subject: [PATCH] [docs] Add the NVIDIA NVFP4 export to the Qwen3.8-27B cookbook (#38611) Co-authored-by: Claude Opus 5 Co-authored-by: Jiminator Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com> --- .../autoregressive/Qwen/Qwen3.8-27B.mdx | 117 +++++--- .../_qwen38_mamba_ratio_calculator.jsx | 12 +- .../src/snippets/configs/Qwen/qwen3.8-27b.jsx | 250 ++++++++++++++---- 3 files changed, 283 insertions(+), 96 deletions(-) diff --git a/docs/cookbook/autoregressive/Qwen/Qwen3.8-27B.mdx b/docs/cookbook/autoregressive/Qwen/Qwen3.8-27B.mdx index 06153c0ea..ac2decb20 100644 --- a/docs/cookbook/autoregressive/Qwen/Qwen3.8-27B.mdx +++ b/docs/cookbook/autoregressive/Qwen/Qwen3.8-27B.mdx @@ -19,9 +19,7 @@ For all methods and hardware platforms, see the [official SGLang installation gu pip install --upgrade pip pip install uv -git clone https://github.com/sgl-project/sglang.git && cd sglang -git checkout 1cf2b8c54d81802abc15dcf23a29b9cc687bc01e -uv pip install --prerelease=allow -e "python[all]" +uv pip install --prerelease=allow sglang ``` Then run the **Python** output of the command panel below in that environment. @@ -31,7 +29,7 @@ Then run the **Python** output of the command panel below in that environment. ```bash Command -docker pull lmsysorg/sglang:dev-qwen38-27b-dflash2 +docker pull lmsysorg/sglang:latest ``` For how to launch the image, see [Install → Method 3: Using Docker](../../../docs/get-started/install#method-3-using-docker). Substitute the inner `sglang serve ...` with what the command generator below produces. @@ -59,14 +57,12 @@ import { Qwen38MambaRatioCalculator } from "/src/snippets/_qwen38_mamba_ratio_ca - The RTX 5090 and RTX PRO 6000 cells above — including every Speculative - Decoding / Serving Strategy / SSM dtype combination — were validated at - ISL 8192 / OSL 1024, concurrency 1; for DFLASH2, to that full standard on - NVFP4 and to boot-and-serve on the RTX PRO 6000 BF16/FP8 cells. The DGX Spark - cells cover the full combination set — DFLASH2 included — re-measured end to - end on `1cf2b8c`, but to a weaker standard: each of the 48 was confirmed to - **boot and serve** at ISL 8192 / OSL 1024, concurrency 1, with no throughput - or acceptance-length numbers taken. + Every cell above — RTX 5090, RTX PRO 6000 and DGX Spark, across all five + checkpoints and every Speculative Decoding / Serving Strategy / SSM dtype + combination — is measured on **v0.5.19**. That is 202 cells, each one served + and scored on the full 1319-question GSM8K (93.18-95.15%). The serving + envelope behind the pins is ISL 8192 / OSL 1024 at concurrency 1; throughput + and acceptance-length numbers were not re-taken in that sweep. ### Mamba ratio calculator @@ -176,18 +172,49 @@ context from earlier messages. Same body, `lm_head` left dense in BF16 RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead + + Qwen3.8-27B-NVFP4 (NVIDIA) + NVIDIA's ModelOpt export of the same W4A4 body, `lm_head` packed to FP4 + nvidia/Qwen3.8-27B-NVFP4 + -The two NVFP4 exports differ only in the `lm_head`: one packs it to FP4, the -other leaves it dense in BF16. The dense head is ~1.7 GB larger on disk and -~3.2 GB larger at runtime, so it is the harder of the two to fit — every +The two RadixArk NVFP4 exports differ only in the `lm_head`: one packs it to +FP4, the other leaves it dense in BF16. The dense head is ~1.7 GB larger on disk +and ~3.2 GB larger at runtime, so it is the harder of the two to fit — every recipe on this page was measured against it, and the FP4-head cells reuse those pins unchanged. -Both NVFP4 checkpoints declare `kv_cache_quant_algo: FP8`; SGLang's default -`--kv-cache-dtype auto` honors it, so the KV pool runs in `fp8_e4m3` with the -checkpoint's calibration scales automatically. +NVIDIA's own export is that same W4A4 body with that same FP4 head: identical +quantized-layer map (FP8 attention and GDN projections, NVFP4 MLPs), identical +tensor set, identical 21.9 GB on disk. On GB300, RTX PRO 6000 and DGX Spark its +cells reuse the FP4-head pins unchanged, and both SM12x grids have been +re-measured against this export on v0.5.19: all 16 overlay combinations per +card serve and score 94.01-95.00% (RTX PRO 6000) and 94.16-95.07% (DGX Spark) +on the full 1319-question GSM8K. + +The RTX 5090 is measured too — all 15 overlay combinations it offers serve and +score 93.93-94.92% — and every winning launch command there is identical to the +FP4-head export's, which is the strongest form of the claim above. What the +32GB card does need is the draft-model rows pinning their own pools: those +recipes pin `--max-running-requests 1`, but nothing caps the pools to match, so +the KV pool sizes itself for 127,332 tokens against the 9,216 one +8192-in/1024-out request needs, and the engine's default split then leaves the +GDN state pool far short of the slots it needs once the draft model's weights +are counted against `--mem-fraction-static`. The DSPARK row therefore pins +`--max-total-tokens` and a measured `--mamba-full-memory-ratio`, as do DFLASH2 +and MTP on the dense-lm_head export at float32 state. Those pins override the +calculator's live value for the selections that carry them. The no-speculation +row needs none of it and runs at the pins shown. + +The two RadixArk checkpoints declare `kv_cache_quant_algo: FP8`, so SGLang's +default `--kv-cache-dtype auto` already puts their KV pool in `fp8_e4m3`. The +NVIDIA export ships no `kv_cache_scheme`, so `auto` would leave its pool in +BF16 instead. Every recipe on this page pins `--kv-cache-dtype fp8_e4m3` +explicitly, so all three run the same `fp8_e4m3` pool regardless; the +difference only shows up if you switch the Playground's **KV Cache Precision** +row back to Auto. ## 2. Configuration Tips @@ -204,25 +231,26 @@ checkpoint's calibration scales automatically. scheduler killed with `exit code -15` and no traceback (`journalctl -u earlyoom` shows the kill). At 0.85, 15 of the 48 cells were killed that way, and which cells is margin noise; at 0.80 every cell served on every attempt. - **Validated on SM121 / aarch64**: all 48 configurations (3 checkpoints x + **Validated on SM121 / aarch64**: all 80 configurations (5 checkpoints x Speculative Decoding x Serving Strategy x Mamba SSM Dtype, DFLASH2 included) - booted and served on GB10 at `1cf2b8c` at ISL 8192 / OSL 1024, concurrency 1. - That is boot-and-serve coverage only — no throughput or acceptance-length - numbers — and it includes the FlashInfer `plan` / `uniform_q_len` path above, - which raised no arity error on that build. Three host quirks when reproducing + served on GB10 on `v0.5.19` at ISL 8192 / OSL 1024, concurrency 1, and each + scored the full 1319-question GSM8K (93.18-95.15%); the float32 and bfloat16 + halves ran on two separate GB10 boxes. No throughput or acceptance-length + numbers were re-taken. The sweep exercises the FlashInfer `plan` / + `uniform_q_len` path above, which raised no arity error on that build. Three host quirks when reproducing on GB10: docker GPU access is CDI-only (`--device nvidia.com/gpu=all`, as no `nvidia` runtime is registered); `nvidia-smi` reports `Not Supported` for memory because it is unified with the CPU — gate a relaunch on `MemAvailable` in `/proc/meminfo` instead; and the BF16 checkpoint takes ~6.5 minutes just to load its 18 shards from NVMe, so budget ~10 minutes to READY before calling a boot hung. -- **H200 (SM90)**: BF16 and FP8 only — the card has no FP4 tensor cores, so the - NVFP4 checkpoint's MLP would fall back to the Marlin W4A16 weight-only path - and its cell is greyed out. The H200 recipes use 32768-token prefill chunks - (SM90 prefill is fast enough that a big chunk barely stalls decode, unlike - the SM120 guidance below), and the FlashInfer GDN prefill backend engages by - default under them. `--attention-backend fa3` is a valid alternative, - measured slightly faster at bs=1. +- **H200 (SM90)**: BF16 and FP8 only — the card has no FP4 tensor cores, so an + NVFP4 checkpoint's MLP would fall back to the Marlin W4A16 weight-only path, + and all three NVFP4 cells are greyed out. The H200 recipes use 32768-token + prefill chunks (SM90 prefill is fast enough that a big chunk barely stalls + decode, unlike the SM120 guidance below), and the FlashInfer GDN prefill + backend engages by default under them. `--attention-backend fa3` is a valid + alternative, measured slightly faster at bs=1. - **MTP**: `--speculative-algorithm EAGLE --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4` uses the in-checkpoint MTP head. (This recipe was originally documented with `NEXTN`, @@ -263,23 +291,24 @@ checkpoint's calibration scales automatically. --mamba-radix-cache-strategy extra_buffer`, and disabled RadixCache for both baseline and DFlash2 to exclude cache warm-up and prefix reuse. The DFlash2 run added the three flags shown above. - Accuracy used zero-shot GSM8K with greedy sampling, `max_new_tokens=2048`, - 128 examples, and concurrency levels 1, 2, 4, 8, and 16. - Validation: NVFP4 measured end-to-end on RTX PRO 6000 and RTX 5090; the - RTX PRO 6000 BF16/FP8 cells boot and serve; all 12 DGX Spark DFLASH2 cells - boot and serve on `1cf2b8c` with the selector folded into the draft CUDA - graph; on H200 and GB300 those cells carry the - **Final Verification In Progress** badge. The - RTX PRO 6000 recipe needs no changes. On the 32GB RTX 5090 every pin is - re-measured against that commit, and the panel applies them automatically: - DFlash2 at `--mem-fraction-static 0.91` with `--chunked-prefill-size 1024` — - the only cell on this page needing a smaller prefill chunk, because at 0.91 - the pools fit but a 2048-token chunk's activations do not — DSpark at 0.88, - EAGLE at 0.93 (bfloat16) and 0.94 (float32), and no-speculation at 0.90. + That comparison's accuracy used zero-shot GSM8K with greedy sampling, + `max_new_tokens=2048`, 128 examples, and concurrency levels 1, 2, 4, 8 + and 16 — a different protocol from this page's own sweep below. + Validation: every SM12x cell on this page is measured end to end on + v0.5.19 — 202 cells over the five checkpoints, four speculative options, two + serving tiers and two GDN state dtypes, full 1319-question GSM8K on each, + 93.18-95.15%. The RTX PRO 6000 and DGX Spark recipes need no changes. On the + 32GB RTX 5090 the panel applies the measured pins automatically: DFlash2 at + `--mem-fraction-static 0.91` with `--chunked-prefill-size 1024` — at 0.91 the + pools fit but a 2048-token chunk's activations do not — DSpark at 0.88 + (bfloat16), 0.91 (float32) and 0.92 on the dense-lm_head export, all three + with their pools pinned and the last two also cutting the prefill chunk to + 1024 and 512, EAGLE at 0.93 (bfloat16) and 0.94 (float32), and + no-speculation at 0.90. Whether float32 is available with a draft model depends on the `lm_head`: on the BF16-head export it is greyed out for both DSpark and DFlash2, since the dense head's ~3.2 GB leave no fp32 state pool that also clears prefill graph - capture. The FP4-head export frees that headroom back — DSpark serves at 0.89 + capture. The FP4-head export frees that headroom back — DSpark serves at 0.91 and DFlash2 High-Throughput at 0.895 with `--mamba-full-memory-ratio 10` overriding the balanced value — and only DFlash2 Low-Latency stays out of reach, where five fp32 slots and a full request's KV never coexist. bfloat16 diff --git a/docs/src/snippets/_qwen38_mamba_ratio_calculator.jsx b/docs/src/snippets/_qwen38_mamba_ratio_calculator.jsx index ba8dfb546..6f87e956c 100644 --- a/docs/src/snippets/_qwen38_mamba_ratio_calculator.jsx +++ b/docs/src/snippets/_qwen38_mamba_ratio_calculator.jsx @@ -89,17 +89,19 @@ export const Qwen38MambaRatioCalculator = () => { // reported as out of range rather than silently mis-computed. const tp = Number(flagArg("--tp")) || Number(flagArg("--tp-size")) || 1; - // The NVFP4 checkpoint declares kv_cache_quant_algo: FP8, so the default - // --kv-cache-dtype auto lands on fp8_e4m3 there with no flag present; the - // BF16 / FP8 checkpoints keep a bf16 KV pool under the same default. An - // explicit flag always wins over the checkpoint's declaration. + // The two RadixArk NVFP4 exports declare kv_cache_quant_algo: FP8, so the + // default --kv-cache-dtype auto lands on fp8_e4m3 there with no flag + // present. The BF16 / FP8 checkpoints keep a bf16 KV pool under the same + // default, and so does NVIDIA's NVFP4 export — it is the same W4A4 body, + // but it ships no kv_cache_scheme, so this cannot key off the `nvfp4` + // prefix. An explicit flag always wins over the checkpoint's declaration. const kvFlag = flagArg("--kv-cache-dtype"); const kvDtype = kvFlag === "fp8_e4m3" ? "fp8_e4m3" : kvFlag === "bfloat16" || kvFlag === "bf16" ? "bfloat16" - : String(quant).startsWith("nvfp4") + : quant === "nvfp4-bf16-head" || quant === "nvfp4-fp4-head" ? "fp8_e4m3" : "bfloat16"; diff --git a/docs/src/snippets/configs/Qwen/qwen3.8-27b.jsx b/docs/src/snippets/configs/Qwen/qwen3.8-27b.jsx index 0ac0d0135..64dded288 100644 --- a/docs/src/snippets/configs/Qwen/qwen3.8-27b.jsx +++ b/docs/src/snippets/configs/Qwen/qwen3.8-27b.jsx @@ -28,10 +28,11 @@ export const config = { ], // Every cell pins `--kv-cache-dtype fp8_e4m3` at the maintainers' direction - // (sign-off recorded in the PR description). NVFP4: a no-op made visible - // (the checkpoint's `kv_cache_quant_algo: FP8` already resolved `auto` to - // fp8_e4m3). BF16/FP8: a real quality/capacity trade — halves - // kv_bytes_per_token but those checkpoints carry no fp8 KV calibration. + // (sign-off recorded in the PR description). The two RadixArk NVFP4 exports: + // a no-op made visible (their `kv_cache_quant_algo: FP8` already resolved + // `auto` to fp8_e4m3). BF16/FP8 and the NVIDIA NVFP4 export: a real + // quality/capacity trade — halves kv_bytes_per_token, and those checkpoints + // declare no KV scheme at all, so `auto` would leave the pool in bf16. // // Speculative decoding and GDN state precision are orthogonal knobs, so // they are overlay rows, not match dims (3 x 2 would turn 12 cells into @@ -51,6 +52,15 @@ export const config = { // BF16-head recipes verbatim. { id: "nvfp4-bf16-head", label: "NVFP4-BF16-Head" }, { id: "nvfp4-fp4-head", label: "NVFP4-FP4-Head" }, + // NVIDIA's own ModelOpt export of the same W4A4 body: identical + // quantized-layer map (401 layers, FP8 attention projections + NVFP4 + // MLPs), identical tensor set, identical 21.9GB on disk, and the same + // FP4-packed lm_head as RadixArk/Qwen3.8-27B-NVFP4 — so every cell + // reuses that checkpoint's recipe verbatim. The one difference is that + // it declares no `kv_cache_scheme`, so `--kv-cache-dtype auto` resolves + // to bf16 here rather than fp8_e4m3; the cells pin fp8_e4m3 explicitly, + // which makes the launch command and the KV pool identical either way. + { id: "nvfp4-nvidia", label: "NVFP4-NVIDIA" }, ] }, { id: "nodes", title: "Nodes", options: [ { id: "single", label: "Single Node" }, @@ -76,7 +86,10 @@ export const config = { // DSpark starves runtime activations and wants it DOWN), so each // option strips the cell's value and re-pins its own. stripPrefixes: (sel) => - sel.hw === "rtx5090" ? ["--mem-fraction-static"] : [], + sel.hw === "rtx5090" + ? ["--mem-fraction-static", "--mamba-full-memory-ratio", + "--max-total-tokens"] + : [], flags: (sel) => [ "--speculative-algorithm EAGLE", "--speculative-num-steps 3", @@ -89,7 +102,7 @@ export const config = { ...(["rtx5090", "rtx6000", "dgx-spark"].includes(sel.hw) ? ["--enable-linear-replayssm-spec"] : []), - // Measured on the 5090 at commit 1cf2b8c: fp32 serves at 0.94, + // Measured on the 5090 on v0.5.19: fp32 serves at 0.94, // bf16 at 0.93. bf16 moved UP from 0.92 with the dense-lm_head // checkpoint -- the heavier weights need a larger static budget // before the state pool fits. @@ -98,6 +111,21 @@ export const config = { ? "--mem-fraction-static 0.94" : "--mem-fraction-static 0.93"] : []), + // The dense-lm_head export is the one case where replayssm's tiny + // state pool still is not enough: its head costs ~3.2GB more at + // runtime, and at fp32 the default split leaves the pool short of + // its slots. Measured on v0.5.19 - the published pins alone, and + // the KV cap alone, both fail to boot here. The pins below are the + // measured pair; pinning the ratio here overrides the calculator's + // live value for this selection. + ...(sel.hw === "rtx5090" && + sel.quant === "nvfp4-bf16-head" && + sel.ssmDtype === "float32" + ? ["--max-total-tokens 16384", + sel.tier === "low-latency" + ? "--mamba-full-memory-ratio 2.67" + : "--mamba-full-memory-ratio 2.14"] + : []), ], }, { @@ -110,26 +138,48 @@ export const config = { disableReason: "On the 32GB RTX 5090 the DSpark draft model only fits on top of the NVFP4 weights", stripPrefixes: (sel) => - sel.hw === "rtx5090" ? ["--mem-fraction-static"] : [], + sel.hw === "rtx5090" + ? ["--mem-fraction-static", "--mamba-full-memory-ratio", + "--chunked-prefill-size", "--max-total-tokens"] + : [], flags: (sel) => [ "--speculative-algorithm DSPARK", "--speculative-draft-model-path RadixArk/Qwen3.8-27B-DSpark", "--speculative-draft-attention-backend flashinfer", - // Measured on the 5090 at commit 1cf2b8c: bf16 serves at 0.88, - // below the 0.90 this recipe carried when it was measured on an - // older build, because a draft model plus the automatic prefill - // CUDA-graph capture no longer fit there. fp32 is greyed out by the - // SSM dtype row. EAGLE and no-speculation are unaffected: replayssm - // keeps EAGLE's state pool tiny and no-spec loads no draft weights. - // Measured on the 5090 at the commit the Install accordion pins: - // bf16 serves at 0.88, and on the FP4-head export fp32 serves at - // 0.89 on the balanced ratio (pool 25,911 / K=6 low-latency, - // 29,490 / K=5 high-throughput). fp32 on the BF16-head export is - // greyed out by the SSM dtype row. + // On the 32GB 5090 this row owns the pools outright. Two things + // fail if it does not: the KV pool sizes itself for concurrency + // --max-running-requests 1 forbids (127,332 tokens against the + // 9,216 one 8192-in/1024-out request needs), and the engine's + // default split leaves the GDN state pool a fraction of what it + // needs once the draft model's 3.64GB is counted against + // --mem-fraction-static, so boot dies with + // `max_mamba_cache_size=0 ... max_num_reqs=0`. + // + // Every value below is measured on v0.5.19 at ISL 8192 / OSL 1024, + // concurrency 1. Pinning the ratio here overrides the calculator's + // live value for these selections. The two FP4-head exports share + // one set of pins; the dense-lm_head export needs its own because + // its head costs ~3.2GB more at runtime. ...(sel.hw === "rtx5090" - ? [sel.ssmDtype === "float32" - ? "--mem-fraction-static 0.89" - : "--mem-fraction-static 0.88"] + ? [ + "--max-total-tokens 16384", + ...(sel.quant === "nvfp4-bf16-head" + ? ["--mem-fraction-static 0.92", + "--chunked-prefill-size 512", + sel.tier === "low-latency" + ? "--mamba-full-memory-ratio 3.38" + : "--mamba-full-memory-ratio 2.71"] + : sel.ssmDtype === "float32" + ? ["--mem-fraction-static 0.91", + "--chunked-prefill-size 1024", + sel.tier === "low-latency" + ? "--mamba-full-memory-ratio 6.94" + : "--mamba-full-memory-ratio 5.56"] + : ["--mem-fraction-static 0.88", + sel.tier === "low-latency" + ? "--mamba-full-memory-ratio 3.38" + : "--mamba-full-memory-ratio 3.12"]), + ] : []), ], }, @@ -157,12 +207,12 @@ export const config = { "--speculative-algorithm DFLASH", "--speculative-draft-model-path incoai/Qwen3.8-27B-DFlash2", "--speculative-num-draft-tokens 8", - // Measured on the 5090 at commit 1cf2b8c, the build the Install - // accordion pins. This is the only cell on the page that also needs - // a prefill chunk smaller than the engine default: at 0.91 the pools - // fit but a 2048-token chunk's activations do not. The pair together - // is the fastest recipe on this card (4.92ms median TPOT, 4.29 - // accept length). fp32 is greyed out by the SSM dtype row. + // Measured on the 5090 on v0.5.19. This cell needs a prefill chunk + // smaller than the engine default (DSPARK is the other row that + // does): at 0.91 the pools fit but a 2048-token chunk's activations + // do not. The pair together is the fastest recipe on this card + // (4.92ms median TPOT, 4.29 accept length). fp32 is greyed out by + // the SSM dtype row. ...(sel.hw === "rtx5090" ? sel.ssmDtype === "float32" // FP4-head export, High-Throughput only (the SSM dtype row @@ -175,6 +225,15 @@ export const config = { : ["--mem-fraction-static 0.91", "--chunked-prefill-size 1024"] : []), + // The dense-lm_head export carries ~3.2GB more weight at runtime, + // which is the difference between serving at the pins above and + // needing the pools pinned outright. Measured on v0.5.19. + ...(sel.hw === "rtx5090" && sel.quant === "nvfp4-bf16-head" + ? ["--max-total-tokens 16384", + sel.tier === "low-latency" + ? "--mamba-full-memory-ratio 3.38" + : "--mamba-full-memory-ratio 3.12"] + : []), ], }, ], @@ -235,7 +294,8 @@ export const config = { // clears prefill CUDA-graph capture, for either draft model. // Measured across 0.86-0.96 at both chunk sizes, plus balanced- // ratio overrides to 20. - // FP4 head — the packed head frees that headroom back: DSpark + // FP4 head — either FP4-head export (RadixArk or NVIDIA; same + // packed head, same footprint) frees that headroom back: DSpark // serves at 0.89 on the balanced ratio and DFlash2 High-Throughput // at 0.895 with the ratio overridden to 10. Only DFlash2 // Low-Latency stays out of reach: S=5 fp32 slots plus a full @@ -268,6 +328,7 @@ export const config = { "default|fp8": "Qwen/Qwen3.8-27B-FP8", "default|nvfp4-bf16-head": "RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead", "default|nvfp4-fp4-head": "RadixArk/Qwen3.8-27B-NVFP4", + "default|nvfp4-nvidia": "nvidia/Qwen3.8-27B-NVFP4", }, placeholders: { @@ -311,16 +372,11 @@ export const config = { ], dockerImages: { - h200: "lmsysorg/sglang:qwen38-27b", - // Both SM120 cards are validated on this image (built from 1cf2b8c, the - // commit every pin on those cards was measured against). - rtx6000: "lmsysorg/sglang:dev-qwen38-27b-dflash2", - rtx5090: "lmsysorg/sglang:dev-qwen38-27b-dflash2", - // Multi-arch: this tag ships both linux/amd64 and linux/arm64, so it pulls - // natively on DGX Spark (GB10 is aarch64). - // Multi-arch (linux/amd64 + linux/arm64), so GB10 pulls it natively. - "dgx-spark": "lmsysorg/sglang:dev-qwen38-27b-dflash2", - gb300: "lmsysorg/sglang:dev", + h200: "lmsysorg/sglang:latest", + rtx6000: "lmsysorg/sglang:latest", + rtx5090: "lmsysorg/sglang:latest", + "dgx-spark": "lmsysorg/sglang:latest", + gb300: "lmsysorg/sglang:latest", }, github: { @@ -462,9 +518,10 @@ export const config = { // is the H200-validated setting: SM90 prefill is fast enough that a big // chunk stalls decode far less than on SM120, and the SM90 FlashInfer GDN // prefill default engages under it (fp32 state pool, chunk <= 32768). - // No NVFP4 cell on this card: SM90 has no FP4 tensor cores, so the W4A4 - // checkpoint's MLP would fall back to the Marlin W4A16 weight-only path — - // runnable, but not a recipe this page ships. + // No NVFP4 cell on this card, for any of the three exports: SM90 has no + // FP4 tensor cores, so a W4A4 checkpoint's MLP would fall back to the + // Marlin W4A16 weight-only path — runnable, but not a recipe this page + // ships. match: { hw: "h200", variant: "default", quant: "fp8", nodes: "single" }, verified: true, // DFLASH2 has not been exercised on this platform; every other overlay @@ -549,6 +606,28 @@ export const config = { "--port {{PORT}}", ], }, + { + // NVIDIA's ModelOpt export of the same W4A4 body and FP4 lm_head as the + // RadixArk FP4-head checkpoint above, so it reuses that recipe verbatim. + // Re-measured against this export on v0.5.19: all 16 overlay + // combinations (spec x tier x state dtype) serve at these pins and score + // 94.01-95.00% on the full 1319-question GSM8K. + match: { hw: "rtx6000", variant: "default", quant: "nvfp4-nvidia", nodes: "single" }, + verified: true, + env: [], + flags: [ + "--trust-remote-code", + "--model-path {{MODEL_NAME}}", + "--kv-cache-dtype fp8_e4m3", + "--mem-fraction-static 0.85", + "--attention-backend flashinfer", + "--chunked-prefill-size 2048", + "--reasoning-parser qwen3", + "--tool-call-parser qwen3_coder", + "--host {{HOST_IP}}", + "--port {{PORT}}", + ], + }, { // FP8 blockwise, ~28.5GB of weights — comfortable on 96GB. match: { hw: "rtx6000", variant: "default", quant: "fp8", nodes: "single" }, @@ -653,6 +732,40 @@ export const config = { "--port {{PORT}}", ], }, + { + // NVIDIA's ModelOpt export: same body, same FP4 lm_head, same 21.9GB of + // weights as the RadixArk FP4-head checkpoint, so the 32GB fit and every + // mem-fraction pin the overlay rows apply carry over unchanged. + match: { hw: "rtx5090", variant: "default", quant: "nvfp4-nvidia", nodes: "single" }, + // Measured on v0.5.19 against this export: all 15 offered overlay + // combinations serve and score 93.93-94.92% on the full 1319-question + // GSM8K. Every winning launch command is identical to the FP4-head + // export's, which is what "reuses that recipe verbatim" above is claiming. + verified: true, + // Rendered with the cell so nobody ships the bs=1 pins into a + // multi-user deployment unaware. + warn: + "This recipe serves ONE request at a time: --max-running-requests 1 " + + "and --cuda-graph-max-bs-decode 1 pin it to the validated single-stream " + + "envelope. To handle more concurrent requests, raise both flags " + + "together and re-derive --mamba-full-memory-ratio (and mem-fraction) " + + "with the [Mamba ratio calculator](#mamba-ratio-calculator) — on this " + + "32GB card the GDN state pool, not KV, is what runs out first.", + env: [], + flags: [ + "--trust-remote-code", + "--model-path {{MODEL_NAME}}", + "--kv-cache-dtype fp8_e4m3", + "--mem-fraction-static 0.9", + "--attention-backend flashinfer", + "--max-running-requests 1", + "--cuda-graph-max-bs-decode 1", + "--reasoning-parser qwen3", + "--tool-call-parser qwen3_coder", + "--host {{HOST_IP}}", + "--port {{PORT}}", + ], + }, // DGX Spark (GB10, SM121): single node, 128GB coherent unified memory // shared with the CPU — every checkpoint fits, so all three quants get a // cell. These cells reuse the RTX PRO 6000 recipe at one lower @@ -665,12 +778,12 @@ export const config = { // leaves ~8GB for the OS — exactly DGX OS earlyoom's SIGTERM threshold — // and the first long prefill or boot-time graph capture dips under it and // gets the scheduler killed (exit code -15, no traceback; check - // `journalctl -u earlyoom`). Re-measured on 1cf2b8c (2026-08-21): at 0.85, + // `journalctl -u earlyoom`). Re-measured on v0.5.19: at 0.85, // 15 of 48 cells were SIGTERMed, and which 15 is margin noise, biased // toward the big-state configs (bfloat16 SSM, DSPARK/DFLASH2 ratios); at // 0.80 every cell served on every attempt. // - // Validated on GB10 (SM121 / aarch64) at 1cf2b8c: all 48 configurations — + // Validated on GB10 (SM121 / aarch64) on v0.5.19: all 80 configurations — // DFLASH2 included — booted and served at ISL 8192 / OSL 1024, // concurrency 1. Boot-and-serve only -- no throughput or acceptance-length // numbers were taken, so this is a weaker standard than the SM120 pair's @@ -680,7 +793,7 @@ export const config = { // 12-cell DFLASH2 pass. { match: { hw: "dgx-spark", variant: "default", quant: "nvfp4-bf16-head", nodes: "single" }, - // All 16 overlay combinations served on GB10 at 1cf2b8c, DFLASH2 + // All 16 overlay combinations served on GB10 on v0.5.19, DFLASH2 // included — its selector folded into the draft CUDA graph in all four // of its cells here. verified: true, @@ -702,7 +815,7 @@ export const config = { // Same recipe as the BF16-head cell above: the FP4 head is smaller, // so anything that fits the bf16 head fits here with room to spare. match: { hw: "dgx-spark", variant: "default", quant: "nvfp4-fp4-head", nodes: "single" }, - // All 16 overlay combinations served on GB10 at 1cf2b8c, DFLASH2 + // All 16 overlay combinations served on GB10 on v0.5.19, DFLASH2 // included — its selector folded into the draft CUDA graph in all four // of its cells here. verified: true, @@ -720,9 +833,32 @@ export const config = { "--port {{PORT}}", ], }, + { + // NVIDIA's ModelOpt export of the same W4A4 body as the RadixArk FP4-head + // checkpoint, on that cell's recipe. + match: { hw: "dgx-spark", variant: "default", quant: "nvfp4-nvidia", nodes: "single" }, + // Re-measured against this export on v0.5.19: all 16 overlay + // combinations serve at these pins and score 94.16-95.07% on the full + // 1319-question GSM8K (float32 and bfloat16 halves run on two separate + // GB10 boxes). + verified: true, + env: [], + flags: [ + "--trust-remote-code", + "--model-path {{MODEL_NAME}}", + "--kv-cache-dtype fp8_e4m3", + "--mem-fraction-static 0.80", + "--attention-backend flashinfer", + "--chunked-prefill-size 2048", + "--reasoning-parser qwen3", + "--tool-call-parser qwen3_coder", + "--host {{HOST_IP}}", + "--port {{PORT}}", + ], + }, { match: { hw: "dgx-spark", variant: "default", quant: "fp8", nodes: "single" }, - // All 16 overlay combinations served on GB10 at 1cf2b8c, DFLASH2 + // All 16 overlay combinations served on GB10 on v0.5.19, DFLASH2 // included. This checkpoint held the sweep's most earlyoom-prone cells // at 0.85 (every bfloat16-SSM pick was killed); all clean at 0.80. verified: true, @@ -742,7 +878,7 @@ export const config = { }, { match: { hw: "dgx-spark", variant: "default", quant: "bf16", nodes: "single" }, - // All 16 overlay combinations served on GB10 at 1cf2b8c, DFLASH2 + // All 16 overlay combinations served on GB10 on v0.5.19, DFLASH2 // included. Heaviest checkpoint (52GB, ~6.5 min to load its 18 shards // from NVMe — budget ~10 min to READY before calling a boot hung). verified: true, @@ -806,6 +942,26 @@ export const config = { "--port {{PORT}}", ], }, + { + match: { hw: "gb300", variant: "default", quant: "nvfp4-nvidia", nodes: "single" }, + verified: true, + // DFLASH2 has not been exercised on this platform; every other overlay + // pick keeps this cell's original validation. + verificationStatus: (sel) => + sel.spec === "dflash" ? "in-progress" : "verified", + env: [], + flags: [ + "--trust-remote-code", + "--model-path {{MODEL_NAME}}", + "--kv-cache-dtype fp8_e4m3", + "--mem-fraction-static 0.85", + "--chunked-prefill-size 2048", + "--reasoning-parser qwen3", + "--tool-call-parser qwen3_coder", + "--host {{HOST_IP}}", + "--port {{PORT}}", + ], + }, { match: { hw: "gb300", variant: "default", quant: "fp8", nodes: "single" }, verified: true,