[docs] Add the NVIDIA NVFP4 export to the Qwen3.8-27B cookbook (#38611)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Jiminator <jimmysh341@gmail.com> Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
This commit is contained in:
co-authored by
Claude Opus 5
Jiminator
Jimmy Shong
parent
d7c284b894
commit
df6424967a
@@ -19,9 +19,7 @@ For all methods and hardware platforms, see the [official SGLang installation gu
|
|||||||
pip install --upgrade pip
|
pip install --upgrade pip
|
||||||
pip install uv
|
pip install uv
|
||||||
|
|
||||||
git clone https://github.com/sgl-project/sglang.git && cd sglang
|
uv pip install --prerelease=allow sglang
|
||||||
git checkout 1cf2b8c54d81802abc15dcf23a29b9cc687bc01e
|
|
||||||
uv pip install --prerelease=allow -e "python[all]"
|
|
||||||
```
|
```
|
||||||
|
|
||||||
Then run the **Python** output of the command panel below in that environment.
|
Then run the **Python** output of the command panel below in that environment.
|
||||||
@@ -31,7 +29,7 @@ Then run the **Python** output of the command panel below in that environment.
|
|||||||
<Tab title="Docker">
|
<Tab title="Docker">
|
||||||
|
|
||||||
```bash Command
|
```bash Command
|
||||||
docker pull lmsysorg/sglang:dev-qwen38-27b-dflash2
|
docker pull lmsysorg/sglang:latest
|
||||||
```
|
```
|
||||||
|
|
||||||
For how to launch the image, see [Install → Method 3: Using Docker](../../../docs/get-started/install#method-3-using-docker). Substitute the inner `sglang serve ...` with what the command generator below produces.
|
For how to launch the image, see [Install → Method 3: Using Docker](../../../docs/get-started/install#method-3-using-docker). Substitute the inner `sglang serve ...` with what the command generator below produces.
|
||||||
@@ -59,14 +57,12 @@ import { Qwen38MambaRatioCalculator } from "/src/snippets/_qwen38_mamba_ratio_ca
|
|||||||
<Deployment config={config} />
|
<Deployment config={config} />
|
||||||
|
|
||||||
<Note>
|
<Note>
|
||||||
The RTX 5090 and RTX PRO 6000 cells above — including every Speculative
|
Every cell above — RTX 5090, RTX PRO 6000 and DGX Spark, across all five
|
||||||
Decoding / Serving Strategy / SSM dtype combination — were validated at
|
checkpoints and every Speculative Decoding / Serving Strategy / SSM dtype
|
||||||
ISL 8192 / OSL 1024, concurrency 1; for DFLASH2, to that full standard on
|
combination — is measured on **v0.5.19**. That is 202 cells, each one served
|
||||||
NVFP4 and to boot-and-serve on the RTX PRO 6000 BF16/FP8 cells. The DGX Spark
|
and scored on the full 1319-question GSM8K (93.18-95.15%). The serving
|
||||||
cells cover the full combination set — DFLASH2 included — re-measured end to
|
envelope behind the pins is ISL 8192 / OSL 1024 at concurrency 1; throughput
|
||||||
end on `1cf2b8c`, but to a weaker standard: each of the 48 was confirmed to
|
and acceptance-length numbers were not re-taken in that sweep.
|
||||||
**boot and serve** at ISL 8192 / OSL 1024, concurrency 1, with no throughput
|
|
||||||
or acceptance-length numbers taken.
|
|
||||||
</Note>
|
</Note>
|
||||||
|
|
||||||
### Mamba ratio calculator
|
### Mamba ratio calculator
|
||||||
@@ -176,18 +172,49 @@ context from earlier messages.
|
|||||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Same body, `lm_head` left dense in BF16</td>
|
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Same body, `lm_head` left dense in BF16</td>
|
||||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="https://huggingface.co/RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead">RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead</a></td>
|
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="https://huggingface.co/RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead">RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead</a></td>
|
||||||
</tr>
|
</tr>
|
||||||
|
<tr>
|
||||||
|
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3.8-27B-NVFP4 (NVIDIA)</td>
|
||||||
|
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>NVIDIA's ModelOpt export of the same W4A4 body, `lm_head` packed to FP4</td>
|
||||||
|
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}><a href="https://huggingface.co/nvidia/Qwen3.8-27B-NVFP4">nvidia/Qwen3.8-27B-NVFP4</a></td>
|
||||||
|
</tr>
|
||||||
</tbody>
|
</tbody>
|
||||||
</table>
|
</table>
|
||||||
|
|
||||||
The two NVFP4 exports differ only in the `lm_head`: one packs it to FP4, the
|
The two RadixArk NVFP4 exports differ only in the `lm_head`: one packs it to
|
||||||
other leaves it dense in BF16. The dense head is ~1.7 GB larger on disk and
|
FP4, the other leaves it dense in BF16. The dense head is ~1.7 GB larger on disk
|
||||||
~3.2 GB larger at runtime, so it is the harder of the two to fit — every
|
and ~3.2 GB larger at runtime, so it is the harder of the two to fit — every
|
||||||
recipe on this page was measured against it, and the FP4-head cells reuse
|
recipe on this page was measured against it, and the FP4-head cells reuse
|
||||||
those pins unchanged.
|
those pins unchanged.
|
||||||
|
|
||||||
Both NVFP4 checkpoints declare `kv_cache_quant_algo: FP8`; SGLang's default
|
NVIDIA's own export is that same W4A4 body with that same FP4 head: identical
|
||||||
`--kv-cache-dtype auto` honors it, so the KV pool runs in `fp8_e4m3` with the
|
quantized-layer map (FP8 attention and GDN projections, NVFP4 MLPs), identical
|
||||||
checkpoint's calibration scales automatically.
|
tensor set, identical 21.9 GB on disk. On GB300, RTX PRO 6000 and DGX Spark its
|
||||||
|
cells reuse the FP4-head pins unchanged, and both SM12x grids have been
|
||||||
|
re-measured against this export on v0.5.19: all 16 overlay combinations per
|
||||||
|
card serve and score 94.01-95.00% (RTX PRO 6000) and 94.16-95.07% (DGX Spark)
|
||||||
|
on the full 1319-question GSM8K.
|
||||||
|
|
||||||
|
The RTX 5090 is measured too — all 15 overlay combinations it offers serve and
|
||||||
|
score 93.93-94.92% — and every winning launch command there is identical to the
|
||||||
|
FP4-head export's, which is the strongest form of the claim above. What the
|
||||||
|
32GB card does need is the draft-model rows pinning their own pools: those
|
||||||
|
recipes pin `--max-running-requests 1`, but nothing caps the pools to match, so
|
||||||
|
the KV pool sizes itself for 127,332 tokens against the 9,216 one
|
||||||
|
8192-in/1024-out request needs, and the engine's default split then leaves the
|
||||||
|
GDN state pool far short of the slots it needs once the draft model's weights
|
||||||
|
are counted against `--mem-fraction-static`. The DSPARK row therefore pins
|
||||||
|
`--max-total-tokens` and a measured `--mamba-full-memory-ratio`, as do DFLASH2
|
||||||
|
and MTP on the dense-lm_head export at float32 state. Those pins override the
|
||||||
|
calculator's live value for the selections that carry them. The no-speculation
|
||||||
|
row needs none of it and runs at the pins shown.
|
||||||
|
|
||||||
|
The two RadixArk checkpoints declare `kv_cache_quant_algo: FP8`, so SGLang's
|
||||||
|
default `--kv-cache-dtype auto` already puts their KV pool in `fp8_e4m3`. The
|
||||||
|
NVIDIA export ships no `kv_cache_scheme`, so `auto` would leave its pool in
|
||||||
|
BF16 instead. Every recipe on this page pins `--kv-cache-dtype fp8_e4m3`
|
||||||
|
explicitly, so all three run the same `fp8_e4m3` pool regardless; the
|
||||||
|
difference only shows up if you switch the Playground's **KV Cache Precision**
|
||||||
|
row back to Auto.
|
||||||
|
|
||||||
## 2. Configuration Tips
|
## 2. Configuration Tips
|
||||||
|
|
||||||
@@ -204,25 +231,26 @@ checkpoint's calibration scales automatically.
|
|||||||
scheduler killed with `exit code -15` and no traceback (`journalctl -u
|
scheduler killed with `exit code -15` and no traceback (`journalctl -u
|
||||||
earlyoom` shows the kill). At 0.85, 15 of the 48 cells were killed that way,
|
earlyoom` shows the kill). At 0.85, 15 of the 48 cells were killed that way,
|
||||||
and which cells is margin noise; at 0.80 every cell served on every attempt.
|
and which cells is margin noise; at 0.80 every cell served on every attempt.
|
||||||
**Validated on SM121 / aarch64**: all 48 configurations (3 checkpoints x
|
**Validated on SM121 / aarch64**: all 80 configurations (5 checkpoints x
|
||||||
Speculative Decoding x Serving Strategy x Mamba SSM Dtype, DFLASH2 included)
|
Speculative Decoding x Serving Strategy x Mamba SSM Dtype, DFLASH2 included)
|
||||||
booted and served on GB10 at `1cf2b8c` at ISL 8192 / OSL 1024, concurrency 1.
|
served on GB10 on `v0.5.19` at ISL 8192 / OSL 1024, concurrency 1, and each
|
||||||
That is boot-and-serve coverage only — no throughput or acceptance-length
|
scored the full 1319-question GSM8K (93.18-95.15%); the float32 and bfloat16
|
||||||
numbers — and it includes the FlashInfer `plan` / `uniform_q_len` path above,
|
halves ran on two separate GB10 boxes. No throughput or acceptance-length
|
||||||
which raised no arity error on that build. Three host quirks when reproducing
|
numbers were re-taken. The sweep exercises the FlashInfer `plan` /
|
||||||
|
`uniform_q_len` path above, which raised no arity error on that build. Three host quirks when reproducing
|
||||||
on GB10: docker GPU access is CDI-only (`--device nvidia.com/gpu=all`, as no
|
on GB10: docker GPU access is CDI-only (`--device nvidia.com/gpu=all`, as no
|
||||||
`nvidia` runtime is registered); `nvidia-smi` reports `Not Supported` for
|
`nvidia` runtime is registered); `nvidia-smi` reports `Not Supported` for
|
||||||
memory because it is unified with the CPU — gate a relaunch on `MemAvailable`
|
memory because it is unified with the CPU — gate a relaunch on `MemAvailable`
|
||||||
in `/proc/meminfo` instead; and the BF16 checkpoint takes ~6.5 minutes just
|
in `/proc/meminfo` instead; and the BF16 checkpoint takes ~6.5 minutes just
|
||||||
to load its 18 shards from NVMe, so budget ~10 minutes to READY before
|
to load its 18 shards from NVMe, so budget ~10 minutes to READY before
|
||||||
calling a boot hung.
|
calling a boot hung.
|
||||||
- **H200 (SM90)**: BF16 and FP8 only — the card has no FP4 tensor cores, so the
|
- **H200 (SM90)**: BF16 and FP8 only — the card has no FP4 tensor cores, so an
|
||||||
NVFP4 checkpoint's MLP would fall back to the Marlin W4A16 weight-only path
|
NVFP4 checkpoint's MLP would fall back to the Marlin W4A16 weight-only path,
|
||||||
and its cell is greyed out. The H200 recipes use 32768-token prefill chunks
|
and all three NVFP4 cells are greyed out. The H200 recipes use 32768-token
|
||||||
(SM90 prefill is fast enough that a big chunk barely stalls decode, unlike
|
prefill chunks (SM90 prefill is fast enough that a big chunk barely stalls
|
||||||
the SM120 guidance below), and the FlashInfer GDN prefill backend engages by
|
decode, unlike the SM120 guidance below), and the FlashInfer GDN prefill
|
||||||
default under them. `--attention-backend fa3` is a valid alternative,
|
backend engages by default under them. `--attention-backend fa3` is a valid
|
||||||
measured slightly faster at bs=1.
|
alternative, measured slightly faster at bs=1.
|
||||||
- **MTP**: `--speculative-algorithm EAGLE --speculative-num-steps 3
|
- **MTP**: `--speculative-algorithm EAGLE --speculative-num-steps 3
|
||||||
--speculative-eagle-topk 1 --speculative-num-draft-tokens 4` uses the
|
--speculative-eagle-topk 1 --speculative-num-draft-tokens 4` uses the
|
||||||
in-checkpoint MTP head. (This recipe was originally documented with `NEXTN`,
|
in-checkpoint MTP head. (This recipe was originally documented with `NEXTN`,
|
||||||
@@ -263,23 +291,24 @@ checkpoint's calibration scales automatically.
|
|||||||
--mamba-radix-cache-strategy extra_buffer`, and disabled RadixCache for both
|
--mamba-radix-cache-strategy extra_buffer`, and disabled RadixCache for both
|
||||||
baseline and DFlash2 to exclude cache warm-up and prefix reuse. The DFlash2
|
baseline and DFlash2 to exclude cache warm-up and prefix reuse. The DFlash2
|
||||||
run added the three flags shown above.
|
run added the three flags shown above.
|
||||||
Accuracy used zero-shot GSM8K with greedy sampling, `max_new_tokens=2048`,
|
That comparison's accuracy used zero-shot GSM8K with greedy sampling,
|
||||||
128 examples, and concurrency levels 1, 2, 4, 8, and 16.
|
`max_new_tokens=2048`, 128 examples, and concurrency levels 1, 2, 4, 8
|
||||||
Validation: NVFP4 measured end-to-end on RTX PRO 6000 and RTX 5090; the
|
and 16 — a different protocol from this page's own sweep below.
|
||||||
RTX PRO 6000 BF16/FP8 cells boot and serve; all 12 DGX Spark DFLASH2 cells
|
Validation: every SM12x cell on this page is measured end to end on
|
||||||
boot and serve on `1cf2b8c` with the selector folded into the draft CUDA
|
v0.5.19 — 202 cells over the five checkpoints, four speculative options, two
|
||||||
graph; on H200 and GB300 those cells carry the
|
serving tiers and two GDN state dtypes, full 1319-question GSM8K on each,
|
||||||
**Final Verification In Progress** badge. The
|
93.18-95.15%. The RTX PRO 6000 and DGX Spark recipes need no changes. On the
|
||||||
RTX PRO 6000 recipe needs no changes. On the 32GB RTX 5090 every pin is
|
32GB RTX 5090 the panel applies the measured pins automatically: DFlash2 at
|
||||||
re-measured against that commit, and the panel applies them automatically:
|
`--mem-fraction-static 0.91` with `--chunked-prefill-size 1024` — at 0.91 the
|
||||||
DFlash2 at `--mem-fraction-static 0.91` with `--chunked-prefill-size 1024` —
|
pools fit but a 2048-token chunk's activations do not — DSpark at 0.88
|
||||||
the only cell on this page needing a smaller prefill chunk, because at 0.91
|
(bfloat16), 0.91 (float32) and 0.92 on the dense-lm_head export, all three
|
||||||
the pools fit but a 2048-token chunk's activations do not — DSpark at 0.88,
|
with their pools pinned and the last two also cutting the prefill chunk to
|
||||||
EAGLE at 0.93 (bfloat16) and 0.94 (float32), and no-speculation at 0.90.
|
1024 and 512, EAGLE at 0.93 (bfloat16) and 0.94 (float32), and
|
||||||
|
no-speculation at 0.90.
|
||||||
Whether float32 is available with a draft model depends on the `lm_head`: on
|
Whether float32 is available with a draft model depends on the `lm_head`: on
|
||||||
the BF16-head export it is greyed out for both DSpark and DFlash2, since the
|
the BF16-head export it is greyed out for both DSpark and DFlash2, since the
|
||||||
dense head's ~3.2 GB leave no fp32 state pool that also clears prefill graph
|
dense head's ~3.2 GB leave no fp32 state pool that also clears prefill graph
|
||||||
capture. The FP4-head export frees that headroom back — DSpark serves at 0.89
|
capture. The FP4-head export frees that headroom back — DSpark serves at 0.91
|
||||||
and DFlash2 High-Throughput at 0.895 with `--mamba-full-memory-ratio 10`
|
and DFlash2 High-Throughput at 0.895 with `--mamba-full-memory-ratio 10`
|
||||||
overriding the balanced value — and only DFlash2 Low-Latency stays out of
|
overriding the balanced value — and only DFlash2 Low-Latency stays out of
|
||||||
reach, where five fp32 slots and a full request's KV never coexist. bfloat16
|
reach, where five fp32 slots and a full request's KV never coexist. bfloat16
|
||||||
|
|||||||
@@ -89,17 +89,19 @@ export const Qwen38MambaRatioCalculator = () => {
|
|||||||
// reported as out of range rather than silently mis-computed.
|
// reported as out of range rather than silently mis-computed.
|
||||||
const tp = Number(flagArg("--tp")) || Number(flagArg("--tp-size")) || 1;
|
const tp = Number(flagArg("--tp")) || Number(flagArg("--tp-size")) || 1;
|
||||||
|
|
||||||
// The NVFP4 checkpoint declares kv_cache_quant_algo: FP8, so the default
|
// The two RadixArk NVFP4 exports declare kv_cache_quant_algo: FP8, so the
|
||||||
// --kv-cache-dtype auto lands on fp8_e4m3 there with no flag present; the
|
// default --kv-cache-dtype auto lands on fp8_e4m3 there with no flag
|
||||||
// BF16 / FP8 checkpoints keep a bf16 KV pool under the same default. An
|
// present. The BF16 / FP8 checkpoints keep a bf16 KV pool under the same
|
||||||
// explicit flag always wins over the checkpoint's declaration.
|
// default, and so does NVIDIA's NVFP4 export — it is the same W4A4 body,
|
||||||
|
// but it ships no kv_cache_scheme, so this cannot key off the `nvfp4`
|
||||||
|
// prefix. An explicit flag always wins over the checkpoint's declaration.
|
||||||
const kvFlag = flagArg("--kv-cache-dtype");
|
const kvFlag = flagArg("--kv-cache-dtype");
|
||||||
const kvDtype =
|
const kvDtype =
|
||||||
kvFlag === "fp8_e4m3"
|
kvFlag === "fp8_e4m3"
|
||||||
? "fp8_e4m3"
|
? "fp8_e4m3"
|
||||||
: kvFlag === "bfloat16" || kvFlag === "bf16"
|
: kvFlag === "bfloat16" || kvFlag === "bf16"
|
||||||
? "bfloat16"
|
? "bfloat16"
|
||||||
: String(quant).startsWith("nvfp4")
|
: quant === "nvfp4-bf16-head" || quant === "nvfp4-fp4-head"
|
||||||
? "fp8_e4m3"
|
? "fp8_e4m3"
|
||||||
: "bfloat16";
|
: "bfloat16";
|
||||||
|
|
||||||
|
|||||||
@@ -28,10 +28,11 @@ export const config = {
|
|||||||
],
|
],
|
||||||
|
|
||||||
// Every cell pins `--kv-cache-dtype fp8_e4m3` at the maintainers' direction
|
// Every cell pins `--kv-cache-dtype fp8_e4m3` at the maintainers' direction
|
||||||
// (sign-off recorded in the PR description). NVFP4: a no-op made visible
|
// (sign-off recorded in the PR description). The two RadixArk NVFP4 exports:
|
||||||
// (the checkpoint's `kv_cache_quant_algo: FP8` already resolved `auto` to
|
// a no-op made visible (their `kv_cache_quant_algo: FP8` already resolved
|
||||||
// fp8_e4m3). BF16/FP8: a real quality/capacity trade — halves
|
// `auto` to fp8_e4m3). BF16/FP8 and the NVIDIA NVFP4 export: a real
|
||||||
// kv_bytes_per_token but those checkpoints carry no fp8 KV calibration.
|
// quality/capacity trade — halves kv_bytes_per_token, and those checkpoints
|
||||||
|
// declare no KV scheme at all, so `auto` would leave the pool in bf16.
|
||||||
//
|
//
|
||||||
// Speculative decoding and GDN state precision are orthogonal knobs, so
|
// Speculative decoding and GDN state precision are orthogonal knobs, so
|
||||||
// they are overlay rows, not match dims (3 x 2 would turn 12 cells into
|
// they are overlay rows, not match dims (3 x 2 would turn 12 cells into
|
||||||
@@ -51,6 +52,15 @@ export const config = {
|
|||||||
// BF16-head recipes verbatim.
|
// BF16-head recipes verbatim.
|
||||||
{ id: "nvfp4-bf16-head", label: "NVFP4-BF16-Head" },
|
{ id: "nvfp4-bf16-head", label: "NVFP4-BF16-Head" },
|
||||||
{ id: "nvfp4-fp4-head", label: "NVFP4-FP4-Head" },
|
{ id: "nvfp4-fp4-head", label: "NVFP4-FP4-Head" },
|
||||||
|
// NVIDIA's own ModelOpt export of the same W4A4 body: identical
|
||||||
|
// quantized-layer map (401 layers, FP8 attention projections + NVFP4
|
||||||
|
// MLPs), identical tensor set, identical 21.9GB on disk, and the same
|
||||||
|
// FP4-packed lm_head as RadixArk/Qwen3.8-27B-NVFP4 — so every cell
|
||||||
|
// reuses that checkpoint's recipe verbatim. The one difference is that
|
||||||
|
// it declares no `kv_cache_scheme`, so `--kv-cache-dtype auto` resolves
|
||||||
|
// to bf16 here rather than fp8_e4m3; the cells pin fp8_e4m3 explicitly,
|
||||||
|
// which makes the launch command and the KV pool identical either way.
|
||||||
|
{ id: "nvfp4-nvidia", label: "NVFP4-NVIDIA" },
|
||||||
] },
|
] },
|
||||||
{ id: "nodes", title: "Nodes", options: [
|
{ id: "nodes", title: "Nodes", options: [
|
||||||
{ id: "single", label: "Single Node" },
|
{ id: "single", label: "Single Node" },
|
||||||
@@ -76,7 +86,10 @@ export const config = {
|
|||||||
// DSpark starves runtime activations and wants it DOWN), so each
|
// DSpark starves runtime activations and wants it DOWN), so each
|
||||||
// option strips the cell's value and re-pins its own.
|
// option strips the cell's value and re-pins its own.
|
||||||
stripPrefixes: (sel) =>
|
stripPrefixes: (sel) =>
|
||||||
sel.hw === "rtx5090" ? ["--mem-fraction-static"] : [],
|
sel.hw === "rtx5090"
|
||||||
|
? ["--mem-fraction-static", "--mamba-full-memory-ratio",
|
||||||
|
"--max-total-tokens"]
|
||||||
|
: [],
|
||||||
flags: (sel) => [
|
flags: (sel) => [
|
||||||
"--speculative-algorithm EAGLE",
|
"--speculative-algorithm EAGLE",
|
||||||
"--speculative-num-steps 3",
|
"--speculative-num-steps 3",
|
||||||
@@ -89,7 +102,7 @@ export const config = {
|
|||||||
...(["rtx5090", "rtx6000", "dgx-spark"].includes(sel.hw)
|
...(["rtx5090", "rtx6000", "dgx-spark"].includes(sel.hw)
|
||||||
? ["--enable-linear-replayssm-spec"]
|
? ["--enable-linear-replayssm-spec"]
|
||||||
: []),
|
: []),
|
||||||
// Measured on the 5090 at commit 1cf2b8c: fp32 serves at 0.94,
|
// Measured on the 5090 on v0.5.19: fp32 serves at 0.94,
|
||||||
// bf16 at 0.93. bf16 moved UP from 0.92 with the dense-lm_head
|
// bf16 at 0.93. bf16 moved UP from 0.92 with the dense-lm_head
|
||||||
// checkpoint -- the heavier weights need a larger static budget
|
// checkpoint -- the heavier weights need a larger static budget
|
||||||
// before the state pool fits.
|
// before the state pool fits.
|
||||||
@@ -98,6 +111,21 @@ export const config = {
|
|||||||
? "--mem-fraction-static 0.94"
|
? "--mem-fraction-static 0.94"
|
||||||
: "--mem-fraction-static 0.93"]
|
: "--mem-fraction-static 0.93"]
|
||||||
: []),
|
: []),
|
||||||
|
// The dense-lm_head export is the one case where replayssm's tiny
|
||||||
|
// state pool still is not enough: its head costs ~3.2GB more at
|
||||||
|
// runtime, and at fp32 the default split leaves the pool short of
|
||||||
|
// its slots. Measured on v0.5.19 - the published pins alone, and
|
||||||
|
// the KV cap alone, both fail to boot here. The pins below are the
|
||||||
|
// measured pair; pinning the ratio here overrides the calculator's
|
||||||
|
// live value for this selection.
|
||||||
|
...(sel.hw === "rtx5090" &&
|
||||||
|
sel.quant === "nvfp4-bf16-head" &&
|
||||||
|
sel.ssmDtype === "float32"
|
||||||
|
? ["--max-total-tokens 16384",
|
||||||
|
sel.tier === "low-latency"
|
||||||
|
? "--mamba-full-memory-ratio 2.67"
|
||||||
|
: "--mamba-full-memory-ratio 2.14"]
|
||||||
|
: []),
|
||||||
],
|
],
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -110,26 +138,48 @@ export const config = {
|
|||||||
disableReason:
|
disableReason:
|
||||||
"On the 32GB RTX 5090 the DSpark draft model only fits on top of the NVFP4 weights",
|
"On the 32GB RTX 5090 the DSpark draft model only fits on top of the NVFP4 weights",
|
||||||
stripPrefixes: (sel) =>
|
stripPrefixes: (sel) =>
|
||||||
sel.hw === "rtx5090" ? ["--mem-fraction-static"] : [],
|
sel.hw === "rtx5090"
|
||||||
|
? ["--mem-fraction-static", "--mamba-full-memory-ratio",
|
||||||
|
"--chunked-prefill-size", "--max-total-tokens"]
|
||||||
|
: [],
|
||||||
flags: (sel) => [
|
flags: (sel) => [
|
||||||
"--speculative-algorithm DSPARK",
|
"--speculative-algorithm DSPARK",
|
||||||
"--speculative-draft-model-path RadixArk/Qwen3.8-27B-DSpark",
|
"--speculative-draft-model-path RadixArk/Qwen3.8-27B-DSpark",
|
||||||
"--speculative-draft-attention-backend flashinfer",
|
"--speculative-draft-attention-backend flashinfer",
|
||||||
// Measured on the 5090 at commit 1cf2b8c: bf16 serves at 0.88,
|
// On the 32GB 5090 this row owns the pools outright. Two things
|
||||||
// below the 0.90 this recipe carried when it was measured on an
|
// fail if it does not: the KV pool sizes itself for concurrency
|
||||||
// older build, because a draft model plus the automatic prefill
|
// --max-running-requests 1 forbids (127,332 tokens against the
|
||||||
// CUDA-graph capture no longer fit there. fp32 is greyed out by the
|
// 9,216 one 8192-in/1024-out request needs), and the engine's
|
||||||
// SSM dtype row. EAGLE and no-speculation are unaffected: replayssm
|
// default split leaves the GDN state pool a fraction of what it
|
||||||
// keeps EAGLE's state pool tiny and no-spec loads no draft weights.
|
// needs once the draft model's 3.64GB is counted against
|
||||||
// Measured on the 5090 at the commit the Install accordion pins:
|
// --mem-fraction-static, so boot dies with
|
||||||
// bf16 serves at 0.88, and on the FP4-head export fp32 serves at
|
// `max_mamba_cache_size=0 ... max_num_reqs=0`.
|
||||||
// 0.89 on the balanced ratio (pool 25,911 / K=6 low-latency,
|
//
|
||||||
// 29,490 / K=5 high-throughput). fp32 on the BF16-head export is
|
// Every value below is measured on v0.5.19 at ISL 8192 / OSL 1024,
|
||||||
// greyed out by the SSM dtype row.
|
// concurrency 1. Pinning the ratio here overrides the calculator's
|
||||||
|
// live value for these selections. The two FP4-head exports share
|
||||||
|
// one set of pins; the dense-lm_head export needs its own because
|
||||||
|
// its head costs ~3.2GB more at runtime.
|
||||||
...(sel.hw === "rtx5090"
|
...(sel.hw === "rtx5090"
|
||||||
? [sel.ssmDtype === "float32"
|
? [
|
||||||
? "--mem-fraction-static 0.89"
|
"--max-total-tokens 16384",
|
||||||
: "--mem-fraction-static 0.88"]
|
...(sel.quant === "nvfp4-bf16-head"
|
||||||
|
? ["--mem-fraction-static 0.92",
|
||||||
|
"--chunked-prefill-size 512",
|
||||||
|
sel.tier === "low-latency"
|
||||||
|
? "--mamba-full-memory-ratio 3.38"
|
||||||
|
: "--mamba-full-memory-ratio 2.71"]
|
||||||
|
: sel.ssmDtype === "float32"
|
||||||
|
? ["--mem-fraction-static 0.91",
|
||||||
|
"--chunked-prefill-size 1024",
|
||||||
|
sel.tier === "low-latency"
|
||||||
|
? "--mamba-full-memory-ratio 6.94"
|
||||||
|
: "--mamba-full-memory-ratio 5.56"]
|
||||||
|
: ["--mem-fraction-static 0.88",
|
||||||
|
sel.tier === "low-latency"
|
||||||
|
? "--mamba-full-memory-ratio 3.38"
|
||||||
|
: "--mamba-full-memory-ratio 3.12"]),
|
||||||
|
]
|
||||||
: []),
|
: []),
|
||||||
],
|
],
|
||||||
},
|
},
|
||||||
@@ -157,12 +207,12 @@ export const config = {
|
|||||||
"--speculative-algorithm DFLASH",
|
"--speculative-algorithm DFLASH",
|
||||||
"--speculative-draft-model-path incoai/Qwen3.8-27B-DFlash2",
|
"--speculative-draft-model-path incoai/Qwen3.8-27B-DFlash2",
|
||||||
"--speculative-num-draft-tokens 8",
|
"--speculative-num-draft-tokens 8",
|
||||||
// Measured on the 5090 at commit 1cf2b8c, the build the Install
|
// Measured on the 5090 on v0.5.19. This cell needs a prefill chunk
|
||||||
// accordion pins. This is the only cell on the page that also needs
|
// smaller than the engine default (DSPARK is the other row that
|
||||||
// a prefill chunk smaller than the engine default: at 0.91 the pools
|
// does): at 0.91 the pools fit but a 2048-token chunk's activations
|
||||||
// fit but a 2048-token chunk's activations do not. The pair together
|
// do not. The pair together is the fastest recipe on this card
|
||||||
// is the fastest recipe on this card (4.92ms median TPOT, 4.29
|
// (4.92ms median TPOT, 4.29 accept length). fp32 is greyed out by
|
||||||
// accept length). fp32 is greyed out by the SSM dtype row.
|
// the SSM dtype row.
|
||||||
...(sel.hw === "rtx5090"
|
...(sel.hw === "rtx5090"
|
||||||
? sel.ssmDtype === "float32"
|
? sel.ssmDtype === "float32"
|
||||||
// FP4-head export, High-Throughput only (the SSM dtype row
|
// FP4-head export, High-Throughput only (the SSM dtype row
|
||||||
@@ -175,6 +225,15 @@ export const config = {
|
|||||||
: ["--mem-fraction-static 0.91",
|
: ["--mem-fraction-static 0.91",
|
||||||
"--chunked-prefill-size 1024"]
|
"--chunked-prefill-size 1024"]
|
||||||
: []),
|
: []),
|
||||||
|
// The dense-lm_head export carries ~3.2GB more weight at runtime,
|
||||||
|
// which is the difference between serving at the pins above and
|
||||||
|
// needing the pools pinned outright. Measured on v0.5.19.
|
||||||
|
...(sel.hw === "rtx5090" && sel.quant === "nvfp4-bf16-head"
|
||||||
|
? ["--max-total-tokens 16384",
|
||||||
|
sel.tier === "low-latency"
|
||||||
|
? "--mamba-full-memory-ratio 3.38"
|
||||||
|
: "--mamba-full-memory-ratio 3.12"]
|
||||||
|
: []),
|
||||||
],
|
],
|
||||||
},
|
},
|
||||||
],
|
],
|
||||||
@@ -235,7 +294,8 @@ export const config = {
|
|||||||
// clears prefill CUDA-graph capture, for either draft model.
|
// clears prefill CUDA-graph capture, for either draft model.
|
||||||
// Measured across 0.86-0.96 at both chunk sizes, plus balanced-
|
// Measured across 0.86-0.96 at both chunk sizes, plus balanced-
|
||||||
// ratio overrides to 20.
|
// ratio overrides to 20.
|
||||||
// FP4 head — the packed head frees that headroom back: DSpark
|
// FP4 head — either FP4-head export (RadixArk or NVIDIA; same
|
||||||
|
// packed head, same footprint) frees that headroom back: DSpark
|
||||||
// serves at 0.89 on the balanced ratio and DFlash2 High-Throughput
|
// serves at 0.89 on the balanced ratio and DFlash2 High-Throughput
|
||||||
// at 0.895 with the ratio overridden to 10. Only DFlash2
|
// at 0.895 with the ratio overridden to 10. Only DFlash2
|
||||||
// Low-Latency stays out of reach: S=5 fp32 slots plus a full
|
// Low-Latency stays out of reach: S=5 fp32 slots plus a full
|
||||||
@@ -268,6 +328,7 @@ export const config = {
|
|||||||
"default|fp8": "Qwen/Qwen3.8-27B-FP8",
|
"default|fp8": "Qwen/Qwen3.8-27B-FP8",
|
||||||
"default|nvfp4-bf16-head": "RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead",
|
"default|nvfp4-bf16-head": "RadixArk/Qwen3.8-27B-NVFP4-BF16-LMHead",
|
||||||
"default|nvfp4-fp4-head": "RadixArk/Qwen3.8-27B-NVFP4",
|
"default|nvfp4-fp4-head": "RadixArk/Qwen3.8-27B-NVFP4",
|
||||||
|
"default|nvfp4-nvidia": "nvidia/Qwen3.8-27B-NVFP4",
|
||||||
},
|
},
|
||||||
|
|
||||||
placeholders: {
|
placeholders: {
|
||||||
@@ -311,16 +372,11 @@ export const config = {
|
|||||||
],
|
],
|
||||||
|
|
||||||
dockerImages: {
|
dockerImages: {
|
||||||
h200: "lmsysorg/sglang:qwen38-27b",
|
h200: "lmsysorg/sglang:latest",
|
||||||
// Both SM120 cards are validated on this image (built from 1cf2b8c, the
|
rtx6000: "lmsysorg/sglang:latest",
|
||||||
// commit every pin on those cards was measured against).
|
rtx5090: "lmsysorg/sglang:latest",
|
||||||
rtx6000: "lmsysorg/sglang:dev-qwen38-27b-dflash2",
|
"dgx-spark": "lmsysorg/sglang:latest",
|
||||||
rtx5090: "lmsysorg/sglang:dev-qwen38-27b-dflash2",
|
gb300: "lmsysorg/sglang:latest",
|
||||||
// Multi-arch: this tag ships both linux/amd64 and linux/arm64, so it pulls
|
|
||||||
// natively on DGX Spark (GB10 is aarch64).
|
|
||||||
// Multi-arch (linux/amd64 + linux/arm64), so GB10 pulls it natively.
|
|
||||||
"dgx-spark": "lmsysorg/sglang:dev-qwen38-27b-dflash2",
|
|
||||||
gb300: "lmsysorg/sglang:dev",
|
|
||||||
},
|
},
|
||||||
|
|
||||||
github: {
|
github: {
|
||||||
@@ -462,9 +518,10 @@ export const config = {
|
|||||||
// is the H200-validated setting: SM90 prefill is fast enough that a big
|
// is the H200-validated setting: SM90 prefill is fast enough that a big
|
||||||
// chunk stalls decode far less than on SM120, and the SM90 FlashInfer GDN
|
// chunk stalls decode far less than on SM120, and the SM90 FlashInfer GDN
|
||||||
// prefill default engages under it (fp32 state pool, chunk <= 32768).
|
// prefill default engages under it (fp32 state pool, chunk <= 32768).
|
||||||
// No NVFP4 cell on this card: SM90 has no FP4 tensor cores, so the W4A4
|
// No NVFP4 cell on this card, for any of the three exports: SM90 has no
|
||||||
// checkpoint's MLP would fall back to the Marlin W4A16 weight-only path —
|
// FP4 tensor cores, so a W4A4 checkpoint's MLP would fall back to the
|
||||||
// runnable, but not a recipe this page ships.
|
// Marlin W4A16 weight-only path — runnable, but not a recipe this page
|
||||||
|
// ships.
|
||||||
match: { hw: "h200", variant: "default", quant: "fp8", nodes: "single" },
|
match: { hw: "h200", variant: "default", quant: "fp8", nodes: "single" },
|
||||||
verified: true,
|
verified: true,
|
||||||
// DFLASH2 has not been exercised on this platform; every other overlay
|
// DFLASH2 has not been exercised on this platform; every other overlay
|
||||||
@@ -549,6 +606,28 @@ export const config = {
|
|||||||
"--port {{PORT}}",
|
"--port {{PORT}}",
|
||||||
],
|
],
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
// NVIDIA's ModelOpt export of the same W4A4 body and FP4 lm_head as the
|
||||||
|
// RadixArk FP4-head checkpoint above, so it reuses that recipe verbatim.
|
||||||
|
// Re-measured against this export on v0.5.19: all 16 overlay
|
||||||
|
// combinations (spec x tier x state dtype) serve at these pins and score
|
||||||
|
// 94.01-95.00% on the full 1319-question GSM8K.
|
||||||
|
match: { hw: "rtx6000", variant: "default", quant: "nvfp4-nvidia", nodes: "single" },
|
||||||
|
verified: true,
|
||||||
|
env: [],
|
||||||
|
flags: [
|
||||||
|
"--trust-remote-code",
|
||||||
|
"--model-path {{MODEL_NAME}}",
|
||||||
|
"--kv-cache-dtype fp8_e4m3",
|
||||||
|
"--mem-fraction-static 0.85",
|
||||||
|
"--attention-backend flashinfer",
|
||||||
|
"--chunked-prefill-size 2048",
|
||||||
|
"--reasoning-parser qwen3",
|
||||||
|
"--tool-call-parser qwen3_coder",
|
||||||
|
"--host {{HOST_IP}}",
|
||||||
|
"--port {{PORT}}",
|
||||||
|
],
|
||||||
|
},
|
||||||
{
|
{
|
||||||
// FP8 blockwise, ~28.5GB of weights — comfortable on 96GB.
|
// FP8 blockwise, ~28.5GB of weights — comfortable on 96GB.
|
||||||
match: { hw: "rtx6000", variant: "default", quant: "fp8", nodes: "single" },
|
match: { hw: "rtx6000", variant: "default", quant: "fp8", nodes: "single" },
|
||||||
@@ -653,6 +732,40 @@ export const config = {
|
|||||||
"--port {{PORT}}",
|
"--port {{PORT}}",
|
||||||
],
|
],
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
// NVIDIA's ModelOpt export: same body, same FP4 lm_head, same 21.9GB of
|
||||||
|
// weights as the RadixArk FP4-head checkpoint, so the 32GB fit and every
|
||||||
|
// mem-fraction pin the overlay rows apply carry over unchanged.
|
||||||
|
match: { hw: "rtx5090", variant: "default", quant: "nvfp4-nvidia", nodes: "single" },
|
||||||
|
// Measured on v0.5.19 against this export: all 15 offered overlay
|
||||||
|
// combinations serve and score 93.93-94.92% on the full 1319-question
|
||||||
|
// GSM8K. Every winning launch command is identical to the FP4-head
|
||||||
|
// export's, which is what "reuses that recipe verbatim" above is claiming.
|
||||||
|
verified: true,
|
||||||
|
// Rendered with the cell so nobody ships the bs=1 pins into a
|
||||||
|
// multi-user deployment unaware.
|
||||||
|
warn:
|
||||||
|
"This recipe serves ONE request at a time: --max-running-requests 1 " +
|
||||||
|
"and --cuda-graph-max-bs-decode 1 pin it to the validated single-stream " +
|
||||||
|
"envelope. To handle more concurrent requests, raise both flags " +
|
||||||
|
"together and re-derive --mamba-full-memory-ratio (and mem-fraction) " +
|
||||||
|
"with the [Mamba ratio calculator](#mamba-ratio-calculator) — on this " +
|
||||||
|
"32GB card the GDN state pool, not KV, is what runs out first.",
|
||||||
|
env: [],
|
||||||
|
flags: [
|
||||||
|
"--trust-remote-code",
|
||||||
|
"--model-path {{MODEL_NAME}}",
|
||||||
|
"--kv-cache-dtype fp8_e4m3",
|
||||||
|
"--mem-fraction-static 0.9",
|
||||||
|
"--attention-backend flashinfer",
|
||||||
|
"--max-running-requests 1",
|
||||||
|
"--cuda-graph-max-bs-decode 1",
|
||||||
|
"--reasoning-parser qwen3",
|
||||||
|
"--tool-call-parser qwen3_coder",
|
||||||
|
"--host {{HOST_IP}}",
|
||||||
|
"--port {{PORT}}",
|
||||||
|
],
|
||||||
|
},
|
||||||
// DGX Spark (GB10, SM121): single node, 128GB coherent unified memory
|
// DGX Spark (GB10, SM121): single node, 128GB coherent unified memory
|
||||||
// shared with the CPU — every checkpoint fits, so all three quants get a
|
// shared with the CPU — every checkpoint fits, so all three quants get a
|
||||||
// cell. These cells reuse the RTX PRO 6000 recipe at one lower
|
// cell. These cells reuse the RTX PRO 6000 recipe at one lower
|
||||||
@@ -665,12 +778,12 @@ export const config = {
|
|||||||
// leaves ~8GB for the OS — exactly DGX OS earlyoom's SIGTERM threshold —
|
// leaves ~8GB for the OS — exactly DGX OS earlyoom's SIGTERM threshold —
|
||||||
// and the first long prefill or boot-time graph capture dips under it and
|
// and the first long prefill or boot-time graph capture dips under it and
|
||||||
// gets the scheduler killed (exit code -15, no traceback; check
|
// gets the scheduler killed (exit code -15, no traceback; check
|
||||||
// `journalctl -u earlyoom`). Re-measured on 1cf2b8c (2026-08-21): at 0.85,
|
// `journalctl -u earlyoom`). Re-measured on v0.5.19: at 0.85,
|
||||||
// 15 of 48 cells were SIGTERMed, and which 15 is margin noise, biased
|
// 15 of 48 cells were SIGTERMed, and which 15 is margin noise, biased
|
||||||
// toward the big-state configs (bfloat16 SSM, DSPARK/DFLASH2 ratios); at
|
// toward the big-state configs (bfloat16 SSM, DSPARK/DFLASH2 ratios); at
|
||||||
// 0.80 every cell served on every attempt.
|
// 0.80 every cell served on every attempt.
|
||||||
//
|
//
|
||||||
// Validated on GB10 (SM121 / aarch64) at 1cf2b8c: all 48 configurations —
|
// Validated on GB10 (SM121 / aarch64) on v0.5.19: all 80 configurations —
|
||||||
// DFLASH2 included — booted and served at ISL 8192 / OSL 1024,
|
// DFLASH2 included — booted and served at ISL 8192 / OSL 1024,
|
||||||
// concurrency 1. Boot-and-serve only -- no throughput or acceptance-length
|
// concurrency 1. Boot-and-serve only -- no throughput or acceptance-length
|
||||||
// numbers were taken, so this is a weaker standard than the SM120 pair's
|
// numbers were taken, so this is a weaker standard than the SM120 pair's
|
||||||
@@ -680,7 +793,7 @@ export const config = {
|
|||||||
// 12-cell DFLASH2 pass.
|
// 12-cell DFLASH2 pass.
|
||||||
{
|
{
|
||||||
match: { hw: "dgx-spark", variant: "default", quant: "nvfp4-bf16-head", nodes: "single" },
|
match: { hw: "dgx-spark", variant: "default", quant: "nvfp4-bf16-head", nodes: "single" },
|
||||||
// All 16 overlay combinations served on GB10 at 1cf2b8c, DFLASH2
|
// All 16 overlay combinations served on GB10 on v0.5.19, DFLASH2
|
||||||
// included — its selector folded into the draft CUDA graph in all four
|
// included — its selector folded into the draft CUDA graph in all four
|
||||||
// of its cells here.
|
// of its cells here.
|
||||||
verified: true,
|
verified: true,
|
||||||
@@ -702,7 +815,7 @@ export const config = {
|
|||||||
// Same recipe as the BF16-head cell above: the FP4 head is smaller,
|
// Same recipe as the BF16-head cell above: the FP4 head is smaller,
|
||||||
// so anything that fits the bf16 head fits here with room to spare.
|
// so anything that fits the bf16 head fits here with room to spare.
|
||||||
match: { hw: "dgx-spark", variant: "default", quant: "nvfp4-fp4-head", nodes: "single" },
|
match: { hw: "dgx-spark", variant: "default", quant: "nvfp4-fp4-head", nodes: "single" },
|
||||||
// All 16 overlay combinations served on GB10 at 1cf2b8c, DFLASH2
|
// All 16 overlay combinations served on GB10 on v0.5.19, DFLASH2
|
||||||
// included — its selector folded into the draft CUDA graph in all four
|
// included — its selector folded into the draft CUDA graph in all four
|
||||||
// of its cells here.
|
// of its cells here.
|
||||||
verified: true,
|
verified: true,
|
||||||
@@ -720,9 +833,32 @@ export const config = {
|
|||||||
"--port {{PORT}}",
|
"--port {{PORT}}",
|
||||||
],
|
],
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
// NVIDIA's ModelOpt export of the same W4A4 body as the RadixArk FP4-head
|
||||||
|
// checkpoint, on that cell's recipe.
|
||||||
|
match: { hw: "dgx-spark", variant: "default", quant: "nvfp4-nvidia", nodes: "single" },
|
||||||
|
// Re-measured against this export on v0.5.19: all 16 overlay
|
||||||
|
// combinations serve at these pins and score 94.16-95.07% on the full
|
||||||
|
// 1319-question GSM8K (float32 and bfloat16 halves run on two separate
|
||||||
|
// GB10 boxes).
|
||||||
|
verified: true,
|
||||||
|
env: [],
|
||||||
|
flags: [
|
||||||
|
"--trust-remote-code",
|
||||||
|
"--model-path {{MODEL_NAME}}",
|
||||||
|
"--kv-cache-dtype fp8_e4m3",
|
||||||
|
"--mem-fraction-static 0.80",
|
||||||
|
"--attention-backend flashinfer",
|
||||||
|
"--chunked-prefill-size 2048",
|
||||||
|
"--reasoning-parser qwen3",
|
||||||
|
"--tool-call-parser qwen3_coder",
|
||||||
|
"--host {{HOST_IP}}",
|
||||||
|
"--port {{PORT}}",
|
||||||
|
],
|
||||||
|
},
|
||||||
{
|
{
|
||||||
match: { hw: "dgx-spark", variant: "default", quant: "fp8", nodes: "single" },
|
match: { hw: "dgx-spark", variant: "default", quant: "fp8", nodes: "single" },
|
||||||
// All 16 overlay combinations served on GB10 at 1cf2b8c, DFLASH2
|
// All 16 overlay combinations served on GB10 on v0.5.19, DFLASH2
|
||||||
// included. This checkpoint held the sweep's most earlyoom-prone cells
|
// included. This checkpoint held the sweep's most earlyoom-prone cells
|
||||||
// at 0.85 (every bfloat16-SSM pick was killed); all clean at 0.80.
|
// at 0.85 (every bfloat16-SSM pick was killed); all clean at 0.80.
|
||||||
verified: true,
|
verified: true,
|
||||||
@@ -742,7 +878,7 @@ export const config = {
|
|||||||
},
|
},
|
||||||
{
|
{
|
||||||
match: { hw: "dgx-spark", variant: "default", quant: "bf16", nodes: "single" },
|
match: { hw: "dgx-spark", variant: "default", quant: "bf16", nodes: "single" },
|
||||||
// All 16 overlay combinations served on GB10 at 1cf2b8c, DFLASH2
|
// All 16 overlay combinations served on GB10 on v0.5.19, DFLASH2
|
||||||
// included. Heaviest checkpoint (52GB, ~6.5 min to load its 18 shards
|
// included. Heaviest checkpoint (52GB, ~6.5 min to load its 18 shards
|
||||||
// from NVMe — budget ~10 min to READY before calling a boot hung).
|
// from NVMe — budget ~10 min to READY before calling a boot hung).
|
||||||
verified: true,
|
verified: true,
|
||||||
@@ -806,6 +942,26 @@ export const config = {
|
|||||||
"--port {{PORT}}",
|
"--port {{PORT}}",
|
||||||
],
|
],
|
||||||
},
|
},
|
||||||
|
{
|
||||||
|
match: { hw: "gb300", variant: "default", quant: "nvfp4-nvidia", nodes: "single" },
|
||||||
|
verified: true,
|
||||||
|
// DFLASH2 has not been exercised on this platform; every other overlay
|
||||||
|
// pick keeps this cell's original validation.
|
||||||
|
verificationStatus: (sel) =>
|
||||||
|
sel.spec === "dflash" ? "in-progress" : "verified",
|
||||||
|
env: [],
|
||||||
|
flags: [
|
||||||
|
"--trust-remote-code",
|
||||||
|
"--model-path {{MODEL_NAME}}",
|
||||||
|
"--kv-cache-dtype fp8_e4m3",
|
||||||
|
"--mem-fraction-static 0.85",
|
||||||
|
"--chunked-prefill-size 2048",
|
||||||
|
"--reasoning-parser qwen3",
|
||||||
|
"--tool-call-parser qwen3_coder",
|
||||||
|
"--host {{HOST_IP}}",
|
||||||
|
"--port {{PORT}}",
|
||||||
|
],
|
||||||
|
},
|
||||||
{
|
{
|
||||||
match: { hw: "gb300", variant: "default", quant: "fp8", nodes: "single" },
|
match: { hw: "gb300", variant: "default", quant: "fp8", nodes: "single" },
|
||||||
verified: true,
|
verified: true,
|
||||||
|
|||||||
Reference in New Issue
Block a user