[Cookbook] Enable DSpark on the DeepSeek-V4 Flash Vision low-latency recipes (#37301)
This commit is contained in:
@@ -225,7 +225,7 @@ For the original Flash and Pro checkpoints:
|
|||||||
- **Verified matrix** — MMMU-Pro via sgl-eval at `temperature 1.0`, `top-p 0.95`, `--reasoning-effort max`.
|
- **Verified matrix** — MMMU-Pro via sgl-eval at `temperature 1.0`, `top-p 0.95`, `--reasoning-effort max`.
|
||||||
- **Engine auto-configuration** — the engine picks the `flashinfer_mxfp4` MoE runner and auto-disables shared-experts fusion for this checkpoint (its HashTopK routing rejects fused shared experts); don't pass `--enforce-shared-experts-fusion`.
|
- **Engine auto-configuration** — the engine picks the `flashinfer_mxfp4` MoE runner and auto-disables shared-experts fusion for this checkpoint (its HashTopK routing rejects fused shared experts); don't pass `--enforce-shared-experts-fusion`.
|
||||||
- **Chunked prefill & radix cache stay enabled** — the scheduler keeps image spans consistent automatically: chunked-prefill truncation points are span-aligned (an image span always prefills within a single extend, overshooting the chunk budget by at most one span), and a radix-cache prefix match ending deep inside an image span is re-issued from the span start.
|
- **Chunked prefill & radix cache stay enabled** — the scheduler keeps image spans consistent automatically: chunked-prefill truncation points are span-aligned (an image span always prefills within a single extend, overshooting the chunk budget by at most one span), and a radix-cache prefix match ending deep inside an image span is re-issued from the span start.
|
||||||
- **Speculative decoding** — the checkpoint bundles a DSpark head, but MTP/DSpark with image batches is not yet verified, so all Flash Vision recipes run target-only. As on the 0731/0813 checkpoints, do not pass the EAGLE flags.
|
- **Speculative decoding** — the checkpoint bundles a DSpark head, and the low-latency recipes enable it with `--speculative-algorithm DSPARK` (verified with image batches on B200 via the MMMU-Pro round; the other hardware rows are pending verification). The balanced and high-throughput recipes run target-only: they use DP Attention, which DSpark is incompatible with on current releases. As on the 0731/0813 checkpoints, do not pass the EAGLE flags.
|
||||||
|
|
||||||
**Shared experts fusion (Blackwell, flashinfer_mxfp4)**
|
**Shared experts fusion (Blackwell, flashinfer_mxfp4)**
|
||||||
|
|
||||||
@@ -625,7 +625,7 @@ For more details, see the [HiCache documentation](../../../docs/advanced_feature
|
|||||||
|
|
||||||
Flash Official (0731) and Pro Official (0813) bundle a DSpark draft head in `deepseek-ai/DeepSeek-V4-Flash-0731` and `deepseek-ai/DeepSeek-V4-Pro-0813`. The target and draft weights therefore come from the same checkpoint: enable DSpark with `--speculative-algorithm DSPARK` and do not set a separate `--speculative-draft-model-path`.
|
Flash Official (0731) and Pro Official (0813) bundle a DSpark draft head in `deepseek-ai/DeepSeek-V4-Flash-0731` and `deepseek-ai/DeepSeek-V4-Pro-0813`. The target and draft weights therefore come from the same checkpoint: enable DSpark with `--speculative-algorithm DSPARK` and do not set a separate `--speculative-draft-model-path`.
|
||||||
|
|
||||||
The experimental [Flash Vision checkpoint](#vision-note) also bundles a DSpark head, but speculative decoding with image batches is not yet verified on it — the Flash Vision recipes run target-only for now, and the Playground greys the DSpark chip out on that variant.
|
The experimental [Flash Vision checkpoint](#vision-note) also bundles a DSpark head, enabled the same way: the Flash Vision low-latency recipes ship with `--speculative-algorithm DSPARK` (verified with image batches on B200 via the MMMU-Pro round; other hardware rows pending). The balanced and high-throughput Flash Vision recipes stay target-only because they run DP Attention.
|
||||||
|
|
||||||
Unlike the EAGLE recipes for the original Flash and Pro checkpoints, this recipe omits `--speculative-num-steps`, `--speculative-eagle-topk`, and `--speculative-num-draft-tokens`. SGLang reads the DSpark shape from the checkpoint.
|
Unlike the EAGLE recipes for the original Flash and Pro checkpoints, this recipe omits `--speculative-num-steps`, `--speculative-eagle-topk`, and `--speculative-num-draft-tokens`. SGLang reads the DSpark shape from the checkpoint.
|
||||||
|
|
||||||
|
|||||||
@@ -640,8 +640,8 @@ export const benchmarks = [
|
|||||||
{
|
{
|
||||||
match: { hw: "b200", variant: "flash-vision", quant: "fp4", strategy: "low-latency", nodes: "single" },
|
match: { hw: "b200", variant: "flash-vision", quant: "fp4", strategy: "low-latency", nodes: "single" },
|
||||||
sglang_version: "dev-dsv4-flash-vision",
|
sglang_version: "dev-dsv4-flash-vision",
|
||||||
accuracy: { mmmu_pro_pct: 74.96 },
|
accuracy: { mmmu_pro_pct: 75.14 },
|
||||||
notes: "MMMU-Pro (standard, 10-option) measured with sgl-eval on 4×B200 (TP=4) at temperature 1.0, top-p 0.95, --reasoning-effort max.",
|
notes: "MMMU-Pro (standard, 10-option) measured with sgl-eval on 4×B200 (TP=4) at temperature 1.0, top-p 0.95, --reasoning-effort max, with the bundled DSpark head enabled (--speculative-algorithm DSPARK).",
|
||||||
},
|
},
|
||||||
{ match: { hw: "b200", variant: "flash-vision", quant: "fp4", strategy: "balanced", nodes: "single" } },
|
{ match: { hw: "b200", variant: "flash-vision", quant: "fp4", strategy: "balanced", nodes: "single" } },
|
||||||
{ match: { hw: "b200", variant: "flash-vision", quant: "fp4", strategy: "high-throughput", nodes: "single" } },
|
{ match: { hw: "b200", variant: "flash-vision", quant: "fp4", strategy: "high-throughput", nodes: "single" } },
|
||||||
|
|||||||
@@ -313,8 +313,6 @@ sgl-eval run mmmu_pro \\
|
|||||||
flags: ["--speculative-algorithm DSPARK"],
|
flags: ["--speculative-algorithm DSPARK"],
|
||||||
hide: { variant: ["flash", "pro"] },
|
hide: { variant: ["flash", "pro"] },
|
||||||
disable: [
|
disable: [
|
||||||
{ when: { variant: ["flash-vision"] },
|
|
||||||
reason: "The Flash Vision checkpoint bundles a DSpark head, but speculative decoding is not yet verified with image inputs — the cookbook recipes run target-only for now." },
|
|
||||||
{ when: { dpAttnOn: [true] },
|
{ when: { dpAttnOn: [true] },
|
||||||
reason: "DSpark is not compatible with DP Attention on the current release." },
|
reason: "DSpark is not compatible with DP Attention on the current release." },
|
||||||
{ when: { hw: ["mi300x", "mi355x"] },
|
{ when: { hw: ["mi300x", "mi355x"] },
|
||||||
@@ -2849,11 +2847,12 @@ sgl-eval run mmmu_pro \\
|
|||||||
//
|
//
|
||||||
// DeepSeek-V4-Flash-Vision-Exp (sgl-project/sglang#37253): the 0731
|
// DeepSeek-V4-Flash-Vision-Exp (sgl-project/sglang#37253): the 0731
|
||||||
// Flash base plus a vision encoder + aligner. The checkpoint bundles a
|
// Flash base plus a vision encoder + aligner. The checkpoint bundles a
|
||||||
// DSpark head, but speculative decoding is not yet verified with image
|
// DSpark head; low-latency recipes enable it (--speculative-algorithm
|
||||||
// batches, so every recipe runs target-only. Low-latency is the serving
|
// DSPARK, no other spec flags — the draft ships in the main checkpoint),
|
||||||
// shape the MMMU-Pro round ran on (4×B200); balanced / high-throughput
|
// verified on B200 via the MMMU-Pro round (4×B200, image batches).
|
||||||
// mirror the Flash Official recipes on the same 4-GPU topology — final
|
// Balanced / high-throughput stay target-only: those recipes run DP
|
||||||
// verification in progress.
|
// attention, which DSpark is incompatible with on the current release.
|
||||||
|
// Non-B200 hardware — final verification in progress.
|
||||||
// ====================================================================
|
// ====================================================================
|
||||||
{
|
{
|
||||||
match: { hw: "b200", variant: "flash-vision", quant: "fp4", strategy: "low-latency", nodes: "single" },
|
match: { hw: "b200", variant: "flash-vision", quant: "fp4", strategy: "low-latency", nodes: "single" },
|
||||||
@@ -2863,6 +2862,7 @@ sgl-eval run mmmu_pro \\
|
|||||||
flags: [
|
flags: [
|
||||||
"--model-path {{MODEL_NAME}}",
|
"--model-path {{MODEL_NAME}}",
|
||||||
"--tp 4",
|
"--tp 4",
|
||||||
|
"--speculative-algorithm DSPARK",
|
||||||
"--mem-fraction-static 0.85",
|
"--mem-fraction-static 0.85",
|
||||||
"--host {{HOST_IP}}",
|
"--host {{HOST_IP}}",
|
||||||
"--port {{PORT}}",
|
"--port {{PORT}}",
|
||||||
@@ -2918,6 +2918,7 @@ sgl-eval run mmmu_pro \\
|
|||||||
flags: [
|
flags: [
|
||||||
"--model-path {{MODEL_NAME}}",
|
"--model-path {{MODEL_NAME}}",
|
||||||
"--tp 4",
|
"--tp 4",
|
||||||
|
"--speculative-algorithm DSPARK",
|
||||||
"--mem-fraction-static 0.85",
|
"--mem-fraction-static 0.85",
|
||||||
"--host {{HOST_IP}}",
|
"--host {{HOST_IP}}",
|
||||||
"--port {{PORT}}",
|
"--port {{PORT}}",
|
||||||
@@ -2969,6 +2970,7 @@ sgl-eval run mmmu_pro \\
|
|||||||
flags: [
|
flags: [
|
||||||
"--model-path {{MODEL_NAME}}",
|
"--model-path {{MODEL_NAME}}",
|
||||||
"--tp 4",
|
"--tp 4",
|
||||||
|
"--speculative-algorithm DSPARK",
|
||||||
"--mem-fraction-static 0.85",
|
"--mem-fraction-static 0.85",
|
||||||
"--host {{HOST_IP}}",
|
"--host {{HOST_IP}}",
|
||||||
"--port {{PORT}}",
|
"--port {{PORT}}",
|
||||||
@@ -3020,6 +3022,7 @@ sgl-eval run mmmu_pro \\
|
|||||||
flags: [
|
flags: [
|
||||||
"--model-path {{MODEL_NAME}}",
|
"--model-path {{MODEL_NAME}}",
|
||||||
"--tp 4",
|
"--tp 4",
|
||||||
|
"--speculative-algorithm DSPARK",
|
||||||
"--mem-fraction-static 0.85",
|
"--mem-fraction-static 0.85",
|
||||||
"--host {{HOST_IP}}",
|
"--host {{HOST_IP}}",
|
||||||
"--port {{PORT}}",
|
"--port {{PORT}}",
|
||||||
@@ -3076,6 +3079,7 @@ sgl-eval run mmmu_pro \\
|
|||||||
"--model-path {{MODEL_NAME}}",
|
"--model-path {{MODEL_NAME}}",
|
||||||
"--tp 4",
|
"--tp 4",
|
||||||
"--moe-runner-backend marlin",
|
"--moe-runner-backend marlin",
|
||||||
|
"--speculative-algorithm DSPARK",
|
||||||
"--mem-fraction-static 0.85",
|
"--mem-fraction-static 0.85",
|
||||||
"--host {{HOST_IP}}",
|
"--host {{HOST_IP}}",
|
||||||
"--port {{PORT}}",
|
"--port {{PORT}}",
|
||||||
|
|||||||
Reference in New Issue
Block a user