[cookbook] Laguna-M.1: add PD disaggregation section (#28737)
This commit is contained in:
@@ -233,3 +233,91 @@ Args: {"location": "Beijing"}
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
### 3.3 Prefill-Decode (PD) Disaggregation
|
||||
|
||||
[PD disaggregation](../../../docs/advanced_features/pd_disaggregation) runs prefill and decode on **separate** SGLang servers linked by an RDMA KV-transfer fabric (mooncake or NIXL), fronted by the PD router. Laguna-M.1 is **global-attention with a standard KV cache** (no sliding window, no sparse "index" side-buffer), so its KV pages transfer with **no model-specific flags** — just the `--disaggregation-*` knobs. Both roles auto-select the same attention backend (FlashAttention-3) and page size because they share the model and flags, so the KV layout lines up for transfer.
|
||||
|
||||
**Supported / validated topology:**
|
||||
|
||||
- **Equal tensor parallelism** — prefill and decode run the same `--tp`.
|
||||
- **Single pipeline stage** — PP = 1 (the default).
|
||||
- **mooncake or NIXL** transfer backend over RDMA / InfiniBand.
|
||||
- Validated on **2 × 8×H200** (TP8 prefill + TP8 decode, BF16), one node each, over an 8× 400 Gb/s NDR InfiniBand fabric.
|
||||
|
||||
Launch the prefill server, then the decode server — the same recipe with `--disaggregation-mode decode` and no bootstrap port. Point `--disaggregation-ib-device` at your RDMA NIC(s).
|
||||
|
||||
```bash Prefill server (node A)
|
||||
sglang serve \
|
||||
--model-path poolside/Laguna-M.1 \
|
||||
--trust-remote-code \
|
||||
--reasoning-parser poolside_v1 \
|
||||
--tool-call-parser poolside_v1 \
|
||||
--tp 8 \
|
||||
--disaggregation-mode prefill \
|
||||
--disaggregation-transfer-backend mooncake \
|
||||
--disaggregation-ib-device mlx5_0,mlx5_1,mlx5_2,mlx5_3,mlx5_4,mlx5_5,mlx5_6,mlx5_7 \
|
||||
--host 0.0.0.0 --port 30000 \
|
||||
--disaggregation-bootstrap-port 8998
|
||||
```
|
||||
|
||||
```bash Decode server (node B)
|
||||
sglang serve \
|
||||
--model-path poolside/Laguna-M.1 \
|
||||
--trust-remote-code \
|
||||
--reasoning-parser poolside_v1 \
|
||||
--tool-call-parser poolside_v1 \
|
||||
--tp 8 \
|
||||
--disaggregation-mode decode \
|
||||
--disaggregation-transfer-backend mooncake \
|
||||
--disaggregation-ib-device mlx5_0,mlx5_1,mlx5_2,mlx5_3,mlx5_4,mlx5_5,mlx5_6,mlx5_7 \
|
||||
--host 0.0.0.0 --port 30001
|
||||
```
|
||||
|
||||
Then start the PD router, pointing it at the prefill bootstrap (URL plus its `--disaggregation-bootstrap-port`) and the decode endpoint:
|
||||
|
||||
```bash PD router
|
||||
python3 -m sglang_router.launch_router \
|
||||
--pd-disaggregation \
|
||||
--prefill http://<prefill-host>:30000 8998 \
|
||||
--decode http://<decode-host>:30001 \
|
||||
--policy round_robin \
|
||||
--host 0.0.0.0 --port 8000
|
||||
```
|
||||
|
||||
Clients hit the router exactly like a single server — it splits each request across the two stages transparently:
|
||||
|
||||
<Accordion title="PD Client Example (Python)">
|
||||
|
||||
```python Example
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(base_url="http://<router-host>:8000/v1", api_key="EMPTY")
|
||||
|
||||
response = client.chat.completions.create(
|
||||
model="poolside/Laguna-M.1",
|
||||
messages=[{"role": "user", "content": "What is 2 + 2?"}],
|
||||
max_tokens=64,
|
||||
)
|
||||
print(response.choices[0].message.content)
|
||||
```
|
||||
|
||||
**Output Example:**
|
||||
|
||||
```text Output
|
||||
2 + 2 = 4
|
||||
```
|
||||
|
||||
</Accordion>
|
||||
|
||||
**Transfer backend — mooncake (recommended).** mooncake honors `--disaggregation-ib-device` and establishes its RDMA connection at registration, so the **first request is already fast** (no cold start). It works with a single NIC or all eight; using **all 8 NICs lowers TTFT** (more aggregate bandwidth for the KV payload — the gap widens at longer context). On 8×H200 (random isl=512 / osl=256, 16 concurrent) it served ≈ **717 tok/s** output (≈ 2.2k tok/s total), mean **TTFT 244 ms**, mean **TPOT 17.7 ms**; with a single `mlx5_0` NIC, ≈ 697 tok/s and TTFT 287 ms (TPOT unchanged — decode is compute-bound).
|
||||
|
||||
**Transfer backend — NIXL (works, with two caveats).**
|
||||
|
||||
<Warning>
|
||||
The NIXL path **ignores `--disaggregation-ib-device`** — that flag is mooncake-only. NIXL uses its UCX backend, whose NIC is selected by the **`UCX_NET_DEVICES`** environment variable. **Set it** (e.g. `export UCX_NET_DEVICES=mlx5_0:1`) on both servers; without it UCX cannot establish a working cross-node path and every KV transfer hangs until it hits the 300 s timeout (`Request … timed out … in KVPoll.WaitingForInput`) and returns a 500.
|
||||
</Warning>
|
||||
|
||||
With `UCX_NET_DEVICES` pinned, NIXL matches mooncake on quality and steady-state speed (≈ 720 tok/s, TTFT 230 ms, TPOT 17.7 ms). One difference: the **first request after launch pays a ~38 s one-time UCX connection cold-start** (a single port or all eight behave the same). Warm the path with one throwaway request after startup, or raise `SGLANG_DISAGGREGATION_WAITING_TIMEOUT` (default 300 s) so the first real request isn't dropped while UCX connects.
|
||||
|
||||
**Validation.** PD disaggregation preserves output quality — disaggregated output matches non-disaggregated serving, and GSM8K (no-thinking, 200-question subset via the router) scored **0.945** (mooncake, 8 NICs) / **0.940** (NIXL) / **0.950** (mooncake, 1 NIC), all with 100% stop-rate and 0% errors — in line with single-node BF16 (≈ 0.93 on the full split). Logs confirm the split: the prefill node logs `Prefill batch` (CUDA graph off), the decode node logs `Decode batch` (CUDA graph on).
|
||||
|
||||
@@ -148,6 +148,46 @@ sgl-eval run gsm8k \\
|
||||
{ id: "toolCall", label: "Tool Call Parser", flag: "--tool-call-parser poolside_v1" },
|
||||
],
|
||||
},
|
||||
|
||||
// Prefill-Decode disaggregation (§3.3). M.1 is standard-KV (global attention, no sparse
|
||||
// index buffer), so it disaggregates with just the --disaggregation-* flags — no model-specific
|
||||
// backend pinning. Verified on 2×8×H200 (TP8+TP8, BF16) over InfiniBand.
|
||||
pdDisagg: {
|
||||
modes: [
|
||||
{ id: "off", label: "Off" },
|
||||
{ id: "prefill", label: "Prefill role" },
|
||||
{ id: "decode", label: "Decode role" },
|
||||
],
|
||||
transferBackends: [
|
||||
// mooncake (recommended): honors --disaggregation-ib-device, no transfer cold-start.
|
||||
// The NCCL/MNNVL env is only needed on NVLink-multinode Grace-Blackwell (GB200/GB300).
|
||||
{ id: "mooncake", label: "Mooncake",
|
||||
env: [
|
||||
"NCCL_MNNVL_ENABLE=1",
|
||||
"NCCL_CUMEM_ENABLE=1",
|
||||
"SGLANG_MOONCAKE_CUSTOM_MEM_POOL=True",
|
||||
"MC_FORCE_MNNVL=1",
|
||||
],
|
||||
envWhen: { hw: ["gb200", "gb300"] } },
|
||||
// NiXL ignores --disaggregation-ib-device; its UCX backend needs the NIC pinned via
|
||||
// UCX_NET_DEVICES or every KV transfer hangs to the 300s timeout (§3.3). Baked in here
|
||||
// for the IB-based HGX platforms; also expect a ~38s one-time UCX cold-start.
|
||||
{ id: "nixl", label: "NiXL",
|
||||
env: ["UCX_NET_DEVICES=mlx5_0:1"],
|
||||
envWhen: { hw: ["h200", "b200", "b300"] } },
|
||||
],
|
||||
ibDevices: [{ id: "auto", label: "Auto" }, "mlx5_0", "mlx5_7"],
|
||||
router: {
|
||||
port: 8000,
|
||||
command:
|
||||
`python3 -m sglang_router.launch_router \\
|
||||
--pd-disaggregation \\
|
||||
--prefill http://<prefill-host>:{{PREFILL_PORT}} \\
|
||||
--decode http://<decode-host>:{{DECODE_PORT}} \\
|
||||
--policy round_robin \\
|
||||
--host 0.0.0.0 --port {{ROUTER_PORT}}`,
|
||||
},
|
||||
},
|
||||
},
|
||||
|
||||
// One Balanced cell per valid (hw × quant): H200×{BF16,FP8}; each Blackwell×{BF16,FP8,NVFP4}.
|
||||
|
||||
Reference in New Issue
Block a user