fix(glm-5.2-nvfp4): bound Mooncake synchronous transfer batches (#32758)

This commit is contained in:
HZY
2026-09-08 22:14:33 +08:00
committed by GitHub
parent 4df5df911b
commit a6b542813f
6 changed files with 205 additions and 2 deletions
@@ -180,6 +180,11 @@ The `SGLANG_MOONCAKE_CUSTOM_MEM_POOL` environment variable enables the custom me
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Sets the number of parallel transfer queues. KVCache transfer requests from multiple decode instances will be sharded into these queues so that they can share the threads and the transfer bandwidth at the same time. If it is set to <code>1</code>, then we transfer requests one by one according to fcfs strategy</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>`4`</td>
</tr>
<tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>**`SGLANG_MOONCAKE_MAX_TRANSFER_BATCH_INDICES`**</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Opt-in limit for the number of KV cache indices represented by one synchronous all-layer Mooncake transfer batch. Set it to a positive value to slice larger index arrays into ordered sub-batches before contiguous address ranges are formed. The custom-memory-pool layerwise path is unchanged.</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}><code>0</code> (disabled)</td>
</tr>
<tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>**`SGLANG_DISAGGREGATION_BOOTSTRAP_TIMEOUT`**</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Timeout (seconds) for receiving destination KV indices during request initialization</td>