[NPU] [DOC] add missing DEEP_NORMAL_MODE_USE_INT8_QUANT for w8a8+deepep scenarios (#29937)
This commit is contained in:
File diff suppressed because it is too large
Load Diff
@@ -76,6 +76,18 @@ This document provides a list of commonly used environment variables and aims to
|
||||
<td>Needs to be enabled when the expert ID to be processed by <br/> DEEPEP contains -1.</td>
|
||||
<td><code>0</code></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>DEEP_NORMAL_MODE_USE_INT8_QUANT</code></td>
|
||||
<td>
|
||||
<strong>Deprecated — will be removed in a future release.</strong>
|
||||
<br/>
|
||||
When set to <code>1</code>, quantizes intermediate activations to INT8 in
|
||||
<br/>the DeepEP dispatch operator during normal mode, reducing communication
|
||||
<br/>volume for W8A8-quantized MoE models. This variable will become a no-op;
|
||||
<br/>the quantization behavior will be inferred automatically.
|
||||
</td>
|
||||
<td><code>0</code></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
|
||||
@@ -98,6 +98,7 @@ export HCCL_OP_EXPANSION_MODE=AIV
|
||||
export HCCL_SOCKET_IFNAME=lo
|
||||
export GLOO_SOCKET_IFNAME=lo
|
||||
# DEEPEP
|
||||
export DEEP_NORMAL_MODE_USE_INT8_QUANT=1
|
||||
export DEEPEP_NORMAL_LONG_SEQ_ROUND=72
|
||||
export DEEPEP_NORMAL_LONG_SEQ_PER_ROUND_TOKENS=1024
|
||||
export DEEPEP_NORMAL_COMBINE_ENABLE_LONG_SEQ=1
|
||||
@@ -219,6 +220,7 @@ export GLOO_SOCKET_IFNAME=lo
|
||||
export DEEPEP_NORMAL_LONG_SEQ_ROUND=72
|
||||
export DEEPEP_NORMAL_LONG_SEQ_PER_ROUND_TOKENS=1024
|
||||
export DEEPEP_NORMAL_COMBINE_ENABLE_LONG_SEQ=1
|
||||
export DEEP_NORMAL_MODE_USE_INT8_QUANT=1
|
||||
|
||||
|
||||
IPS=('your ip1' 'your ip2')
|
||||
|
||||
@@ -760,6 +760,10 @@ Breaks large prefill computations into smaller, fixed-size chunks for better sch
|
||||
|
||||
Captures the compute graph on NPU and replays it to eliminate kernel launch overhead, analogous to CUDA Graph on NVIDIA GPUs. Internally uses `torch.npu.NPUGraph`. Most effective for decode with stable, predictable batch sizes (via `--cuda-graph-bs`). Has partial compatibility with most features because graph capture requires fixed control flow and tensor shapes.
|
||||
|
||||
<Note>
|
||||
`--enable-torch-compile` is incompatible with NPU Graph. When torch.compile is enabled, NPU Graph must be disabled via `--disable-cuda-graph`.
|
||||
</Note>
|
||||
|
||||
### Speculative Decoding (`--speculative-algorithm`)
|
||||
|
||||
Reduces per-token latency by predicting multiple future tokens in a single forward pass, then verifying them against the model. Ascend supports `NEXTN` (for DeepSeek models, using the model's own hidden states) and `EAGLE3` (for Qwen MoE models, using a separate draft model). Controlled by `--speculative-num-steps`, `--speculative-eagle-topk`, and `--speculative-num-draft-tokens`. On Ascend, `SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1` enables the optimized overlap scheduler for speculative decoding.
|
||||
|
||||
@@ -1662,7 +1662,7 @@ click [Server Arguments](../../advanced_features/server_arguments).
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>`--init-expert-location`</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>`trivial`</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>`trivial`,<br/> `<path.pt>`,<br/> `<path.json>`,<br/> `<json_string>`</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>`trivial`,<br/> `<path.pt>`,<br/> `<path.json>`,<br/> `<json_string>`</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>A2, A3</td>
|
||||
</tr>
|
||||
<tr>
|
||||
|
||||
@@ -206,6 +206,7 @@ export HCCL_BUFFSIZE=1500
|
||||
export DEEPEP_NORMAL_LONG_SEQ_PER_ROUND_TOKENS=1024
|
||||
export DEEPEP_NORMAL_LONG_SEQ_ROUND=128
|
||||
export DEEPEP_NORMAL_COMBINE_ENABLE_LONG_SEQ=1
|
||||
export DEEP_NORMAL_MODE_USE_INT8_QUANT=1
|
||||
|
||||
python3 -m sglang.launch_server \
|
||||
--model-path ${MODEL_PATH} \
|
||||
|
||||
Reference in New Issue
Block a user