Update GLM-5.2 NVFP4 B200/B300 for AgentX HiCache (#35368)

This commit is contained in:
Faradawn Yang
2026-09-02 14:39:09 -07:00
committed by GitHub
parent 3a855b050a
commit 9c70d22721
@@ -269,3 +269,33 @@ When deploying with PD Disaggregation, the prefill node can choose to enable [La
--cp-strategy interleave \
```
With LayerSplit, the kv cache on each rank can be sharded over the CP attention group, and prefetched when necessary. This can reduce kv cache memory by up to 75%, thus increasing the throughput on prefill side.
### 3.6 Agentic Long-Context with HiCache DRAM Offload (NVFP4, MTP)
**B300 (TP8):**
```bash Command
python3 -m sglang.launch_server \
--model-path nvidia/GLM-5.2-NVFP4 \
--trust-remote-code \
--tp 8 \
--ep-size 1 \
--quantization modelopt_fp4 \
--kv-cache-dtype fp8_e4m3 \
--bf16-gemm-backend cutedsl \
--max-prefill-tokens 8192 \
--chunked-prefill-size 8192 \
--mem-fraction-static 0.85 \
--tool-call-parser glm47 \
--reasoning-parser glm45 \
--speculative-algorithm EAGLE \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4 \
--enable-hierarchical-cache \
--hicache-size 270 \
--hicache-write-policy write_back \
--hicache-io-backend direct \
--hicache-mem-layout page_first_direct
```
**B200 (TP8):** same command, with `--mem-fraction-static 0.83` and `--hicache-size 169`. Lower concurrency (c1/c4/c8) uses `--hicache-ratio 0.75` instead.