[Docs] Sync docs_new with legacy docs and update migration redirects (#23337)
Co-authored-by: Mingyi <wisclmy0611@gmail.com>
This commit is contained in:
@@ -142,7 +142,7 @@ python3 -m sglang.bench_serving \
|
||||
|
||||
- `--output-file FILE.jsonl`: append JSONL results to file; auto-named if unspecified
|
||||
- `--output-details`: include per-request arrays (generated texts, errors, ttfts, itls, input/output lens)
|
||||
- `--extra-request-body '{"top_p":0.9,"temperature":0.6}'`: merged into payload (sampling params, etc.)
|
||||
- `--extra-request-body '{"top_p":0.9,"temperature":0.6}'`: merged into payload (sampling params, etc.)
|
||||
- `--disable-ignore-eos`: pass through EOS behavior (varies by backend)
|
||||
- `--warmup-requests N`: run warmup requests with short output first (default 1)
|
||||
- `--flush-cache`: call `/flush_cache` (sglang) before main run
|
||||
@@ -335,7 +335,7 @@ python3 -m sglang.bench_serving \
|
||||
python3 -m sglang.bench_serving \
|
||||
--backend sglang \
|
||||
--host 127.0.0.1 --port 30000 \
|
||||
--model mode-name \
|
||||
--model model-name \
|
||||
--dataset-name mooncake \
|
||||
--mooncake-slowdown-factor 1.0 \
|
||||
--mooncake-num-rounds 1000 \
|
||||
@@ -344,6 +344,41 @@ python3 -m sglang.bench_serving \
|
||||
--random-output-len 256
|
||||
```
|
||||
|
||||
10) Fake decode stress testing (PD disaggregation, decode-only):
|
||||
|
||||
When benchmarking pure decode performance in a PD disaggregation setup, you can bypass the prefill node entirely by using `--fake-prefill`. This requires the decode server to be started with `--disaggregation-transfer-backend fake`:
|
||||
|
||||
```bash Command
|
||||
# Step 1: Start a decode-only server with fake transfer backend
|
||||
python -m sglang.launch_server \
|
||||
--model-path meta-llama/Llama-3.1-8B-Instruct \
|
||||
--disaggregation-mode decode \
|
||||
--disaggregation-transfer-backend fake \
|
||||
--port 30001
|
||||
|
||||
# Step 2: Run bench_serving with --fake-prefill
|
||||
python3 -m sglang.bench_serving \
|
||||
--backend sglang \
|
||||
--host 127.0.0.1 --port 30001 \
|
||||
--model meta-llama/Llama-3.1-8B-Instruct \
|
||||
--dataset-name random \
|
||||
--num-prompts 500 \
|
||||
--random-input-len 1024 --random-output-len 256 \
|
||||
--fake-prefill
|
||||
```
|
||||
|
||||
Similarly, `bench_one_batch_server` also supports `--fake-prefill`:
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.bench_one_batch_server \
|
||||
--base-url http://127.0.0.1:30001 \
|
||||
--model-path meta-llama/Llama-3.1-8B-Instruct \
|
||||
--batch-size 32 --input-len 1024 --output-len 256 \
|
||||
--fake-prefill
|
||||
```
|
||||
|
||||
The `--fake-prefill` flag automatically injects special sentinel values into each request, telling the decode server to skip real KV transfer and generate fake KV data locally.
|
||||
|
||||
### Troubleshooting
|
||||
|
||||
- All requests failed: verify `--backend`, server URL/port, `--model`, and authentication. Check warmup errors printed by the script.
|
||||
@@ -355,4 +390,4 @@ python3 -m sglang.bench_serving \
|
||||
### Notes
|
||||
|
||||
- The script raises the file descriptor soft limit (`RLIMIT_NOFILE`) to help with many concurrent connections.
|
||||
- For sglang, `/get_server_info` is queried post-run to report speculative decoding accept length when available.
|
||||
- For sglang, `/server_info` is queried post-run to report speculative decoding accept length when available.
|
||||
|
||||
Reference in New Issue
Block a user