[NPU][DOCS]Add best practice and benchmark result parameter description (#25875)
This commit is contained in:
@@ -264,7 +264,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>BF16</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-32b-4k-1_5k-11ms-on-a2-8-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-32B</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td>
|
||||
@@ -274,7 +274,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-32b-1k-0_3k-12ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-32B</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td>
|
||||
@@ -284,7 +284,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-32b-6k-1_5k-17ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-8B</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
|
||||
@@ -294,7 +294,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-8b-1k-0_3k-7ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-8B</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
|
||||
@@ -304,7 +304,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-8b-6k-1_5k-12ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-8B</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
|
||||
@@ -314,7 +314,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-8b-3_5k-1_5k-5ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-30B-A3B</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
|
||||
@@ -324,7 +324,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-30b-a3b-6k-1_5k-10ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-30B-A3B</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
|
||||
@@ -334,7 +334,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-30b-a3b-1k-0_3k-7ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-Next-A3B-Instruct</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td>
|
||||
@@ -344,7 +344,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-next-1k-0_3k-14_21ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-Next-A3B-Instruct</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td>
|
||||
@@ -354,7 +354,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-next-6k-1_5k-15_62ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-Next-A3B-Instruct</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
|
||||
@@ -364,7 +364,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-next-3_5k-1_5k-20ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-14B</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
|
||||
@@ -374,6 +374,36 @@ you encounter issues or have any questions, please [open an issue](https://githu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-14b-3_5k-1_5k-9ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-27B</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>PD Mixed</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>3.5K+1.5K</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>20ms</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-27b-3_5k-1_5k-20ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-27B</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>PD Mixed</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>16K+1K</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>20ms</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-27b-16k-1k-20ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-27B</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>PD Mixed</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>64K+1K</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>20ms</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-27b-64k-1k-20ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
@@ -533,7 +563,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-32b-2k-2k-50ms-on-a2-8-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-14B</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
|
||||
@@ -543,7 +573,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-14b-3_5k-1_5k-50ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-8B</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
|
||||
@@ -553,6 +583,36 @@ you encounter issues or have any questions, please [open an issue](https://githu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-8b-3_5k-1_5k-50ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-27B</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>PD Mixed</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>3.5K+1.5K</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>50ms</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-27b-3_5k-1_5k-50ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-27B</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>PD Mixed</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>16K+1K</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>50ms</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-27b-16k-1k-50ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-27B</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>PD Mixed</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>64K+1K</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>50ms</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-27b-64k-1k-50ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
@@ -4154,3 +4214,447 @@ We tested it based on the `RANDOM` dataset.
|
||||
```bash Command
|
||||
python3 -m sglang.bench_serving --dataset-name random --backend sglang --host 127.0.0.1 --port 6699 --random-range-ratio 1 --max-concurrency 1 --random-output-len 1500 --random-input-len 3500 --num-prompts 1
|
||||
```
|
||||
|
||||
### Qwen3-27B 3_5K-1_5K 20ms on A3 2 Cards Mixed Mode
|
||||
|
||||
Model: Eco-Tech/Qwen3.5-27B-w8a8-mtp
|
||||
|
||||
Hardware: Atlas 800I A3 2Card
|
||||
|
||||
DeployMode: PD Mixed
|
||||
|
||||
Dataset: random
|
||||
|
||||
Input Output Length: 3.5K+1.5K
|
||||
|
||||
TPOT: 20ms
|
||||
|
||||
#### Model Deployment
|
||||
|
||||
```bash Command
|
||||
# high performance cpu
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl -w kernel.sched_migration_cost_ns=50000
|
||||
# bind cpu
|
||||
export SGLANG_SET_CPU_AFFINITY=1
|
||||
|
||||
unset https_proxy
|
||||
unset http_proxy
|
||||
unset HTTPS_PROXY
|
||||
unset HTTP_PROXY
|
||||
|
||||
# on-demand set device
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
|
||||
export ASCEND_LAUNCH_BLOCKING=1
|
||||
export STREAMS_PER_DEVICE=32
|
||||
export HCCL_BUFFSIZE=3000
|
||||
export HCCL_OP_EXPANSION_MODE="AIV"
|
||||
export HCCL_SOCKET_IFNAME=lo
|
||||
export GLOO_SOCKET_IFNAME=lo
|
||||
export SGLANG_NPU_PROFILING=0
|
||||
export SGLANG_DISAGGEGATION_WAITING_TIMEOUT=3600
|
||||
export SGLANG_ENABLE_SPEC_V2=1
|
||||
export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=0
|
||||
|
||||
MODEL_PATH=xxx
|
||||
|
||||
python -m sglang.launch_server --model-path ${MODEL_PATH} \
|
||||
--attention-backend ascend \
|
||||
--host 127.0.0.1 --port 6699 \
|
||||
--device npu \
|
||||
--tp-size 4\
|
||||
--trust-remote-code \
|
||||
--watchdog-timeout 9000 \
|
||||
--chunked-prefill-size -1 \
|
||||
--max-prefill-tokens 186000 \
|
||||
--enable-prefill-delayer \
|
||||
--prefill-delayer-max-delay-passes 200 \
|
||||
--disable-radix-cache \
|
||||
--mem-fraction-static 0.94 \
|
||||
--max-total-tokens 700000 \
|
||||
--max-running-requests 38 \
|
||||
--max-mamba-cache-size 200 \
|
||||
--quantization modelslim \
|
||||
--dtype bfloat16 \
|
||||
--mamba-ssm-dtype bfloat16 \
|
||||
--enable-multimodal \
|
||||
--mm-attention-backend ascend_attn \
|
||||
--cuda-graph-bs 1 2 4 8 12 18 24 32 34 36 38 \
|
||||
--speculative-algorithm NEXTN \
|
||||
--speculative-num-steps 3 \
|
||||
--speculative-eagle-topk 1 \
|
||||
--speculative-num-draft-tokens 4
|
||||
```
|
||||
|
||||
#### Benchmark
|
||||
|
||||
We tested it based on the `RANDOM` dataset.
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.bench_serving --backend sglang --host 127.0.0.1 --port 6699 --dataset-name random --max-concurrency 38 --num-prompts 152 --random-range-ratio 1 --random-output-len 1500 --random-input-len 3500
|
||||
```
|
||||
|
||||
### Qwen3-27B 16K-1K 20ms on A3 1 Cards Mixed Mode
|
||||
|
||||
Model: Eco-Tech/Qwen3.5-27B-w8a8-mtp
|
||||
|
||||
Hardware: Atlas 800I A3 1Card
|
||||
|
||||
DeployMode: PD Mixed
|
||||
|
||||
Dataset: random
|
||||
|
||||
Input Output Length: 16K+1K
|
||||
|
||||
TPOT: 20ms
|
||||
|
||||
#### Model Deployment
|
||||
|
||||
```bash Command
|
||||
# high performance cpu
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl -w kernel.sched_migration_cost_ns=50000
|
||||
# bind cpu
|
||||
export SGLANG_SET_CPU_AFFINITY=1
|
||||
|
||||
unset https_proxy
|
||||
unset http_proxy
|
||||
unset HTTPS_PROXY
|
||||
unset HTTP_PROXY
|
||||
# cann
|
||||
source /usr/local/Ascend/ascend-toolkit/set_env.sh
|
||||
source /usr/local/Ascend/nnal/atb/set_env.sh
|
||||
|
||||
export STREAMS_PER_DEVICE=32
|
||||
export HCCL_OP_EXPANSION_MODE=AIV
|
||||
export HCCL_SOCKET_IFNAME=lo
|
||||
export GLOO_SOCKET_IFNAME=lo
|
||||
export SGLANG_ENABLE_SPEC_V2=1
|
||||
export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=0
|
||||
export SGLANG_SCHEDULER_DECREASE_PREFILL_IDLE=1
|
||||
export SGLANG_PREFILL_DELAYER_MAX_DELAY_PASSES=100
|
||||
|
||||
# on-demand set device
|
||||
export ASCEND_RT_VISIBLE_DEVICES=8,9
|
||||
|
||||
MODEL_PATH=xxx
|
||||
|
||||
sglang serve --model-path ${MODEL_PATH} \
|
||||
--attention-backend ascend \
|
||||
--device npu \
|
||||
--tp-size 2 --nnodes 1 --node-rank 0 \
|
||||
--chunked-prefill-size -1 --max-prefill-tokens 65000 \
|
||||
--disable-radix-cache \
|
||||
--trust-remote-code \
|
||||
--host 127.0.0.1 --max-running-requests 32 --max-mamba-cache-size 32 \
|
||||
--mem-fraction-static 0.85 \
|
||||
--port 8001 \
|
||||
--cuda-graph-bs 2 3 4 5 6 \
|
||||
--enable-multimodal \
|
||||
--quantization modelslim \
|
||||
--mm-attention-backend ascend_attn \
|
||||
--dtype bfloat16 --mamba-ssm-dtype bfloat16 --max-total-tokens 310000 \
|
||||
--speculative-algorithm NEXTN --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4
|
||||
```
|
||||
|
||||
#### Benchmark
|
||||
|
||||
We tested it based on the `RANDOM` dataset.
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.bench_serving --backend sglang --host 127.0.0.1 --port 8001 --dataset-name random --max-concurrency 32 --num-prompts 128 --random-range-ratio 1 --random-output-len 1000 --random-input-len 16000
|
||||
```
|
||||
|
||||
### Qwen3-27B 64K-1K 20ms on A3 1 Cards Mixed Mode
|
||||
|
||||
Model: Eco-Tech/Qwen3.5-27B-w8a8-mtp
|
||||
|
||||
Hardware: Atlas 800I A3 1Card
|
||||
|
||||
DeployMode: PD Mixed
|
||||
|
||||
Dataset: random
|
||||
|
||||
Input Output Length: 64K+1K
|
||||
|
||||
TPOT: 20ms
|
||||
|
||||
#### Model Deployment
|
||||
|
||||
```bash Command
|
||||
# high performance cpu
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl -w kernel.sched_migration_cost_ns=50000
|
||||
# bind cpu
|
||||
export SGLANG_SET_CPU_AFFINITY=1
|
||||
|
||||
unset https_proxy
|
||||
unset http_proxy
|
||||
unset HTTPS_PROXY
|
||||
unset HTTP_PROXY
|
||||
|
||||
# cann
|
||||
source /usr/local/Ascend/ascend-toolkit/set_env.sh
|
||||
source /usr/local/Ascend/nnal/atb/set_env.sh
|
||||
|
||||
export STREAMS_PER_DEVICE=32
|
||||
export HCCL_OP_EXPANSION_MODE=AIV
|
||||
export HCCL_SOCKET_IFNAME=lo
|
||||
export GLOO_SOCKET_IFNAME=lo
|
||||
export SGLANG_NPU_PROFILING=1
|
||||
export SGLANG_ENABLE_SPEC_V2=1
|
||||
export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=0
|
||||
# on-demand set device
|
||||
export ASCEND_RT_VISIBLE_DEVICES=4,5
|
||||
|
||||
MODEL_PATH=xxx
|
||||
|
||||
python -m sglang.launch_server --model-path ${MODEL_PATH} \
|
||||
--attention-backend ascend \
|
||||
--device npu \
|
||||
--tp-size 2 --nnodes 1 --node-rank 0 \
|
||||
--chunked-prefill-size -1 --max-prefill-tokens 130000 \
|
||||
--disable-radix-cache \
|
||||
--trust-remote-code \
|
||||
--host 127.0.0.1 --max-running-requests 32 --max-mamba-cache-size 18 \
|
||||
--mem-fraction-static 0.5 \
|
||||
--port 8004 \
|
||||
--cuda-graph-bs 2 3 4 \
|
||||
--enable-multimodal \
|
||||
--quantization modelslim \
|
||||
--mm-attention-backend ascend_attn \
|
||||
--dtype bfloat16 --mamba-ssm-dtype bfloat16 --max-total-tokens 280000 \
|
||||
--speculative-algorithm NEXTN --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4
|
||||
```
|
||||
|
||||
#### Benchmark
|
||||
|
||||
We tested it based on the `RANDOM` dataset.
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.bench_serving --backend sglang --host 127.0.0.1 --port 8004 --dataset-name random --max-concurrency 9 --num-prompts 36 --random-range-ratio 1 --random-output-len 1000 --random-input-len 64000
|
||||
```
|
||||
|
||||
### Qwen3-27B 3_5K-1_5K 50ms on A3 1 Cards Mixed Mode
|
||||
|
||||
Model: Eco-Tech/Qwen3.5-27B-w8a8-mtp
|
||||
|
||||
Hardware: Atlas 800I A3 1Card
|
||||
|
||||
DeployMode: PD Mixed
|
||||
|
||||
Dataset: random
|
||||
|
||||
Input Output Length: 3.5K+1.5K
|
||||
|
||||
TPOT: 50ms
|
||||
|
||||
#### Model Deployment
|
||||
|
||||
```bash Command
|
||||
# high performance cpu
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl -w kernel.sched_migration_cost_ns=50000
|
||||
# bind cpu
|
||||
export SGLANG_SET_CPU_AFFINITY=1
|
||||
|
||||
unset https_proxy
|
||||
unset http_proxy
|
||||
unset HTTPS_PROXY
|
||||
unset HTTP_PROXY
|
||||
|
||||
# cann
|
||||
source /usr/local/Ascend/ascend-toolkit/set_env.sh
|
||||
source /usr/local/Ascend/nnal/atb/set_env.sh
|
||||
|
||||
export STREAMS_PER_DEVICE=32
|
||||
export HCCL_OP_EXPANSION_MODE=AIV
|
||||
export HCCL_SOCKET_IFNAME=lo
|
||||
export GLOO_SOCKET_IFNAME=lo
|
||||
export SGLANG_ENABLE_SPEC_V2=1
|
||||
export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=0
|
||||
export SGLANG_SCHEDULER_DECREASE_PREFILL_IDLE=1
|
||||
export SGLANG_PREFILL_DELAYER_MAX_DELAY_PASSES=100
|
||||
|
||||
MODEL_PATH=xxx
|
||||
|
||||
python -m sglang.launch_server --model-path ${MODEL_PATH} \
|
||||
--attention-backend ascend \
|
||||
--device npu \
|
||||
--tp-size 2 --nnodes 1 --node-rank 0 \
|
||||
--chunked-prefill-size -1 --max-prefill-tokens 60000 \
|
||||
--disable-radix-cache \
|
||||
--trust-remote-code \
|
||||
--host 127.0.0.1 --max-running-requests 48 --max-mamba-cache-size 60 \
|
||||
--mem-fraction-static 0.7 \
|
||||
--port 8000 \
|
||||
--cuda-graph-bs 2 8 16 32 48 \
|
||||
--enable-multimodal \
|
||||
--quantization modelslim \
|
||||
--mm-attention-backend ascend_attn \
|
||||
--dtype bfloat16 --mamba-ssm-dtype bfloat16 \
|
||||
--speculative-algorithm NEXTN --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4
|
||||
```
|
||||
|
||||
#### Benchmark
|
||||
|
||||
We tested it based on the `RANDOM` dataset.
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.bench_serving --backend sglang --host 127.0.0.1 --port 8000 --dataset-name random --max-concurrency 48 --num-prompts 192 --random-range-ratio 1 --random-output-len 1500 --random-input-len 3500
|
||||
```
|
||||
|
||||
### Qwen3-27B 16K-1K 50ms on A3 2 Cards Mixed Mode
|
||||
|
||||
Model: Eco-Tech/Qwen3.5-27B-w8a8-mtp
|
||||
|
||||
Hardware: Atlas 800I A3 2Card
|
||||
|
||||
DeployMode: PD Mixed
|
||||
|
||||
Dataset: random
|
||||
|
||||
Input Output Length: 16K+1K
|
||||
|
||||
TPOT: 50ms
|
||||
|
||||
#### Model Deployment
|
||||
|
||||
```bash Command
|
||||
# high performance cpu
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl -w kernel.sched_migration_cost_ns=50000
|
||||
# bind cpu
|
||||
export SGLANG_SET_CPU_AFFINITY=1
|
||||
|
||||
unset https_proxy
|
||||
unset http_proxy
|
||||
unset HTTPS_PROXY
|
||||
unset HTTP_PROXY
|
||||
|
||||
# cann
|
||||
source /usr/local/Ascend/ascend-toolkit/set_env.sh
|
||||
source /usr/local/Ascend/nnal/atb/set_env.sh
|
||||
|
||||
export STREAMS_PER_DEVICE=32
|
||||
export HCCL_OP_EXPANSION_MODE=AIV
|
||||
export HCCL_SOCKET_IFNAME=lo
|
||||
export GLOO_SOCKET_IFNAME=lo
|
||||
export SGLANG_ENABLE_SPEC_V2=1
|
||||
export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1
|
||||
export SGLANG_SCHEDULER_DECREASE_PREFILL_IDLE=1
|
||||
export SGLANG_PREFILL_DELAYER_MAX_DELAY_PASSES=30
|
||||
# on-demand set device
|
||||
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
|
||||
|
||||
MODEL_PATH=xxx
|
||||
|
||||
python3 -m sglang.launch_server --model-path ${MODEL_PATH} \
|
||||
--attention-backend ascend \
|
||||
--device npu \
|
||||
--tp-size 4 --nnodes 1 --node-rank 0 \
|
||||
--chunked-prefill-size -1 --max-prefill-tokens 50000 \
|
||||
--disable-radix-cache \
|
||||
--trust-remote-code \
|
||||
--host 127.0.0.1 --max-running-requests 28 --max-mamba-cache-size 50 \
|
||||
--mem-fraction-static 0.7 \
|
||||
--port 8001 \
|
||||
--cuda-graph-bs 2 8 12 16 20 24 28\
|
||||
--enable-multimodal \
|
||||
--quantization modelslim \
|
||||
--mm-attention-backend ascend_attn \
|
||||
--dtype bfloat16 --mamba-ssm-dtype bfloat16 \
|
||||
--speculative-algorithm NEXTN --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4
|
||||
```
|
||||
|
||||
#### Benchmark
|
||||
|
||||
We tested it based on the `RANDOM` dataset.
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.bench_serving --backend sglang --host 127.0.0.1 --port 8001 --dataset-name random --max-concurrency 28 --num-prompts 152 --random-range-ratio 1 --random-output-len 1000 --random-input-len 16000
|
||||
```
|
||||
|
||||
### Qwen3-27B 64K-1K 50ms on A3 2 Cards Mixed Mode
|
||||
|
||||
Model: Eco-Tech/Qwen3.5-27B-w8a8-mtp
|
||||
|
||||
Hardware: Atlas 800I A3 2Card
|
||||
|
||||
DeployMode: PD Mixed
|
||||
|
||||
Dataset: random
|
||||
|
||||
Input Output Length: 64K+1K
|
||||
|
||||
TPOT: 50ms
|
||||
|
||||
#### Model Deployment
|
||||
|
||||
```bash Command
|
||||
# high performance cpu
|
||||
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
|
||||
sysctl -w vm.swappiness=0
|
||||
sysctl -w kernel.numa_balancing=0
|
||||
sysctl -w kernel.sched_migration_cost_ns=50000
|
||||
# bind cpu
|
||||
export SGLANG_SET_CPU_AFFINITY=1
|
||||
|
||||
unset https_proxy
|
||||
unset http_proxy
|
||||
unset HTTPS_PROXY
|
||||
unset HTTP_PROXY
|
||||
|
||||
# cann
|
||||
source /usr/local/Ascend/ascend-toolkit/set_env.sh
|
||||
source /usr/local/Ascend/nnal/atb/set_env.sh
|
||||
|
||||
export STREAMS_PER_DEVICE=32
|
||||
export HCCL_OP_EXPANSION_MODE=AIV
|
||||
export HCCL_SOCKET_IFNAME=lo
|
||||
export GLOO_SOCKET_IFNAME=lo
|
||||
export SGLANG_ENABLE_SPEC_V2=1
|
||||
export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=0
|
||||
export SGLANG_SCHEDULER_DECREASE_PREFILL_IDLE=1
|
||||
export SGLANG_PREFILL_DELAYER_MAX_DELAY_PASSES=100
|
||||
# on-demand set device
|
||||
export ASCEND_RT_VISIBLE_DEVICES=4,5,6,7
|
||||
|
||||
MODEL_PATH=xxx
|
||||
|
||||
python3 -m sglang.launch_server --model-path ${MODEL_PATH} \
|
||||
--attention-backend ascend \
|
||||
--device npu \
|
||||
--tp-size 4 --nnodes 1 --node-rank 0 \
|
||||
--chunked-prefill-size -1 --max-prefill-tokens 200000 \
|
||||
--disable-radix-cache \
|
||||
--trust-remote-code \
|
||||
--host 127.0.0.1 --max-running-requests 32 --max-mamba-cache-size 22 \
|
||||
--mem-fraction-static 0.5 \
|
||||
--port 9000 \
|
||||
--cuda-graph-bs 2 4 8 11 12 13 \
|
||||
--enable-multimodal \
|
||||
--quantization modelslim \
|
||||
--mm-attention-backend ascend_attn \
|
||||
--dtype bfloat16 --mamba-ssm-dtype bfloat16 --max-total-tokens 850000 \
|
||||
--speculative-algorithm NEXTN --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4
|
||||
```
|
||||
|
||||
#### Benchmark
|
||||
|
||||
We tested it based on the `RANDOM` dataset.
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.bench_serving --backend sglang --host 127.0.0.1 --port 9000 --dataset-name random --max-concurrency 9 --num-prompts 36 --random-range-ratio 1 --random-output-len 1000 --random-input-len 64000
|
||||
```
|
||||
|
||||
@@ -351,6 +351,111 @@ Max ITL (ms): 2229.30
|
||||
==================================================
|
||||
```
|
||||
|
||||
#### SGLang Serving Benchmark Result — Complete Reference
|
||||
|
||||
The output format is **hardcoded in `bench_serving.py`. All formatting decisions — including column widths, alignment, and decimal precision — are statically defined in the source and cannot be changed via command-line arguments.
|
||||
|
||||
##### Test Configuration
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Parameter</th>
|
||||
<th>Description</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td><code>Backend</code></td>
|
||||
<td>The serving backend under test (e.g., <code>sglang</code>, <code>vllm</code>).</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>Traffic request rate</code></td>
|
||||
<td>Request generation rate in req/s. <code>inf</code> means maximum rate (concurrency-bounded). <code>trace</code> indicates trace timestamp mode. A fixed value enforces constant inter-arrival time.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>Max request concurrency</code></td>
|
||||
<td>Maximum number of concurrent requests from the client side. Displays <code>not set</code> when unspecified.</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
##### Core Statistics & Throughput Metrics
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Parameter</th>
|
||||
<th>Description</th>
|
||||
<th>Format Specification</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr><td><code>Successful requests</code></td><td>Total number of successfully completed requests (HTTP 200, no generation errors).</td><td>Integer, no decimal places</td></tr>
|
||||
<tr><td><code>Benchmark duration (s)</code></td><td>Total elapsed time from first request sent to last response fully received (seconds).</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>Total input tokens</code></td><td>Total number of input (prompt) tokens across all requests, counted by server-side tokenizer.</td><td>Integer, no decimal places</td></tr>
|
||||
<tr><td><code>Total input text tokens</code></td><td>Same as <code>Total input tokens</code>. For multimodal inputs, this may differ.</td><td>Integer, no decimal places</td></tr>
|
||||
<tr><td><code>Total generated tokens</code></td><td>Total number of output tokens actually generated by the server (server-side tokenizer count).</td><td>Integer, no decimal places</td></tr>
|
||||
<tr><td><code>Total generated tokens (retokenized)</code></td><td>Output text re-tokenized by the client using its own tokenizer. A large discrepancy indicates tokenizer mismatch or special tokens in output.</td><td>Integer, no decimal places</td></tr>
|
||||
<tr><td><code>Request throughput (req/s)</code></td><td>Number of successful requests processed per second. Formula: <code>Successful requests / Benchmark duration (s)</code>.</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>Input token throughput (tok/s)</code></td><td>Number of input tokens processed per second. Formula: <code>Total input tokens / Benchmark duration (s)</code>.</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>Output token throughput (tok/s)</code></td><td>Number of output tokens generated per second. Formula: <code>Total generated tokens / Benchmark duration (s)</code>.</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>Peak output token throughput (tok/s)</code></td><td>Observed instantaneous peak output token generation rate during the test (computed over a sliding window).</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>Peak concurrent requests</code></td><td>Maximum number of requests being processed simultaneously on the server side. May exceed client-side <code>Max request concurrency</code> due to queueing.</td><td>Integer, no decimal places</td></tr>
|
||||
<tr><td><code>Total token throughput (tok/s)</code></td><td>Sum of input and output token throughputs. Formula: <code>Input token throughput + Output token throughput</code>.</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>Concurrency</code></td><td>Average number of concurrent requests during the test (Little's Law). Formula: <code>Sum of all E2E latencies / Benchmark duration</code>.</td><td>2 decimal places</td></tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
##### End-to-End Latency (E2E Latency)
|
||||
|
||||
<table>
|
||||
<thead><tr><th>Statistic</th><th>Description</th><th>Format</th></tr></thead>
|
||||
<tbody>
|
||||
<tr><td><code>Mean E2E Latency (ms)</code></td><td>Arithmetic mean</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>Median E2E Latency (ms)</code></td><td>50th percentile</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>P90 E2E Latency (ms)</code></td><td>90th percentile (90% of requests have latency ≤ this value)</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>P99 E2E Latency (ms)</code></td><td>99th percentile</td><td>2 decimal places</td></tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
##### Time to First Token (TTFT)
|
||||
|
||||
<table>
|
||||
<thead><tr><th>Statistic</th><th>Description</th><th>Format</th></tr></thead>
|
||||
<tbody>
|
||||
<tr><td><code>Mean TTFT (ms)</code></td><td>Arithmetic mean</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>Median TTFT (ms)</code></td><td>50th percentile</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>P99 TTFT (ms)</code></td><td>99th percentile</td><td>2 decimal places</td></tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
##### Time per Output Token (TPOT) – Excluding First Token
|
||||
Formula: <code>(E2E Latency - TTFT) / (Number of output tokens - 1)</code>
|
||||
|
||||
<table>
|
||||
<thead><tr><th>Statistic</th><th>Description</th><th>Format</th></tr></thead>
|
||||
<tbody>
|
||||
<tr><td><code>Mean TPOT (ms)</code></td><td>Arithmetic mean</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>Median TPOT (ms)</code></td><td>50th percentile</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>P99 TPOT (ms)</code></td><td>99th percentile</td><td>2 decimal places</td></tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
##### Inter-Token Latency (ITL)
|
||||
|
||||
<table>
|
||||
<thead><tr><th>Statistic</th><th>Description</th><th>Format</th></tr></thead>
|
||||
<tbody>
|
||||
<tr><td><code>Mean ITL (ms)</code></td><td>Average inter-token interval</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>Median ITL (ms)</code></td><td>50th percentile inter-token interval</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>P95 ITL (ms)</code></td><td>95th percentile (used to detect stalls)</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>P99 ITL (ms)</code></td><td>99th percentile</td><td>2 decimal places</td></tr>
|
||||
<tr><td><code>Max ITL (ms)</code></td><td>Maximum observed inter-token interval; useful for identifying severe blocking events</td><td>2 decimal places</td></tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
|
||||
## 3. Online Service: Multimodal Model
|
||||
|
||||
Test `Qwen/Qwen2.5-VL-7B-Instruct` for vision-language tasks.
|
||||
|
||||
Reference in New Issue
Block a user