[NPU][DOCS]Add best practice and benchmark result parameter description (#25875)

This commit is contained in:
loading66
2026-05-21 19:08:10 +08:00
committed by GitHub
parent 64f21b1589
commit 2e0d2d4c18
2 changed files with 622 additions and 13 deletions
@@ -264,7 +264,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>BF16</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>BF16</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-32b-4k-1_5k-11ms-on-a2-8-cards-mixed-mode">Optimal Configuration</a></td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-32b-4k-1_5k-11ms-on-a2-8-cards-mixed-mode">Optimal Configuration</a></td>
</tr> </tr>
<tr> <tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-32B</td> <td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-32B</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td>
@@ -274,7 +274,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-32b-1k-0_3k-12ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-32b-1k-0_3k-12ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td>
</tr> </tr>
<tr> <tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-32B</td> <td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-32B</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td>
@@ -284,7 +284,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-32b-6k-1_5k-17ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-32b-6k-1_5k-17ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td>
</tr> </tr>
<tr> <tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-8B</td> <td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-8B</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
@@ -294,7 +294,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-8b-1k-0_3k-7ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-8b-1k-0_3k-7ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
</tr> </tr>
<tr> <tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-8B</td> <td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-8B</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
@@ -304,7 +304,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-8b-6k-1_5k-12ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-8b-6k-1_5k-12ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
</tr> </tr>
<tr> <tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-8B</td> <td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-8B</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
@@ -314,7 +314,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-8b-3_5k-1_5k-5ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-8b-3_5k-1_5k-5ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
</tr> </tr>
<tr> <tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-30B-A3B</td> <td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-30B-A3B</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
@@ -324,7 +324,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-30b-a3b-6k-1_5k-10ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-30b-a3b-6k-1_5k-10ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
</tr> </tr>
<tr> <tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-30B-A3B</td> <td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-30B-A3B</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
@@ -334,7 +334,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-30b-a3b-1k-0_3k-7ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-30b-a3b-1k-0_3k-7ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
</tr> </tr>
<tr> <tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-Next-A3B-Instruct</td> <td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-Next-A3B-Instruct</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td>
@@ -344,7 +344,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-next-1k-0_3k-14_21ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-next-1k-0_3k-14_21ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td>
</tr> </tr>
<tr> <tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-Next-A3B-Instruct</td> <td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-Next-A3B-Instruct</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td>
@@ -354,7 +354,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-next-6k-1_5k-15_62ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-next-6k-1_5k-15_62ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td>
</tr> </tr>
<tr> <tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-Next-A3B-Instruct</td> <td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-Next-A3B-Instruct</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
@@ -364,7 +364,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-next-3_5k-1_5k-20ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-next-3_5k-1_5k-20ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
</tr> </tr>
<tr> <tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-14B</td> <td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-14B</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
@@ -374,6 +374,36 @@ you encounter issues or have any questions, please [open an issue](https://githu
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-14b-3_5k-1_5k-9ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-14b-3_5k-1_5k-9ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
</tr> </tr>
<tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-27B</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>PD Mixed</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>3.5K+1.5K</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>20ms</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-27b-3_5k-1_5k-20ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td>
</tr>
<tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-27B</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>PD Mixed</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>16K+1K</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>20ms</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-27b-16k-1k-20ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
</tr>
<tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-27B</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>PD Mixed</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>64K+1K</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>20ms</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-27b-64k-1k-20ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
</tr>
</tbody> </tbody>
</table> </table>
@@ -533,7 +563,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-32b-2k-2k-50ms-on-a2-8-cards-mixed-mode">Optimal Configuration</a></td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-32b-2k-2k-50ms-on-a2-8-cards-mixed-mode">Optimal Configuration</a></td>
</tr> </tr>
<tr> <tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-14B</td> <td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-14B</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
@@ -543,7 +573,7 @@ you encounter issues or have any questions, please [open an issue](https://githu
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-14b-3_5k-1_5k-50ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-14b-3_5k-1_5k-50ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
</tr> </tr>
<tr> <tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-8B</td> <td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-8B</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
@@ -553,6 +583,36 @@ you encounter issues or have any questions, please [open an issue](https://githu
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-8b-3_5k-1_5k-50ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td> <td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-8b-3_5k-1_5k-50ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
</tr> </tr>
<tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-27B</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>1</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>PD Mixed</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>3.5K+1.5K</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>50ms</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-27b-3_5k-1_5k-50ms-on-a3-1-cards-mixed-mode">Optimal Configuration</a></td>
</tr>
<tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-27B</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>PD Mixed</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>16K+1K</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>50ms</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-27b-16k-1k-50ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td>
</tr>
<tr>
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>Qwen3-27B</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Atlas 800I A3</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>2</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>PD Mixed</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>64K+1K</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>50ms</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>W8A8 INT8</td>
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="#qwen3-27b-64k-1k-50ms-on-a3-2-cards-mixed-mode">Optimal Configuration</a></td>
</tr>
</tbody> </tbody>
</table> </table>
@@ -4154,3 +4214,447 @@ We tested it based on the `RANDOM` dataset.
```bash Command ```bash Command
python3 -m sglang.bench_serving --dataset-name random --backend sglang --host 127.0.0.1 --port 6699 --random-range-ratio 1 --max-concurrency 1 --random-output-len 1500 --random-input-len 3500 --num-prompts 1 python3 -m sglang.bench_serving --dataset-name random --backend sglang --host 127.0.0.1 --port 6699 --random-range-ratio 1 --max-concurrency 1 --random-output-len 1500 --random-input-len 3500 --num-prompts 1
``` ```
### Qwen3-27B 3_5K-1_5K 20ms on A3 2 Cards Mixed Mode
Model: Eco-Tech/Qwen3.5-27B-w8a8-mtp
Hardware: Atlas 800I A3 2Card
DeployMode: PD Mixed
Dataset: random
Input Output Length: 3.5K+1.5K
TPOT: 20ms
#### Model Deployment
```bash Command
# high performance cpu
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
sysctl -w vm.swappiness=0
sysctl -w kernel.numa_balancing=0
sysctl -w kernel.sched_migration_cost_ns=50000
# bind cpu
export SGLANG_SET_CPU_AFFINITY=1
unset https_proxy
unset http_proxy
unset HTTPS_PROXY
unset HTTP_PROXY
# on-demand set device
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
export ASCEND_LAUNCH_BLOCKING=1
export STREAMS_PER_DEVICE=32
export HCCL_BUFFSIZE=3000
export HCCL_OP_EXPANSION_MODE="AIV"
export HCCL_SOCKET_IFNAME=lo
export GLOO_SOCKET_IFNAME=lo
export SGLANG_NPU_PROFILING=0
export SGLANG_DISAGGEGATION_WAITING_TIMEOUT=3600
export SGLANG_ENABLE_SPEC_V2=1
export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=0
MODEL_PATH=xxx
python -m sglang.launch_server --model-path ${MODEL_PATH} \
--attention-backend ascend \
--host 127.0.0.1 --port 6699 \
--device npu \
--tp-size 4\
--trust-remote-code \
--watchdog-timeout 9000 \
--chunked-prefill-size -1 \
--max-prefill-tokens 186000 \
--enable-prefill-delayer \
--prefill-delayer-max-delay-passes 200 \
--disable-radix-cache \
--mem-fraction-static 0.94 \
--max-total-tokens 700000 \
--max-running-requests 38 \
--max-mamba-cache-size 200 \
--quantization modelslim \
--dtype bfloat16 \
--mamba-ssm-dtype bfloat16 \
--enable-multimodal \
--mm-attention-backend ascend_attn \
--cuda-graph-bs 1 2 4 8 12 18 24 32 34 36 38 \
--speculative-algorithm NEXTN \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4
```
#### Benchmark
We tested it based on the `RANDOM` dataset.
```bash Command
python3 -m sglang.bench_serving --backend sglang --host 127.0.0.1 --port 6699 --dataset-name random --max-concurrency 38 --num-prompts 152 --random-range-ratio 1 --random-output-len 1500 --random-input-len 3500
```
### Qwen3-27B 16K-1K 20ms on A3 1 Cards Mixed Mode
Model: Eco-Tech/Qwen3.5-27B-w8a8-mtp
Hardware: Atlas 800I A3 1Card
DeployMode: PD Mixed
Dataset: random
Input Output Length: 16K+1K
TPOT: 20ms
#### Model Deployment
```bash Command
# high performance cpu
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
sysctl -w vm.swappiness=0
sysctl -w kernel.numa_balancing=0
sysctl -w kernel.sched_migration_cost_ns=50000
# bind cpu
export SGLANG_SET_CPU_AFFINITY=1
unset https_proxy
unset http_proxy
unset HTTPS_PROXY
unset HTTP_PROXY
# cann
source /usr/local/Ascend/ascend-toolkit/set_env.sh
source /usr/local/Ascend/nnal/atb/set_env.sh
export STREAMS_PER_DEVICE=32
export HCCL_OP_EXPANSION_MODE=AIV
export HCCL_SOCKET_IFNAME=lo
export GLOO_SOCKET_IFNAME=lo
export SGLANG_ENABLE_SPEC_V2=1
export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=0
export SGLANG_SCHEDULER_DECREASE_PREFILL_IDLE=1
export SGLANG_PREFILL_DELAYER_MAX_DELAY_PASSES=100
# on-demand set device
export ASCEND_RT_VISIBLE_DEVICES=8,9
MODEL_PATH=xxx
sglang serve --model-path ${MODEL_PATH} \
--attention-backend ascend \
--device npu \
--tp-size 2 --nnodes 1 --node-rank 0 \
--chunked-prefill-size -1 --max-prefill-tokens 65000 \
--disable-radix-cache \
--trust-remote-code \
--host 127.0.0.1 --max-running-requests 32 --max-mamba-cache-size 32 \
--mem-fraction-static 0.85 \
--port 8001 \
--cuda-graph-bs 2 3 4 5 6 \
--enable-multimodal \
--quantization modelslim \
--mm-attention-backend ascend_attn \
--dtype bfloat16 --mamba-ssm-dtype bfloat16 --max-total-tokens 310000 \
--speculative-algorithm NEXTN --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4
```
#### Benchmark
We tested it based on the `RANDOM` dataset.
```bash Command
python3 -m sglang.bench_serving --backend sglang --host 127.0.0.1 --port 8001 --dataset-name random --max-concurrency 32 --num-prompts 128 --random-range-ratio 1 --random-output-len 1000 --random-input-len 16000
```
### Qwen3-27B 64K-1K 20ms on A3 1 Cards Mixed Mode
Model: Eco-Tech/Qwen3.5-27B-w8a8-mtp
Hardware: Atlas 800I A3 1Card
DeployMode: PD Mixed
Dataset: random
Input Output Length: 64K+1K
TPOT: 20ms
#### Model Deployment
```bash Command
# high performance cpu
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
sysctl -w vm.swappiness=0
sysctl -w kernel.numa_balancing=0
sysctl -w kernel.sched_migration_cost_ns=50000
# bind cpu
export SGLANG_SET_CPU_AFFINITY=1
unset https_proxy
unset http_proxy
unset HTTPS_PROXY
unset HTTP_PROXY
# cann
source /usr/local/Ascend/ascend-toolkit/set_env.sh
source /usr/local/Ascend/nnal/atb/set_env.sh
export STREAMS_PER_DEVICE=32
export HCCL_OP_EXPANSION_MODE=AIV
export HCCL_SOCKET_IFNAME=lo
export GLOO_SOCKET_IFNAME=lo
export SGLANG_NPU_PROFILING=1
export SGLANG_ENABLE_SPEC_V2=1
export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=0
# on-demand set device
export ASCEND_RT_VISIBLE_DEVICES=4,5
MODEL_PATH=xxx
python -m sglang.launch_server --model-path ${MODEL_PATH} \
--attention-backend ascend \
--device npu \
--tp-size 2 --nnodes 1 --node-rank 0 \
--chunked-prefill-size -1 --max-prefill-tokens 130000 \
--disable-radix-cache \
--trust-remote-code \
--host 127.0.0.1 --max-running-requests 32 --max-mamba-cache-size 18 \
--mem-fraction-static 0.5 \
--port 8004 \
--cuda-graph-bs 2 3 4 \
--enable-multimodal \
--quantization modelslim \
--mm-attention-backend ascend_attn \
--dtype bfloat16 --mamba-ssm-dtype bfloat16 --max-total-tokens 280000 \
--speculative-algorithm NEXTN --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4
```
#### Benchmark
We tested it based on the `RANDOM` dataset.
```bash Command
python3 -m sglang.bench_serving --backend sglang --host 127.0.0.1 --port 8004 --dataset-name random --max-concurrency 9 --num-prompts 36 --random-range-ratio 1 --random-output-len 1000 --random-input-len 64000
```
### Qwen3-27B 3_5K-1_5K 50ms on A3 1 Cards Mixed Mode
Model: Eco-Tech/Qwen3.5-27B-w8a8-mtp
Hardware: Atlas 800I A3 1Card
DeployMode: PD Mixed
Dataset: random
Input Output Length: 3.5K+1.5K
TPOT: 50ms
#### Model Deployment
```bash Command
# high performance cpu
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
sysctl -w vm.swappiness=0
sysctl -w kernel.numa_balancing=0
sysctl -w kernel.sched_migration_cost_ns=50000
# bind cpu
export SGLANG_SET_CPU_AFFINITY=1
unset https_proxy
unset http_proxy
unset HTTPS_PROXY
unset HTTP_PROXY
# cann
source /usr/local/Ascend/ascend-toolkit/set_env.sh
source /usr/local/Ascend/nnal/atb/set_env.sh
export STREAMS_PER_DEVICE=32
export HCCL_OP_EXPANSION_MODE=AIV
export HCCL_SOCKET_IFNAME=lo
export GLOO_SOCKET_IFNAME=lo
export SGLANG_ENABLE_SPEC_V2=1
export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=0
export SGLANG_SCHEDULER_DECREASE_PREFILL_IDLE=1
export SGLANG_PREFILL_DELAYER_MAX_DELAY_PASSES=100
MODEL_PATH=xxx
python -m sglang.launch_server --model-path ${MODEL_PATH} \
--attention-backend ascend \
--device npu \
--tp-size 2 --nnodes 1 --node-rank 0 \
--chunked-prefill-size -1 --max-prefill-tokens 60000 \
--disable-radix-cache \
--trust-remote-code \
--host 127.0.0.1 --max-running-requests 48 --max-mamba-cache-size 60 \
--mem-fraction-static 0.7 \
--port 8000 \
--cuda-graph-bs 2 8 16 32 48 \
--enable-multimodal \
--quantization modelslim \
--mm-attention-backend ascend_attn \
--dtype bfloat16 --mamba-ssm-dtype bfloat16 \
--speculative-algorithm NEXTN --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4
```
#### Benchmark
We tested it based on the `RANDOM` dataset.
```bash Command
python3 -m sglang.bench_serving --backend sglang --host 127.0.0.1 --port 8000 --dataset-name random --max-concurrency 48 --num-prompts 192 --random-range-ratio 1 --random-output-len 1500 --random-input-len 3500
```
### Qwen3-27B 16K-1K 50ms on A3 2 Cards Mixed Mode
Model: Eco-Tech/Qwen3.5-27B-w8a8-mtp
Hardware: Atlas 800I A3 2Card
DeployMode: PD Mixed
Dataset: random
Input Output Length: 16K+1K
TPOT: 50ms
#### Model Deployment
```bash Command
# high performance cpu
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
sysctl -w vm.swappiness=0
sysctl -w kernel.numa_balancing=0
sysctl -w kernel.sched_migration_cost_ns=50000
# bind cpu
export SGLANG_SET_CPU_AFFINITY=1
unset https_proxy
unset http_proxy
unset HTTPS_PROXY
unset HTTP_PROXY
# cann
source /usr/local/Ascend/ascend-toolkit/set_env.sh
source /usr/local/Ascend/nnal/atb/set_env.sh
export STREAMS_PER_DEVICE=32
export HCCL_OP_EXPANSION_MODE=AIV
export HCCL_SOCKET_IFNAME=lo
export GLOO_SOCKET_IFNAME=lo
export SGLANG_ENABLE_SPEC_V2=1
export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1
export SGLANG_SCHEDULER_DECREASE_PREFILL_IDLE=1
export SGLANG_PREFILL_DELAYER_MAX_DELAY_PASSES=30
# on-demand set device
export ASCEND_RT_VISIBLE_DEVICES=0,1,2,3
MODEL_PATH=xxx
python3 -m sglang.launch_server --model-path ${MODEL_PATH} \
--attention-backend ascend \
--device npu \
--tp-size 4 --nnodes 1 --node-rank 0 \
--chunked-prefill-size -1 --max-prefill-tokens 50000 \
--disable-radix-cache \
--trust-remote-code \
--host 127.0.0.1 --max-running-requests 28 --max-mamba-cache-size 50 \
--mem-fraction-static 0.7 \
--port 8001 \
--cuda-graph-bs 2 8 12 16 20 24 28\
--enable-multimodal \
--quantization modelslim \
--mm-attention-backend ascend_attn \
--dtype bfloat16 --mamba-ssm-dtype bfloat16 \
--speculative-algorithm NEXTN --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4
```
#### Benchmark
We tested it based on the `RANDOM` dataset.
```bash Command
python3 -m sglang.bench_serving --backend sglang --host 127.0.0.1 --port 8001 --dataset-name random --max-concurrency 28 --num-prompts 152 --random-range-ratio 1 --random-output-len 1000 --random-input-len 16000
```
### Qwen3-27B 64K-1K 50ms on A3 2 Cards Mixed Mode
Model: Eco-Tech/Qwen3.5-27B-w8a8-mtp
Hardware: Atlas 800I A3 2Card
DeployMode: PD Mixed
Dataset: random
Input Output Length: 64K+1K
TPOT: 50ms
#### Model Deployment
```bash Command
# high performance cpu
echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
sysctl -w vm.swappiness=0
sysctl -w kernel.numa_balancing=0
sysctl -w kernel.sched_migration_cost_ns=50000
# bind cpu
export SGLANG_SET_CPU_AFFINITY=1
unset https_proxy
unset http_proxy
unset HTTPS_PROXY
unset HTTP_PROXY
# cann
source /usr/local/Ascend/ascend-toolkit/set_env.sh
source /usr/local/Ascend/nnal/atb/set_env.sh
export STREAMS_PER_DEVICE=32
export HCCL_OP_EXPANSION_MODE=AIV
export HCCL_SOCKET_IFNAME=lo
export GLOO_SOCKET_IFNAME=lo
export SGLANG_ENABLE_SPEC_V2=1
export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=0
export SGLANG_SCHEDULER_DECREASE_PREFILL_IDLE=1
export SGLANG_PREFILL_DELAYER_MAX_DELAY_PASSES=100
# on-demand set device
export ASCEND_RT_VISIBLE_DEVICES=4,5,6,7
MODEL_PATH=xxx
python3 -m sglang.launch_server --model-path ${MODEL_PATH} \
--attention-backend ascend \
--device npu \
--tp-size 4 --nnodes 1 --node-rank 0 \
--chunked-prefill-size -1 --max-prefill-tokens 200000 \
--disable-radix-cache \
--trust-remote-code \
--host 127.0.0.1 --max-running-requests 32 --max-mamba-cache-size 22 \
--mem-fraction-static 0.5 \
--port 9000 \
--cuda-graph-bs 2 4 8 11 12 13 \
--enable-multimodal \
--quantization modelslim \
--mm-attention-backend ascend_attn \
--dtype bfloat16 --mamba-ssm-dtype bfloat16 --max-total-tokens 850000 \
--speculative-algorithm NEXTN --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4
```
#### Benchmark
We tested it based on the `RANDOM` dataset.
```bash Command
python3 -m sglang.bench_serving --backend sglang --host 127.0.0.1 --port 9000 --dataset-name random --max-concurrency 9 --num-prompts 36 --random-range-ratio 1 --random-output-len 1000 --random-input-len 64000
```
@@ -351,6 +351,111 @@ Max ITL (ms): 2229.30
================================================== ==================================================
``` ```
#### SGLang Serving Benchmark Result — Complete Reference
The output format is **hardcoded in `bench_serving.py`. All formatting decisions — including column widths, alignment, and decimal precision — are statically defined in the source and cannot be changed via command-line arguments.
##### Test Configuration
<table>
<thead>
<tr>
<th>Parameter</th>
<th>Description</th>
</tr>
</thead>
<tbody>
<tr>
<td><code>Backend</code></td>
<td>The serving backend under test (e.g., <code>sglang</code>, <code>vllm</code>).</td>
</tr>
<tr>
<td><code>Traffic request rate</code></td>
<td>Request generation rate in req/s. <code>inf</code> means maximum rate (concurrency-bounded). <code>trace</code> indicates trace timestamp mode. A fixed value enforces constant inter-arrival time.</td>
</tr>
<tr>
<td><code>Max request concurrency</code></td>
<td>Maximum number of concurrent requests from the client side. Displays <code>not set</code> when unspecified.</td>
</tr>
</tbody>
</table>
##### Core Statistics & Throughput Metrics
<table>
<thead>
<tr>
<th>Parameter</th>
<th>Description</th>
<th>Format Specification</th>
</tr>
</thead>
<tbody>
<tr><td><code>Successful requests</code></td><td>Total number of successfully completed requests (HTTP 200, no generation errors).</td><td>Integer, no decimal places</td></tr>
<tr><td><code>Benchmark duration (s)</code></td><td>Total elapsed time from first request sent to last response fully received (seconds).</td><td>2 decimal places</td></tr>
<tr><td><code>Total input tokens</code></td><td>Total number of input (prompt) tokens across all requests, counted by server-side tokenizer.</td><td>Integer, no decimal places</td></tr>
<tr><td><code>Total input text tokens</code></td><td>Same as <code>Total input tokens</code>. For multimodal inputs, this may differ.</td><td>Integer, no decimal places</td></tr>
<tr><td><code>Total generated tokens</code></td><td>Total number of output tokens actually generated by the server (server-side tokenizer count).</td><td>Integer, no decimal places</td></tr>
<tr><td><code>Total generated tokens (retokenized)</code></td><td>Output text re-tokenized by the client using its own tokenizer. A large discrepancy indicates tokenizer mismatch or special tokens in output.</td><td>Integer, no decimal places</td></tr>
<tr><td><code>Request throughput (req/s)</code></td><td>Number of successful requests processed per second. Formula: <code>Successful requests / Benchmark duration (s)</code>.</td><td>2 decimal places</td></tr>
<tr><td><code>Input token throughput (tok/s)</code></td><td>Number of input tokens processed per second. Formula: <code>Total input tokens / Benchmark duration (s)</code>.</td><td>2 decimal places</td></tr>
<tr><td><code>Output token throughput (tok/s)</code></td><td>Number of output tokens generated per second. Formula: <code>Total generated tokens / Benchmark duration (s)</code>.</td><td>2 decimal places</td></tr>
<tr><td><code>Peak output token throughput (tok/s)</code></td><td>Observed instantaneous peak output token generation rate during the test (computed over a sliding window).</td><td>2 decimal places</td></tr>
<tr><td><code>Peak concurrent requests</code></td><td>Maximum number of requests being processed simultaneously on the server side. May exceed client-side <code>Max request concurrency</code> due to queueing.</td><td>Integer, no decimal places</td></tr>
<tr><td><code>Total token throughput (tok/s)</code></td><td>Sum of input and output token throughputs. Formula: <code>Input token throughput + Output token throughput</code>.</td><td>2 decimal places</td></tr>
<tr><td><code>Concurrency</code></td><td>Average number of concurrent requests during the test (Little's Law). Formula: <code>Sum of all E2E latencies / Benchmark duration</code>.</td><td>2 decimal places</td></tr>
</tbody>
</table>
##### End-to-End Latency (E2E Latency)
<table>
<thead><tr><th>Statistic</th><th>Description</th><th>Format</th></tr></thead>
<tbody>
<tr><td><code>Mean E2E Latency (ms)</code></td><td>Arithmetic mean</td><td>2 decimal places</td></tr>
<tr><td><code>Median E2E Latency (ms)</code></td><td>50th percentile</td><td>2 decimal places</td></tr>
<tr><td><code>P90 E2E Latency (ms)</code></td><td>90th percentile (90% of requests have latency ≤ this value)</td><td>2 decimal places</td></tr>
<tr><td><code>P99 E2E Latency (ms)</code></td><td>99th percentile</td><td>2 decimal places</td></tr>
</tbody>
</table>
##### Time to First Token (TTFT)
<table>
<thead><tr><th>Statistic</th><th>Description</th><th>Format</th></tr></thead>
<tbody>
<tr><td><code>Mean TTFT (ms)</code></td><td>Arithmetic mean</td><td>2 decimal places</td></tr>
<tr><td><code>Median TTFT (ms)</code></td><td>50th percentile</td><td>2 decimal places</td></tr>
<tr><td><code>P99 TTFT (ms)</code></td><td>99th percentile</td><td>2 decimal places</td></tr>
</tbody>
</table>
##### Time per Output Token (TPOT) – Excluding First Token
Formula: <code>(E2E Latency - TTFT) / (Number of output tokens - 1)</code>
<table>
<thead><tr><th>Statistic</th><th>Description</th><th>Format</th></tr></thead>
<tbody>
<tr><td><code>Mean TPOT (ms)</code></td><td>Arithmetic mean</td><td>2 decimal places</td></tr>
<tr><td><code>Median TPOT (ms)</code></td><td>50th percentile</td><td>2 decimal places</td></tr>
<tr><td><code>P99 TPOT (ms)</code></td><td>99th percentile</td><td>2 decimal places</td></tr>
</tbody>
</table>
##### Inter-Token Latency (ITL)
<table>
<thead><tr><th>Statistic</th><th>Description</th><th>Format</th></tr></thead>
<tbody>
<tr><td><code>Mean ITL (ms)</code></td><td>Average inter-token interval</td><td>2 decimal places</td></tr>
<tr><td><code>Median ITL (ms)</code></td><td>50th percentile inter-token interval</td><td>2 decimal places</td></tr>
<tr><td><code>P95 ITL (ms)</code></td><td>95th percentile (used to detect stalls)</td><td>2 decimal places</td></tr>
<tr><td><code>P99 ITL (ms)</code></td><td>99th percentile</td><td>2 decimal places</td></tr>
<tr><td><code>Max ITL (ms)</code></td><td>Maximum observed inter-token interval; useful for identifying severe blocking events</td><td>2 decimal places</td></tr>
</tbody>
</table>
## 3. Online Service: Multimodal Model ## 3. Online Service: Multimodal Model
Test `Qwen/Qwen2.5-VL-7B-Instruct` for vision-language tasks. Test `Qwen/Qwen2.5-VL-7B-Instruct` for vision-language tasks.