886 lines
32 KiB
Plaintext
886 lines
32 KiB
Plaintext
---
|
|
title: Qwen3
|
|
metatags:
|
|
description: "Deploy Qwen3 series models with SGLang - featuring advanced reasoning, 256K context, and flexible Dense/MoE architectures for edge to cloud."
|
|
---
|
|
|
|
|
|
## 1. Model Introduction
|
|
|
|
[Qwen3 series](https://github.com/QwenLM/Qwen3) are the most powerful vision-language models in the Qwen series to date, featuring advanced capabilities in multi-modal understanding, reasoning, and agentic applications.
|
|
|
|
This generation delivers comprehensive upgrades across the board:
|
|
|
|
- **Stronger general intelligence**: Significant improvements in instruction following, logical reasoning, text comprehension, mathematics, science, coding, and tool usage.
|
|
- **Broader multilingual knowledge**: Substantial gains in long-tail knowledge coverage across multiple languages.
|
|
- **More helpful & aligned responses**: Markedly better alignment with user preferences in subjective and open-ended tasks, enabling higher-quality, more useful text generation.
|
|
- **Extended context length**: Enhanced capabilities in understanding and reasoning over 256K-token long contexts.
|
|
- **Stronger agent interaction capabilities**: Improved tool use and search-based agent performance.
|
|
- **Flexible deployment options**: Available in Dense and MoE architectures that scale from edge to cloud, with Instruct and reasoning-enhanced Thinking editions.
|
|
|
|
For more details, please refer to the [official Qwen3 GitHub Repository](https://github.com/QwenLM/Qwen3).
|
|
|
|
## 2. SGLang Installation
|
|
|
|
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
|
|
|
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
|
|
|
## 3. Model Deployment
|
|
|
|
This section provides deployment configurations optimized for different hardware platforms and use cases.
|
|
|
|
### 3.1 Basic Configuration
|
|
|
|
The Qwen3 series offers models in various sizes and architectures, optimized for different hardware platforms including NVIDIA GPUs, AMD GPUs, Intel Arc Pro B-Series GPUs(codename: BMG (Battlemage)), and Intel Xeon CPUs. The recommended launch configurations vary by hardware and model size.
|
|
|
|
**Interactive Command Generator**: Use the configuration selector below to automatically generate the appropriate deployment command for your hardware platform, model size, quantization method, and thinking capabilities.
|
|
|
|
import { Qwen3Deployment } from "/src/snippets/autoregressive/qwen3-deployment.jsx";
|
|
|
|
<Qwen3Deployment />
|
|
|
|
### 3.2 Configuration Tips
|
|
|
|
- **Memory Management:** Set lower `--context-length` to conserve memory. A value of `128000` is sufficient for most scenarios, down from the default 262K.
|
|
- **Expert Parallelism:** SGLang supports Expert Parallelism (EP) via `--ep`, allowing experts in MoE models to be deployed on separate GPUs for better throughput. One thing to note is that, for quantized models, you need to set `--ep` to a value that satisfies the requirement: `(moe_intermediate_size / moe_tp_size) % weight_block_size_n == 0, where moe_tp_size is equal to tp_size divided by ep_size.` Note that EP may perform worse in low concurrency scenarios due to additional communication overhead. Check out [Expert Parallelism Deployment](../../../docs/advanced_features/expert_parallelism) for more details.
|
|
- **Kernel Tuning:** For MoE Triton kernel tuning on your specific hardware, refer to [fused_moe_triton](https://github.com/sgl-project/sglang/tree/main/benchmark/kernels/fused_moe_triton).
|
|
- **Speculative Decoding:** Using Speculative Decoding for latency-sensitive scenarios.
|
|
- `--speculative-algorithm EAGLE3`: Speculative decoding algorithm
|
|
- `--speculative-num-steps 3`: Number of speculative verification rounds
|
|
- `--speculative-eagle-topk 1`: Top-k sampling for draft tokens
|
|
- `--speculative-num-draft-tokens 4`: Number of draft tokens per step
|
|
- `--speculative-draft-model-path`: The path of the draft model weights. This can be a local folder or a Hugging Face repo ID such as [`lmsys/SGLang-EAGLE3-Qwen3-235B-A22B-Instruct-2507-SpecForge-Meituan`](https://huggingface.co/lmsys/SGLang-EAGLE3-Qwen3-235B-A22B-Instruct-2507-SpecForge-Meituan).
|
|
- **Xeon CPU service configuration:** Please refer to the `Notes` part in the serving engine launching section in [the SGLang CPU server document](../../../docs/hardware-platforms/cpu_server#launch-of-the-serving-engine) to better understand how to configure the arguments, especially for TP (tensor parallel) and NUMA binding settings.
|
|
|
|
## 4. Model Invocation
|
|
|
|
### 4.1 Basic Usage
|
|
|
|
For basic API usage and request examples, please refer to:
|
|
|
|
- [SGLang Basic Usage Guide](../../../docs/basic_usage/send_request)
|
|
- [SGLang OpenAI Vision API Guide](../../../docs/basic_usage/openai_api_vision)
|
|
|
|
### 4.2 Advanced Usage
|
|
|
|
#### 4.2.1 Reasoning Parser
|
|
|
|
Qwen3-235B-A22B supports reasoning mode. Enable the reasoning parser during deployment to separate the thinking and content sections:
|
|
|
|
```shell Command
|
|
python -m sglang.launch_server \
|
|
--model Qwen/Qwen3-235B-A22B-Thinking-2507 \
|
|
--reasoning-parser qwen3 \
|
|
--tp 8 \
|
|
--host 0.0.0.0 \
|
|
--port 8000
|
|
```
|
|
|
|
**Streaming with Thinking Process:**
|
|
|
|
```python Example
|
|
from openai import OpenAI
|
|
|
|
client = OpenAI(
|
|
base_url="http://localhost:8000/v1",
|
|
api_key="EMPTY"
|
|
)
|
|
|
|
# Enable streaming to see the thinking process in real-time
|
|
response = client.chat.completions.create(
|
|
model="Qwen/Qwen3-235B-A22B-Thinking-2507",
|
|
messages=[
|
|
{"role": "user", "content": "Solve this problem step by step: What is 15% of 240?"}
|
|
],
|
|
temperature=0.7,
|
|
max_tokens=2048,
|
|
stream=True
|
|
)
|
|
|
|
# Process the stream
|
|
has_thinking = False
|
|
has_answer = False
|
|
thinking_started = False
|
|
|
|
for chunk in response:
|
|
if chunk.choices and len(chunk.choices) > 0:
|
|
delta = chunk.choices[0].delta
|
|
|
|
# Print thinking process
|
|
if hasattr(delta, 'reasoning_content') and delta.reasoning_content:
|
|
if not thinking_started:
|
|
print("=============== Thinking =================", flush=True)
|
|
thinking_started = True
|
|
has_thinking = True
|
|
print(delta.reasoning_content, end="", flush=True)
|
|
|
|
# Print answer content
|
|
if delta.content:
|
|
# Close thinking section and add content header
|
|
if has_thinking and not has_answer:
|
|
print("\n=============== Content =================", flush=True)
|
|
has_answer = True
|
|
print(delta.content, end="", flush=True)
|
|
|
|
print()
|
|
```
|
|
|
|
**Output Example:**
|
|
|
|
```text Output
|
|
=============== Thinking =================
|
|
|
|
Okay, so I need to figure out what 15% of 240 is. Hmm, percentages can sometimes trip me up, but I think I remember some basics. Let me start by recalling that "percent" means "per hundred," so 15% is the same as 15 per 100, or 15/100. So, maybe I can convert 15% into a decimal first? Yeah, I think that's a common method.
|
|
...
|
|
So conclusion: The answer is 36.
|
|
|
|
=============== Content =================
|
|
|
|
|
|
To determine what 15% of 240 is, we can follow a systematic approach that involves converting the percentage to a decimal and then performing multiplication. Here's a step-by-step breakdown of the solution:
|
|
|
|
....
|
|
|
|
### Final Answer:
|
|
|
|
$$
|
|
\boxed{36}
|
|
$$
|
|
|
|
Thus, 15% of 240 is **36**.
|
|
```
|
|
|
|
**Note:** The reasoning parser captures the model's step-by-step thinking process, allowing you to see how the model arrives at its conclusions.
|
|
|
|
#### 4.2.3 Tool Calling
|
|
|
|
Qwen3 supports tool calling capabilities. Enable the tool call parser:
|
|
|
|
```shell Command
|
|
python -m sglang.launch_server \
|
|
--model Qwen/Qwen3-235B-A22B-Thinking-2507 \
|
|
--reasoning-parser qwen3 \
|
|
--tool-call-parser qwen25 \
|
|
--tp 8 \
|
|
--host 0.0.0.0 \
|
|
--port 8000
|
|
```
|
|
|
|
**Python Example (with Thinking Process):**
|
|
|
|
```python Example
|
|
from openai import OpenAI
|
|
|
|
client = OpenAI(
|
|
base_url="http://localhost:8000/v1",
|
|
api_key="EMPTY"
|
|
)
|
|
|
|
# Define available tools
|
|
tools = [
|
|
{
|
|
"type": "function",
|
|
"function": {
|
|
"name": "get_weather",
|
|
"description": "Get the current weather for a location",
|
|
"parameters": {
|
|
"type": "object",
|
|
"properties": {
|
|
"location": {
|
|
"type": "string",
|
|
"description": "The city name"
|
|
},
|
|
"unit": {
|
|
"type": "string",
|
|
"enum": ["celsius", "fahrenheit"],
|
|
"description": "Temperature unit"
|
|
}
|
|
},
|
|
"required": ["location"]
|
|
}
|
|
}
|
|
}
|
|
]
|
|
|
|
# Make request with streaming to see thinking process
|
|
response = client.chat.completions.create(
|
|
model="Qwen/Qwen3-235B-A22B-Thinking-2507",
|
|
messages=[
|
|
{"role": "user", "content": "What's the weather in Beijing?"}
|
|
],
|
|
tools=tools,
|
|
temperature=0.7,
|
|
stream=True
|
|
)
|
|
|
|
# Process streaming response
|
|
thinking_started = False
|
|
has_thinking = False
|
|
tool_calls_accumulator = {}
|
|
|
|
for chunk in response:
|
|
if chunk.choices and len(chunk.choices) > 0:
|
|
delta = chunk.choices[0].delta
|
|
|
|
# Print thinking process
|
|
if hasattr(delta, 'reasoning_content') and delta.reasoning_content:
|
|
if not thinking_started:
|
|
print("=============== Thinking =================", flush=True)
|
|
thinking_started = True
|
|
has_thinking = True
|
|
print(delta.reasoning_content, end="", flush=True)
|
|
|
|
# Accumulate tool calls
|
|
if hasattr(delta, 'tool_calls') and delta.tool_calls:
|
|
# Close thinking section if needed
|
|
if has_thinking and thinking_started:
|
|
print("\n=============== Content =================\n", flush=True)
|
|
thinking_started = False
|
|
|
|
for tool_call in delta.tool_calls:
|
|
index = tool_call.index
|
|
if index not in tool_calls_accumulator:
|
|
tool_calls_accumulator[index] = {
|
|
'name': None,
|
|
'arguments': ''
|
|
}
|
|
|
|
if tool_call.function:
|
|
if tool_call.function.name:
|
|
tool_calls_accumulator[index]['name'] = tool_call.function.name
|
|
if tool_call.function.arguments:
|
|
tool_calls_accumulator[index]['arguments'] += tool_call.function.arguments
|
|
|
|
# Print content
|
|
if delta.content:
|
|
print(delta.content, end="", flush=True)
|
|
|
|
# Print accumulated tool calls
|
|
for index, tool_call in sorted(tool_calls_accumulator.items()):
|
|
print(f"🔧 Tool Call: {tool_call['name']}")
|
|
print(f" Arguments: {tool_call['arguments']}")
|
|
|
|
print()
|
|
```
|
|
|
|
**Output Example:**
|
|
|
|
```text Output
|
|
=============== Thinking =================
|
|
|
|
Okay, the user is asking for the weather in Beijing. Let me check the tools available. There's a function called get_weather that takes location and unit parameters. The location is required, so I need to specify Beijing as the location. The unit is optional and can be either celsius or fahrenheit. Since the user didn't specify the unit, maybe I should default to a common one. In China, they usually use celsius, so I'll set unit to celsius. I'll call the get_weather function with location: Beijing and unit: celsius. That should get the current weather for them.
|
|
|
|
|
|
|
|
=============== Content =================
|
|
|
|
🔧 Tool Call: get_weather
|
|
Arguments: {"location": "Beijing", "unit": "celsius"}
|
|
```
|
|
|
|
**Note:**
|
|
|
|
- The reasoning parser shows how the model decides to use a tool
|
|
- Tool calls are clearly marked with the function name and arguments
|
|
- You can then execute the function and send the result back to continue the conversation
|
|
|
|
**Handling Tool Call Results:**
|
|
|
|
```python Example
|
|
# After getting the tool call, execute the function
|
|
def get_weather(location, unit="celsius"):
|
|
# Your actual weather API call here
|
|
return f"The weather in {location} is 22°{unit[0].upper()} and sunny."
|
|
|
|
# Send tool result back to the model
|
|
messages = [
|
|
{"role": "user", "content": "What's the weather in Beijing?"},
|
|
{
|
|
"role": "assistant",
|
|
"content": None,
|
|
"tool_calls": [{
|
|
"id": "call_123",
|
|
"type": "function",
|
|
"function": {
|
|
"name": "get_weather",
|
|
"arguments": '{"location": "Beijing", "unit": "celsius"}'
|
|
}
|
|
}]
|
|
},
|
|
{
|
|
"role": "tool",
|
|
"tool_call_id": "call_123",
|
|
"content": get_weather("Beijing", "celsius")
|
|
}
|
|
]
|
|
|
|
final_response = client.chat.completions.create(
|
|
model="Qwen/Qwen3-235B-A22B-Thinking-2507",
|
|
messages=messages,
|
|
temperature=0.7
|
|
)
|
|
|
|
print(final_response.choices[0].message.content)
|
|
# Output: "The current weather in Beijing is **22°C** and **sunny**. A perfect day to enjoy outdoor activities! 🌞"
|
|
```
|
|
|
|
## 5. Benchmark
|
|
|
|
### 5.1 Speed Benchmark
|
|
|
|
**Test Environment:**
|
|
|
|
- Hardware: NVIDIA B200 GPU (8x)
|
|
- Model: Qwen3-235B-A22B-Instruct-2507
|
|
- Tensor Parallelism: 8
|
|
- sglang version: 0.5.6
|
|
|
|
We use SGLang's built-in benchmarking tool to conduct performance evaluation on the [ShareGPT_Vicuna_unfiltered](https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered) dataset. This dataset contains real conversation data and can better reflect performance in actual use scenarios.
|
|
|
|
#### 5.1.1 Standard Scenario Benchmark
|
|
|
|
- Model Deployment Command:
|
|
|
|
```shell Command
|
|
python -m sglang.launch_server \
|
|
--model Qwen/Qwen3-235B-A22B-Instruct-2507 \
|
|
--tp 8
|
|
```
|
|
|
|
##### 5.1.1.1 Low Concurrency
|
|
|
|
- Benchmark Command:
|
|
|
|
```shell Command
|
|
python3 -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model Qwen/Qwen3-235B-A22B-Instruct-2507 \
|
|
--dataset-name random \
|
|
--random-input-len 1000 \
|
|
--random-output-len 1000 \
|
|
--num-prompts 10 \
|
|
--max-concurrency 1
|
|
```
|
|
|
|
- Test Results:
|
|
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 1
|
|
Successful requests: 10
|
|
Benchmark duration (s): 43.56
|
|
Total input tokens: 6101
|
|
Total input text tokens: 6101
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 4210
|
|
Total generated tokens (retokenized): 4206
|
|
Request throughput (req/s): 0.23
|
|
Input token throughput (tok/s): 140.07
|
|
Output token throughput (tok/s): 96.65
|
|
Peak output token throughput (tok/s): 100.00
|
|
Peak concurrent requests: 2
|
|
Total token throughput (tok/s): 236.72
|
|
Concurrency: 1.00
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 4353.63
|
|
Median E2E Latency (ms): 3475.79
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 99.03
|
|
Median TTFT (ms): 92.18
|
|
P99 TTFT (ms): 166.05
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 10.12
|
|
Median TPOT (ms): 10.12
|
|
P99 TPOT (ms): 10.15
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 10.13
|
|
Median ITL (ms): 10.12
|
|
P95 ITL (ms): 10.49
|
|
P99 ITL (ms): 10.70
|
|
Max ITL (ms): 13.45
|
|
==================================================
|
|
```
|
|
|
|
##### 5.1.1.2 Medium Concurrency
|
|
|
|
- Benchmark Command:
|
|
|
|
```shell Command
|
|
python3 -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model Qwen/Qwen3-235B-A22B-Instruct-2507 \
|
|
--dataset-name random \
|
|
--random-input-len 1000 \
|
|
--random-output-len 1000 \
|
|
--num-prompts 80 \
|
|
--max-concurrency 16
|
|
```
|
|
|
|
- Test Results:
|
|
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 16
|
|
Successful requests: 80
|
|
Benchmark duration (s): 48.95
|
|
Total input tokens: 39668
|
|
Total input text tokens: 39668
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 40725
|
|
Total generated tokens (retokenized): 40716
|
|
Request throughput (req/s): 1.63
|
|
Input token throughput (tok/s): 810.44
|
|
Output token throughput (tok/s): 832.04
|
|
Peak output token throughput (tok/s): 1151.00
|
|
Peak concurrent requests: 21
|
|
Total token throughput (tok/s): 1642.48
|
|
Concurrency: 13.61
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 8326.72
|
|
Median E2E Latency (ms): 8827.86
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 215.70
|
|
Median TTFT (ms): 88.82
|
|
P99 TTFT (ms): 727.08
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 16.36
|
|
Median TPOT (ms): 16.12
|
|
P99 TPOT (ms): 24.09
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 15.96
|
|
Median ITL (ms): 14.52
|
|
P95 ITL (ms): 16.04
|
|
P99 ITL (ms): 67.69
|
|
Max ITL (ms): 457.52
|
|
==================================================
|
|
```
|
|
|
|
##### 5.1.1.3 High Concurrency
|
|
|
|
- Benchmark Command:
|
|
|
|
```shell Command
|
|
python3 -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model Qwen/Qwen3-235B-A22B-Instruct-2507 \
|
|
--dataset-name random \
|
|
--random-input-len 1000 \
|
|
--random-output-len 1000 \
|
|
--num-prompts 500 \
|
|
--max-concurrency 100
|
|
```
|
|
|
|
- Test Results:
|
|
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 100
|
|
Successful requests: 500
|
|
Benchmark duration (s): 92.07
|
|
Total input tokens: 249831
|
|
Total input text tokens: 249831
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 252162
|
|
Total generated tokens (retokenized): 251124
|
|
Request throughput (req/s): 5.43
|
|
Input token throughput (tok/s): 2713.46
|
|
Output token throughput (tok/s): 2738.78
|
|
Peak output token throughput (tok/s): 4400.00
|
|
Peak concurrent requests: 110
|
|
Total token throughput (tok/s): 5452.24
|
|
Concurrency: 90.50
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 16665.09
|
|
Median E2E Latency (ms): 16060.10
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 260.55
|
|
Median TTFT (ms): 122.68
|
|
P99 TTFT (ms): 863.11
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 32.94
|
|
Median TPOT (ms): 34.04
|
|
P99 TPOT (ms): 41.19
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 32.59
|
|
Median ITL (ms): 23.54
|
|
P95 ITL (ms): 69.79
|
|
P99 ITL (ms): 119.09
|
|
Max ITL (ms): 577.70
|
|
==================================================
|
|
```
|
|
|
|
#### 5.1.2 Reasoning Scenario Benchmark
|
|
|
|
- Model Deployment Command:
|
|
|
|
```shell Command
|
|
python -m sglang.launch_server \
|
|
--model Qwen/Qwen3-235B-A22B-Instruct-2507 \
|
|
--tp 8
|
|
```
|
|
|
|
##### 5.1.2.1 Low Concurrency
|
|
|
|
- Benchmark Command:
|
|
|
|
```shell Command
|
|
python3 -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model Qwen/Qwen3-235B-A22B-Instruct-2507 \
|
|
--dataset-name random \
|
|
--random-input-len 1000 \
|
|
--random-output-len 8000 \
|
|
--num-prompts 10 \
|
|
--max-concurrency 1
|
|
```
|
|
|
|
- Test Results:
|
|
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 1
|
|
Successful requests: 10
|
|
Benchmark duration (s): 457.45
|
|
Total input tokens: 6101
|
|
Total input text tokens: 6101
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 44452
|
|
Total generated tokens (retokenized): 44059
|
|
Request throughput (req/s): 0.02
|
|
Input token throughput (tok/s): 13.34
|
|
Output token throughput (tok/s): 97.17
|
|
Peak output token throughput (tok/s): 100.00
|
|
Peak concurrent requests: 2
|
|
Total token throughput (tok/s): 110.51
|
|
Concurrency: 1.00
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 45742.42
|
|
Median E2E Latency (ms): 49266.87
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 110.60
|
|
Median TTFT (ms): 109.36
|
|
P99 TTFT (ms): 167.43
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 10.23
|
|
Median TPOT (ms): 10.24
|
|
P99 TPOT (ms): 10.32
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 10.27
|
|
Median ITL (ms): 10.26
|
|
P95 ITL (ms): 10.71
|
|
P99 ITL (ms): 10.97
|
|
Max ITL (ms): 15.79
|
|
==================================================
|
|
```
|
|
|
|
##### 5.1.2.2 Medium Concurrency
|
|
|
|
- Benchmark Command:
|
|
|
|
```shell Command
|
|
python3 -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model Qwen/Qwen3-235B-A22B-Instruct-2507 \
|
|
--dataset-name random \
|
|
--random-input-len 1000 \
|
|
--random-output-len 8000 \
|
|
--num-prompts 80 \
|
|
--max-concurrency 16
|
|
```
|
|
|
|
- Test Results:
|
|
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 16
|
|
Successful requests: 80
|
|
Benchmark duration (s): 340.17
|
|
Total input tokens: 39668
|
|
Total input text tokens: 39668
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 318226
|
|
Total generated tokens (retokenized): 318104
|
|
Request throughput (req/s): 0.24
|
|
Input token throughput (tok/s): 116.61
|
|
Output token throughput (tok/s): 935.49
|
|
Peak output token throughput (tok/s): 1120.00
|
|
Peak concurrent requests: 19
|
|
Total token throughput (tok/s): 1052.10
|
|
Concurrency: 13.85
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 58885.30
|
|
Median E2E Latency (ms): 59238.70
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 169.71
|
|
Median TTFT (ms): 101.61
|
|
P99 TTFT (ms): 455.71
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 14.82
|
|
Median TPOT (ms): 14.91
|
|
P99 TPOT (ms): 15.20
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 14.76
|
|
Median ITL (ms): 14.63
|
|
P95 ITL (ms): 15.46
|
|
P99 ITL (ms): 16.62
|
|
Max ITL (ms): 104.94
|
|
==================================================
|
|
```
|
|
|
|
##### 5.1.2.3 High Concurrency
|
|
|
|
- Benchmark Command:
|
|
|
|
```shell Command
|
|
python3 -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model Qwen/Qwen3-235B-A22B-Instruct-2507 \
|
|
--dataset-name random \
|
|
--random-input-len 1000 \
|
|
--random-output-len 8000 \
|
|
--num-prompts 320 \
|
|
--max-concurrency 64
|
|
```
|
|
|
|
- Test Results:
|
|
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 64
|
|
Successful requests: 320
|
|
Benchmark duration (s): 544.83
|
|
Total input tokens: 158939
|
|
Total input text tokens: 158939
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 1300705
|
|
Total generated tokens (retokenized): 1293015
|
|
Request throughput (req/s): 0.59
|
|
Input token throughput (tok/s): 291.72
|
|
Output token throughput (tok/s): 2387.34
|
|
Peak output token throughput (tok/s): 3008.00
|
|
Peak concurrent requests: 68
|
|
Total token throughput (tok/s): 2679.06
|
|
Concurrency: 56.35
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 95937.70
|
|
Median E2E Latency (ms): 99362.32
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 265.03
|
|
Median TTFT (ms): 129.11
|
|
P99 TTFT (ms): 823.85
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 23.66
|
|
Median TPOT (ms): 24.07
|
|
P99 TPOT (ms): 24.97
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 23.54
|
|
Median ITL (ms): 23.07
|
|
P95 ITL (ms): 25.92
|
|
P99 ITL (ms): 63.87
|
|
Max ITL (ms): 408.30
|
|
==================================================
|
|
```
|
|
|
|
#### 5.1.3 Summarization Scenario Benchmark
|
|
|
|
##### 5.1.3.1 Low Concurrency
|
|
|
|
- Benchmark Command:
|
|
|
|
```shell Command
|
|
python3 -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model Qwen/Qwen3-235B-A22B-Instruct-2507 \
|
|
--dataset-name random \
|
|
--random-input-len 8000 \
|
|
--random-output-len 1000 \
|
|
--num-prompts 10 \
|
|
--max-concurrency 1
|
|
```
|
|
|
|
- Test Results:
|
|
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 1
|
|
Successful requests: 10
|
|
Benchmark duration (s): 44.82
|
|
Total input tokens: 41941
|
|
Total input text tokens: 41941
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 4210
|
|
Total generated tokens (retokenized): 4210
|
|
Request throughput (req/s): 0.22
|
|
Input token throughput (tok/s): 935.86
|
|
Output token throughput (tok/s): 93.94
|
|
Peak output token throughput (tok/s): 99.00
|
|
Peak concurrent requests: 2
|
|
Total token throughput (tok/s): 1029.80
|
|
Concurrency: 1.00
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 4479.60
|
|
Median E2E Latency (ms): 3622.99
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 139.90
|
|
Median TTFT (ms): 114.85
|
|
P99 TTFT (ms): 225.17
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 10.31
|
|
Median TPOT (ms): 10.33
|
|
P99 TPOT (ms): 10.51
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 10.33
|
|
Median ITL (ms): 10.33
|
|
P95 ITL (ms): 10.73
|
|
P99 ITL (ms): 10.93
|
|
Max ITL (ms): 14.48
|
|
==================================================
|
|
```
|
|
|
|
##### 5.1.3.2 Medium Concurrency
|
|
|
|
- Benchmark Command:
|
|
|
|
```shell Command
|
|
python3 -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model Qwen/Qwen3-235B-A22B-Instruct-2507 \
|
|
--dataset-name random \
|
|
--random-input-len 8000 \
|
|
--random-output-len 1000 \
|
|
--num-prompts 80 \
|
|
--max-concurrency 16
|
|
```
|
|
|
|
- Test Results:
|
|
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 16
|
|
Successful requests: 80
|
|
Benchmark duration (s): 50.68
|
|
Total input tokens: 300020
|
|
Total input text tokens: 300020
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 41589
|
|
Total generated tokens (retokenized): 41578
|
|
Request throughput (req/s): 1.58
|
|
Input token throughput (tok/s): 5920.41
|
|
Output token throughput (tok/s): 820.69
|
|
Peak output token throughput (tok/s): 1200.00
|
|
Peak concurrent requests: 20
|
|
Total token throughput (tok/s): 6741.10
|
|
Concurrency: 13.90
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 8805.54
|
|
Median E2E Latency (ms): 9368.79
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 284.29
|
|
Median TTFT (ms): 168.48
|
|
P99 TTFT (ms): 1027.21
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 16.81
|
|
Median TPOT (ms): 16.66
|
|
P99 TPOT (ms): 27.18
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 16.42
|
|
Median ITL (ms): 13.68
|
|
P95 ITL (ms): 17.23
|
|
P99 ITL (ms): 90.75
|
|
Max ITL (ms): 574.64
|
|
==================================================
|
|
```
|
|
|
|
##### 5.1.3.3 High Concurrency
|
|
|
|
- Benchmark Command:
|
|
|
|
```shell Command
|
|
python3 -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model Qwen/Qwen3-235B-A22B-Instruct-2507 \
|
|
--dataset-name random \
|
|
--random-input-len 8000 \
|
|
--random-output-len 1000 \
|
|
--num-prompts 320 \
|
|
--max-concurrency 64
|
|
```
|
|
|
|
- Test Results:
|
|
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 64
|
|
Successful requests: 320
|
|
Benchmark duration (s): 94.77
|
|
Total input tokens: 1273893
|
|
Total input text tokens: 1273893
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 169680
|
|
Total generated tokens (retokenized): 169640
|
|
Request throughput (req/s): 3.38
|
|
Input token throughput (tok/s): 13441.86
|
|
Output token throughput (tok/s): 1790.43
|
|
Peak output token throughput (tok/s): 2687.00
|
|
Peak concurrent requests: 70
|
|
Total token throughput (tok/s): 15232.28
|
|
Concurrency: 58.63
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 17364.14
|
|
Median E2E Latency (ms): 17495.95
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 238.22
|
|
Median TTFT (ms): 203.27
|
|
P99 TTFT (ms): 510.48
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 32.50
|
|
Median TPOT (ms): 34.27
|
|
P99 TPOT (ms): 40.59
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 32.36
|
|
Median ITL (ms): 22.50
|
|
P95 ITL (ms): 97.81
|
|
P99 ITL (ms): 151.55
|
|
Max ITL (ms): 352.79
|
|
==================================================
|
|
```
|
|
|
|
### 5.2 Accuracy Benchmark
|
|
|
|
#### 5.2.1 GSM8K Benchmark
|
|
|
|
- **Benchmark Command:**
|
|
|
|
```shell Command
|
|
python3 -m sglang.test.few_shot_gsm8k --num-questions 200
|
|
```
|
|
|
|
- **Results**:
|
|
|
|
- Qwen/Qwen3-235B-A22B-Instruct-2507
|
|
```text Output
|
|
Accuracy: 0.945
|
|
Invalid: 0.000
|
|
Latency: 11.980 s
|
|
Output throughput: 2358.105 token/s
|
|
```
|