830 lines
28 KiB
Plaintext
830 lines
28 KiB
Plaintext
---
|
|
title: GLM Glyph
|
|
metatags:
|
|
description: "Deploy GLM-Glyph with SGLang - community contribution guide for Zhipu AI's GLM Glyph model deployment."
|
|
---
|
|
|
|
import { GLMGlyphDeployment } from '/src/snippets/autoregressive/glm-glyph-deployment.jsx';
|
|
|
|
## 1. Model Introduction
|
|
|
|
[Glyph](https://huggingface.co/zai-org/Glyph) is a powerful language model developed by Zhipu AI, featuring advanced capabilities in reasoning, function calling, and multi-modal understanding.
|
|
|
|
**Hardware Support:** NVIDIA B200/H100/H200, AMD MI300X/MI325X/MI355X
|
|
|
|
**Key Features:**
|
|
|
|
- **Advanced Reasoning**: Built-in reasoning capabilities for complex problem-solving
|
|
- **Multiple Quantizations**: BF16 and FP8 variants for different performance/memory trade-offs
|
|
- **High Performance**: Optimized for both throughput and latency scenarios
|
|
|
|
**Available Models:**
|
|
|
|
- **BF16 (Full precision)**: [zai-org/Glyph](https://huggingface.co/zai-org/Glyph)
|
|
- **FP8 (8-bit quantized)**: [zai-org/Glyph-FP8](https://huggingface.co/zai-org/Glyph-FP8)
|
|
|
|
**License:**
|
|
|
|
Please refer to the [official Glyph model card](https://huggingface.co/zai-org/Glyph) for license details.
|
|
|
|
## 2. SGLang Installation
|
|
|
|
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
|
|
|
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
|
|
|
## 3. Model Deployment
|
|
|
|
This section provides deployment configurations optimized for different hardware platforms and use cases.
|
|
|
|
### 3.1 Basic Configuration
|
|
|
|
**Interactive Command Generator**: Use the configuration selector below to automatically generate the appropriate deployment command for your hardware platform, quantization method, and other options.
|
|
|
|
<GLMGlyphDeployment />
|
|
|
|
### 3.2 Configuration Tips
|
|
|
|
- **Thinking Budget:** Use `--enable-custom-logit-processor` flag and pass `Glm4MoeThinkingBudgetLogitProcessor` in requests to cap the model's thinking token count. See the [GLM-4.5 cookbook page](/cookbook/autoregressive/GLM/GLM-4.5) for the full Thinking Budget usage example.
|
|
|
|
## 4. Model Invocation
|
|
|
|
### 4.1 Basic Usage
|
|
|
|
For basic API usage and request examples, please refer to:
|
|
|
|
- [SGLang Basic Usage Guide](../../../docs/basic_usage/send_request)
|
|
|
|
### 4.2 Advanced Usage
|
|
|
|
#### 4.2.1 Thinking Mode
|
|
|
|
Glyph supports thinking mode for enhanced reasoning. Enable the reasoning parser during deployment to separate the thinking and content sections:
|
|
|
|
```shell Command
|
|
python -m sglang.launch_server \
|
|
--model-path zai-org/Glyph \
|
|
--reasoning-parser glm45 \
|
|
--tp 4
|
|
```
|
|
|
|
**Streaming with Thinking Process:**
|
|
|
|
```python Example
|
|
from openai import OpenAI
|
|
|
|
client = OpenAI(
|
|
base_url="http://localhost:30000/v1",
|
|
api_key="EMPTY"
|
|
)
|
|
|
|
# Enable streaming to see the thinking process in real-time
|
|
response = client.chat.completions.create(
|
|
model="zai-org/Glyph",
|
|
messages=[
|
|
{"role": "user", "content": "Solve this problem step by step: What is 15% of 240?"}
|
|
],
|
|
temperature=0.7,
|
|
max_tokens=2048,
|
|
stream=True
|
|
)
|
|
|
|
# Process the stream
|
|
has_thinking = False
|
|
has_answer = False
|
|
thinking_started = False
|
|
|
|
for chunk in response:
|
|
if chunk.choices and len(chunk.choices) > 0:
|
|
delta = chunk.choices[0].delta
|
|
|
|
# Print thinking process
|
|
if hasattr(delta, 'reasoning_content') and delta.reasoning_content:
|
|
if not thinking_started:
|
|
print("=============== Thinking =================", flush=True)
|
|
thinking_started = True
|
|
has_thinking = True
|
|
print(delta.reasoning_content, end="", flush=True)
|
|
|
|
# Print answer content
|
|
if delta.content:
|
|
# Close thinking section and add content header
|
|
if has_thinking and not has_answer:
|
|
print("\n=============== Content =================", flush=True)
|
|
has_answer = True
|
|
print(delta.content, end="", flush=True)
|
|
|
|
print()
|
|
```
|
|
|
|
**Note:** The reasoning parser captures the model's step-by-step thinking process, allowing you to see how the model arrives at its conclusions.
|
|
|
|
**Disable Thinking Mode:**
|
|
|
|
To disable thinking mode for a specific request:
|
|
|
|
```python Example
|
|
response = client.chat.completions.create(
|
|
model="zai-org/Glyph",
|
|
messages=[{"role": "user", "content": "What is the capital of France?"}],
|
|
extra_body={"chat_template_kwargs": {"enable_thinking": False}}
|
|
)
|
|
```
|
|
|
|
#### 4.2.2 Tool Calling
|
|
|
|
Glyph supports tool calling capabilities. Enable the tool call parser:
|
|
|
|
```shell Command
|
|
python -m sglang.launch_server \
|
|
--model-path zai-org/Glyph \
|
|
--reasoning-parser glm45 \
|
|
--tool-call-parser glm45 \
|
|
--tp 4
|
|
```
|
|
|
|
**Python Example (with Thinking Process):**
|
|
|
|
```python Example
|
|
from openai import OpenAI
|
|
|
|
client = OpenAI(
|
|
base_url="http://localhost:30000/v1",
|
|
api_key="EMPTY"
|
|
)
|
|
|
|
# Define available tools
|
|
tools = [
|
|
{
|
|
"type": "function",
|
|
"function": {
|
|
"name": "get_weather",
|
|
"description": "Get the current weather for a location",
|
|
"parameters": {
|
|
"type": "object",
|
|
"properties": {
|
|
"location": {
|
|
"type": "string",
|
|
"description": "The city name"
|
|
},
|
|
"unit": {
|
|
"type": "string",
|
|
"enum": ["celsius", "fahrenheit"],
|
|
"description": "Temperature unit"
|
|
}
|
|
},
|
|
"required": ["location"]
|
|
}
|
|
}
|
|
}
|
|
]
|
|
|
|
# Make request with streaming to see thinking process
|
|
response = client.chat.completions.create(
|
|
model="zai-org/Glyph",
|
|
messages=[
|
|
{"role": "user", "content": "What's the weather in Beijing?"}
|
|
],
|
|
tools=tools,
|
|
temperature=0.7,
|
|
stream=True
|
|
)
|
|
|
|
# Process streaming response
|
|
thinking_started = False
|
|
has_thinking = False
|
|
tool_calls_accumulator = {}
|
|
|
|
for chunk in response:
|
|
if chunk.choices and len(chunk.choices) > 0:
|
|
delta = chunk.choices[0].delta
|
|
|
|
# Print thinking process
|
|
if hasattr(delta, 'reasoning_content') and delta.reasoning_content:
|
|
if not thinking_started:
|
|
print("=============== Thinking =================", flush=True)
|
|
thinking_started = True
|
|
has_thinking = True
|
|
print(delta.reasoning_content, end="", flush=True)
|
|
|
|
# Accumulate tool calls
|
|
if hasattr(delta, 'tool_calls') and delta.tool_calls:
|
|
# Close thinking section if needed
|
|
if has_thinking and thinking_started:
|
|
print("\n=============== Content =================\n", flush=True)
|
|
thinking_started = False
|
|
|
|
for tool_call in delta.tool_calls:
|
|
index = tool_call.index
|
|
if index not in tool_calls_accumulator:
|
|
tool_calls_accumulator[index] = {
|
|
'name': None,
|
|
'arguments': ''
|
|
}
|
|
|
|
if tool_call.function:
|
|
if tool_call.function.name:
|
|
tool_calls_accumulator[index]['name'] = tool_call.function.name
|
|
if tool_call.function.arguments:
|
|
tool_calls_accumulator[index]['arguments'] += tool_call.function.arguments
|
|
|
|
# Print content
|
|
if delta.content:
|
|
print(delta.content, end="", flush=True)
|
|
|
|
# Print accumulated tool calls
|
|
for index, tool_call in sorted(tool_calls_accumulator.items()):
|
|
print(f"Tool Call: {tool_call['name']}")
|
|
print(f" Arguments: {tool_call['arguments']}")
|
|
|
|
print()
|
|
```
|
|
|
|
**Output Example:**
|
|
|
|
```text Output
|
|
=============== Thinking =================
|
|
The user is asking about the weather in Beijing. I need to use the get_weather function to retrieve this information.
|
|
I should call the function with location="Beijing".
|
|
=============== Content =================
|
|
|
|
Tool Call: get_weather
|
|
Arguments: {"location": "Beijing", "unit": "celsius"}
|
|
```
|
|
|
|
**Note:**
|
|
|
|
- The reasoning parser shows how the model decides to use a tool
|
|
- Tool calls are clearly marked with the function name and arguments
|
|
- You can then execute the function and send the result back to continue the conversation
|
|
|
|
**Handling Tool Call Results:**
|
|
|
|
```python Example
|
|
# After getting the tool call, execute the function
|
|
def get_weather(location, unit="celsius"):
|
|
# Your actual weather API call here
|
|
return f"The weather in {location} is 22°{unit[0].upper()} and sunny."
|
|
|
|
# Send tool result back to the model
|
|
messages = [
|
|
{"role": "user", "content": "What's the weather in Beijing?"},
|
|
{
|
|
"role": "assistant",
|
|
"content": None,
|
|
"tool_calls": [{
|
|
"id": "call_123",
|
|
"type": "function",
|
|
"function": {
|
|
"name": "get_weather",
|
|
"arguments": '{"location": "Beijing", "unit": "celsius"}'
|
|
}
|
|
}]
|
|
},
|
|
{
|
|
"role": "tool",
|
|
"tool_call_id": "call_123",
|
|
"content": get_weather("Beijing", "celsius")
|
|
}
|
|
]
|
|
|
|
final_response = client.chat.completions.create(
|
|
model="zai-org/Glyph",
|
|
messages=messages,
|
|
temperature=0.7
|
|
)
|
|
|
|
print(final_response.choices[0].message.content)
|
|
# Output: "The weather in Beijing is currently 22°C and sunny."
|
|
```
|
|
|
|
## 5. Benchmark
|
|
|
|
This section uses **industry-standard configurations** for comparable benchmark results.
|
|
|
|
### 5.1 Speed Benchmark
|
|
|
|
**Test Environment:**
|
|
|
|
- Model: Glyph
|
|
- SGLang Version: 0.5.6.post1
|
|
|
|
**Benchmark Methodology:**
|
|
|
|
We use industry-standard benchmark configurations to ensure results are comparable across frameworks and hardware platforms.
|
|
|
|
#### 5.1.1 Standard Scenario Benchmark
|
|
|
|
- **Model Deployment**
|
|
```bash Command
|
|
python -m sglang.launch_server \
|
|
--model zai-org/Glyph \
|
|
--tp 2
|
|
```
|
|
|
|
##### 5.1.1.1 Low Concurrency
|
|
- **Benchmark Command**:
|
|
```bash Command
|
|
python -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model zai-org/Glyph \
|
|
--dataset-name random \
|
|
--random-input-len 1000 \
|
|
--random-output-len 1000 \
|
|
--num-prompts 10 \
|
|
--max-concurrency 1 \
|
|
--request-rate inf
|
|
```
|
|
|
|
- **Test Results**:
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 1
|
|
Successful requests: 10
|
|
Benchmark duration (s): 17.03
|
|
Total input tokens: 6101
|
|
Total input text tokens: 6101
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 4220
|
|
Total generated tokens (retokenized): 4220
|
|
Request throughput (req/s): 0.59
|
|
Input token throughput (tok/s): 358.17
|
|
Output token throughput (tok/s): 247.74
|
|
Peak output token throughput (tok/s): 251.00
|
|
Peak concurrent requests: 3
|
|
Total token throughput (tok/s): 605.91
|
|
Concurrency: 1.00
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 1702.14
|
|
Median E2E Latency (ms): 1361.72
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 22.35
|
|
Median TTFT (ms): 22.61
|
|
P99 TTFT (ms): 23.76
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 3.99
|
|
Median TPOT (ms): 3.99
|
|
P99 TPOT (ms): 4.01
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 3.99
|
|
Median ITL (ms): 3.99
|
|
P95 ITL (ms): 4.03
|
|
P99 ITL (ms): 4.12
|
|
Max ITL (ms): 7.46
|
|
==================================================
|
|
```
|
|
|
|
##### 5.1.1.2 Medium Concurrency
|
|
- **Benchmark Command**:
|
|
```bash Command
|
|
python -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model zai-org/Glyph \
|
|
--dataset-name random \
|
|
--random-input-len 1000 \
|
|
--random-output-len 1000 \
|
|
--num-prompts 80 \
|
|
--max-concurrency 16 \
|
|
--request-rate inf
|
|
```
|
|
|
|
- **Test Results**:
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 16
|
|
Successful requests: 80
|
|
Benchmark duration (s): 16.27
|
|
Total input tokens: 39668
|
|
Total input text tokens: 39668
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 40805
|
|
Total generated tokens (retokenized): 40804
|
|
Request throughput (req/s): 4.92
|
|
Input token throughput (tok/s): 2438.06
|
|
Output token throughput (tok/s): 2507.94
|
|
Peak output token throughput (tok/s): 3069.00
|
|
Peak concurrent requests: 26
|
|
Total token throughput (tok/s): 4946.00
|
|
Concurrency: 13.44
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 2733.43
|
|
Median E2E Latency (ms): 2892.98
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 33.10
|
|
Median TTFT (ms): 27.73
|
|
P99 TTFT (ms): 49.34
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 5.33
|
|
Median TPOT (ms): 5.39
|
|
P99 TPOT (ms): 5.86
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 5.30
|
|
Median ITL (ms): 4.89
|
|
P95 ITL (ms): 5.54
|
|
P99 ITL (ms): 21.17
|
|
Max ITL (ms): 25.14
|
|
==================================================
|
|
```
|
|
|
|
##### 5.1.1.3 High Concurrency
|
|
- **Benchmark Command**:
|
|
```bash Command
|
|
python -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model zai-org/Glyph \
|
|
--dataset-name random \
|
|
--random-input-len 1000 \
|
|
--random-output-len 1000 \
|
|
--num-prompts 500 \
|
|
--max-concurrency 100 \
|
|
--request-rate inf
|
|
```
|
|
- **Test Results**:
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 100
|
|
Successful requests: 500
|
|
Benchmark duration (s): 25.67
|
|
Total input tokens: 249831
|
|
Total input text tokens: 249831
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 252662
|
|
Total generated tokens (retokenized): 252657
|
|
Request throughput (req/s): 19.48
|
|
Input token throughput (tok/s): 9733.69
|
|
Output token throughput (tok/s): 9843.99
|
|
Peak output token throughput (tok/s): 13398.00
|
|
Peak concurrent requests: 127
|
|
Total token throughput (tok/s): 19577.68
|
|
Concurrency: 89.49
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 4593.75
|
|
Median E2E Latency (ms): 4431.03
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 48.66
|
|
Median TTFT (ms): 35.88
|
|
P99 TTFT (ms): 120.61
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 9.10
|
|
Median TPOT (ms): 9.55
|
|
P99 TPOT (ms): 11.00
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 9.01
|
|
Median ITL (ms): 6.51
|
|
P95 ITL (ms): 23.19
|
|
P99 ITL (ms): 25.54
|
|
Max ITL (ms): 52.93
|
|
==================================================
|
|
```
|
|
|
|
#### 5.1.2 Reasoning Scenario Benchmark
|
|
|
|
##### 5.1.2.1 Low Concurrency
|
|
- **Benchmark Command**:
|
|
```bash Command
|
|
python -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model zai-org/Glyph \
|
|
--dataset-name random \
|
|
--random-input-len 1000 \
|
|
--random-output-len 8000 \
|
|
--num-prompts 10 \
|
|
--max-concurrency 1 \
|
|
--request-rate inf
|
|
```
|
|
- **Test Results**:
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 1
|
|
Successful requests: 10
|
|
Benchmark duration (s): 201.53
|
|
Total input tokens: 6101
|
|
Total input text tokens: 6101
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 44462
|
|
Total generated tokens (retokenized): 44455
|
|
Request throughput (req/s): 0.05
|
|
Input token throughput (tok/s): 30.27
|
|
Output token throughput (tok/s): 220.63
|
|
Peak output token throughput (tok/s): 251.00
|
|
Peak concurrent requests: 2
|
|
Total token throughput (tok/s): 250.90
|
|
Concurrency: 1.00
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 20151.45
|
|
Median E2E Latency (ms): 21576.31
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 2362.23
|
|
Median TTFT (ms): 23.03
|
|
P99 TTFT (ms): 21310.14
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 4.00
|
|
Median TPOT (ms): 4.00
|
|
P99 TPOT (ms): 4.01
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 4.00
|
|
Median ITL (ms): 4.00
|
|
P95 ITL (ms): 4.05
|
|
P99 ITL (ms): 4.08
|
|
Max ITL (ms): 5.67
|
|
==================================================
|
|
```
|
|
|
|
##### 5.1.2.2 Medium Concurrency
|
|
- **Benchmark Command**:
|
|
```bash Command
|
|
python -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model zai-org/Glyph \
|
|
--dataset-name random \
|
|
--random-input-len 1000 \
|
|
--random-output-len 8000 \
|
|
--num-prompts 80 \
|
|
--max-concurrency 16 \
|
|
--request-rate inf
|
|
```
|
|
- **Test Results**:
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 16
|
|
Successful requests: 80
|
|
Benchmark duration (s): 118.67
|
|
Total input tokens: 39668
|
|
Total input text tokens: 39668
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 318306
|
|
Total generated tokens (retokenized): 318270
|
|
Request throughput (req/s): 0.67
|
|
Input token throughput (tok/s): 334.27
|
|
Output token throughput (tok/s): 2682.26
|
|
Peak output token throughput (tok/s): 3264.00
|
|
Peak concurrent requests: 19
|
|
Total token throughput (tok/s): 3016.53
|
|
Concurrency: 13.74
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 20387.23
|
|
Median E2E Latency (ms): 20466.09
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 132.47
|
|
Median TTFT (ms): 27.19
|
|
P99 TTFT (ms): 583.15
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 5.09
|
|
Median TPOT (ms): 5.13
|
|
P99 TPOT (ms): 5.19
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 5.09
|
|
Median ITL (ms): 5.08
|
|
P95 ITL (ms): 5.18
|
|
P99 ITL (ms): 5.57
|
|
Max ITL (ms): 522.26
|
|
==================================================
|
|
```
|
|
|
|
##### 5.1.2.3 High Concurrency
|
|
- **Benchmark Command**:
|
|
```bash Command
|
|
python -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model zai-org/Glyph \
|
|
--dataset-name random \
|
|
--random-input-len 1000 \
|
|
--random-output-len 8000 \
|
|
--num-prompts 320 \
|
|
--max-concurrency 64 \
|
|
--request-rate inf
|
|
```
|
|
- **Test Results**:
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 64
|
|
Successful requests: 320
|
|
Benchmark duration (s): 150.00
|
|
Total input tokens: 158939
|
|
Total input text tokens: 158939
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 1301025
|
|
Total generated tokens (retokenized): 1300901
|
|
Request throughput (req/s): 2.13
|
|
Input token throughput (tok/s): 1059.59
|
|
Output token throughput (tok/s): 8673.49
|
|
Peak output token throughput (tok/s): 11899.00
|
|
Peak concurrent requests: 71
|
|
Total token throughput (tok/s): 9733.09
|
|
Concurrency: 54.71
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 25645.42
|
|
Median E2E Latency (ms): 26913.26
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 163.75
|
|
Median TTFT (ms): 93.67
|
|
P99 TTFT (ms): 426.19
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 6.27
|
|
Median TPOT (ms): 6.39
|
|
P99 TPOT (ms): 6.59
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 6.27
|
|
Median ITL (ms): 0.17
|
|
P95 ITL (ms): 32.94
|
|
P99 ITL (ms): 67.89
|
|
Max ITL (ms): 136.00
|
|
==================================================
|
|
```
|
|
|
|
#### 5.1.3 Summarization Scenario Benchmark
|
|
|
|
#### 5.1.3.1 Low Concurrency
|
|
- **Benchmark Command**:
|
|
```bash Command
|
|
python -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model zai-org/Glyph \
|
|
--dataset-name random \
|
|
--random-input-len 8000 \
|
|
--random-output-len 1000 \
|
|
--num-prompts 10 \
|
|
--max-concurrency 1 \
|
|
--request-rate inf
|
|
```
|
|
- **Test Results**:
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 1
|
|
Successful requests: 10
|
|
Benchmark duration (s): 17.44
|
|
Total input tokens: 41941
|
|
Total input text tokens: 41941
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 4220
|
|
Total generated tokens (retokenized): 4220
|
|
Request throughput (req/s): 0.57
|
|
Input token throughput (tok/s): 2405.19
|
|
Output token throughput (tok/s): 242.00
|
|
Peak output token throughput (tok/s): 250.00
|
|
Peak concurrent requests: 2
|
|
Total token throughput (tok/s): 2647.19
|
|
Concurrency: 1.00
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 1742.54
|
|
Median E2E Latency (ms): 1412.47
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 53.48
|
|
Median TTFT (ms): 45.05
|
|
P99 TTFT (ms): 98.57
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 4.01
|
|
Median TPOT (ms): 4.01
|
|
P99 TPOT (ms): 4.03
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 4.01
|
|
Median ITL (ms): 4.01
|
|
P95 ITL (ms): 4.06
|
|
P99 ITL (ms): 4.09
|
|
Max ITL (ms): 4.95
|
|
==================================================
|
|
```
|
|
|
|
##### 5.1.3.2 Medium Concurrency
|
|
- **Benchmark Command**:
|
|
```bash Command
|
|
python -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model zai-org/Glyph \
|
|
--dataset-name random \
|
|
--random-input-len 8000 \
|
|
--random-output-len 1000 \
|
|
--num-prompts 80 \
|
|
--max-concurrency 16 \
|
|
--request-rate inf
|
|
```
|
|
- **Test Results**:
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 16
|
|
Successful requests: 80
|
|
Benchmark duration (s): 16.90
|
|
Total input tokens: 300020
|
|
Total input text tokens: 300020
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 41669
|
|
Total generated tokens (retokenized): 41668
|
|
Request throughput (req/s): 4.73
|
|
Input token throughput (tok/s): 17753.58
|
|
Output token throughput (tok/s): 2465.75
|
|
Peak output token throughput (tok/s): 3005.00
|
|
Peak concurrent requests: 25
|
|
Total token throughput (tok/s): 20219.33
|
|
Concurrency: 13.68
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 2890.33
|
|
Median E2E Latency (ms): 3069.55
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 41.46
|
|
Median TTFT (ms): 31.75
|
|
P99 TTFT (ms): 93.18
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 5.52
|
|
Median TPOT (ms): 5.58
|
|
P99 TPOT (ms): 6.14
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 5.48
|
|
Median ITL (ms): 5.13
|
|
P95 ITL (ms): 5.93
|
|
P99 ITL (ms): 20.76
|
|
Max ITL (ms): 36.01
|
|
==================================================
|
|
```
|
|
|
|
##### 5.1.3.3 High Concurrency
|
|
|
|
- **Benchmark Command**:
|
|
```bash Command
|
|
python -m sglang.bench_serving \
|
|
--backend sglang \
|
|
--model zai-org/Glyph \
|
|
--dataset-name random \
|
|
--random-input-len 8000 \
|
|
--random-output-len 1000 \
|
|
--num-prompts 320 \
|
|
--max-concurrency 64 \
|
|
--request-rate inf
|
|
```
|
|
- **Test Results**:
|
|
```text Output
|
|
============ Serving Benchmark Result ============
|
|
Backend: sglang
|
|
Traffic request rate: inf
|
|
Max request concurrency: 64
|
|
Successful requests: 320
|
|
Benchmark duration (s): 35.54
|
|
Total input tokens: 1273893
|
|
Total input text tokens: 1273893
|
|
Total input vision tokens: 0
|
|
Total generated tokens: 170000
|
|
Total generated tokens (retokenized): 169994
|
|
Request throughput (req/s): 9.01
|
|
Input token throughput (tok/s): 35848.57
|
|
Output token throughput (tok/s): 4783.96
|
|
Peak output token throughput (tok/s): 8396.00
|
|
Peak concurrent requests: 80
|
|
Total token throughput (tok/s): 40632.53
|
|
Concurrency: 59.26
|
|
----------------End-to-End Latency----------------
|
|
Mean E2E Latency (ms): 6580.96
|
|
Median E2E Latency (ms): 6248.74
|
|
---------------Time to First Token----------------
|
|
Mean TTFT (ms): 345.27
|
|
Median TTFT (ms): 96.06
|
|
P99 TTFT (ms): 2823.92
|
|
-----Time per Output Token (excl. 1st token)------
|
|
Mean TPOT (ms): 12.26
|
|
Median TPOT (ms): 12.53
|
|
P99 TPOT (ms): 23.58
|
|
---------------Inter-Token Latency----------------
|
|
Mean ITL (ms): 11.76
|
|
Median ITL (ms): 6.57
|
|
P95 ITL (ms): 27.66
|
|
P99 ITL (ms): 91.24
|
|
Max ITL (ms): 2609.64
|
|
==================================================
|
|
```
|
|
|
|
### 5.2 Accuracy Benchmark
|
|
|
|
Document model accuracy on standard benchmarks:
|
|
|
|
#### 5.2.1 GSM8K Benchmark
|
|
|
|
- Benchmark Command
|
|
|
|
```bash Command
|
|
python -m sglang.test.few_shot_gsm8k \
|
|
--num-questions 200
|
|
```
|
|
|
|
- Test Result
|
|
|
|
```text Output
|
|
Accuracy: 0.890
|
|
Invalid: 0.000
|
|
Latency: 3.718 s
|
|
Output throughput: 5245.606 token/s
|
|
```
|