Add new Mintlify documentation site (docs_new/) (#23001)
Co-authored-by: AdityaVKochar <adityavardhankochar@gmail.com> Co-authored-by: mintlify[bot] <109931778+mintlify[bot]@users.noreply.github.com> Co-authored-by: adhyan-jain <adhyanjain2006@gmail.com> Co-authored-by: Adhyan Jain <71976554+adhyan-jain@users.noreply.github.com> Co-authored-by: Maitri-shah29 <maitrirajivshah@gmail.com> Co-authored-by: Adarsh Shirawalmath <114558126+adarshxs@users.noreply.github.com> Co-authored-by: Maitri Shah <shah29maitri@gmail.com> Co-authored-by: Aditya Vardhan Kochar <80113212+AdityaVKochar@users.noreply.github.com> Co-authored-by: Rishit Shivam <164783543+pokymono@users.noreply.github.com> Co-authored-by: Rishitshivam <164783543+Rishitshivam@users.noreply.github.com> Co-authored-by: IshhanKheria <ishhankheria06@gmail.com> Co-authored-by: Ishita Joshi <ishitata.joshi@gmail.com> Co-authored-by: Richard Chen <104477092+Richardczl98@users.noreply.github.com> Co-authored-by: longGGGGGG <553746008@qq.com> Co-authored-by: Richard <richardchen@radixark.ai> Co-authored-by: Nakul Sinha <nakul.new4socials@gmail.com> Co-authored-by: Divyam Agrawal <ludicrouslytrue@gmail.com> Co-authored-by: Richardczl98 <Zhenlinc@stanford.edu> Co-authored-by: Krishang Zinzuwadia <krishangzinzuwadia@gmail.com> Co-authored-by: nimeshas <nimesha.s106@gmail.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Jignas Paturu <86356085+JignasP@users.noreply.github.com> Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
This commit is contained in:
co-authored by
AdityaVKochar
mintlify[bot]
adhyan-jain
Adhyan Jain
Maitri-shah29
Adarsh Shirawalmath
Maitri Shah
Aditya Vardhan Kochar
Rishit Shivam
Rishitshivam
IshhanKheria
Ishita Joshi
Richard Chen
longGGGGGG
Richard
Nakul Sinha
Divyam Agrawal
Richardczl98
Krishang Zinzuwadia
nimeshas
Claude Opus 4.6
github-actions[bot]
Jignas Paturu
zijiexia
parent
575fdc2c4c
commit
a3291b5654
@@ -0,0 +1,695 @@
|
||||
---
|
||||
title: Step3-VL-10B
|
||||
metatags:
|
||||
description: "Deploy Step3-VL-10B multimodal model with SGLang - compact 10B dense model with frontier-level vision understanding, complex reasoning, and tool calling capabilities."
|
||||
tag: NEW
|
||||
---
|
||||
|
||||
import { Step3VL10BDeployment } from '/src/snippets/autoregressive/step-3vl-10b-deployment.jsx';
|
||||
|
||||
## 1. Model Introduction
|
||||
|
||||
[Step3-VL-10B](https://huggingface.co/stepfun-ai/Step3-VL-10B) is a lightweight open-source multimodal model developed by StepFun, designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence. Despite its compact 10B parameter footprint, Step3-VL-10B excels in visual perception, complex reasoning, and human-centric alignment.
|
||||
|
||||
Key highlights of Step3-VL-10B include:
|
||||
|
||||
- **STEM Reasoning**: Achieves 94.43% on AIME 2025 and 75.95% on MathVision (with PaCoRe), demonstrating exceptional complex reasoning capabilities that outperform models 10×–20× larger.
|
||||
- **Visual Perception**: Records 92.05% on MMBench and 80.11% on MMMU, establishing strong general visual understanding and multimodal reasoning.
|
||||
- **GUI & OCR**: Delivers state-of-the-art performance on ScreenSpot-V2 (92.61%), ScreenSpot-Pro (51.55%), and OCRBench (86.75%), optimized for agentic and document understanding tasks.
|
||||
- **Spatial Understanding**: Demonstrates emergent spatial awareness with 66.79% on BLINK and 57.21% on All-Angles-Bench, establishing strong potential for embodied intelligence applications.
|
||||
|
||||
For more details, please refer to the [Step3-VL-10B model card on Hugging Face](https://huggingface.co/stepfun-ai/Step3-VL-10B).
|
||||
|
||||
## 2. SGLang Installation
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
This section provides deployment configurations optimized for different hardware platforms and use cases.
|
||||
|
||||
### 3.1 Basic Configuration
|
||||
|
||||
Step3-VL-10B is a compact 10B dense model that can run on a single GPU. Recommended starting configurations vary depending on hardware.
|
||||
|
||||
**Interactive Command Generator**: Use the configuration selector below to automatically generate the appropriate deployment command for your hardware platform and quantization method. SGLang supports serving Step3-VL-10B on NVIDIA B200, H200, H100, and AMD MI355X, MI325X, MI300X GPUs.
|
||||
|
||||
<Step3VL10BDeployment />
|
||||
|
||||
### 3.2 Configuration Tips
|
||||
|
||||
- **Single GPU Deployment**: Step3-VL-10B fits comfortably on a single GPU with BF16 precision, no tensor parallelism required.
|
||||
- **Memory Management**: Set lower `--context-length` to conserve memory if needed. A value of `32768` is sufficient for most scenarios.
|
||||
- **FP8 Quantization**: Use FP8 quantization to further reduce memory usage while maintaining quality.
|
||||
|
||||
## 4. Model Invocation
|
||||
|
||||
### 4.1 Basic Usage
|
||||
|
||||
For basic API usage and request examples, please refer to:
|
||||
|
||||
- [SGLang Basic Usage Guide](../../../docs/basic_usage/send_request)
|
||||
- [SGLang OpenAI Vision API Guide](../../../docs/basic_usage/openai_api_vision)
|
||||
|
||||
### 4.2 Advanced Usage
|
||||
|
||||
#### 4.2.1 Multi-Modal Inputs
|
||||
|
||||
Step3-VL-10B supports image inputs. Here's a basic example with image input:
|
||||
|
||||
```python Example
|
||||
import time
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(
|
||||
api_key="EMPTY",
|
||||
base_url="http://localhost:30000/v1",
|
||||
timeout=3600
|
||||
)
|
||||
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "image_url",
|
||||
"image_url": {
|
||||
"url": "https://ofasys-multimodal-wlcb-3-toshanghai.oss-accelerate.aliyuncs.com/wpf272043/keepme/image/receipt.png"
|
||||
}
|
||||
},
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Read all the text in the image."
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
|
||||
start = time.time()
|
||||
response = client.chat.completions.create(
|
||||
model="stepfun-ai/Step3-VL-10B",
|
||||
messages=messages,
|
||||
max_tokens=2048,
|
||||
extra_body={"top_k": -1}
|
||||
)
|
||||
print(f"Response costs: {time.time() - start:.2f}s")
|
||||
print(f"Generated text: {response.choices[0].message.content}")
|
||||
```
|
||||
|
||||
**Example output:**
|
||||
|
||||
```text Output
|
||||
Response costs: 5.89s
|
||||
Generated text: Auntie Anne's
|
||||
|
||||
CINNAMON SUGAR
|
||||
1 × 17,000 17,000
|
||||
|
||||
SUB TOTAL 17,000
|
||||
|
||||
GRAND TOTAL 17,000
|
||||
|
||||
CASH IDR 20,000
|
||||
|
||||
CHANGE DUE 3,000
|
||||
```
|
||||
|
||||
**Multi-Image Input Example:**
|
||||
|
||||
Step3-VL-10B can process multiple images in a single request for comparison or analysis:
|
||||
|
||||
```python Example
|
||||
import time
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(
|
||||
api_key="EMPTY",
|
||||
base_url="http://localhost:30000/v1",
|
||||
timeout=3600
|
||||
)
|
||||
|
||||
messages = [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{
|
||||
"type": "image_url",
|
||||
"image_url": {
|
||||
"url": "https://www.civitatis.com/f/china/hong-kong/guia/taxi.jpg"
|
||||
}
|
||||
},
|
||||
{
|
||||
"type": "image_url",
|
||||
"image_url": {
|
||||
"url": "https://cdn.cheapoguides.com/wp-content/uploads/sites/7/2025/05/GettyImages-509614603-1280x600.jpg"
|
||||
}
|
||||
},
|
||||
{
|
||||
"type": "text",
|
||||
"text": "Compare these two images and describe the differences in 100 words or less."
|
||||
}
|
||||
]
|
||||
}
|
||||
]
|
||||
|
||||
start = time.time()
|
||||
response = client.chat.completions.create(
|
||||
model="stepfun-ai/Step3-VL-10B",
|
||||
messages=messages,
|
||||
max_tokens=2048,
|
||||
extra_body={"top_k": -1}
|
||||
)
|
||||
print(f"Response costs: {time.time() - start:.2f}s")
|
||||
print(f"Generated text: {response.choices[0].message.content}")
|
||||
```
|
||||
|
||||
**Example Output:**
|
||||
|
||||
```text Output
|
||||
Response costs: 3.24s
|
||||
Generated text: First image: Single red Hong Kong taxi close - up, clear license plate (RX 5004), “4 SEATS” sticker, urban street with shops behind. Second image: Aerial view of many taxis (red, green) on a highway with a viaduct, some hoods open, dense arrangement. Differences: Scale (single vs many), perspective (close - up vs aerial), context (street shops vs highway), and taxi conditions (normal vs some open hoods).
|
||||
```
|
||||
|
||||
#### 4.2.2 Reasoning Parser
|
||||
|
||||
Step3-VL-10B supports reasoning mode. Enable the reasoning parser during deployment to separate the thinking and content sections:
|
||||
|
||||
```shell Command
|
||||
python -m sglang.launch_server \
|
||||
--model stepfun-ai/Step3-VL-10B \
|
||||
--reasoning-parser deepseek-r1 \
|
||||
--host 0.0.0.0 \
|
||||
--port 30000 \
|
||||
--trust-remote-code
|
||||
```
|
||||
|
||||
**Streaming with Thinking Process:**
|
||||
|
||||
```python Example
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(
|
||||
base_url="http://localhost:30000/v1",
|
||||
api_key="EMPTY"
|
||||
)
|
||||
|
||||
# Enable streaming to see the thinking process in real-time
|
||||
response = client.chat.completions.create(
|
||||
model="stepfun-ai/Step3-VL-10B",
|
||||
messages=[
|
||||
{"role": "user", "content": "Solve this problem step by step: What is 15% of 240?"}
|
||||
],
|
||||
temperature=0.7,
|
||||
max_tokens=2048,
|
||||
stream=True,
|
||||
extra_body={"top_k": -1}
|
||||
)
|
||||
|
||||
# Process the stream
|
||||
has_thinking = False
|
||||
has_answer = False
|
||||
thinking_started = False
|
||||
|
||||
for chunk in response:
|
||||
if chunk.choices and len(chunk.choices) > 0:
|
||||
delta = chunk.choices[0].delta
|
||||
|
||||
# Print thinking process
|
||||
if hasattr(delta, 'reasoning_content') and delta.reasoning_content:
|
||||
if not thinking_started:
|
||||
print("=============== Thinking =================", flush=True)
|
||||
thinking_started = True
|
||||
has_thinking = True
|
||||
print(delta.reasoning_content, end="", flush=True)
|
||||
|
||||
# Print answer content
|
||||
if delta.content:
|
||||
# Close thinking section and add content header
|
||||
if has_thinking and not has_answer:
|
||||
print("\n=============== Content =================", flush=True)
|
||||
has_answer = True
|
||||
print(delta.content, end="", flush=True)
|
||||
|
||||
print()
|
||||
```
|
||||
|
||||
**Example Output:**
|
||||
```text Output
|
||||
=============== Thinking =================
|
||||
Okay, let's see. The problem is asking for 15% of 240. Hmm, I need to remember how to calculate percentages. So, percentage means "per hundred," right? So, 15% is the same as 15 per 100, or 15/100. To find a percentage of a number, I think you convert the percentage to a decimal and then multiply it by the number. Let me check that.
|
||||
|
||||
First, 15% as a decimal. To convert a percentage to a decimal, you divide by 100. So 15 divided by 100 is 0.15. Yeah, that's right. So 15% is 0.15 in decimal form. Then, to find 15% of 24
|
||||
0, I need to multiply 0.15 by 240. Let me do that calculation.
|
||||
|
||||
Let me write it out: 0.15 * 240. Let's compute that. Maybe break it down. 0.1 is 10%, and 0.05 is 5%, so 10% of 240 is 24, and 5% of 240 is 12. Then 10% + 5% is 15%, so 24 + 12 is 36. Oh,
|
||||
that's a good way to check. So 15% is 10% plus 5%, which adds up to 36. Let me verify with the decimal method. 0.15 * 240. Let's multiply 240 by 0.1 first: 24. Then 240 by 0.05: 12. Adding them gives 36. Yep, that matches. Alternatively, 240 * 15 = 3600, and then divide by 100 (since it's per hundred), so 3600 / 100 = 36. That's another way. So 15% of 240 is 36. Let me make sure I didn't make a mistake. Let's check with another method. 10% of 240 is 24, 20% would be 48, so 15% is halfway between 10% and 20%, which is (24 + 48)/2 = 36. Yep, that works too
|
||||
. So all methods point to 36. I think that's the answer.
|
||||
|
||||
=============== Content =================
|
||||
|
||||
To solve the problem "What is 15% of 240?" step by step:
|
||||
|
||||
---
|
||||
|
||||
### **Step 1: Understand the concept of percentage**
|
||||
A percentage represents a portion of a whole. Specifically, "percent" means "per hundred." So, 15% means **15 out of 100**, or **15/100**.
|
||||
|
||||
---
|
||||
|
||||
### **Step 2: Convert the percentage to a decimal**
|
||||
To use percentages in calculations, convert them to decimals by dividing by 100:
|
||||
$$
|
||||
15\% = \frac{15}{100} = 0.15
|
||||
$$
|
||||
|
||||
---
|
||||
|
||||
### **Step 3: Multiply the decimal by the given number**
|
||||
Now, multiply 0.15 (the decimal form of 15%) by 240:
|
||||
$$
|
||||
0.15 \times 240 = 36
|
||||
$$
|
||||
|
||||
---
|
||||
|
||||
### **Alternative Verification Methods**
|
||||
|
||||
#### **Method A: Break into parts**
|
||||
- 10% of 240 = $ 0.10 \times 240 = 24 $
|
||||
- 5% of 240 = $ 0.05 \times 240 = 12 $
|
||||
- Add them: $ 24 + 12 = 36 $
|
||||
|
||||
#### **Method B: Use direct multiplication**
|
||||
- $ 15\% \text{ of } 240 = \frac{15}{100} \times 240 = \frac{3600}{100} = 36 $
|
||||
|
||||
#### **Method C: Estimate using known percentages**
|
||||
- 20% of 240 = $ 0.20 \times 240 = 48 $
|
||||
- 10% of 240 = $ 0.10 \times 240 = 24 $
|
||||
- 15% is halfway between 10% and 20%: $ \frac{24 + 48}{2} = 36 $
|
||||
|
||||
---
|
||||
|
||||
### **Final Answer**
|
||||
$$
|
||||
\boxed{36}
|
||||
$$
|
||||
```
|
||||
|
||||
**Note:** The reasoning parser captures the model's step-by-step thinking process, allowing you to see how the model arrives at its conclusions.
|
||||
|
||||
#### 4.2.3 Tool Calling
|
||||
|
||||
Step3-VL-10B supports tool calling capabilities. Enable the tool call parser:
|
||||
|
||||
```shell Command
|
||||
python -m sglang.launch_server \
|
||||
--model stepfun-ai/Step3-VL-10B \
|
||||
--reasoning-parser deepseek-r1 \
|
||||
--tool-call-parser hermes \
|
||||
--host 0.0.0.0 \
|
||||
--port 30000 \
|
||||
--trust-remote-code
|
||||
```
|
||||
|
||||
**Python Example (with Thinking Process):**
|
||||
|
||||
```python Example
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(
|
||||
base_url="http://localhost:30000/v1",
|
||||
api_key="EMPTY"
|
||||
)
|
||||
|
||||
# Define available tools
|
||||
tools = [
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "get_weather",
|
||||
"description": "Get the current weather for a location",
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"location": {
|
||||
"type": "string",
|
||||
"description": "The city name"
|
||||
},
|
||||
"unit": {
|
||||
"type": "string",
|
||||
"enum": ["celsius", "fahrenheit"],
|
||||
"description": "Temperature unit"
|
||||
}
|
||||
},
|
||||
"required": ["location"]
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
|
||||
# Make request with streaming to see thinking process
|
||||
response = client.chat.completions.create(
|
||||
model="stepfun-ai/Step3-VL-10B",
|
||||
messages=[
|
||||
{"role": "user", "content": "What's the weather in Beijing?"}
|
||||
],
|
||||
tools=tools,
|
||||
temperature=0.7,
|
||||
stream=True,
|
||||
extra_body={"top_k": -1}
|
||||
)
|
||||
|
||||
# Process streaming response
|
||||
thinking_started = False
|
||||
has_thinking = False
|
||||
tool_calls_accumulator = {}
|
||||
|
||||
for chunk in response:
|
||||
if chunk.choices and len(chunk.choices) > 0:
|
||||
delta = chunk.choices[0].delta
|
||||
|
||||
# Print thinking process
|
||||
if hasattr(delta, 'reasoning_content') and delta.reasoning_content:
|
||||
if not thinking_started:
|
||||
print("=============== Thinking =================", flush=True)
|
||||
thinking_started = True
|
||||
has_thinking = True
|
||||
print(delta.reasoning_content, end="", flush=True)
|
||||
|
||||
# Accumulate tool calls
|
||||
if hasattr(delta, 'tool_calls') and delta.tool_calls:
|
||||
# Close thinking section if needed
|
||||
if has_thinking and thinking_started:
|
||||
print("\n=============== Content =================\n", flush=True)
|
||||
thinking_started = False
|
||||
|
||||
for tool_call in delta.tool_calls:
|
||||
index = tool_call.index
|
||||
if index not in tool_calls_accumulator:
|
||||
tool_calls_accumulator[index] = {
|
||||
'name': None,
|
||||
'arguments': ''
|
||||
}
|
||||
|
||||
if tool_call.function:
|
||||
if tool_call.function.name:
|
||||
tool_calls_accumulator[index]['name'] = tool_call.function.name
|
||||
if tool_call.function.arguments:
|
||||
tool_calls_accumulator[index]['arguments'] += tool_call.function.arguments
|
||||
|
||||
# Print content
|
||||
if delta.content:
|
||||
print(delta.content, end="", flush=True)
|
||||
|
||||
# Print accumulated tool calls
|
||||
for index, tool_call in sorted(tool_calls_accumulator.items()):
|
||||
print(f"Tool Call: {tool_call['name']}")
|
||||
print(f" Arguments: {tool_call['arguments']}")
|
||||
|
||||
print()
|
||||
```
|
||||
|
||||
**Example Output:**
|
||||
```text Output
|
||||
=============== Thinking =================
|
||||
The user is asking about the weather in Beijing. I have a function called "get_weather" that can provide weather information for a location. Let me check the parameters:
|
||||
|
||||
- location: required (string) - "Beijing"
|
||||
- unit: optional (string, enum: ["celsius", "fahrenheit"]) - not specified by the user, so I won't include it
|
||||
|
||||
I should call the function with location="Beijing".
|
||||
|
||||
<tool_calls>
|
||||
|
||||
=============== Content =================
|
||||
|
||||
</tool_calls>Tool Call: get_weather
|
||||
Arguments: {"location": "Beijing"}
|
||||
```
|
||||
|
||||
**Handling Tool Call Results:**
|
||||
|
||||
```python Example
|
||||
# After getting the tool call, execute the function
|
||||
def get_weather(location, unit="celsius"):
|
||||
# Your actual weather API call here
|
||||
return f"The weather in {location} is 22°{unit[0].upper()} and sunny."
|
||||
|
||||
# Send tool result back to the model
|
||||
messages = [
|
||||
{"role": "user", "content": "What's the weather in Beijing?"},
|
||||
{
|
||||
"role": "assistant",
|
||||
"content": None,
|
||||
"tool_calls": [{
|
||||
"id": "call_123",
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "get_weather",
|
||||
"arguments": '{"location": "Beijing", "unit": "celsius"}'
|
||||
}
|
||||
}]
|
||||
},
|
||||
{
|
||||
"role": "tool",
|
||||
"tool_call_id": "call_123",
|
||||
"content": get_weather("Beijing", "celsius")
|
||||
}
|
||||
]
|
||||
|
||||
final_response = client.chat.completions.create(
|
||||
model="stepfun-ai/Step3-VL-10B",
|
||||
messages=messages,
|
||||
temperature=0.7,
|
||||
extra_body={"top_k": -1}
|
||||
)
|
||||
|
||||
print(final_response.choices[0].message.content)
|
||||
```
|
||||
|
||||
**Note:**
|
||||
|
||||
- The reasoning parser shows how the model decides to use a tool
|
||||
- Tool calls are clearly marked with the function name and arguments
|
||||
- You can then execute the function and send the result back to continue the conversation
|
||||
|
||||
## 5. Benchmark
|
||||
|
||||
### 5.1 Speed Benchmark
|
||||
|
||||
**Test Environment:**
|
||||
|
||||
- Hardware: NVIDIA B200 GPU (1x)
|
||||
- Model: stepfun-ai/Step3-VL-10B
|
||||
- Tensor Parallelism: 1
|
||||
- sglang version: 0.5.8+
|
||||
|
||||
We use SGLang's built-in benchmarking tool to conduct performance evaluation with random images.
|
||||
|
||||
#### 5.1.1 Latency-Sensitive Benchmark
|
||||
|
||||
- Model Deployment Command:
|
||||
|
||||
```shell Command
|
||||
python -m sglang.launch_server \
|
||||
--model stepfun-ai/Step3-VL-10B \
|
||||
--host 0.0.0.0 \
|
||||
--port 30000 \
|
||||
--trust-remote-code
|
||||
```
|
||||
|
||||
- Benchmark Command:
|
||||
|
||||
```shell Command
|
||||
python3 -m sglang.bench_serving \
|
||||
--backend sglang-oai-chat \
|
||||
--host 127.0.0.1 \
|
||||
--port 30000 \
|
||||
--model stepfun-ai/Step3-VL-10B \
|
||||
--dataset-name image \
|
||||
--image-count 2 \
|
||||
--image-resolution 720p \
|
||||
--random-input-len 128 \
|
||||
--random-output-len 1024 \
|
||||
--num-prompts 10 \
|
||||
--max-concurrency 1
|
||||
```
|
||||
|
||||
- Result:
|
||||
|
||||
```text Output
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang-oai-chat
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 1
|
||||
Successful requests: 10
|
||||
Benchmark duration (s): 30.85
|
||||
Total input tokens: 14120
|
||||
Total input text tokens: 720
|
||||
Total input vision tokens: 13400
|
||||
Total generated tokens: 4220
|
||||
Total generated tokens (retokenized): 4217
|
||||
Request throughput (req/s): 0.32
|
||||
Input token throughput (tok/s): 457.71
|
||||
Output token throughput (tok/s): 136.79
|
||||
Peak output token throughput (tok/s): 240.00
|
||||
Peak concurrent requests: 2
|
||||
Total token throughput (tok/s): 594.50
|
||||
Concurrency: 1.00
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 3083.40
|
||||
Median E2E Latency (ms): 2747.00
|
||||
P90 E2E Latency (ms): 4574.50
|
||||
P99 E2E Latency (ms): 5462.49
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 1327.69
|
||||
Median TTFT (ms): 1341.01
|
||||
P99 TTFT (ms): 1486.11
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 4.16
|
||||
Median TPOT (ms): 4.17
|
||||
P99 TPOT (ms): 4.18
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 4.17
|
||||
Median ITL (ms): 4.18
|
||||
P95 ITL (ms): 4.30
|
||||
P99 ITL (ms): 4.38
|
||||
Max ITL (ms): 8.24
|
||||
==================================================
|
||||
```
|
||||
|
||||
#### 5.1.2 Throughput-Sensitive Benchmark
|
||||
|
||||
- Benchmark Command:
|
||||
|
||||
```shell Command
|
||||
python3 -m sglang.bench_serving \
|
||||
--backend sglang-oai-chat \
|
||||
--host 127.0.0.1 \
|
||||
--port 30000 \
|
||||
--model stepfun-ai/Step3-VL-10B \
|
||||
--dataset-name image \
|
||||
--image-count 2 \
|
||||
--image-resolution 720p \
|
||||
--random-input-len 128 \
|
||||
--random-output-len 1024 \
|
||||
--num-prompts 1000 \
|
||||
--max-concurrency 100
|
||||
```
|
||||
|
||||
- Result:
|
||||
|
||||
```text Output
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang-oai-chat
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 100
|
||||
Successful requests: 1000
|
||||
Benchmark duration (s): 976.52
|
||||
Total input tokens: 1416949
|
||||
Total input text tokens: 76949
|
||||
Total input vision tokens: 1340000
|
||||
Total generated tokens: 510855
|
||||
Total generated tokens (retokenized): 510526
|
||||
Request throughput (req/s): 1.02
|
||||
Input token throughput (tok/s): 1451.02
|
||||
Output token throughput (tok/s): 523.14
|
||||
Peak output token throughput (tok/s): 20429.00
|
||||
Peak concurrent requests: 103
|
||||
Total token throughput (tok/s): 1974.16
|
||||
Concurrency: 99.81
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 97463.22
|
||||
Median E2E Latency (ms): 91872.75
|
||||
P90 E2E Latency (ms): 118553.42
|
||||
P99 E2E Latency (ms): 198445.56
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 94379.07
|
||||
Median TTFT (ms): 87163.09
|
||||
P99 TTFT (ms): 194871.41
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 5.89
|
||||
Median TPOT (ms): 5.72
|
||||
P99 TPOT (ms): 23.58
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 6.05
|
||||
Median ITL (ms): 0.13
|
||||
P95 ITL (ms): 0.56
|
||||
P99 ITL (ms): 3.99
|
||||
Max ITL (ms): 97551.06
|
||||
==================================================
|
||||
```
|
||||
|
||||
### 5.2 Accuracy Benchmark
|
||||
|
||||
#### 5.2.1 MMMU Benchmark
|
||||
|
||||
You can evaluate the model's accuracy using the MMMU dataset:
|
||||
|
||||
- Model Deployment Command:
|
||||
|
||||
```shell Command
|
||||
python -m sglang.launch_server \
|
||||
--model stepfun-ai/Step3-VL-10B \
|
||||
--host 0.0.0.0 \
|
||||
--port 30000 \
|
||||
--trust-remote-code
|
||||
```
|
||||
|
||||
- Benchmark Command:
|
||||
|
||||
```shell Command
|
||||
python3 benchmark/mmmu/bench_sglang.py \
|
||||
--port 30000 \
|
||||
--concurrency 64
|
||||
```
|
||||
|
||||
- Result:
|
||||
|
||||
```text Output
|
||||
Benchmark time: 934.6179109360091
|
||||
answers saved to: ./answer_sglang.json
|
||||
Evaluating...
|
||||
answers saved to: ./answer_sglang.json
|
||||
{'Accounting': {'acc': 0.667, 'num': 30},
|
||||
'Agriculture': {'acc': 0.367, 'num': 30},
|
||||
'Architecture_and_Engineering': {'acc': 0.4, 'num': 30},
|
||||
'Art': {'acc': 0.467, 'num': 30},
|
||||
'Art_Theory': {'acc': 0.5, 'num': 30},
|
||||
'Basic_Medical_Science': {'acc': 0.367, 'num': 30},
|
||||
'Biology': {'acc': 0.3, 'num': 30},
|
||||
'Chemistry': {'acc': 0.467, 'num': 30},
|
||||
'Clinical_Medicine': {'acc': 0.567, 'num': 30},
|
||||
'Computer_Science': {'acc': 0.467, 'num': 30},
|
||||
'Design': {'acc': 0.567, 'num': 30},
|
||||
'Diagnostics_and_Laboratory_Medicine': {'acc': 0.3, 'num': 30},
|
||||
'Economics': {'acc': 0.6, 'num': 30},
|
||||
'Electronics': {'acc': 0.567, 'num': 30},
|
||||
'Energy_and_Power': {'acc': 0.633, 'num': 30},
|
||||
'Finance': {'acc': 0.733, 'num': 30},
|
||||
'Geography': {'acc': 0.333, 'num': 30},
|
||||
'History': {'acc': 0.533, 'num': 30},
|
||||
'Literature': {'acc': 0.533, 'num': 30},
|
||||
'Manage': {'acc': 0.6, 'num': 30},
|
||||
'Marketing': {'acc': 0.767, 'num': 30},
|
||||
'Materials': {'acc': 0.6, 'num': 30},
|
||||
'Math': {'acc': 0.7, 'num': 30},
|
||||
'Mechanical_Engineering': {'acc': 0.333, 'num': 30},
|
||||
'Music': {'acc': 0.4, 'num': 30},
|
||||
'Overall': {'acc': 0.523, 'num': 900},
|
||||
'Overall-Art and Design': {'acc': 0.483, 'num': 120},
|
||||
'Overall-Business': {'acc': 0.673, 'num': 150},
|
||||
'Overall-Health and Medicine': {'acc': 0.513, 'num': 150},
|
||||
'Overall-Humanities and Social Science': {'acc': 0.492, 'num': 120},
|
||||
'Overall-Science': {'acc': 0.5, 'num': 150},
|
||||
'Overall-Tech and Engineering': {'acc': 0.481, 'num': 210},
|
||||
'Pharmacy': {'acc': 0.6, 'num': 30},
|
||||
'Physics': {'acc': 0.7, 'num': 30},
|
||||
'Psychology': {'acc': 0.467, 'num': 30},
|
||||
'Public_Health': {'acc': 0.733, 'num': 30},
|
||||
'Sociology': {'acc': 0.433, 'num': 30}}
|
||||
eval out saved to ./val_sglang.json
|
||||
Overall accuracy: 0.523
|
||||
```
|
||||
@@ -0,0 +1,529 @@
|
||||
---
|
||||
title: Step-3.5
|
||||
metatags:
|
||||
description: "Deploy Step-3.5 reasoning engine with SGLang. "
|
||||
tag: NEW
|
||||
---
|
||||
|
||||
import { Step35Deployment } from '/src/snippets/autoregressive/step-35-deployment.jsx';
|
||||
|
||||
## 1. Model Introduction
|
||||
|
||||
[Step-3.5-Flash](https://huggingface.co/stepfun-ai/Step-3.5-Flash) is StepFun's production-grade reasoning engine built to decouple elite intelligence from heavy compute, and cuts attention cost for low-latency, cost-effective long-context inference—purpose-built for autonomous agents in real-world workflows. The model is available in multiple quantization formats optimized for different hardware platforms.
|
||||
|
||||
This generation delivers comprehensive upgrades across the board:
|
||||
- **Hybrid Attention Architecture**: Interleaves Sliding Window Attention (SWA) and Global Attention (GA) with a 3:1 ratio and an aggressive 128-token window. This hybrid approach ensures consistent performance across massive datasets or long codebases while significantly reducing the computational overhead typical of standard long-context models.
|
||||
- **Sparse Mixture-of-Experts**: Only 11B active parameters out of 196B parameters.
|
||||
- **Multi-Layer Multi-Token Prediction (MTP)**: Equipped with a 3-way Multi-Token Prediction (MTP-3). This allows for complex, multi-step reasoning chains with immediate responsiveness.
|
||||
|
||||
## 2.SGLang Installation
|
||||
|
||||
Step-3.5-Flash is currently available in SGLang via Docker image install.
|
||||
|
||||
### Docker (NVIDIA)
|
||||
```bash Command
|
||||
# Pull the docker image
|
||||
docker pull lmsysorg/sglang:dev-pr-18084
|
||||
|
||||
# Launch the container
|
||||
docker run -it --gpus all \
|
||||
--shm-size=32g \
|
||||
--ipc=host \
|
||||
--network=host \
|
||||
lmsysorg/sglang:dev-pr-18084 bash
|
||||
```
|
||||
|
||||
### Docker (AMD ROCm)
|
||||
```bash Command
|
||||
# For MI300X/MI325X
|
||||
docker pull lmsysorg/sglang:v0.5.9-rocm700-mi30x
|
||||
|
||||
# For MI350X/MI355X
|
||||
docker pull lmsysorg/sglang:v0.5.9-rocm700-mi35x
|
||||
|
||||
docker run -it \
|
||||
--device=/dev/kfd --device=/dev/dri \
|
||||
--shm-size=32g \
|
||||
--ipc=host \
|
||||
--network=host \
|
||||
--group-add video --cap-add=SYS_PTRACE \
|
||||
--security-opt seccomp=unconfined \
|
||||
lmsysorg/sglang:v0.5.9-rocm700-mi30x bash # or mi35x for MI350X/MI355X
|
||||
```
|
||||
|
||||
## 3.Model Deployment
|
||||
|
||||
This section provides deployment configurations optimized for different hardware platforms and use cases.
|
||||
|
||||
### 3.1 Basic Configuration
|
||||
|
||||
The Step-3.5-Flash series comes in only one sizes. Recommended starting configurations vary depending on hardware.
|
||||
|
||||
**Interactive Command Generator**: Use the configuration selector below to automatically generate the appropriate deployment command for your hardware platform, model size, quantization method, and thinking capabilities.
|
||||
|
||||
<Step35Deployment />
|
||||
|
||||
### 3.2 Configuration Tips
|
||||
|
||||
- **Memory**: Requires GPUs with high VRAM capacity. Supported platforms: H200 (4×, TP=4), MI300X/MI325X/MI350X/MI355X (4×, TP=4 EP=4).
|
||||
- **AMD Docker Image**: Use `lmsysorg/sglang:v0.5.9-rocm700-mi30x` for MI300X/MI325X and `lmsysorg/sglang:v0.5.9-rocm700-mi35x` for MI350X/MI355X.
|
||||
- **AMD Expert Parallelism Required**: On AMD GPUs, always use `--ep 4` with `--tp 4`. Both BF16 and FP8 models require expert parallelism. Without EP, the MoE intermediate dimension is split across GPUs (N=320), which triggers an AITER CK GEMM incompatibility. With EP=4, each GPU handles 72 full experts (N=1280), which works correctly with cuda graph enabled.
|
||||
- **AITER JIT Compilation**: First inference on AMD may take 30-40 seconds for AITER kernel JIT compilation. Subsequent requests use cached kernels.
|
||||
|
||||
## 4.Model Invocation
|
||||
|
||||
### 4.1 Basic Usage
|
||||
|
||||
For basic API usage and request examples, please refer to:
|
||||
|
||||
- [SGLang Basic Usage Guide](../../../docs/basic_usage/send_request)
|
||||
|
||||
### 4.2 Advanced Usage
|
||||
|
||||
#### 4.2.1 Reasoning Parser
|
||||
|
||||
Step-3.5-Flash only supports reasoning mode. Enable the reasoning parser during deployment to separate the thinking and content sections:
|
||||
|
||||
```shell Command
|
||||
sglang serve \
|
||||
--model-path stepfun-ai/Step-3.5-Flash \
|
||||
--tp 4 \
|
||||
--ep 4 \
|
||||
--reasoning-parser step3p5
|
||||
```
|
||||
|
||||
```python Example
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(
|
||||
base_url="http://localhost:30000/v1",
|
||||
api_key="EMPTY"
|
||||
)
|
||||
|
||||
# Enable streaming to see the thinking process in real-time
|
||||
response = client.chat.completions.create(
|
||||
model="stepfun-ai/Step-3.5-Flash",
|
||||
messages=[
|
||||
{"role": "user", "content": "Solve this problem step by step: What is 15% of 240?"}
|
||||
],
|
||||
temperature=0.7,
|
||||
max_tokens=2048,
|
||||
stream=True
|
||||
)
|
||||
|
||||
# Process the stream
|
||||
has_thinking = False
|
||||
has_answer = False
|
||||
thinking_started = False
|
||||
|
||||
for chunk in response:
|
||||
if chunk.choices and len(chunk.choices) > 0:
|
||||
delta = chunk.choices[0].delta
|
||||
|
||||
# Print thinking process
|
||||
if hasattr(delta, 'reasoning_content') and delta.reasoning_content:
|
||||
if not thinking_started:
|
||||
print("=============== Thinking =================", flush=True)
|
||||
thinking_started = True
|
||||
has_thinking = True
|
||||
print(delta.reasoning_content, end="", flush=True)
|
||||
|
||||
# Print answer content
|
||||
if delta.content:
|
||||
# Close thinking section and add content header
|
||||
if has_thinking and not has_answer:
|
||||
print("\n=============== Content =================", flush=True)
|
||||
has_answer = True
|
||||
print(delta.content, end="", flush=True)
|
||||
|
||||
print()
|
||||
```
|
||||
|
||||
**Output Example:**
|
||||
|
||||
```text Output
|
||||
=============== Thinking =================
|
||||
We are asked: "What is 15% of 240?" We need to solve step by step.
|
||||
|
||||
Step 1: Understand that "15% of 240" means we need to calculate 15 percent of 240. In mathematical terms, it is (15/100) * 240.
|
||||
|
||||
Step 2: Simplify the calculation. We can compute 15% of 240 by first finding 10% of 240 and then 5% of 240, and adding them. Alternatively, we can multiply directly.
|
||||
|
||||
Method 1:
|
||||
10% of 240 = 240 * 0.10 = 24.
|
||||
5% is half of 10%, so 5% of 240 = 24 / 2 = 12.
|
||||
Then 15% = 10% + 5% = 24 + 12 = 36.
|
||||
|
||||
Method 2: Direct multiplication: 15% = 15/100 = 0.15, so 0.15 * 240 = 36.
|
||||
|
||||
We can also compute fractionally: (15/100)*240 = (15*240)/100. 15*240 = 3600, divided by 100 gives 36.
|
||||
|
||||
Thus, the answer is 36.
|
||||
|
||||
We'll present the solution step by step.
|
||||
|
||||
=============== Content =================
|
||||
|
||||
To find 15% of 240, follow these steps:
|
||||
|
||||
1. **Convert the percentage to a decimal**:
|
||||
\( 15\% = \frac{15}{100} = 0.15 \)
|
||||
|
||||
2. **Multiply by the number**:
|
||||
\( 0.15 \times 240 = 36 \)
|
||||
|
||||
Alternatively, break it down:
|
||||
- \( 10\% \text{ of } 240 = 240 \times 0.10 = 24 \)
|
||||
- \( 5\% \text{ of } 240 = \frac{24}{2} = 12 \) (since 5% is half of 10%)
|
||||
- \( 15\% = 10\% + 5\% = 24 + 12 = 36 \)
|
||||
|
||||
**Answer:** 36
|
||||
```
|
||||
|
||||
#### 4.2.2 Tool Calling
|
||||
|
||||
Step-3.5 supports tool calling capabilities. Enable the tool call parser:
|
||||
|
||||
**Python Example:**
|
||||
|
||||
Start sglang server:
|
||||
|
||||
```shell Command
|
||||
sglang serve \
|
||||
--model-path stepfun-ai/Step-3.5-Flash \
|
||||
--tp 4 \
|
||||
--ep 4 \
|
||||
--reasoning-parser step3p5 \
|
||||
--tool-call-parser step3p5
|
||||
```
|
||||
|
||||
```python Example
|
||||
from openai import OpenAI
|
||||
import json
|
||||
|
||||
client = OpenAI(
|
||||
base_url="http://localhost:30000/v1",
|
||||
api_key="EMPTY"
|
||||
)
|
||||
|
||||
# 1. define tools
|
||||
tools = [
|
||||
{
|
||||
"type": "function",
|
||||
"function": {
|
||||
"name": "get_weather",
|
||||
"description": "Get the current weather for a location",
|
||||
"parameters": {
|
||||
"type": "object",
|
||||
"properties": {
|
||||
"location": {"type": "string", "description": "The city name"},
|
||||
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"], "description": "Temperature unit"}
|
||||
},
|
||||
"required": ["location"]
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
|
||||
# 2. tool run
|
||||
def get_weather(location, unit="celsius"):
|
||||
return f"The weather in {location} is 22°{unit[0].upper()} and sunny."
|
||||
|
||||
# 3. send first request
|
||||
print("--- Sending first request ---")
|
||||
response = client.chat.completions.create(
|
||||
model="stepfun-ai/Step-3.5-Flash",
|
||||
messages=[
|
||||
{"role": "user", "content": "What's the weather in Beijing?"}
|
||||
],
|
||||
tools=tools,
|
||||
temperature=1.0,
|
||||
stream=False
|
||||
)
|
||||
|
||||
message = response.choices[0].message
|
||||
|
||||
# 4. Handle Reasoning Content
|
||||
reasoning = getattr(message, 'reasoning_content', None)
|
||||
if reasoning:
|
||||
print("=============== Thinking =================")
|
||||
print(reasoning)
|
||||
print("==========================================")
|
||||
|
||||
# 5. Handle Tool Calls
|
||||
if message.tool_calls:
|
||||
print("\n🔧 Tool Calls detected:")
|
||||
history_messages = [
|
||||
{"role": "user", "content": "What's the weather in Beijing?"},
|
||||
message
|
||||
]
|
||||
|
||||
for tool_call in message.tool_calls:
|
||||
print(f" Tool: {tool_call.function.name}")
|
||||
print(f" Args: {tool_call.function.arguments}")
|
||||
|
||||
args = json.loads(tool_call.function.arguments)
|
||||
tool_result = get_weather(args.get("location"), args.get("unit", "celsius"))
|
||||
|
||||
history_messages.append({
|
||||
"role": "tool",
|
||||
"tool_call_id": tool_call.id,
|
||||
"content": tool_result
|
||||
})
|
||||
|
||||
print("\n--- Sending tool results ---")
|
||||
final_response = client.chat.completions.create(
|
||||
model="stepfun-ai/Step-3.5-Flash",
|
||||
messages=history_messages,
|
||||
temperature=1.0,
|
||||
stream=False
|
||||
)
|
||||
|
||||
print("=============== Final Content =================")
|
||||
print(final_response.choices[0].message.content)
|
||||
|
||||
else:
|
||||
if message.content:
|
||||
print("=============== Content =================")
|
||||
print(message.content)
|
||||
```
|
||||
|
||||
**Output Example:**
|
||||
|
||||
```text Output
|
||||
--- Sending first request ---
|
||||
=============== Thinking =================
|
||||
The user is asking for the weather in Beijing. I should use the get_weather function with location="Beijing". The unit parameter is optional and the user didn't specify a preference, so I'll leave it out (the default should be fine).
|
||||
|
||||
==========================================
|
||||
|
||||
🔧 Tool Calls detected:
|
||||
Tool: get_weather
|
||||
Args: {"location": "Beijing"}
|
||||
|
||||
--- Sending tool results ---
|
||||
=============== Final Content =================
|
||||
The weather in Beijing is 22°C and sunny.
|
||||
```
|
||||
|
||||
**Note:**
|
||||
|
||||
- The reasoning parser shows how the model decides to use a tool
|
||||
- Tool calls are clearly marked with the function name and arguments
|
||||
- You can then execute the function and send the result back to continue the conversation
|
||||
|
||||
## 5. Benchmark
|
||||
|
||||
### 5.1 Speed Benchmark
|
||||
|
||||
**Test Environment:**
|
||||
|
||||
- Hardware: NVIDIA H200 GPU (4x)
|
||||
- Model: Step-3.5-Flash
|
||||
- Tensor Parallelism: 4
|
||||
- Expert Parallelism: 4
|
||||
- sglang version: 0.5.8
|
||||
|
||||
We use SGLang's built-in benchmarking tool to conduct performance evaluation on the [ShareGPT_Vicuna_unfiltered](https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered) dataset. This dataset contains real conversation data and can better reflect performance in actual use scenarios.
|
||||
|
||||
#### 5.1.1 Standard Scenario Benchmark
|
||||
|
||||
- Model Deployment Command:
|
||||
|
||||
```shell Command
|
||||
sglang serve \
|
||||
--model-path stepfun-ai/Step-3.5-Flash \
|
||||
--tp 4 \
|
||||
--ep 4
|
||||
```
|
||||
|
||||
##### 5.1.1.1 Low Concurrency
|
||||
|
||||
- Benchmark Command:
|
||||
|
||||
```shell Command
|
||||
python3 -m sglang.bench_serving \
|
||||
--backend sglang \
|
||||
--model stepfun-ai/Step-3.5-Flash \
|
||||
--dataset-name random \
|
||||
--random-input-len 1000 \
|
||||
--random-output-len 1000 \
|
||||
--num-prompts 10 \
|
||||
--max-concurrency 1
|
||||
```
|
||||
|
||||
- Test Results:
|
||||
|
||||
```text Output
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 1
|
||||
Successful requests: 10
|
||||
Benchmark duration (s): 35.30
|
||||
Total input tokens: 6091
|
||||
Total input text tokens: 6091
|
||||
Total generated tokens: 4220
|
||||
Total generated tokens (retokenized): 4212
|
||||
Request throughput (req/s): 0.28
|
||||
Input token throughput (tok/s): 172.57
|
||||
Output token throughput (tok/s): 119.56
|
||||
Peak output token throughput (tok/s): 124.00
|
||||
Peak concurrent requests: 2
|
||||
Total token throughput (tok/s): 292.14
|
||||
Concurrency: 1.00
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 3527.94
|
||||
Median E2E Latency (ms): 2884.72
|
||||
P90 E2E Latency (ms): 6350.38
|
||||
P99 E2E Latency (ms): 7858.53
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 107.53
|
||||
Median TTFT (ms): 80.93
|
||||
P99 TTFT (ms): 269.52
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 8.12
|
||||
Median TPOT (ms): 8.13
|
||||
P99 TPOT (ms): 8.14
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 8.12
|
||||
Median ITL (ms): 8.11
|
||||
P95 ITL (ms): 8.61
|
||||
P99 ITL (ms): 8.91
|
||||
Max ITL (ms): 20.77
|
||||
==================================================
|
||||
```
|
||||
|
||||
##### 5.1.1.2 Medium Concurrency
|
||||
|
||||
- Benchmark Command:
|
||||
|
||||
```shell Command
|
||||
python3 -m sglang.bench_serving \
|
||||
--backend sglang \
|
||||
--model stepfun-ai/Step-3.5-Flash \
|
||||
--dataset-name random \
|
||||
--random-input-len 1000 \
|
||||
--random-output-len 1000 \
|
||||
--num-prompts 80 \
|
||||
--max-concurrency 16
|
||||
```
|
||||
|
||||
- Test Results:
|
||||
|
||||
```text Output
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 16
|
||||
Successful requests: 80
|
||||
Benchmark duration (s): 54.06
|
||||
Total input tokens: 39588
|
||||
Total input text tokens: 39588
|
||||
Total generated tokens: 40805
|
||||
Total generated tokens (retokenized): 40479
|
||||
Request throughput (req/s): 1.48
|
||||
Input token throughput (tok/s): 732.33
|
||||
Output token throughput (tok/s): 754.84
|
||||
Peak output token throughput (tok/s): 928.00
|
||||
Peak concurrent requests: 21
|
||||
Total token throughput (tok/s): 1487.17
|
||||
Concurrency: 14.06
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 9501.23
|
||||
Median E2E Latency (ms): 10010.71
|
||||
P90 E2E Latency (ms): 15655.09
|
||||
P99 E2E Latency (ms): 18803.63
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 198.34
|
||||
Median TTFT (ms): 89.50
|
||||
P99 TTFT (ms): 984.66
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 18.97
|
||||
Median TPOT (ms): 18.80
|
||||
P99 TPOT (ms): 35.67
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 18.27
|
||||
Median ITL (ms): 17.48
|
||||
P95 ITL (ms): 18.44
|
||||
P99 ITL (ms): 62.47
|
||||
Max ITL (ms): 460.85
|
||||
==================================================
|
||||
```
|
||||
|
||||
##### 5.1.1.3 High Concurrency
|
||||
|
||||
- Benchmark Command:
|
||||
|
||||
```shell Command
|
||||
python3 -m sglang.bench_serving \
|
||||
--backend sglang \
|
||||
--model stepfun-ai/Step-3.5-Flash \
|
||||
--dataset-name random \
|
||||
--random-input-len 1000 \
|
||||
--random-output-len 1000 \
|
||||
--num-prompts 500 \
|
||||
--max-concurrency 100
|
||||
```
|
||||
|
||||
- Test Results:
|
||||
|
||||
```text Output
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 100
|
||||
Successful requests: 500
|
||||
Benchmark duration (s): 125.88
|
||||
Total input tokens: 249331
|
||||
Total input text tokens: 249331
|
||||
Total generated tokens: 252662
|
||||
Total generated tokens (retokenized): 251323
|
||||
Request throughput (req/s): 3.97
|
||||
Input token throughput (tok/s): 1980.77
|
||||
Output token throughput (tok/s): 2007.23
|
||||
Peak output token throughput (tok/s): 2500.00
|
||||
Peak concurrent requests: 109
|
||||
Total token throughput (tok/s): 3987.99
|
||||
Concurrency: 92.25
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 23223.31
|
||||
Median E2E Latency (ms): 22631.90
|
||||
P90 E2E Latency (ms): 42269.38
|
||||
P99 E2E Latency (ms): 47637.53
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 372.13
|
||||
Median TTFT (ms): 127.26
|
||||
P99 TTFT (ms): 1880.42
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 46.06
|
||||
Median TPOT (ms): 47.61
|
||||
P99 TPOT (ms): 51.34
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 45.31
|
||||
Median ITL (ms): 39.86
|
||||
P95 ITL (ms): 72.49
|
||||
P99 ITL (ms): 117.05
|
||||
Max ITL (ms): 1359.81
|
||||
==================================================
|
||||
```
|
||||
|
||||
### 5.2 Accuracy Benchmark
|
||||
|
||||
#### 5.2.1 GSM8K Benchmark
|
||||
|
||||
- **Benchmark Command:**
|
||||
|
||||
```shell Command
|
||||
python3 -m sglang.test.few_shot_gsm8k --num-questions 200
|
||||
```
|
||||
|
||||
- **Results**:
|
||||
|
||||
- Step-3.5-Flash
|
||||
```
|
||||
Accuracy: 0.885
|
||||
Invalid: 0.005
|
||||
Latency: 9.986 s
|
||||
Output throughput: 1972.911 token/s
|
||||
```
|
||||
Reference in New Issue
Block a user