Files
sglang/docs/cookbook/autoregressive/Qwen/Qwen3-Coder.mdx
T
2026-08-03 16:51:00 -07:00

789 lines
28 KiB
Plaintext

---
title: Qwen3-Coder
metatags:
description: "Deploy Qwen3-Coder(480B, 30B) MoE coding model with SGLang on AMD MI300X (MI325X, MI355X)"
---
import { Qwen3CoderDeployment } from '/src/snippets/autoregressive/qwen3-coder-deployment.jsx';
## 1. Model Introduction
[Qwen3-Coder](https://huggingface.co/collections/Qwen/qwen3-coder) is the latest code-focused large language model series from the Qwen team. Built on the foundation of Qwen3, Qwen3-Coder delivers exceptional performance in code generation, understanding, and reasoning tasks.
**Key Features:**
- **State-of-the-art Coding Performance**: Achieves top-tier results on HumanEval, MBPP, LiveCodeBench, and other major coding benchmarks.
- **Tool Calling Support**: Native support for function calling and tool use, enabling seamless integration with external APIs and services.
- **Extended Context Length**: Supports up to 256K tokens for processing large codebases and long documents.
- **Multilingual Code Support**: Proficient in Python, JavaScript, TypeScript, Java, C++, Go, Rust, and many other programming languages.
- **MoE Architecture**: Efficient Mixture-of-Experts design for optimal performance-to-cost ratio.
- **ROCm Support**: Compatible with AMD MI300X, MI325X and MI355X GPUs via SGLang (verified).
- **NVIDIA GPU Support**: Compatible with NVIDIA GB200 and B200 GPUs via SGLang (verified).
For more details, please refer to the [official Qwen3-Coder GitHub Repository](https://github.com/QwenLM/Qwen3-Coder).
## 2. SGLang Installation
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
For SGLang CPU installation, please refer to the [CPU version installation guide](../../../docs/hardware-platforms/cpu_server#installation).
## 3. Model Deployment
This section provides deployment configurations verified on AMD MI300X, MI325X, MI355X, NVIDIA B200, GB200, and Intel Xeon CPU hardware platforms.
### 3.1 Configuration
**Interactive Command Generator**: Use the configuration selector below to automatically generate the appropriate deployment command for your hardware platform, model size, and quantization method.
<Qwen3CoderDeployment />
### 3.2 Configuration Tips
**AMD (MI300X/MI325X/MI355X):**
* **Memory Management**: We have verified successful deployment on MI300X/MI325X/MI355X with `--context-length 8192`. Larger context lengths may be supported but require additional memory.
* **Expert Parallelism**: For 480B-A35B with FP8 quantization, `--ep 2` is required to satisfy the dimension alignment requirement.
* **Page Size**: `--page-size 32` is recommended for MoE models to optimize memory usage.
* **Environment Variable**: If you encounter aiter-related issues, try setting `SGLANG_USE_AITER=0`.
**NVIDIA (B200/GB200):**
* **GB200 Parallelism**: Use `--tp 4 --ep 4` on GB200. B200 uses the default NVIDIA settings generated above.
* **NVFP4 Quantization**: Requires `--quantization modelopt_fp4` and uses a different model path (`nvidia/Qwen3-Coder-...`).
* **DP Attention**: NVFP4 configuration supports `--enable-dp-attention` for improved throughput.
**Intel Xeon CPU:**
* Please refer to the `Notes` part in the serving engine launching section in [the SGLang CPU server document](../../../docs/hardware-platforms/cpu_server#launch-of-the-serving-engine) to better understand how to configure the arguments, especially for TP (tensor parallel) and NUMA binding settings.
**General:**
* **Tool Use**: To enable tool calling capabilities, add `--tool-call-parser qwen3_coder` to the launch command.
## 4. Model Invocation
### 4.1 Basic Usage
For basic API usage and request examples, please refer to:
- [SGLang Basic Usage Guide](../../../docs/basic_usage/send_request)
### 4.2 Advanced Usage
#### 4.2.1 Code Generation Example
```python Example
from openai import OpenAI
client = OpenAI(
api_key="EMPTY",
base_url="http://localhost:30000/v1",
timeout=3600
)
messages = [
{
"role": "user",
"content": "Write a Python function that implements binary search on a sorted list. Include docstring and type hints."
}
]
response = client.chat.completions.create(
model="Qwen/Qwen3-Coder-480B-A35B-Instruct",
messages=messages,
max_tokens=2048,
temperature=0.7
)
print(response.choices[0].message.content)
```
**Example Output:**
````text Output
```python
from typing import List, Optional, TypeVar
T = TypeVar('T')
def binary_search(arr: List[T], target: T) -> Optional[int]:
"""
Perform binary search on a sorted list to find the index of a target element.
This function implements the binary search algorithm, which efficiently finds
a target value in a sorted array by repeatedly dividing the search interval
in half.
Args:
arr (List[T]): A sorted list of elements to search through.
target (T): The element to search for in the list.
Returns:
Optional[int]: The index of the target element if found, None otherwise.
Time Complexity:
O(log n) where n is the number of elements in the array.
Space Complexity:
O(1) - iterative implementation uses constant extra space.
Examples:
>>> binary_search([1, 2, 3, 4, 5], 3)
2
>>> binary_search([1, 2, 3, 4, 5], 6)
None
>>> binary_search(['a', 'b', 'c', 'd'], 'b')
1
>>> binary_search([], 1)
None
"""
if not arr:
return None
left: int = 0
right: int = len(arr) - 1
while left <= right:
mid: int = (left + right) // 2
if arr[mid] == target:
return mid
elif arr[mid] < target:
left = mid + 1
else:
right = mid - 1
return None
# Alternative recursive implementation
def binary_search_recursive(arr: List[T], target: T, left: int = 0, right: Optional[int] = None) -> Optional[int]:
"""
Perform binary search recursively on a sorted list to find the index of a target element.
Args:
arr (List[T]): A sorted list of elements to search through.
target (T): The element to search for in the list.
left (int): Left boundary of the search range (inclusive).
right (Optional[int]): Right boundary of the search range (inclusive).
Returns:
Optional[int]: The index of the target element if found, None otherwise.
Time Complexity:
O(log n) where n is the number of elements in the array.
Space Complexity:
O(log n) due to recursive call stack.
Examples:
>>> binary_search_recursive([1, 2, 3, 4, 5], 3)
2
>>> binary_search_recursive([1, 2, 3, 4, 5], 6)
None
"""
if not arr:
return None
if right is None:
right = len(arr) - 1
if left > right:
return None
mid: int = (left + right) // 2
if arr[mid] == target:
return mid
elif arr[mid] < target:
return binary_search_recursive(arr, target, mid + 1, right)
else:
return binary_search_recursive(arr, target, left, mid - 1)
```
This implementation provides:
1. **Main function** (`binary_search`): An iterative implementation that's more memory-efficient
2. **Alternative function** (`binary_search_recursive`): A recursive implementation for educational purposes
3. **Type hints**: Using generics (`TypeVar`) to work with any comparable type
4. **Comprehensive docstring**: Including description, parameters, return value, complexity analysis, and examples
5. **Edge case handling**: Empty lists, elements not found, etc.
6. **Clear variable names**: Self-documenting code
7. **Examples**: Doctest-style examples in the docstring
The function works with any sorted list of comparable elements (integers, strings, etc.) and returns the index of the target element if found, or `None` if not found.
````
#### 4.2.2 Tool Calling Example
Qwen3-Coder supports tool calling capabilities. Enable the tool call parser during deployment. The following example uses 30B-A3B model:
```shell Command
SGLANG_USE_AITER=0 python -m sglang.launch_server \
--model Qwen/Qwen3-Coder-30B-A3B-Instruct \
--tp 1 \
--context-length 8192 \
--page-size 32 \
--tool-call-parser qwen3_coder
```
**Python Example:**
```python Example
from openai import OpenAI
client = OpenAI(
api_key="EMPTY",
base_url="http://localhost:30000/v1",
timeout=3600
)
# Define available tools
tools = [
{
"type": "function",
"function": {
"name": "execute_code",
"description": "Execute Python code and return the result",
"parameters": {
"type": "object",
"properties": {
"code": {
"type": "string",
"description": "The Python code to execute"
}
},
"required": ["code"]
}
}
}
]
response = client.chat.completions.create(
model="Qwen/Qwen3-Coder-30B-A3B-Instruct",
messages=[
{"role": "user", "content": "Calculate the factorial of 10 using Python"}
],
tools=tools,
temperature=0.7
)
# Check if the model wants to call a tool
if response.choices[0].message.tool_calls:
tool_call = response.choices[0].message.tool_calls[0]
print(f"Tool: {tool_call.function.name}")
print(f"Arguments: {tool_call.function.arguments}")
else:
# Model may return tool call in content format
print(response.choices[0].message.content)
```
**Example Output:**
```text Output
Tool: execute_code
Arguments: {"code": "def factorial(n):\n if n == 0 or n == 1:\n return 1\n else:\n return n * factorial(n-1)\n\nresult = factorial(10)\nresult"}
```
## 5. Benchmark
### 5.1 Speed Benchmark
**Test Environment:**
- Hardware: AMD MI300X GPU (8x)
- Model: Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8
- Tensor Parallelism: 8
- Expert Parallelism: 2
- sglang version: 0.5.7
We use SGLang's built-in benchmarking tool to conduct performance evaluation with random dataset.
#### 5.1.1 AMD Standard Scenario Benchmark
- Model Deployment Command:
```shell Command
SGLANG_USE_AITER=0 python -m sglang.launch_server \
--model Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8 \
--tp 8 \
--ep 2 \
--context-length 8192 \
--page-size 32 \
--trust-remote-code
```
##### 5.1.1.1 Low Concurrency
- Benchmark Command:
```shell Command
python3 -m sglang.bench_serving \
--backend sglang \
--model Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8 \
--dataset-name random \
--random-input-len 1000 \
--random-output-len 1000 \
--num-prompts 10 \
--max-concurrency 1
```
- Test Results:
```text Output
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 1
Successful requests: 10
Benchmark duration (s): 73.79
Total input tokens: 6101
Total input text tokens: 6101
Total generated tokens: 4220
Total generated tokens (retokenized): 4104
Request throughput (req/s): 0.14
Input token throughput (tok/s): 82.68
Output token throughput (tok/s): 57.19
Peak output token throughput (tok/s): 59.00
Peak concurrent requests: 2
Total token throughput (tok/s): 139.86
Concurrency: 1.00
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 7376.26
Median E2E Latency (ms): 5851.51
P90 E2E Latency (ms): 13351.89
P99 E2E Latency (ms): 16908.32
---------------Time to First Token----------------
Mean TTFT (ms): 191.93
Median TTFT (ms): 126.06
P99 TTFT (ms): 662.15
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 17.06
Median TPOT (ms): 17.07
P99 TPOT (ms): 17.08
---------------Inter-Token Latency----------------
Mean ITL (ms): 17.06
Median ITL (ms): 17.06
P95 ITL (ms): 17.14
P99 ITL (ms): 17.19
Max ITL (ms): 18.53
==================================================
```
##### 5.1.1.2 Medium Concurrency
- Benchmark Command:
```shell Command
python3 -m sglang.bench_serving \
--backend sglang \
--model Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8 \
--dataset-name random \
--random-input-len 1000 \
--random-output-len 1000 \
--num-prompts 80 \
--max-concurrency 16
```
- Test Results:
```text Output
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 16
Successful requests: 80
Benchmark duration (s): 87.04
Total input tokens: 39668
Total input text tokens: 39668
Total generated tokens: 40805
Total generated tokens (retokenized): 40364
Request throughput (req/s): 0.92
Input token throughput (tok/s): 455.77
Output token throughput (tok/s): 468.83
Peak output token throughput (tok/s): 608.00
Peak concurrent requests: 20
Total token throughput (tok/s): 924.59
Concurrency: 13.76
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 14966.88
Median E2E Latency (ms): 15871.93
P90 E2E Latency (ms): 24983.41
P99 E2E Latency (ms): 29504.85
---------------Time to First Token----------------
Mean TTFT (ms): 388.94
Median TTFT (ms): 157.49
P99 TTFT (ms): 1318.63
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 29.41
Median TPOT (ms): 29.22
P99 TPOT (ms): 43.48
---------------Inter-Token Latency----------------
Mean ITL (ms): 28.64
Median ITL (ms): 26.42
P95 ITL (ms): 27.51
P99 ITL (ms): 131.63
Max ITL (ms): 995.11
==================================================
```
##### 5.1.1.3 High Concurrency
- Benchmark Command:
```shell Command
python3 -m sglang.bench_serving \
--backend sglang \
--model Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8 \
--dataset-name random \
--random-input-len 1000 \
--random-output-len 1000 \
--num-prompts 320 \
--max-concurrency 64
```
- Test Results:
```text Output
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 64
Successful requests: 320
Benchmark duration (s): 177.82
Total input tokens: 158939
Total input text tokens: 158939
Total generated tokens: 170134
Total generated tokens (retokenized): 168387
Request throughput (req/s): 1.80
Input token throughput (tok/s): 893.84
Output token throughput (tok/s): 956.80
Peak output token throughput (tok/s): 1728.00
Peak concurrent requests: 70
Total token throughput (tok/s): 1850.64
Concurrency: 58.88
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 32716.53
Median E2E Latency (ms): 30896.37
P90 E2E Latency (ms): 65605.24
P99 E2E Latency (ms): 80970.63
---------------Time to First Token----------------
Mean TTFT (ms): 372.97
Median TTFT (ms): 181.67
P99 TTFT (ms): 529.01
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 62.98
Median TPOT (ms): 50.44
P99 TPOT (ms): 204.24
---------------Inter-Token Latency----------------
Mean ITL (ms): 60.95
Median ITL (ms): 37.87
P95 ITL (ms): 143.98
P99 ITL (ms): 148.02
Max ITL (ms): 36863.32
==================================================
```
#### 5.1.2 NVIDIA (B200/GB200) Standard Scenario Benchmark
The following runs use the same random dataset benchmark client commands as the AMD section. On B200, launch the server with the following command:
```bash
sglang serve --model Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8 --tp 8 --ep 8 --context-length 8192 --page-size 32 --trust-remote-code
##### 5.1.2.1 FP8 Model
- Low Concurrency:
```text Output
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 1
Successful requests: 10
Benchmark duration (s): 42.68
Total input tokens: 6101
Total input text tokens: 6101
Total generated tokens: 4220
Total generated tokens (retokenized): 4204
Request throughput (req/s): 0.23
Input token throughput (tok/s): 142.95
Output token throughput (tok/s): 98.88
Peak output token throughput (tok/s): 102.00
Peak concurrent requests: 2
Total token throughput (tok/s): 241.83
Concurrency: 1.00
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 4266.06
Median E2E Latency (ms): 3420.24
P90 E2E Latency (ms): 7717.19
P99 E2E Latency (ms): 9504.50
---------------Time to First Token----------------
Mean TTFT (ms): 112.03
Median TTFT (ms): 112.70
P99 TTFT (ms): 115.35
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 9.87
Median TPOT (ms): 9.86
P99 TPOT (ms): 9.92
---------------Inter-Token Latency----------------
Mean ITL (ms): 9.87
Median ITL (ms): 9.87
P95 ITL (ms): 10.06
P99 ITL (ms): 10.18
Max ITL (ms): 14.80
==================================================
```
- Medium Concurrency:
```text Output
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 16
Successful requests: 80
Benchmark duration (s): 60.80
Total input tokens: 39668
Total input text tokens: 39668
Total generated tokens: 40805
Total generated tokens (retokenized): 40543
Request throughput (req/s): 1.32
Input token throughput (tok/s): 652.43
Output token throughput (tok/s): 671.13
Peak output token throughput (tok/s): 864.00
Peak concurrent requests: 20
Total token throughput (tok/s): 1323.57
Concurrency: 13.93
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 10587.26
Median E2E Latency (ms): 11486.18
P90 E2E Latency (ms): 17374.75
P99 E2E Latency (ms): 21107.18
---------------Time to First Token----------------
Mean TTFT (ms): 155.27
Median TTFT (ms): 121.57
P99 TTFT (ms): 294.31
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 20.77
Median TPOT (ms): 21.13
P99 TPOT (ms): 23.62
---------------Inter-Token Latency----------------
Mean ITL (ms): 20.49
Median ITL (ms): 18.73
P95 ITL (ms): 19.65
P99 ITL (ms): 98.85
Max ITL (ms): 536.87
==================================================
```
- High Concurrency:
```text Output
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 64
Successful requests: 320
Benchmark duration (s): 100.07
Total input tokens: 158939
Total input text tokens: 158939
Total generated tokens: 170134
Total generated tokens (retokenized): 169119
Request throughput (req/s): 3.20
Input token throughput (tok/s): 1588.32
Output token throughput (tok/s): 1700.19
Peak output token throughput (tok/s): 2303.00
Peak concurrent requests: 71
Total token throughput (tok/s): 3288.51
Concurrency: 57.93
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 18114.01
Median E2E Latency (ms): 18279.15
P90 E2E Latency (ms): 30557.22
P99 E2E Latency (ms): 35889.84
---------------Time to First Token----------------
Mean TTFT (ms): 346.40
Median TTFT (ms): 129.75
P99 TTFT (ms): 1370.20
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 33.76
Median TPOT (ms): 34.62
P99 TPOT (ms): 39.97
---------------Inter-Token Latency----------------
Mean ITL (ms): 33.48
Median ITL (ms): 25.70
P95 ITL (ms): 99.36
P99 ITL (ms): 132.30
Max ITL (ms): 1132.39
==================================================
```
##### 5.1.2.2 NVFP4 Model
- Low Concurrency:
```text Output
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 1
Successful requests: 10
Benchmark duration (s): 34.49
Total input tokens: 6101
Total input text tokens: 6101
Total generated tokens: 4220
Total generated tokens (retokenized): 4218
Request throughput (req/s): 0.29
Input token throughput (tok/s): 176.87
Output token throughput (tok/s): 122.34
Peak output token throughput (tok/s): 127.00
Peak concurrent requests: 2
Total token throughput (tok/s): 299.21
Concurrency: 1.00
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 3448.01
Median E2E Latency (ms): 2768.11
P90 E2E Latency (ms): 6225.73
P99 E2E Latency (ms): 7668.26
---------------Time to First Token----------------
Mean TTFT (ms): 104.55
Median TTFT (ms): 105.38
P99 TTFT (ms): 105.63
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 7.94
Median TPOT (ms): 7.95
P99 TPOT (ms): 7.97
---------------Inter-Token Latency----------------
Mean ITL (ms): 7.94
Median ITL (ms): 7.94
P95 ITL (ms): 8.05
P99 ITL (ms): 8.11
Max ITL (ms): 24.64
==================================================
```
- Medium Concurrency:
```text Output
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 16
Successful requests: 80
Benchmark duration (s): 43.30
Total input tokens: 39668
Total input text tokens: 39668
Total generated tokens: 40805
Total generated tokens (retokenized): 39975
Request throughput (req/s): 1.85
Input token throughput (tok/s): 916.16
Output token throughput (tok/s): 942.42
Peak output token throughput (tok/s): 1264.00
Peak concurrent requests: 21
Total token throughput (tok/s): 1858.57
Concurrency: 13.90
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 7521.95
Median E2E Latency (ms): 8246.89
P90 E2E Latency (ms): 12370.93
P99 E2E Latency (ms): 15023.96
---------------Time to First Token----------------
Mean TTFT (ms): 137.27
Median TTFT (ms): 109.59
P99 TTFT (ms): 208.78
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 14.69
Median TPOT (ms): 14.87
P99 TPOT (ms): 17.63
---------------Inter-Token Latency----------------
Mean ITL (ms): 14.51
Median ITL (ms): 12.75
P95 ITL (ms): 13.33
P99 ITL (ms): 92.85
Max ITL (ms): 113.70
==================================================
```
- High Concurrency:
```text Output
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 64
Successful requests: 320
Benchmark duration (s): 73.93
Total input tokens: 158939
Total input text tokens: 158939
Total generated tokens: 170134
Total generated tokens (retokenized): 168841
Request throughput (req/s): 4.33
Input token throughput (tok/s): 2149.98
Output token throughput (tok/s): 2301.42
Peak output token throughput (tok/s): 3497.00
Peak concurrent requests: 71
Total token throughput (tok/s): 4451.40
Concurrency: 58.28
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 13463.58
Median E2E Latency (ms): 13498.74
P90 E2E Latency (ms): 22957.10
P99 E2E Latency (ms): 26656.95
---------------Time to First Token----------------
Mean TTFT (ms): 239.00
Median TTFT (ms): 113.42
P99 TTFT (ms): 713.87
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 25.13
Median TPOT (ms): 26.02
P99 TPOT (ms): 30.90
---------------Inter-Token Latency----------------
Mean ITL (ms): 24.92
Median ITL (ms): 16.68
P95 ITL (ms): 93.33
P99 ITL (ms): 119.26
Max ITL (ms): 548.82
==================================================
```
### 5.2 Accuracy Benchmark
#### 5.2.1 GSM8K Benchmark
- **Benchmark Command:**
```shell Command
python3 -m sglang.test.few_shot_gsm8k --num-questions 200
```
##### AMD (MI300X/MI325X/MI355X)
- **Results**:
- Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8
```
Accuracy: 0.965
Invalid: 0.000
Latency: 23.084 s
Output throughput: 1148.425 token/s
```
##### NVIDIA (B200/GB200)
For deployment commands, see [Section 3.1](#3-1-configuration).
- Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8
```
Accuracy: 0.965
Invalid: 0.000
Latency: 14.870 s
Output throughput: 1777.726 token/s
```
- nvidia/Qwen3-Coder-480B-A35B-Instruct-NVFP (NVFP4)
```
Accuracy: 0.960
Invalid: 0.000
Latency: 13.948 s
Output throughput: 1988.548 token/s
```