[Docs] Sync docs_new with legacy docs and update migration redirects (#23337)
Co-authored-by: Mingyi <wisclmy0611@gmail.com>
This commit is contained in:
@@ -93,6 +93,13 @@ repos:
|
||||
language: system
|
||||
files: ^test/registered/.*\.py$
|
||||
pass_filenames: false
|
||||
- id: check-no-docs-changes
|
||||
name: reject changes under legacy docs/
|
||||
entry: python3 scripts/ci/check_no_docs_changes.py
|
||||
language: system
|
||||
pass_filenames: false
|
||||
always_run: true
|
||||
stages: [pre-commit]
|
||||
- repo: https://github.com/lycheeverse/lychee.git
|
||||
rev: lychee-v0.22.0
|
||||
hooks:
|
||||
|
||||
@@ -26,7 +26,7 @@ To use DeepSeek-Math-V2, you must agree to DeepSeek's Community License. See [LI
|
||||
|
||||
## 2. SGLang Installation
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -28,7 +28,7 @@ For more details, please refer to the [official DeepSeek-OCR-2 repository](https
|
||||
|
||||
## 2. SGLang Installation
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -24,7 +24,7 @@ For more details, please refer to the [official DeepSeek-OCR repository](https:/
|
||||
|
||||
## 2. SGLang Installation
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -30,7 +30,7 @@ For more details, please refer to the [official DeepSeek-R1 repository](https://
|
||||
|
||||
## 2. SGLang Installation
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -20,7 +20,7 @@ Key highlights include:
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -16,7 +16,7 @@ metatags:
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -22,7 +22,7 @@ The DeepSeek-V3.2 series includes three model variants, each optimized for diffe
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -21,7 +21,7 @@ ERNIE-4.5 delivers advanced features as below:
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -2,7 +2,6 @@
|
||||
title: Chroma-1.0
|
||||
metatags:
|
||||
description: "Deploy Chroma-1.0 end-to-end speech conversation model with SGLang - real-time speech generation, voice cloning, and speech reasoning."
|
||||
tag: NEW
|
||||
---
|
||||
|
||||
## 1. Model Introduction
|
||||
|
||||
@@ -28,7 +28,7 @@ Please refer to the [official GLM-4.5 model card](https://huggingface.co/zai-org
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -22,7 +22,7 @@ GLM-4.5V introduces several key features:
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -22,7 +22,7 @@ For more details, please refer to the [official GLM-4.6 documentation](https://d
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -54,7 +54,7 @@ sudo apt install ffmpeg
|
||||
- Want to use the latest development features
|
||||
- Participate in SGLang project development
|
||||
|
||||
For general installation instructions, you can also refer to the [official SGLang installation guide](../../../docs/get-started/installation).
|
||||
For general installation instructions, you can also refer to the [official SGLang installation guide](../../../docs/get-started/install).
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -37,7 +37,7 @@ Please refer to the [official GLM-4.7-Flash model card](https://huggingface.co/z
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -38,7 +38,7 @@ Please refer to the [official GLM-4.7 model card](https://huggingface.co/zai-org
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -16,7 +16,7 @@ tag: NEW
|
||||
|
||||
## 2. SGLang Installation
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -2,7 +2,6 @@
|
||||
title: GLM-5
|
||||
metatags:
|
||||
description: "Deploy GLM-5 with SGLang on NVIDIA H100/H200/B200 and AMD MI300X/MI325X/MI355X — state-of-the-art reasoning, enhanced coding, and robust tool calling capabilities."
|
||||
tag: NEW
|
||||
---
|
||||
|
||||
## 1. Model Introduction
|
||||
@@ -28,7 +27,7 @@ With advances in both pre-training (28.5T tokens) and post-training via [slime](
|
||||
|
||||
## 2. SGLang Installation
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -31,7 +31,7 @@ Please refer to the [official Glyph model card](https://huggingface.co/zai-org/G
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -25,7 +25,7 @@ For more details, please refer to the [official GLM-OCR model card](https://hugg
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -2,6 +2,7 @@
|
||||
title: Gemma 4
|
||||
metatags:
|
||||
description: "Deploy Gemma 4 with SGLang - Google's next-generation open models with MoE variants and multimodal support for text, vision, and audio."
|
||||
tag: NEW
|
||||
---
|
||||
|
||||
import { Gemma4Deployment } from '/src/snippets/autoregressive/gemma4-deployment.jsx';
|
||||
@@ -79,7 +80,7 @@ docker pull lmsysorg/sglang:dev-gemma4 # CUDA 12.9
|
||||
docker pull lmsysorg/sglang:dev-cu13-gemma4 # CUDA 13
|
||||
```
|
||||
|
||||
For the full Docker setup and other installation methods, please refer to the [official SGLang installation guide](../../../docs/get-started/installation).
|
||||
For the full Docker setup and other installation methods, please refer to the [official SGLang installation guide](../../../docs/get-started/install).
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -2,7 +2,6 @@
|
||||
title: LLaDA 2.1
|
||||
metatags:
|
||||
description: "Deploy LLaDA 2.1 with SGLang - large-scale discrete diffusion language model with parallel token generation, iterative denoising, MoE architecture, and reinforcement learning for reasoning."
|
||||
tag: NEW
|
||||
---
|
||||
|
||||
import { LLaDA21Deployment } from '/src/snippets/autoregressive/llada-21-deployment.jsx';
|
||||
@@ -64,7 +63,7 @@ Apache 2.0. Please refer to the [official LLaDA2.X repository](https://github.co
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -2,7 +2,6 @@
|
||||
title: Ling-2.5-1T
|
||||
metatags:
|
||||
description: "Deploy Ling-2.5-1T with SGLang - 1T parameter MoE model with 63B active parameters, trillion-scale context length up to 1M tokens, and agentic tool calling capabilities."
|
||||
tag: NEW
|
||||
---
|
||||
|
||||
## 1. Model Introduction
|
||||
@@ -35,9 +34,9 @@ docker pull lmsysorg/sglang:nightly-dev-20260213-a0ebaa64
|
||||
docker pull lmsysorg/sglang:nightly-dev-cu13-20260213-a0ebaa64
|
||||
```
|
||||
|
||||
For other installation methods, please refer to the [official SGLang installation guide](../../../docs/get-started/installation).
|
||||
For other installation methods, please refer to the [official SGLang installation guide](../../../docs/get-started/install).
|
||||
|
||||
Ling-2.5-1T is also supported via the **nightly PyPI builds**. See the [SGLang Installation (PyPI)](../../../docs/get-started/installation) guide for setup instructions.
|
||||
Ling-2.5-1T is also supported via the **nightly PyPI builds**. See the [SGLang Installation (PyPI)](../../../docs/get-started/install) guide for setup instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -2,7 +2,6 @@
|
||||
title: Ring-2.5-1T
|
||||
metatags:
|
||||
description: "Deploy Ring-2.5-1T with SGLang - world's first open-source 1T parameter reasoning model with hybrid linear attention, deep reasoning, and agentic tool calling capabilities."
|
||||
tag: NEW
|
||||
---
|
||||
|
||||
## 1. Model Introduction
|
||||
@@ -42,7 +41,7 @@ docker pull lmsysorg/sglang:v0.5.9-rocm700-mi30x
|
||||
docker pull lmsysorg/sglang:v0.5.9-rocm700-mi35x
|
||||
```
|
||||
|
||||
For other installation methods, please refer to the [official SGLang installation guide](../../../docs/get-started/installation).
|
||||
For other installation methods, please refer to the [official SGLang installation guide](../../../docs/get-started/install).
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -20,7 +20,7 @@ For further details, please refer to the [Llama 3.1 blog](https://ai.meta.com/bl
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -22,7 +22,7 @@ For more details, please refer to the [official Llama models repository](https:/
|
||||
|
||||
## 2. SGLang Installation
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -27,7 +27,7 @@ For more details, please refer to the official llama4 Repository:https://www.lla
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -2,7 +2,6 @@
|
||||
title: MiniMax-M2.5
|
||||
metatags:
|
||||
description: "Deploy MiniMax-M2.5 with SGLang - community contribution guide for MiniMax M2.5 model deployment."
|
||||
tag: NEW
|
||||
---
|
||||
|
||||
import { MiniMaxM25Deployment } from '/src/snippets/autoregressive/minimax-m25-deployment.jsx';
|
||||
@@ -24,7 +23,7 @@ For more details, please refer to the [official MiniMax-M2.5 announcement](https
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
**For AMD MI300X/MI325X/MI355X GPUs:**
|
||||
|
||||
|
||||
@@ -24,7 +24,7 @@ For more details, see the [official MiniMax-M2.7 blog post](https://www.minimax.
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
**Docker Images by Hardware Platform:**
|
||||
|
||||
|
||||
@@ -26,7 +26,7 @@ For more details, please refer to the [official Minimax GitHub Repository](https
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions. The AMD environment is currently available in SGLang via Docker image install.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions. The AMD environment is currently available in SGLang via Docker image install.
|
||||
|
||||
### 2.1 AMD Docker
|
||||
#### 2.1.1 Launch docker
|
||||
|
||||
@@ -35,7 +35,7 @@ For enterprises requiring specialized capabilities (increased context, domain-sp
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
<Warning title="Transformers version requirement">
|
||||
Devstral 2 requires a recent `transformers`. Please verify `transformers >= 5.0.0.rc`:
|
||||
|
||||
@@ -23,7 +23,7 @@ For further details, please refer to the [official documentation](https://github
|
||||
|
||||
## 2. SGLang Installation
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -43,10 +43,10 @@ With its multimodal capabilities, efficient MoE architecture, and flexible mode
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
<Info>
|
||||
Mistral Small 4 support landed in [sgl-project/sglang#20708](https://github.com/sgl-project/sglang/pull/20708) and has been merged into `main`. A model-specific Docker image is no longer required. Use the standard SGLang installation methods from the [official installation guide](../../../docs/get-started/installation).
|
||||
Mistral Small 4 support landed in [sgl-project/sglang#20708](https://github.com/sgl-project/sglang/pull/20708) and has been merged into `main`. A model-specific Docker image is no longer required. Use the standard SGLang installation methods from the [official installation guide](../../../docs/get-started/install).
|
||||
</Info>
|
||||
|
||||
---
|
||||
|
||||
@@ -23,7 +23,7 @@ For details, see [official documentation](https://huggingface.co/moonshotai/Kimi
|
||||
|
||||
## 2. SGLang Installation
|
||||
|
||||
Refer to the [official SGLang installation guide](../../../docs/get-started/installation).
|
||||
Refer to the [official SGLang installation guide](../../../docs/get-started/install).
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -72,7 +72,7 @@ For details, see [official documentation](https://huggingface.co/moonshotai/Kimi
|
||||
|
||||
## 2. SGLang Installation
|
||||
|
||||
Refer to the [official SGLang installation guide](../../../docs/get-started/installation).
|
||||
Refer to the [official SGLang installation guide](../../../docs/get-started/install).
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -20,7 +20,7 @@ For details, see [official documentation](https://github.com/MoonshotAI/Kimi-K2)
|
||||
|
||||
## 2. SGLang Installation
|
||||
|
||||
Refer to the [official SGLang installation guide](../../../docs/get-started/installation).
|
||||
Refer to the [official SGLang installation guide](../../../docs/get-started/install).
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -24,7 +24,7 @@ For more details, please refer to the [official Kimi Linear GitHub Repository]:
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -20,7 +20,7 @@ At a high level:
|
||||
|
||||
## 2. SGLang Installation
|
||||
|
||||
Refer to the [official SGLang installation guide](../../../docs/get-started/installation), or install nightly wheel through:
|
||||
Refer to the [official SGLang installation guide](../../../docs/get-started/install), or install nightly wheel through:
|
||||
```bash Command
|
||||
uv pip install sglang==0.5.6.post3.dev1278+gad1b4e472 --extra-index-url https://sgl-project.github.io/whl/nightly/
|
||||
```
|
||||
|
||||
@@ -31,7 +31,7 @@ uv pip install 'git+https://github.com/sgl-project/sglang.git#subdirectory=pytho
|
||||
docker pull lmsysorg/sglang:nightly-dev-20260310-0fd9a57d
|
||||
```
|
||||
|
||||
For the full Docker setup and other installation methods, please refer to the [official SGLang installation guide](../../../docs/get-started/installation).
|
||||
For the full Docker setup and other installation methods, please refer to the [official SGLang installation guide](../../../docs/get-started/install).
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -23,7 +23,7 @@ GPT-OSS introduces several groundbreaking innovations:
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3.Model Deployment
|
||||
|
||||
|
||||
@@ -27,7 +27,7 @@ For more details, please refer to the [official Qwen2.5-VL GitHub Repository](ht
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -24,7 +24,7 @@ For more details, please refer to the [Qwen3-Coder-Next model card](https://hugg
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
**Note:** Qwen3-Coder-Next requires SGLang v0.5.8 or later.
|
||||
|
||||
|
||||
@@ -26,7 +26,7 @@ For more details, please refer to the [official Qwen3-Coder GitHub Repository](h
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -30,7 +30,7 @@ For more details, please refer to the [official Qwen3-Next blog](https://qwen.ai
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -24,7 +24,7 @@ For more details, please refer to the [official Qwen3-VL GitHub Repository](http
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -104,7 +104,7 @@ docker pull lmsysorg/sglang:v0.5.9-rocm720-mi30x
|
||||
docker pull lmsysorg/sglang:v0.5.9-rocm720-mi35x
|
||||
```
|
||||
|
||||
For the full Docker setup and other installation methods, please refer to the [official SGLang installation guide](../../../docs/get-started/installation).
|
||||
For the full Docker setup and other installation methods, please refer to the [official SGLang installation guide](../../../docs/get-started/install).
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -61,7 +61,7 @@ uv pip install 'git+https://github.com/sgl-project/sglang.git#subdirectory=pytho
|
||||
docker pull lmsysorg/sglang:latest
|
||||
```
|
||||
|
||||
For the full Docker setup and other installation methods, please refer to the [official SGLang installation guide](../../../docs/get-started/installation).
|
||||
For the full Docker setup and other installation methods, please refer to the [official SGLang installation guide](../../../docs/get-started/install).
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -24,7 +24,7 @@ For more details, please refer to the [official Qwen3 GitHub Repository](https:/
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -2,7 +2,6 @@
|
||||
title: Step3-VL-10B
|
||||
metatags:
|
||||
description: "Deploy Step3-VL-10B multimodal model with SGLang - compact 10B dense model with frontier-level vision understanding, complex reasoning, and tool calling capabilities."
|
||||
tag: NEW
|
||||
---
|
||||
|
||||
import { Step3VL10BDeployment } from '/src/snippets/autoregressive/step-3vl-10b-deployment.jsx';
|
||||
@@ -24,7 +23,7 @@ For more details, please refer to the [Step3-VL-10B model card on Hugging Face](
|
||||
|
||||
SGLang offers multiple installation methods. You can choose the most suitable installation method based on your hardware platform and requirements.
|
||||
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/installation) for installation instructions.
|
||||
Please refer to the [official SGLang installation guide](../../../docs/get-started/install) for installation instructions.
|
||||
|
||||
## 3. Model Deployment
|
||||
|
||||
|
||||
@@ -2,7 +2,6 @@
|
||||
title: Step-3.5
|
||||
metatags:
|
||||
description: "Deploy Step-3.5 reasoning engine with SGLang. "
|
||||
tag: NEW
|
||||
---
|
||||
|
||||
import { Step35Deployment } from '/src/snippets/autoregressive/step-35-deployment.jsx';
|
||||
|
||||
@@ -32,7 +32,7 @@ Qwen-Image is a text-to-image model. The recommended launch configurations vary
|
||||
|
||||
### 3.2 Configuration Tips
|
||||
|
||||
Current supported optimization all listed [here](../../../docs/sglang-diffusion/attention-backends#platform-support-matrix).
|
||||
Current supported optimization all listed [here](../../../docs/sglang-diffusion/attention_backends#platform-support-matrix).
|
||||
|
||||
- `--vae-path`: Path to a custom VAE model or HuggingFace model ID (e.g., fal/FLUX.2-Tiny-AutoEncoder). If not specified, the VAE will be loaded from the main model path.
|
||||
- `--num-gpus`: Number of GPUs to use
|
||||
@@ -45,7 +45,7 @@ Current supported optimization all listed [here](../../../docs/sglang-diffusion/
|
||||
|
||||
## 4. API Usage
|
||||
|
||||
For complete API documentation, please refer to the [official API usage guide](../../../docs/sglang-diffusion/api/openai-api).
|
||||
For complete API documentation, please refer to the [official API usage guide](../../../docs/sglang-diffusion/api/openai_api).
|
||||
|
||||
### 4.1 Generate an Image
|
||||
|
||||
@@ -72,7 +72,7 @@ with open("output.png", "wb") as f:
|
||||
|
||||
#### 4.2.1 Cache-DiT Acceleration
|
||||
|
||||
SGLang integrates [Cache-DiT](https://github.com/vipshop/cache-dit), a caching acceleration engine for Diffusion Transformers (DiT), to achieve up to 7.4x inference speedup with minimal quality loss. You can set `SGLANG_CACHE_DIT_ENABLED=True` to enable it. For more details, please refer to the SGLang Cache-DiT [documentation](../../../docs/sglang-diffusion/cache-dit).
|
||||
SGLang integrates [Cache-DiT](https://github.com/vipshop/cache-dit), a caching acceleration engine for Diffusion Transformers (DiT), to achieve up to 7.4x inference speedup with minimal quality loss. You can set `SGLANG_CACHE_DIT_ENABLED=True` to enable it. For more details, please refer to the SGLang Cache-DiT [documentation](../../../docs/sglang-diffusion/cache_dit).
|
||||
|
||||
**Basic Usage**
|
||||
|
||||
|
||||
@@ -43,7 +43,7 @@ The Wan2.1 series offers models in multiple sizes and resolutions, optimized for
|
||||
|
||||
### 3.2 Configuration Tips
|
||||
|
||||
Current supported optimization options are listed in the [SGLang diffusion support matrix](../../../docs/sglang-diffusion/attention-backends#platform-support-matrix).
|
||||
Current supported optimization options are listed in the [SGLang diffusion support matrix](../../../docs/sglang-diffusion/attention_backends#platform-support-matrix).
|
||||
|
||||
- `--vae-path`: Path to a custom VAE model or HuggingFace model ID. If not specified, the VAE will be loaded from the main model path.
|
||||
- `--num-gpus {NUM_GPUS}`: Number of GPUs to use.
|
||||
@@ -58,7 +58,7 @@ Current supported optimization options are listed in the [SGLang diffusion suppo
|
||||
### 4.1 Basic Usage
|
||||
|
||||
For more API usage and request examples, please refer to:
|
||||
[SGLang Diffusion OpenAI API](../../../docs/sglang-diffusion/api/openai-api)
|
||||
[SGLang Diffusion OpenAI API](../../../docs/sglang-diffusion/api/openai_api)
|
||||
|
||||
#### 4.1.1 Launch a server and then send requests
|
||||
|
||||
@@ -104,7 +104,7 @@ sglang generate "${SERVER_ARGS[@]}" "${SAMPLING_ARGS[@]}"
|
||||
|
||||
#### 4.2.1 Cache-DiT Acceleration
|
||||
|
||||
SGLang integrates [Cache-DiT](https://github.com/vipshop/cache-dit), a caching acceleration engine for Diffusion Transformers (DiT), to achieve significant inference speedups with minimal quality loss. You can set `SGLANG_CACHE_DIT_ENABLED=True` to enable it. For more details, please refer to the SGLang Cache-DiT [documentation](../../../docs/sglang-diffusion/cache-dit).
|
||||
SGLang integrates [Cache-DiT](https://github.com/vipshop/cache-dit), a caching acceleration engine for Diffusion Transformers (DiT), to achieve significant inference speedups with minimal quality loss. You can set `SGLANG_CACHE_DIT_ENABLED=True` to enable it. For more details, please refer to the SGLang Cache-DiT [documentation](../../../docs/sglang-diffusion/cache_dit).
|
||||
|
||||
**Basic Usage**
|
||||
|
||||
|
||||
@@ -149,7 +149,7 @@ Each recipe provides step-by-step instructions to help you quickly implement SGL
|
||||
|
||||
## Reference
|
||||
|
||||
- [Installation (PyPI)](../docs/get-started/installation) - Install SGLang via pip or uv (stable and nightly)
|
||||
- [Installation (PyPI)](../docs/get-started/install) - Install SGLang via pip or uv (stable and nightly)
|
||||
- [Server arguments](./base/reference/server_arguments) - Understanding all the arguments
|
||||
|
||||
## 🚀 Quick Start
|
||||
|
||||
+115
-93
@@ -18,7 +18,7 @@
|
||||
},
|
||||
{
|
||||
"source": "/advanced_features/adaptive_speculative_decoding.html",
|
||||
"destination": "/docs/advanced_features/speculative_decoding"
|
||||
"destination": "/docs/advanced_features/adaptive_speculative_decoding"
|
||||
},
|
||||
{
|
||||
"source": "/advanced_features/attention_backend.html",
|
||||
@@ -78,7 +78,7 @@
|
||||
},
|
||||
{
|
||||
"source": "/advanced_features/hisparse_guide.html",
|
||||
"destination": "/docs/advanced_features/overview"
|
||||
"destination": "/docs/advanced_features/hisparse_guide"
|
||||
},
|
||||
{
|
||||
"source": "/advanced_features/hyperparameter_tuning.html",
|
||||
@@ -158,7 +158,7 @@
|
||||
},
|
||||
{
|
||||
"source": "/basic_usage/deepseek_ocr.html",
|
||||
"destination": "/docs/basic_usage/overview"
|
||||
"destination": "/docs/basic_usage/deepseek_ocr"
|
||||
},
|
||||
{
|
||||
"source": "/basic_usage/deepseek_v3.html",
|
||||
@@ -226,7 +226,7 @@
|
||||
},
|
||||
{
|
||||
"source": "/basic_usage/qwen3_5.html",
|
||||
"destination": "/docs/basic_usage/qwen3"
|
||||
"destination": "/docs/basic_usage/qwen3_5"
|
||||
},
|
||||
{
|
||||
"source": "/basic_usage/qwen3_vl.html",
|
||||
@@ -258,7 +258,7 @@
|
||||
},
|
||||
{
|
||||
"source": "/developer_guide/development_jit_kernel_guide.html",
|
||||
"destination": "/docs/developer_guide/JIT_kernels"
|
||||
"destination": "/docs/developer_guide/development_jit_kernel_guide"
|
||||
},
|
||||
{
|
||||
"source": "/developer_guide/evaluating_new_models.html",
|
||||
@@ -278,23 +278,23 @@
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/api/openai_api.html",
|
||||
"destination": "/docs/sglang-diffusion/api/openai-api"
|
||||
"destination": "/docs/sglang-diffusion/api/openai_api"
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/api/post_processing.html",
|
||||
"destination": "/docs/sglang-diffusion/installation"
|
||||
"destination": "/docs/sglang-diffusion/api/post_processing"
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/ci_perf.html",
|
||||
"destination": "/docs/sglang-diffusion/ci-performance"
|
||||
"destination": "/docs/sglang-diffusion/ci_perf"
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/compatibility_matrix.html",
|
||||
"destination": "/docs/sglang-diffusion/installation"
|
||||
"destination": "/docs/sglang-diffusion/compatibility_matrix"
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/contributing.html",
|
||||
"destination": "/docs/sglang-diffusion/installation"
|
||||
"destination": "/docs/sglang-diffusion/contributing"
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/development.html",
|
||||
@@ -302,15 +302,15 @@
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/disaggregation.html",
|
||||
"destination": "/docs/sglang-diffusion/installation"
|
||||
"destination": "/docs/sglang-diffusion/disaggregation"
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/environment_variables.html",
|
||||
"destination": "/docs/sglang-diffusion/environment-variables"
|
||||
"destination": "/docs/sglang-diffusion/environment_variables"
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/index.html",
|
||||
"destination": "/docs/sglang-diffusion/installation"
|
||||
"destination": "/docs/sglang-diffusion/index"
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/installation.html",
|
||||
@@ -318,11 +318,11 @@
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/performance/attention_backends.html",
|
||||
"destination": "/docs/sglang-diffusion/attention-backends"
|
||||
"destination": "/docs/sglang-diffusion/attention_backends"
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/performance/cache/cache_dit.html",
|
||||
"destination": "/docs/sglang-diffusion/cache-dit"
|
||||
"destination": "/docs/sglang-diffusion/cache_dit"
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/performance/cache/index.html",
|
||||
@@ -330,7 +330,7 @@
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/performance/cache/teacache.html",
|
||||
"destination": "/docs/sglang-diffusion/tea-cache"
|
||||
"destination": "/docs/sglang-diffusion/teacache"
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/performance/index.html",
|
||||
@@ -342,11 +342,11 @@
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/performance/ring_sp_performance.html",
|
||||
"destination": "/docs/sglang-diffusion/performance-optimization"
|
||||
"destination": "/docs/sglang-diffusion/ring_sp_performance"
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/quantization.html",
|
||||
"destination": "/docs/sglang-diffusion/installation"
|
||||
"destination": "/docs/sglang-diffusion/quantization"
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/reference.html",
|
||||
@@ -354,7 +354,7 @@
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/support_new_models.html",
|
||||
"destination": "/docs/sglang-diffusion/installation"
|
||||
"destination": "/docs/sglang-diffusion/support_new_models"
|
||||
},
|
||||
{
|
||||
"source": "/diffusion/usage.html",
|
||||
@@ -362,87 +362,91 @@
|
||||
},
|
||||
{
|
||||
"source": "/get_started/install.html",
|
||||
"destination": "/docs/get-started/installation"
|
||||
"destination": "/docs/get-started/install"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/amd_gpu.html",
|
||||
"destination": "/docs/hardware-platforms/amd-gpus"
|
||||
"destination": "/docs/hardware-platforms/amd_gpu"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/apple_metal.html",
|
||||
"destination": "/docs/hardware-platforms/overview"
|
||||
"destination": "/docs/hardware-platforms/apple_metal"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/ascend/ascend_contribution_guide.html",
|
||||
"destination": "/docs/hardware-platforms/overview"
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/ascend_contribution_guide"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/ascend/ascend_npu.html",
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/SGLang-installation-with-NPUs-support"
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/ascend_npu"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/ascend/ascend_npu_best_practice.html",
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/Best-Practice-on-Ascend-NPU"
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/ascend_npu_best_practice"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/ascend/ascend_npu_deepseek_example.html",
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/DeepSeek-Examples"
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/ascend_npu_deepseek_example"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/ascend/ascend_npu_environment_variables.html",
|
||||
"destination": "/docs/hardware-platforms/overview"
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/ascend_npu_environment_variables"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/ascend/ascend_npu_glm5_examples.html",
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/GLM-5"
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/ascend_npu_glm5_examples"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/ascend/ascend_npu_quantization.html",
|
||||
"destination": "/docs/hardware-platforms/overview"
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/ascend_npu_quantization"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/ascend/ascend_npu_quick_start.html",
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/ascend_npu_quick_start"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/ascend/ascend_npu_qwen3_5_examples.html",
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/Qwen3.5"
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/ascend_npu_qwen3_5_examples"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/ascend/ascend_npu_qwen3_examples.html",
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/Qwen3-Examples"
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/ascend_npu_qwen3_examples"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/ascend/ascend_npu_support.html",
|
||||
"destination": "/docs/hardware-platforms/overview"
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/ascend_npu_quick_start"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/ascend/ascend_npu_support_features.html",
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/Support-Features-on-Ascend-NPU"
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/ascend_npu_support_features"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/ascend/ascend_npu_support_models.html",
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/Support-Models-on-Ascend-NPU"
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/ascend_npu_support_models"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/ascend/mindspore_backend.html",
|
||||
"destination": "/docs/hardware-platforms/overview"
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/mindspore_backend"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/ascend_npu_ring_sp_performance.html",
|
||||
"destination": "/docs/hardware-platforms/overview"
|
||||
"destination": "/docs/hardware-platforms/ascend-npus/ascend_npu_ring_sp_performance"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/cpu_server.html",
|
||||
"destination": "/docs/hardware-platforms/cpu-server"
|
||||
"destination": "/docs/hardware-platforms/cpu_server"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/mthreads_gpu.html",
|
||||
"destination": "/docs/hardware-platforms/overview"
|
||||
"destination": "/docs/hardware-platforms/mthreads_gpu"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/nvidia_jetson.html",
|
||||
"destination": "/docs/hardware-platforms/overview"
|
||||
"destination": "/docs/hardware-platforms/nvidia_jetson"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/plugin.html",
|
||||
"destination": "/docs/hardware-platforms/overview"
|
||||
"destination": "/docs/hardware-platforms/plugin"
|
||||
},
|
||||
{
|
||||
"source": "/platforms/tpu.html",
|
||||
@@ -526,7 +530,7 @@
|
||||
},
|
||||
{
|
||||
"source": "/supported_models/extending/mindspore_models.html",
|
||||
"destination": "/docs/supported-models/mindspore-models"
|
||||
"destination": "/docs/supported-models/mindspore_models"
|
||||
},
|
||||
{
|
||||
"source": "/supported_models/extending/modelscope.html",
|
||||
@@ -534,11 +538,11 @@
|
||||
},
|
||||
{
|
||||
"source": "/supported_models/extending/support_new_models.html",
|
||||
"destination": "/docs/supported-models/new-model-support"
|
||||
"destination": "/docs/supported-models/support_new_models"
|
||||
},
|
||||
{
|
||||
"source": "/supported_models/extending/transformers_fallback.html",
|
||||
"destination": "/docs/supported-models/transformers-fallback"
|
||||
"destination": "/docs/supported-models/transformers_fallback"
|
||||
},
|
||||
{
|
||||
"source": "/supported_models/index.html",
|
||||
@@ -546,11 +550,11 @@
|
||||
},
|
||||
{
|
||||
"source": "/supported_models/retrieval_ranking/classify_models.html",
|
||||
"destination": "/docs/supported-models/classification-models"
|
||||
"destination": "/docs/supported-models/classify_models"
|
||||
},
|
||||
{
|
||||
"source": "/supported_models/retrieval_ranking/embedding_models.html",
|
||||
"destination": "/docs/supported-models/embedding-models"
|
||||
"destination": "/docs/supported-models/embedding_models"
|
||||
},
|
||||
{
|
||||
"source": "/supported_models/retrieval_ranking/index.html",
|
||||
@@ -558,7 +562,7 @@
|
||||
},
|
||||
{
|
||||
"source": "/supported_models/retrieval_ranking/rerank_models.html",
|
||||
"destination": "/docs/supported-models/rerank-models"
|
||||
"destination": "/docs/supported-models/rerank_models"
|
||||
},
|
||||
{
|
||||
"source": "/supported_models/specialized/index.html",
|
||||
@@ -566,15 +570,15 @@
|
||||
},
|
||||
{
|
||||
"source": "/supported_models/specialized/reward_models.html",
|
||||
"destination": "/docs/supported-models/reward-models"
|
||||
"destination": "/docs/supported-models/reward_models"
|
||||
},
|
||||
{
|
||||
"source": "/supported_models/text_generation/diffusion_language_models.html",
|
||||
"destination": "/docs/supported-models/diffusion-language-models"
|
||||
"destination": "/docs/supported-models/diffusion_language_models"
|
||||
},
|
||||
{
|
||||
"source": "/supported_models/text_generation/generative_models.html",
|
||||
"destination": "/docs/supported-models/large-language-models"
|
||||
"destination": "/docs/supported-models/generative_models"
|
||||
},
|
||||
{
|
||||
"source": "/supported_models/text_generation/index.html",
|
||||
@@ -582,7 +586,7 @@
|
||||
},
|
||||
{
|
||||
"source": "/supported_models/text_generation/multimodal_language_models.html",
|
||||
"destination": "/docs/supported-models/vision-language-models"
|
||||
"destination": "/docs/supported-models/multimodal_language_models"
|
||||
},
|
||||
{
|
||||
"source": "/supported_models.html",
|
||||
@@ -590,7 +594,7 @@
|
||||
},
|
||||
{
|
||||
"source": "/diffusion.html",
|
||||
"destination": "/docs/sglang-diffusion/installation"
|
||||
"destination": "/docs/sglang-diffusion/index"
|
||||
}
|
||||
],
|
||||
"colors": {
|
||||
@@ -626,7 +630,7 @@
|
||||
"icon": "play",
|
||||
"pages": [
|
||||
"index",
|
||||
"docs/get-started/installation",
|
||||
"docs/get-started/install",
|
||||
"docs/get-started/quickstart",
|
||||
"docs/basic_usage/send_request"
|
||||
]
|
||||
@@ -660,12 +664,14 @@
|
||||
"docs/basic_usage/popular_model_usage",
|
||||
"docs/basic_usage/deepseek_v3",
|
||||
"docs/basic_usage/deepseek_v32",
|
||||
"docs/basic_usage/deepseek_ocr",
|
||||
"docs/basic_usage/glm45",
|
||||
"docs/basic_usage/glmv",
|
||||
"docs/basic_usage/gpt_oss",
|
||||
"docs/basic_usage/kimi_k2_5",
|
||||
"docs/basic_usage/minimax_m2",
|
||||
"docs/basic_usage/qwen3",
|
||||
"docs/basic_usage/qwen3_5",
|
||||
"docs/basic_usage/qwen3_vl",
|
||||
"docs/basic_usage/llama4"
|
||||
]
|
||||
@@ -681,7 +687,9 @@
|
||||
"docs/advanced_features/object_storage",
|
||||
"docs/advanced_features/hyperparameter_tuning",
|
||||
"docs/advanced_features/attention_backend",
|
||||
"docs/advanced_features/hisparse_guide",
|
||||
"docs/advanced_features/speculative_decoding",
|
||||
"docs/advanced_features/adaptive_speculative_decoding",
|
||||
"docs/advanced_features/structured_outputs",
|
||||
"docs/advanced_features/structured_outputs_for_reasoning_models",
|
||||
"docs/advanced_features/tool_parser",
|
||||
@@ -723,32 +731,32 @@
|
||||
{
|
||||
"group": "Text Generation",
|
||||
"pages": [
|
||||
"docs/supported-models/large-language-models",
|
||||
"docs/supported-models/vision-language-models",
|
||||
"docs/supported-models/diffusion-language-models"
|
||||
"docs/supported-models/generative_models",
|
||||
"docs/supported-models/multimodal_language_models",
|
||||
"docs/supported-models/diffusion_language_models"
|
||||
]
|
||||
},
|
||||
{
|
||||
"group": "Retrieval and Ranking",
|
||||
"pages": [
|
||||
"docs/supported-models/embedding-models",
|
||||
"docs/supported-models/rerank-models",
|
||||
"docs/supported-models/classification-models"
|
||||
"docs/supported-models/embedding_models",
|
||||
"docs/supported-models/rerank_models",
|
||||
"docs/supported-models/classify_models"
|
||||
]
|
||||
},
|
||||
{
|
||||
"group": "Specialized Models",
|
||||
"pages": [
|
||||
"docs/supported-models/reward-models"
|
||||
"docs/supported-models/reward_models"
|
||||
]
|
||||
},
|
||||
{
|
||||
"group": "Extending SGLang",
|
||||
"pages": [
|
||||
"docs/supported-models/new-model-support",
|
||||
"docs/supported-models/transformers-fallback",
|
||||
"docs/supported-models/support_new_models",
|
||||
"docs/supported-models/transformers_fallback",
|
||||
"docs/supported-models/modelscope",
|
||||
"docs/supported-models/mindspore-models"
|
||||
"docs/supported-models/mindspore_models"
|
||||
]
|
||||
}
|
||||
]
|
||||
@@ -763,7 +771,7 @@
|
||||
"group": "Development",
|
||||
"pages": [
|
||||
"docs/developer_guide/development_guide_using_docker",
|
||||
"docs/developer_guide/JIT_kernels"
|
||||
"docs/developer_guide/development_jit_kernel_guide"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -824,30 +832,38 @@
|
||||
"pages": [
|
||||
"docs/hardware-platforms/overview",
|
||||
"docs/hardware-platforms/nvidia-gpus",
|
||||
"docs/hardware-platforms/amd-gpus",
|
||||
"docs/hardware-platforms/amd_gpu",
|
||||
"docs/hardware-platforms/apple_metal",
|
||||
{
|
||||
"group": "Ascend NPUs",
|
||||
"pages": [
|
||||
"docs/hardware-platforms/ascend-npus/Best-Practice-on-Ascend-NPU",
|
||||
"docs/hardware-platforms/ascend-npus/DeepSeek-Examples",
|
||||
"docs/hardware-platforms/ascend-npus/GLM-5",
|
||||
"docs/hardware-platforms/ascend-npus/MindSpore-Models",
|
||||
"docs/hardware-platforms/ascend-npus/Qwen3-Examples",
|
||||
"docs/hardware-platforms/ascend-npus/Qwen3.5",
|
||||
"docs/hardware-platforms/ascend-npus/SGLang-installation-with-NPUs-support",
|
||||
"docs/hardware-platforms/ascend-npus/Support-Features-on-Ascend-NPU",
|
||||
"docs/hardware-platforms/ascend-npus/Support-Models-on-Ascend-NPU"
|
||||
"docs/hardware-platforms/ascend-npus/ascend_npu_quick_start",
|
||||
"docs/hardware-platforms/ascend-npus/ascend_npu",
|
||||
"docs/hardware-platforms/ascend-npus/ascend_npu_support_features",
|
||||
"docs/hardware-platforms/ascend-npus/ascend_npu_support_models",
|
||||
"docs/hardware-platforms/ascend-npus/ascend_npu_quantization",
|
||||
"docs/hardware-platforms/ascend-npus/ascend_npu_deepseek_example",
|
||||
"docs/hardware-platforms/ascend-npus/ascend_npu_qwen3_examples",
|
||||
"docs/hardware-platforms/ascend-npus/mindspore_backend",
|
||||
"docs/hardware-platforms/ascend-npus/ascend_contribution_guide",
|
||||
"docs/hardware-platforms/ascend-npus/ascend_npu_best_practice",
|
||||
"docs/hardware-platforms/ascend-npus/ascend_npu_ring_sp_performance",
|
||||
"docs/hardware-platforms/ascend-npus/ascend_npu_qwen3_5_examples",
|
||||
"docs/hardware-platforms/ascend-npus/ascend_npu_glm5_examples",
|
||||
"docs/hardware-platforms/ascend-npus/ascend_npu_environment_variables"
|
||||
]
|
||||
},
|
||||
"docs/hardware-platforms/cpu-server",
|
||||
"docs/hardware-platforms/cpu_server",
|
||||
{
|
||||
"group": "Edge & Embedded",
|
||||
"pages": [
|
||||
"docs/hardware-platforms/nvidia"
|
||||
"docs/hardware-platforms/nvidia_jetson"
|
||||
]
|
||||
},
|
||||
"docs/hardware-platforms/mthreads_gpu",
|
||||
"docs/hardware-platforms/tpu",
|
||||
"docs/hardware-platforms/xpu"
|
||||
"docs/hardware-platforms/xpu",
|
||||
"docs/hardware-platforms/plugin"
|
||||
]
|
||||
}
|
||||
]
|
||||
@@ -925,14 +941,14 @@
|
||||
]
|
||||
},
|
||||
{
|
||||
"group": "Moonshotai",
|
||||
"pages": [
|
||||
"cookbook/autoregressive/Moonshotai/Kimi-K2.6",
|
||||
"cookbook/autoregressive/Moonshotai/Kimi-K2.5",
|
||||
"cookbook/autoregressive/Moonshotai/Kimi-K2",
|
||||
"cookbook/autoregressive/Moonshotai/Kimi-Linear"
|
||||
]
|
||||
},
|
||||
"group": "Moonshotai",
|
||||
"pages": [
|
||||
"cookbook/autoregressive/Moonshotai/Kimi-K2.6",
|
||||
"cookbook/autoregressive/Moonshotai/Kimi-K2.5",
|
||||
"cookbook/autoregressive/Moonshotai/Kimi-K2",
|
||||
"cookbook/autoregressive/Moonshotai/Kimi-Linear"
|
||||
]
|
||||
},
|
||||
{
|
||||
"group": "MiniMax",
|
||||
"pages": [
|
||||
@@ -1073,38 +1089,44 @@
|
||||
"group": "SGLang Diffusion",
|
||||
"icon": "sparkles",
|
||||
"pages": [
|
||||
"sglang-diffusion/intro",
|
||||
"docs/sglang-diffusion/index",
|
||||
"docs/sglang-diffusion/installation",
|
||||
"docs/sglang-diffusion/supported-models",
|
||||
"docs/sglang-diffusion/compatibility_matrix",
|
||||
"docs/sglang-diffusion/disaggregation",
|
||||
"docs/sglang-diffusion/quantization",
|
||||
{
|
||||
"group": "Usage",
|
||||
"pages": [
|
||||
"docs/sglang-diffusion/api/cli",
|
||||
"docs/sglang-diffusion/api/openai-api"
|
||||
"docs/sglang-diffusion/api/openai_api",
|
||||
"docs/sglang-diffusion/api/post_processing"
|
||||
]
|
||||
},
|
||||
{
|
||||
"group": "Performance Optimization",
|
||||
"pages": [
|
||||
"docs/sglang-diffusion/performance-optimization",
|
||||
"docs/sglang-diffusion/attention-backends",
|
||||
"docs/sglang-diffusion/ring_sp_performance",
|
||||
"docs/sglang-diffusion/attention_backends",
|
||||
"docs/sglang-diffusion/profiling",
|
||||
"docs/sglang-diffusion/ci-performance"
|
||||
"docs/sglang-diffusion/ci_perf"
|
||||
]
|
||||
},
|
||||
{
|
||||
"group": "Caching Strategies",
|
||||
"pages": [
|
||||
"docs/sglang-diffusion/caching-acceleration",
|
||||
"docs/sglang-diffusion/cache-dit",
|
||||
"docs/sglang-diffusion/tea-cache"
|
||||
"docs/sglang-diffusion/cache_dit",
|
||||
"docs/sglang-diffusion/teacache"
|
||||
]
|
||||
},
|
||||
{
|
||||
"group": "References",
|
||||
"pages": [
|
||||
"docs/sglang-diffusion/environment-variables",
|
||||
"docs/sglang-diffusion/supported-models"
|
||||
"docs/sglang-diffusion/environment_variables",
|
||||
"docs/sglang-diffusion/compatibility_matrix",
|
||||
"docs/sglang-diffusion/support_new_models",
|
||||
"docs/sglang-diffusion/contributing"
|
||||
]
|
||||
}
|
||||
]
|
||||
|
||||
@@ -0,0 +1,197 @@
|
||||
---
|
||||
title: "Adaptive Speculative Decoding"
|
||||
metatags:
|
||||
description: "Configure adaptive speculative decoding so SGLang can adjust speculative steps and draft tokens at runtime based on acceptance behavior."
|
||||
---
|
||||
|
||||
Adaptive speculative decoding lets SGLang adjust `speculative_num_steps/speculative_num_draft_tokens` at runtime instead of keeping a single fixed value for the whole server lifetime.
|
||||
It is designed for workloads whose accept length changes over time, where one static step count is rarely optimal.
|
||||
|
||||
## Current support
|
||||
|
||||
- Only `--speculative-algorithm EAGLE`
|
||||
- Only `--speculative-eagle-topk 1`
|
||||
- If either condition is not met, SGLang falls back to static speculative settings
|
||||
|
||||
## Why adaptive steps help
|
||||
|
||||
`speculative_num_steps` controls how many draft-model autoregressive steps run in each speculative round. In practice, the best value depends on the current workload.
|
||||
|
||||
- If `num_steps` is too small, the draft model could have produced more accepted tokens, but the round stops too early.
|
||||
- If `num_steps` is too large, the draft model produces many candidate tokens that the target model rejects, so extra draft work is wasted.
|
||||
|
||||
Real traffic often moves between high-acceptance and low-acceptance phases, so one fixed step count is usually a compromise. Adaptive mode tries to follow the workload instead of hard-coding a single global `num_steps`.
|
||||
|
||||
## Design overview
|
||||
|
||||
The adaptive mechanism has three pieces:
|
||||
|
||||
- `AdaptiveSpeculativeParams`: the EMA-based policy
|
||||
- `SpecRuntimeState`: the per-tier runtime state bundle
|
||||
- `AdaptiveController`: the coordinator that chooses a tier and activates the matching runtime state
|
||||
|
||||
At startup, SGLang pre-builds one runtime state per candidate tier. By default, the candidate tiers are `candidate_steps = [1, 3, 7]`.
|
||||
|
||||
```mermaid
|
||||
---
|
||||
title: "SpecRuntimeState — speculative_num_steps / speculative_num_draft_tokens"
|
||||
---
|
||||
graph LR
|
||||
subgraph SR[" "]
|
||||
direction LR
|
||||
subgraph D["Draft stage"]
|
||||
direction TB
|
||||
d1[attn_backend]
|
||||
d2[cuda_graph]
|
||||
end
|
||||
subgraph V["Verify stage"]
|
||||
direction TB
|
||||
v1[attn_backend]
|
||||
v2[cuda_graph]
|
||||
end
|
||||
subgraph E["Extend stage"]
|
||||
direction TB
|
||||
e1[attn_backend]
|
||||
e2[cuda_graph]
|
||||
end
|
||||
end
|
||||
```
|
||||
|
||||
This matters because `CudaGraphRunner` is shape-dependent. Each candidate tier owns its own graph and backend state, so runtime switching is a reference swap, not an online graph recapture.
|
||||
|
||||
## Runtime flow
|
||||
|
||||
The adaptive update happens after verify and affects the next round, not the current one:
|
||||
|
||||
```mermaid
|
||||
---
|
||||
title: "EAGLEWorker.forward_batch_generation() — decode path"
|
||||
---
|
||||
flowchart TD
|
||||
A["① draft(batch)<br/>draft model multi-step generation with current tier"]
|
||||
B["② verify(batch, spec_info)<br/>target model tree verification → produces accept_length_per_req"]
|
||||
C["③ forward_draft_extend_after_decode(batch)<br/>draft model KV-cache catch-up"]
|
||||
D["④ adaptive_controller.on_verify_complete(accept_lengths)<br/>update EMA, apply warmup / interval / hysteresis gates<br/>if tier changed, select a pre-built state from pool"]
|
||||
E["worker.apply_runtime_state(state)"]
|
||||
A --> B --> C --> D --> E
|
||||
```
|
||||
|
||||
> Tier switch happens after the current round completes. Backends and CUDA graphs are never swapped mid-round.
|
||||
|
||||
## How the policy decides
|
||||
|
||||
After each verify pass, SGLang reads the accepted draft length per request, computes the batch average, smooths it with an exponential moving average (EMA), and switches among the pre-built candidate tiers `[1, 3, 7]` by default.
|
||||
|
||||
The decision logic is intentionally conservative:
|
||||
|
||||
- `warmup_batches` skips the first few batches
|
||||
- `update_interval` avoids switching every batch
|
||||
- `down_hysteresis` and `up_hysteresis` reduce oscillation
|
||||
|
||||
Conceptually, the policy probes one step beyond the observed acceptance:
|
||||
|
||||
```text
|
||||
target_steps ≈ clamp(round(ema_accept_len) + 1, min(candidate_steps), max(candidate_steps))
|
||||
```
|
||||
|
||||
So if recent requests consistently accept more drafted tokens, the policy tends to move up. If they start rejecting earlier, it tends to move down.
|
||||
|
||||
## Usage
|
||||
|
||||
`--speculative-adaptive-config` is optional, but the speculative setup still needs to be valid for adaptive mode.
|
||||
|
||||
```bash
|
||||
python3 -m sglang.launch_server \
|
||||
--model meta-llama/Llama-2-7b-chat-hf \
|
||||
--speculative-algorithm EAGLE \
|
||||
--speculative-draft-model-path lmsys/sglang-EAGLE-llama2-chat-7B \
|
||||
--speculative-eagle-topk 1 \
|
||||
--speculative-num-steps 3 \
|
||||
--speculative-num-draft-tokens 4 \
|
||||
--speculative-adaptive
|
||||
```
|
||||
|
||||
If you want to override the defaults, add `--speculative-adaptive-config /path/to/adaptive_spec.json`.
|
||||
|
||||
Example config:
|
||||
|
||||
```json
|
||||
{
|
||||
"candidate_steps": [1, 3, 7],
|
||||
"ema_alpha": 0.2,
|
||||
"warmup_batches": 10,
|
||||
"update_interval": 5
|
||||
}
|
||||
```
|
||||
|
||||
## Config file reference
|
||||
|
||||
The config file is optional. Any omitted keys use defaults.
|
||||
|
||||
<table style={{width: "100%", borderCollapse: "collapse", tableLayout: "fixed"}}>
|
||||
<colgroup>
|
||||
<col style={{width: "33.33%"}} />
|
||||
<col style={{width: "33.33%"}} />
|
||||
<col style={{width: "33.33%"}} />
|
||||
</colgroup>
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Key</th>
|
||||
<th>Default</th>
|
||||
<th>Meaning</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td><code>candidate_steps</code></td>
|
||||
<td><code>[1, 3, 7]</code></td>
|
||||
<td>Discrete <code>speculative_num_steps</code> tiers that adaptive mode can switch between</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>ema_alpha</code></td>
|
||||
<td><code>0.2</code></td>
|
||||
<td>EMA smoothing factor for accepted draft length</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>update_interval</code></td>
|
||||
<td><code>5</code></td>
|
||||
<td>Recompute interval, in verify batches, after warmup</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>warmup_batches</code></td>
|
||||
<td><code>10</code></td>
|
||||
<td>Number of verify batches to observe before switching</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>down_hysteresis</code></td>
|
||||
<td><code>-0.25</code></td>
|
||||
<td>Extra margin before moving to a smaller step</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>up_hysteresis</code></td>
|
||||
<td><code>0.0</code></td>
|
||||
<td>Extra margin before moving to a larger step</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
The initial `--speculative-num-steps` is snapped to the nearest value in `candidate_steps`.
|
||||
|
||||
## Monitoring
|
||||
|
||||
You can inspect the active tier and acceptance metric via `/server_info`:
|
||||
|
||||
```bash
|
||||
curl -s http://127.0.0.1:30000/server_info | jq '.internal_states[0] | {speculative_num_steps, avg_spec_accept_length}'
|
||||
```
|
||||
|
||||
- `speculative_num_steps` is the current active tier
|
||||
- `avg_spec_accept_length` helps explain whether the server is likely to move up or down
|
||||
|
||||
## Tuning tips
|
||||
|
||||
- Start with the default candidate tiers `[1, 3, 7]`
|
||||
- Use fewer tiers if you want lower startup and graph-memory overhead
|
||||
- Increase `ema_alpha` to react faster, or lower it for more stability
|
||||
- Increase `warmup_batches` or `update_interval` if tier switching is too noisy
|
||||
- If your workload is already stable and one static setting is well tuned, adaptive mode may not help much
|
||||
@@ -14,7 +14,7 @@ If you don't specify `--attention-backend`, SGLang makes a best effort to automa
|
||||
|
||||
## Support Matrix
|
||||
|
||||
The support matrix is split into two parts: MHA (standard attention) and MLA (multi-head latent attention). For an explanation of the key differences between MHA and MLA, please see the [SGLang documentation on DeepSeek MLA](../basic_usage/deepseek_v3.md#multi-head-latent-attention-mla-throughput-optimizations) and the original [DeepSeek MLA paper](https://arxiv.org/pdf/2405.04434).
|
||||
The support matrix is split into two parts: MHA (standard attention) and MLA (multi-head latent attention). For an explanation of the key differences between MHA and MLA, please see the [SGLang documentation on DeepSeek MLA](../basic_usage/deepseek_v3#multi-head-latent-attention-mla-throughput-optimizations) and the original [DeepSeek MLA paper](https://arxiv.org/pdf/2405.04434).
|
||||
|
||||
### MHA Backends
|
||||
|
||||
@@ -67,15 +67,15 @@ The support matrix is split into two parts: MHA (standard attention) and MLA (mu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>128</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>**Triton**</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
@@ -129,7 +129,7 @@ The support matrix is split into two parts: MHA (standard attention) and MLA (mu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
</tr>
|
||||
<tr>
|
||||
@@ -147,9 +147,9 @@ The support matrix is split into two parts: MHA (standard attention) and MLA (mu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
</tr>
|
||||
<tr>
|
||||
@@ -258,7 +258,7 @@ The support matrix is split into two parts: MHA (standard attention) and MLA (mu
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>1</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>❌</td>
|
||||
</tr>
|
||||
@@ -279,10 +279,14 @@ Multimodal attention is selected by `--mm-attention-backend`. The "MultiModal" c
|
||||
</Note>
|
||||
|
||||
<Note>
|
||||
- FlashAttention 4 is prefill-only for now.
|
||||
- NSA is specifically designed for [DeepSeek V3.2 DSA](https://lmsys.org/blog/2025-09-29-deepseek-V32/).
|
||||
- FlashAttention 4 supports both prefill and decode on SM90 (Hopper) and SM100 (Blackwell). FA4 MLA supports `page_size = 1`; FA4 MHA requires `page_size = 128`. On SM100, this is auto-enforced by the server; on SM90, users must set `--page-size 128` manually.
|
||||
- NSA is specifically designed for [DeepSeek V3.2 DSA](https://lmsys.org/blog/2025-09-29-deepseek-V32/). See the [DSA Attention Backend (NSA)](#dsa-attention-backend-nsa) section and [DeepSeek V3.2 deployment guide](../basic_usage/deepseek_v32) for details.
|
||||
</Note>
|
||||
|
||||
<Warning>
|
||||
**FA4 on Hopper (SM90):** FA4 decode speed decreases as sequence length grows due to lack of SplitKV support. At batch=1 compared to FA3 on H100: ~-10% at 2K tokens, ~-18% at 4K, ~-31% at 8K, ~-49% at 16K. Larger batch sizes reduce the gap (e.g., batch=8: ~-2% at 2K, ~-8% at 4K). Blackwell (SM100) is not affected.
|
||||
</Warning>
|
||||
|
||||
<Note>
|
||||
For the KV4 FA4 scenario, FA4 requires using a different --decode-attention-backend to run. Except for trtllm_mha being incompatible with FA4, all other decode backends behave as shown in the table.
|
||||
</Note>
|
||||
@@ -291,8 +295,16 @@ For the KV4 FA4 scenario, FA4 requires using a different --decode-attention-back
|
||||
Speculative decoding topk: `topk` is the number of draft tokens sampled per step from the draft model. `topk = 1` follows classic EAGLE; `topk > 1` explores multiple branches and requires backend support in both draft and verification paths.
|
||||
</Tip>
|
||||
|
||||
<Note>
|
||||
**Speculative Decoding V2 (Spec V2):** Spec V2 uses overlap scheduling (`SGLANG_ENABLE_SPEC_V2=True`) that benefits various attention backends. Requires `--speculative-eagle-topk 1` and currently applies to EAGLE and EAGLE3.
|
||||
|
||||
**Verified backends:** TRTLLM MLA, TRTLLM MHA, FA3, Ascend (NPU), Triton.
|
||||
|
||||
**Limited support:** FlashInfer can run under Spec V2, but its plan stream (used for split-KV optimization) introduces a synchronization point that limits overlap benefits.
|
||||
</Note>
|
||||
|
||||
<Tip>
|
||||
Page size controls how many tokens are grouped into a KV cache block. For the prefix cache to take effect, the number of tokens must fill at least one complete page. For example, if your prompt is only 32 tokens and `page_size = 64`, it won't fill a complete page and cannot be matched in the prefix cache (pages cannot be padded). With 65 tokens and `page_size = 64`, only the first page of 64 tokens will be cached and matched; the remaining 1 token is discarded. Use `page_size = 1` for maximum prefix reuse (token-level matching).
|
||||
Page size controls how many tokens are grouped into a KV cache block. For the prefix cache to take effect, the number of tokens must fill at least one complete page. For example, if your prompt is only 32 tokens and `page_size = 64`, it won't fill a complete page and cannot be matched in the prefix cache (pages cannot be padded). With 65 tokens and `page_size = 64`, only the first page of 64 tokens will be cached and matched; the remaining 1 token is discarded. Use `page_size = 1` for maximum prefix reuse (token-level matching). Note that higher page sizes generally improve attention kernel performance, so prefer `page_size > 1` when prefix cache reuse is not critical.
|
||||
</Tip>
|
||||
|
||||
Many backends that do not natively operate on pages can emulate `page_size > 1` at the wrapper layer by expanding page tables to per-token indices. The "Page Size > 1 (native)" column indicates true in-kernel paging. Some backends require fixed native page sizes and cannot be reduced/emulated differently: TRTLLM MHA (16/32/64), TRTLLM MLA (32/64), FlashMLA (64), Cutlass MLA (128), Ascend (128).
|
||||
@@ -303,6 +315,138 @@ MLA page-size constraints:
|
||||
- Cutlass MLA: page_size = 128.
|
||||
- TRTLLM MLA: page_size ∈ {32, 64}.
|
||||
|
||||
### GDN Attention Backends
|
||||
|
||||
GDN (Gated Delta Network) is a linear attention mechanism with O(n) complexity, used in hybrid models that alternate GDN linear attention layers with standard full attention layers. GDN is **not** selected via `--attention-backend`; it is automatically activated when the model architecture requires it (e.g., Qwen 3.5, Qwen 3 Next, Jet Nemotron, Jet VLM).
|
||||
|
||||
The GDN linear attention layers have their own kernel backends, selected via `--linear-attn-backend` (default: `triton`). You can override the kernel per phase with `--linear-attn-decode-backend` and `--linear-attn-prefill-backend`.
|
||||
|
||||
<table style={{width: "100%", borderCollapse: "collapse", tableLayout: "fixed"}}>
|
||||
<colgroup>
|
||||
<col style={{width: "28%"}} />
|
||||
<col style={{width: "16%"}} />
|
||||
<col style={{width: "24%"}} />
|
||||
<col style={{width: "32%"}} />
|
||||
</colgroup>
|
||||
<thead>
|
||||
<tr style={{borderBottom: "2px solid #d55816"}}>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.02)"}}>Backend</th>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.05)"}}>Decode</th>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.02)"}}>Prefill / Extend</th>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.05)"}}>Spec Decoding (Target Verify)</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong>Triton (CUDA)</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong>Triton (AMD/ROCm)</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong>Triton (NPU)</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>❌</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong>Triton (CPU)</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>❌</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong>CuTe DSL (CUDA only)</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>❌</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
<Warning>
|
||||
GDN models are hybrid: the full-attention layers still require a standard `--attention-backend`. Platform constraints for the full-attention backend on hybrid GDN models:
|
||||
- **Blackwell (e.g., B200)**: `triton`, `trtllm_mha`, or `fa4` only.
|
||||
- **NPU (Ascend)**: `ascend` only.
|
||||
- **AMD (ROCm)**: `triton` recommended.
|
||||
- **Other CUDA (Hopper, Ampere, etc.)**: auto-selection works; no special constraints.
|
||||
</Warning>
|
||||
|
||||
### DSA Attention Backend (NSA)
|
||||
|
||||
DSA (Deepseek Sparse Attention) is a native sparse attention mechanism used by [DeepSeek V3.2](https://lmsys.org/blog/2025-09-29-deepseek-V32/). It is activated automatically when the model architecture requires it and is selected via `--attention-backend nsa`.
|
||||
|
||||
Internally, the NSA backend dispatches to different sub-backends for prefill and decode phases. You can override these with `--nsa-prefill-backend` and `--nsa-decode-backend`:
|
||||
|
||||
<table style={{width: "100%", borderCollapse: "collapse", tableLayout: "fixed"}}>
|
||||
<colgroup>
|
||||
<col style={{width: "26%"}} />
|
||||
<col style={{width: "16%"}} />
|
||||
<col style={{width: "16%"}} />
|
||||
<col style={{width: "42%"}} />
|
||||
</colgroup>
|
||||
<thead>
|
||||
<tr style={{borderBottom: "2px solid #d55816"}}>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.02)"}}>Sub-backend</th>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.05)"}}>Prefill</th>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.02)"}}>Decode</th>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.05)"}}>Notes</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong>flashmla_sparse</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Default prefill on Hopper and Blackwell (bf16)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong>flashmla_kv</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Default decode for FP8 on Blackwell with DP</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong>flashmla_auto</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>❌</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Auto-selects flashmla_sparse or flashmla_kv based on kv_cache_dtype</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong>fa3</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Default decode on Hopper (bf16)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong>trtllm</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Default decode on Blackwell (bf16); default for both on Blackwell without DP</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong>tilelang</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Default on AMD (ROCm)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong>aiter</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>✅</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>AMD-specific kernel library (requires aiter package)</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
For deployment examples, see the [DeepSeek V3.2 deployment guide](../basic_usage/deepseek_v32).
|
||||
|
||||
### Hybrid attention (different backends for prefill vs decode) (Experimental)
|
||||
|
||||
<Warning>
|
||||
@@ -354,7 +498,7 @@ If the `--attention-backend` argument is not specified, SGLang automatically sel
|
||||
|
||||
**2. MLA Models (e.g., DeepSeek V3)**
|
||||
- **Hopper**: Defaults to `fa3` (requires CUDA 12.3+).
|
||||
- **Blackwell**: Defaults to `trtllm_mla`.
|
||||
- **Blackwell**: Defaults to `flashinfer`; `trtllm_mla` is auto-selected for DeepSeek V3 models specifically.
|
||||
- **Other Architectures**: Defaults to `triton`.
|
||||
|
||||
|
||||
@@ -432,8 +576,34 @@ python3 -m sglang.launch_server \
|
||||
--trust-remote-code
|
||||
```
|
||||
|
||||
- TRTLLM MHA (Optimized for Blackwell Architecture, e.g., B200)
|
||||
```bash Command
|
||||
python3 -m sglang.launch_server \
|
||||
--tp 4 \
|
||||
--model Qwen/Qwen3.5-35B-A3B-FP8 \
|
||||
--attention-backend trtllm_mha \
|
||||
--trust-remote-code
|
||||
```
|
||||
|
||||
- TRTLLM MHA (XQA backend) (Optimized for SM90 and SM120, e.g., H20, H200, 5090)
|
||||
Note that TRTLLM XQA backend only works well for pagesize 64.
|
||||
```bash Command
|
||||
python3 -m sglang.launch_server \
|
||||
--tp 4 \
|
||||
--model Qwen/Qwen3.5-35B-A3B-FP8 \
|
||||
--decode-attention-backend trtllm_mha \
|
||||
--trust-remote-code
|
||||
```
|
||||
|
||||
- FlashAttention 4 (MHA & MLA)
|
||||
```bash Command
|
||||
# FA4 for both prefill and decode on SM90/SM100
|
||||
python3 -m sglang.launch_server \
|
||||
--model-path Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 \
|
||||
--attention-backend fa4 \
|
||||
--page-size 128 \
|
||||
--trust-remote-code
|
||||
|
||||
python3 -m sglang.launch_server \
|
||||
--tp 8 \
|
||||
--model deepseek-ai/DeepSeek-R1 \
|
||||
@@ -497,6 +667,10 @@ To add a new attention backend, you can learn from the existing backends
|
||||
(`python/sglang/srt/layers/attention/triton_backend.py`, `python/sglang/srt/layers/attention/flashattention_backend.py`)
|
||||
and follow the steps below.
|
||||
|
||||
<Note>
|
||||
Linear attention kernel backends (GDN, KDA) follow a different pattern. They implement `LinearAttnKernelBase` in `python/sglang/srt/layers/attention/linear/kernels/` and are dispatched by `GDNKernelDispatcher` / `KDAKernelDispatcher` rather than registered via `@register_attention_backend`.
|
||||
</Note>
|
||||
|
||||
1. Run without cuda graph. Support the two forward functions
|
||||
- forward_extend
|
||||
- Will be used for prefill, prefill with KV cache, and target verification
|
||||
|
||||
@@ -19,6 +19,81 @@ When launching a language-only model, you must additionally specify the encoder
|
||||
|
||||
We support multiple encoder transfer backends, including zmq_to_scheduler, zmq_to_tokenizer, and mooncake (the default is zmq_to_scheduler). The backend can be selected using `--encoder-transfer-backend`.
|
||||
|
||||
### Encoder transfer with Mooncake
|
||||
|
||||
`--encoder-transfer-backend mooncake` controls **how encoder outputs are transferred** between encoder and language/prefill services. It is an encoder transfer option and can be used independently of the global multimodal embedding cache.
|
||||
|
||||
Example:
|
||||
|
||||
```bash Command
|
||||
# encoder
|
||||
python -m sglang.launch_server \
|
||||
--model-path Qwen/Qwen3-VL-8B-Instruct \
|
||||
--encoder-only \
|
||||
--encoder-transfer-backend mooncake \
|
||||
--port 30000
|
||||
|
||||
# language-only server
|
||||
python -m sglang.launch_server \
|
||||
--model-path Qwen/Qwen3-VL-8B-Instruct \
|
||||
--language-only \
|
||||
--encoder-urls http://127.0.0.1:30000 \
|
||||
--encoder-transfer-backend mooncake \
|
||||
--port 30002
|
||||
```
|
||||
|
||||
### Global multimodal embedding cache with Mooncake
|
||||
|
||||
SGLang also supports a Mooncake-backed **global multimodal embedding cache** for EPD workloads. When enabled on encoder servers, repeated image inputs can reuse previously computed ViT embeddings across instances instead of running the vision encoder again.
|
||||
|
||||
This feature is useful when:
|
||||
|
||||
- the deployment serves repeated or overlapping image inputs,
|
||||
- encoder compute is the bottleneck, and
|
||||
- Mooncake is already available in the cluster.
|
||||
|
||||
At a high level, the encoder checks whether the image embedding already exists in Mooncake. Cache hits are prefetched from the global store, while misses are encoded normally and inserted into the cache in the background.
|
||||
|
||||
To enable it:
|
||||
|
||||
- install and configure Mooncake in the same way as other SGLang Mooncake integrations,
|
||||
- add `--enable-mm-global-cache` on the encoder server.
|
||||
|
||||
`--enable-mm-global-cache` controls **whether multimodal embeddings are looked up and stored in the global Mooncake cache**. It is separate from `--encoder-transfer-backend`, which only controls encoder output transport.
|
||||
|
||||
For Mooncake deployment and configuration details, see [HiCache best practices](./hicache_best_practices#deployment-with-mooncake) and the [Mooncake backend README](https://github.com/sgl-project/sglang/blob/main/python/sglang/srt/mem_cache/storage/mooncake_store/README.md).
|
||||
|
||||
Example:
|
||||
|
||||
```bash Command
|
||||
# Shared Mooncake configuration
|
||||
export MOONCAKE_TE_META_DATA_SERVER="http://127.0.0.1:8080/metadata"
|
||||
export MOONCAKE_MASTER="127.0.0.1:50051"
|
||||
export MOONCAKE_PROTOCOL="rdma"
|
||||
export MOONCAKE_GLOBAL_SEGMENT_SIZE="4gb"
|
||||
|
||||
# encoder with global multimodal cache enabled
|
||||
python -m sglang.launch_server \
|
||||
--model-path Qwen/Qwen3-VL-8B-Instruct \
|
||||
--encoder-only \
|
||||
--enable-mm-global-cache \
|
||||
--port 30000
|
||||
|
||||
# language-only server
|
||||
python -m sglang.launch_server \
|
||||
--model-path Qwen/Qwen3-VL-8B-Instruct \
|
||||
--language-only \
|
||||
--encoder-urls http://127.0.0.1:30000 \
|
||||
--port 30002
|
||||
```
|
||||
|
||||
Notes:
|
||||
|
||||
- This cache is for **multimodal encoder embeddings**, not the language model KV cache.
|
||||
- The feature currently uses Mooncake as the shared backing store.
|
||||
- It can be enabled regardless of which `--encoder-transfer-backend` you use.
|
||||
- It is most relevant for EPD or encoder-disaggregated VLM deployments where the same images are likely to appear across requests or instances.
|
||||
|
||||
#### Qwen VL
|
||||
|
||||
- EP Disaggregation
|
||||
@@ -81,3 +156,42 @@ python -m sglang_router.launch_router \
|
||||
--port 8000
|
||||
|
||||
```
|
||||
|
||||
#### gRPC Encoder (EPD)
|
||||
|
||||
You can run the encoder as a gRPC server while keeping prefill/decode as HTTP.
|
||||
When using gRPC encoders, set `SGLANG_ENCODER_MM_RECEIVER_MODE=grpc` for the
|
||||
prefill process so it uses the gRPC receiver.
|
||||
|
||||
```bash Command
|
||||
# gRPC encoder
|
||||
python -m sglang.launch_server \
|
||||
--model-path Qwen/Qwen3-VL-8B-Instruct \
|
||||
--encoder-only \
|
||||
--grpc-mode \
|
||||
--encoder-transfer-backend zmq_to_scheduler \
|
||||
--port 30000
|
||||
|
||||
# prefill (HTTP) - tell it to use gRPC receiver
|
||||
SGLANG_ENCODER_MM_RECEIVER_MODE=grpc \
|
||||
python -m sglang.launch_server \
|
||||
--model-path Qwen/Qwen3-VL-8B-Instruct \
|
||||
--disaggregation-mode prefill \
|
||||
--language-only \
|
||||
--encoder-urls grpc://127.0.0.1:30000 \
|
||||
--encoder-transfer-backend zmq_to_scheduler \
|
||||
--port 30002
|
||||
|
||||
# decode (HTTP)
|
||||
python -m sglang.launch_server \
|
||||
--model-path Qwen/Qwen3-VL-8B-Instruct \
|
||||
--disaggregation-mode decode \
|
||||
--port 30003
|
||||
|
||||
# router
|
||||
python -m sglang_router.launch_router \
|
||||
--pd-disaggregation \
|
||||
--prefill http://$PREFILL_HOST:30002 \
|
||||
--decode http://$DECODE_HOST:30003 \
|
||||
--port 8000
|
||||
```
|
||||
|
||||
@@ -42,6 +42,16 @@ SGLang's EP integrates diverse, highly efficient backends for different use case
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>An extension of DeepEP for elastic inference, leveraging RDMA for high-performance data transfers.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>Elastic EP serving.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><code>nixl</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><a href="https://github.com/ai-dynamo/nixl/tree/main/examples/device/ep">NIXL-EP</a>, an elastic EP communication library built on NVIDIA's <a href="https://github.com/ai-dynamo/nixl">NIXL</a> framework with native RDMA and NVLink support.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>Elastic EP serving with fault tolerance and dynamic scaling.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><code>mori</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>MORI-EP, AMD's native all-to-all communication implementation optimized for ROCm.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>AMD GPU deployments.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>`flashinfer`</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Flashinfer implementation of all-to-all.</td>
|
||||
@@ -55,9 +65,9 @@ SGLang's EP integrates diverse, highly efficient backends for different use case
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
DeepEP and Mooncake backends support two modes for token dispatch: `normal` mode (optimized for prefill workloads with high throughput) and `low_latency` mode (optimized for decode workloads with low latency and CUDA Graph compatibility). Users are recommended to set `--deepep-mode auto` to enable automatic dispatch mode switching during runtime. Setting `--deepep-mode normal` or `--deepep-mode low_latency` is useful for debugging or development purposes.
|
||||
DeepEP and Mooncake backends support two modes for token dispatch: `normal` mode (optimized for prefill workloads with high throughput) and `low_latency` mode (optimized for decode workloads with low latency and CUDA Graph compatibility). MORI backend only supports `normal` mode now. NIXL-EP currently operates in low-latency mode with CUDA Graph support. Users are recommended to set `--deepep-mode auto` to enable automatic dispatch mode switching during runtime. Setting `--deepep-mode normal` or `--deepep-mode low_latency` is useful for debugging or development purposes.
|
||||
|
||||
Currently, DeepEP and Mooncake only support cases where `ep_size = tp_size`. For hybrid EP and TP (i.e., `ep_size < tp_size`), only the `none` backend (All-Reduce or All-Gather-based dispatching) is supported.
|
||||
Currently, DeepEP, Mooncake, NIXL-EP, `ascend_fuseep` and MORI only support cases where `ep_size = tp_size`. For hybrid EP and TP (i.e., `ep_size < tp_size`), only the `none` backend (All-Reduce or All-Gather-based dispatching) is supported.
|
||||
|
||||
### Backends for MoE Computation
|
||||
|
||||
@@ -82,7 +92,7 @@ Currently, DeepEP and Mooncake only support cases where `ep_size = tp_size`. For
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>`triton`</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Triton-based implementation for grouped GEMMs. To achieve higher performance, it's highly recommended to create [tuned configurations](https://github.com/sgl-project/sglang/blob/main/benchmark/kernels/fused_moe_triton/README).</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Triton-based implementation for grouped GEMMs. To achieve higher performance, it's highly recommended to create <a href="https://github.com/sgl-project/sglang/blob/main/benchmark/kernels/fused_moe_triton/README.md">tuned configurations</a>.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>Custom kernel development or scenarios requiring high extensibility with Torch compilation support.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
@@ -100,6 +110,11 @@ Currently, DeepEP and Mooncake only support cases where `ep_size = tp_size`. For
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>FlashInfer integrated with TensorRT-LLM for accelerated MoE computations, supporting FP4 communication operators and high-performance GEMMs.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>Blackwell with TRT-LLM.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><code>flashinfer_trtllm_routed</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>FlashInfer integrated with TensorRT-LLM for accelerated routed MoE computations, consuming SGLang-computed top-k expert assignments and weights.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>Blackwell with TRT-LLM.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>`flashinfer_cutlass`</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>FlashInfer combined with CUTLASS for high-performance grouped GEMMs in MoE layers, handling FP4/FP8 quantization efficiently.</td>
|
||||
@@ -242,7 +257,7 @@ For model like `nvidia/DeepSeek-R1-0528-NVFP4-v2`, the target model uses NVFP4 p
|
||||
|
||||
## Ascend NPU Guidance
|
||||
### Guidance on SGLang configuration in Ascend NPU
|
||||
- `--moe-a2a-backend` only supports deepep and ascend_fuseep backends,
|
||||
- `--moe-a2a-backend` only supports `deepep` and `ascend_fuseep` backends,
|
||||
|
||||
- `deepep`: The mechanism is consistent with the above description.
|
||||
|
||||
@@ -252,12 +267,13 @@ For model like `nvidia/DeepSeek-R1-0528-NVFP4-v2`, the target model uses NVFP4 p
|
||||
|
||||
- `--deepep-mode`:
|
||||
|
||||
- In PD mixed mode, please set `--deepep-mode` auto.
|
||||
- In PD mixed mode, please set `--deepep-mode auto`.
|
||||
|
||||
- In PD Disaggregation Mode, prefill instance sets `--deepep-mode` normal, and decode instance sets `--deepep-mode` low_latency.
|
||||
- In PD Disaggregation Mode, prefill instance sets `--deepep-mode normal`, and decode instance sets `--deepep-mode low_latency`.
|
||||
|
||||
### DeepEP Ascend Introduction
|
||||
DeepEP Ascend is the adapted version of the DeepEP communication library for Huawei Ascend NPUs, specifically designed for Mixture-of-Experts (MoE) model Expert Parallelism (EP). It supports the Ant-moving Function (Split the sequence length into rounds for streaming batch transmission) to optimize the buffer size occupied during collective communication in prefill stage, especially for long sequences.
|
||||
DeepEP Ascend is the adapted version of the DeepEP communication library for Huawei Ascend NPUs, specifically designed for Mixture-of-Experts (MoE) model Expert Parallelism (EP).
|
||||
It supports the Ant-moving Function (Split the sequence length into rounds for streaming batch transmission) to optimize the buffer size occupied during collective communication in prefill stage, especially for long sequences.
|
||||
|
||||
Ant-moving Function can be enabled for both the dispatch and combine phases via the following environment variables:
|
||||
|
||||
|
||||
@@ -42,6 +42,23 @@ Notes:
|
||||
- `page_first`: Only compatible with `kernel` I/O backend, automatically switches to `layer_first` with `direct` backend
|
||||
- `page_first_direct`: Specifically designed for `direct` I/O backend with optimized memory organization
|
||||
|
||||
### Heterogeneous TP Support (GQA/MHA models)
|
||||
|
||||
HiCache storage supports cross-cluster KV reuse when different deployments use different TP sizes (for example, `tp=4` and `tp=8`) and share the same storage backend namespace.
|
||||
|
||||
Use `tp_lcm_size` in `--hicache-storage-backend-extra-config`:
|
||||
|
||||
```bash Command
|
||||
# Example: heterogeneous TP = {4, 8}, so lcm = 8
|
||||
--hicache-storage-backend-extra-config '{"tp_lcm_size": 8}'
|
||||
```
|
||||
|
||||
Guidelines:
|
||||
|
||||
- Set `tp_lcm_size` to the least common multiple (LCM) of all TP sizes that will share the same HiCache storage.
|
||||
- For MHA models with Mooncake and `page_head` layout, HiCache will split head shards based on `tp_lcm_size` to make keys reusable across heterogeneous TP deployments.
|
||||
- If all clusters use the same TP size, this option is not needed.
|
||||
|
||||
### Prefetch Policies
|
||||
|
||||
```bash Command
|
||||
@@ -108,7 +125,7 @@ python3 -m sglang.launch_server \
|
||||
|
||||
### Deployment with HF3FS
|
||||
|
||||
Here is an example of deploying DeepSeek-R1 with HiCache-HF3FS. For more details, see the [HF3FS Documentation](https://github.com/sgl-project/sglang/tree/main/python/sglang/srt/mem_cache/storage/hf3fs/docs).
|
||||
Here is an example of deploying DeepSeek-R1 with HiCache-HF3FS. For more details, see the [HF3FS Documentation](https://github.com/sgl-project/sglang/blob/main/python/sglang/srt/mem_cache/storage/hf3fs/docs/README.md).
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.launch_server \
|
||||
@@ -133,7 +150,7 @@ python3 -m sglang.launch_server \
|
||||
|
||||
### Deployment with Mooncake
|
||||
|
||||
Here is an example of deploying Qwen3-235B-A22B-Instruct-2507 with Mooncake. For more details, see the [Mooncake Documentation](https://github.com/sgl-project/sglang/tree/main/python/sglang/srt/mem_cache/storage/mooncake_store).
|
||||
Here is an example of deploying Qwen3-235B-A22B-Instruct-2507 with Mooncake. For more details, see the [Mooncake Documentation](https://github.com/sgl-project/sglang/blob/main/python/sglang/srt/mem_cache/storage/mooncake_store/README.md).
|
||||
|
||||
```bash Command
|
||||
# Set Mooncake environment variables
|
||||
|
||||
@@ -21,8 +21,8 @@ The control path is:
|
||||
|
||||
1. **HTTP Server** (`python/sglang/srt/entrypoints/http_server.py`)
|
||||
- Exposes `PUT /hicache/storage-backend`, `DELETE /hicache/storage-backend`, `GET /hicache/storage-backend`
|
||||
2. **TokenizerManager** (`python/sglang/srt/managers/tokenizer_communicator_mixin.py`)
|
||||
- Sends the request to the Scheduler via `_Communicator`
|
||||
2. **TokenizerManager** (`python/sglang/srt/managers/tokenizer_control_mixin.py`)
|
||||
- Sends the request to the Scheduler via `FanOutCommunicator`
|
||||
3. **Scheduler** (`python/sglang/srt/managers/scheduler.py`)
|
||||
- Performs a **strict idle check**
|
||||
- Calls `tree_cache.attach_storage_backend(...)` / `detach_storage_backend(...)`
|
||||
@@ -36,11 +36,11 @@ The control path is:
|
||||
***
|
||||
## 2. Idle-state requirement (strict)
|
||||
|
||||
The Scheduler uses a stricter `_is_idle_for_hicache_storage_op()`:
|
||||
The Scheduler uses `is_fully_idle()` which checks:
|
||||
|
||||
- `_is_no_request()` is true (covers running/overlap/pp/disagg and other active states)
|
||||
- `waiting_queue` is empty
|
||||
- `grammar_queue` is empty (if the grammar backend is enabled)
|
||||
- No running batches (including chunked prefill, overlap, pipeline-parallel, and disaggregation paths)
|
||||
- No waiting requests in any queue (waiting, grammar, disagg bootstrap/prealloc/transfer/inflight)
|
||||
- No DLLM staging requests
|
||||
|
||||
If the condition is not met, attach/detach returns an error like:
|
||||
|
||||
|
||||
@@ -0,0 +1,187 @@
|
||||
---
|
||||
title: "HiSparse: Hierarchical Sparse Attention"
|
||||
metatags:
|
||||
description: "Use HiSparse hierarchical sparse attention to reduce decode GPU KV memory with CPU pinned host storage and PD disaggregation."
|
||||
---
|
||||
|
||||
HiSparse reduces per-request GPU memory consumption during the decode phase by maintaining only a small "hot" KV buffer on GPU while keeping complete KV data in CPU pinned memory. Combined with PD disaggregation, it enables significantly higher decode concurrency.
|
||||
|
||||
> **Prerequisites**: HiSparse only works with models that use **DeepSeek Sparse Attention (DSA)** architectures (e.g., DeepSeek-V3.2, GLM-5). These models natively select a subset of tokens for attention, making it possible to keep only the top-k KV on GPU while storing the full KV in host memory — without accuracy loss. Additionally, HiSparse currently requires **PD disaggregation mode** and is enabled on the **decode instance** only.
|
||||
|
||||
## Why HiSparse?
|
||||
|
||||
In long-context LLM inference, each decoding request holds a full-length KV cache on GPU, limiting the number of concurrent requests a decode instance can serve. HiSparse addresses this by:
|
||||
|
||||
- **Reducing GPU memory per request**: Each request occupies only a fixed-size device buffer (e.g., 4KB tokens) instead of the full sequence length.
|
||||
- **On-demand swap-in**: A CUDA kernel dynamically loads the top-k most relevant KV entries from host memory based on attention scores.
|
||||
- **Transparent to prefill**: HiSparse is entirely a decode-side optimization; the prefill instance requires no changes.
|
||||
|
||||
## Design Overview
|
||||
|
||||
### Decode Workflow
|
||||
|
||||
Each decode step follows this flow:
|
||||
|
||||
1. **Forward decode** — generate the next token
|
||||
2. **Top-k selection** — select the most relevant token positions via attention scores
|
||||
3. **Swap-in** — the CUDA kernel loads top-k KV entries from host to device buffer:
|
||||
- *Short sequences* (`seq_len ≤ device_buffer_size`): fast path, all KV already in buffer
|
||||
- *Long sequences*: hit detection → LRU reordering → miss handling (host → device copy)
|
||||
4. **Decode attention** — compute attention using the top-k device locations
|
||||
5. **Eager backup** — asynchronously copy the previous token's KV from device to host
|
||||
|
||||
### PD Disaggregation Integration (Direct-to-Host)
|
||||
|
||||
In PD disaggregation mode, the prefill instance transfers KV cache directly into the decode instance's host pool via RDMA, bypassing the GPU entirely on the decode side. This eliminates the transient GPU memory spike during KV transfer and removes the staging DMA step.
|
||||
|
||||
```
|
||||
Prefill GPU ──RDMA──▶ Decode Host Pool (CPU pinned memory)
|
||||
│
|
||||
▼
|
||||
alloc device buffer (4KB)
|
||||
│
|
||||
▼
|
||||
swap-in kernel (on-demand top-k)
|
||||
```
|
||||
|
||||
## Server Arguments
|
||||
|
||||
<table style={{width: "100%", borderCollapse: "collapse", tableLayout: "fixed"}}>
|
||||
<colgroup>
|
||||
<col style={{width: "33.33%"}} />
|
||||
<col style={{width: "33.33%"}} />
|
||||
<col style={{width: "33.33%"}} />
|
||||
</colgroup>
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Argument</th>
|
||||
<th>Type / Default</th>
|
||||
<th>Description</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td><code>--enable-hisparse</code></td>
|
||||
<td>flag; default: disabled</td>
|
||||
<td>Enable HiSparse on the decode instance</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>--hisparse-config</code></td>
|
||||
<td>JSON string</td>
|
||||
<td>Configuration for HiSparse (see below)</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
### HiSparse Config Parameters
|
||||
|
||||
Pass as a JSON string via `--hisparse-config`:
|
||||
|
||||
<table style={{width: "100%", borderCollapse: "collapse", tableLayout: "fixed"}}>
|
||||
<colgroup>
|
||||
<col style={{width: "33.33%"}} />
|
||||
<col style={{width: "33.33%"}} />
|
||||
<col style={{width: "33.33%"}} />
|
||||
</colgroup>
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Parameter</th>
|
||||
<th>Type / Default</th>
|
||||
<th>Description</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td><code>top_k</code></td>
|
||||
<td>int</td>
|
||||
<td>Number of topk entries</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>device_buffer_size</code></td>
|
||||
<td>int</td>
|
||||
<td>Number of token slots in the per-request GPU device buffer</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>host_to_device_ratio</code></td>
|
||||
<td>int</td>
|
||||
<td>Ratio of logical pool size to device pool size, determining host memory capacity</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
Example: `--hisparse-config='{"top_k": 2048, "device_buffer_size": 6144, "host_to_device_ratio": 10}'`
|
||||
|
||||
## Deployment
|
||||
|
||||
HiSparse currently requires **PD disaggregation mode** and is enabled only on the **decode instance**.
|
||||
|
||||
### Prefill Instance
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.launch_server \
|
||||
--model-path /path/to/model \
|
||||
--trust-remote-code \
|
||||
--port 8000 --host 0.0.0.0 \
|
||||
--context-length 81920 \
|
||||
--chunked-prefill-size 65536 \
|
||||
--tp-size 8 --dp-size 8 --enable-dp-attention \
|
||||
--mem-fraction-static 0.85 \
|
||||
--disaggregation-mode prefill \
|
||||
--disaggregation-ib-device mlx5_0,mlx5_1,mlx5_2,mlx5_3 \
|
||||
--nnodes 1 --node-rank 0
|
||||
```
|
||||
|
||||
### Decode Instance (with HiSparse)
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.launch_server \
|
||||
--model-path /path/to/model \
|
||||
--trust-remote-code \
|
||||
--port 8000 --host 0.0.0.0 \
|
||||
--context-length 81920 \
|
||||
--tp-size 8 --dp-size 8 --enable-dp-attention \
|
||||
--mem-fraction-static 0.85 \
|
||||
--kv-cache-dtype bfloat16 \
|
||||
--nsa-decode-backend flashmla_sparse \
|
||||
--disaggregation-mode decode \
|
||||
--disaggregation-ib-device mlx5_0,mlx5_1,mlx5_2,mlx5_3 \
|
||||
--dist-init-addr 127.0.0.1:5757 \
|
||||
--nnodes 1 --node-rank 0 \
|
||||
--enable-hisparse \
|
||||
--hisparse-config='{"top_k": 2048, "device_buffer_size": 6144, "host_to_device_ratio": 10}'
|
||||
```
|
||||
|
||||
### Benchmark
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.bench_serving \
|
||||
--backend sglang \
|
||||
--dataset-path /path/to/ShareGPT_V3_unfiltered_cleaned_split.json \
|
||||
--dataset-name random \
|
||||
--random-input 40000 \
|
||||
--random-output 20000 \
|
||||
--num-prompts 200 \
|
||||
--max-concurrency 200 \
|
||||
--request-rate 40 \
|
||||
--random-range-ratio 1.0 \
|
||||
--host 127.0.0.1 \
|
||||
--port 20000 \
|
||||
--model /path/to/model \
|
||||
--flush-cache \
|
||||
```
|
||||
|
||||
### Key Notes
|
||||
|
||||
- The prefill instance does not need `--enable-hisparse`; it is unaware of HiSparse.
|
||||
- On the decode instance, the following flags are **required** for HiSparse:
|
||||
- `--kv-cache-dtype bfloat16` — currently only bfloat16 KV cache is supported (more dtypes planned).
|
||||
- `--nsa-decode-backend flashmla_sparse` — currently only `flashmla_sparse` backend is supported.
|
||||
- `--enable-hisparse` — enables HiSparse.
|
||||
- `--hisparse-config` — HiSparse configuration (top_k, device_buffer_size, host_to_device_ratio).
|
||||
- `host_to_device_ratio` should be configured based on the host machine's available memory. For example:
|
||||
- **~1 TB** host memory → `host_to_device_ratio: 5`
|
||||
- **~2 TB** host memory → `host_to_device_ratio: 10`
|
||||
|
||||
## Acknowledgments
|
||||
|
||||
We would like to thank the SGLang team and community for the implementation and generous support, especially Zhiqiang Xie, Zhangheng Huang, Tingwei Huang, Shangming Cai, Teng Ma, and many others. We also thank the Alibaba Cloud TairKVCache team and the AntGroup SCT Inference team for their valuable contributions.
|
||||
@@ -102,7 +102,7 @@
|
||||
"\"\"\"\n",
|
||||
")\n",
|
||||
"\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")"
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=server_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -155,12 +155,12 @@
|
||||
"python3 -m sglang.launch_server --model-path meta-llama/Meta-Llama-3.1-8B-Instruct \\\n",
|
||||
" --enable-lora \\\n",
|
||||
" --lora-paths lora0=algoprog/fact-generation-llama-3.1-8b-instruct-lora \\\n",
|
||||
" lora1=Nutanix/Meta-Llama-3.1-8B-Instruct_lora_4_alpha_16 \\\n",
|
||||
" lora1=Nutanix/Meta-Llama-3.1-8B-Instruct_SFT_lora_4_alpha_16_humaneval_raw_json \\\n",
|
||||
" --max-loras-per-batch 2 \\\n",
|
||||
" --log-level warning \\\n",
|
||||
"\"\"\")\n",
|
||||
"\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")"
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=server_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -218,7 +218,7 @@
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"lora0 = \"Nutanix/Meta-Llama-3.1-8B-Instruct_lora_4_alpha_16\" # rank - 4, target modules - q_proj, k_proj, v_proj, o_proj, gate_proj\n",
|
||||
"lora0 = \"Nutanix/Meta-Llama-3.1-8B-Instruct_SFT_lora_4_alpha_16_humaneval_raw_json\" # rank - 4, target modules - q_proj, k_proj, v_proj, o_proj, gate_proj\n",
|
||||
"lora1 = \"algoprog/fact-generation-llama-3.1-8b-instruct-lora\" # rank - 64, target modules - q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj\n",
|
||||
"lora0_new = \"philschmid/code-llama-3-1-8b-text-to-sql-lora\" # rank - 256, target modules - q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj\n",
|
||||
"\n",
|
||||
@@ -236,7 +236,7 @@
|
||||
" \"\"\")\n",
|
||||
"\n",
|
||||
"url = f\"http://127.0.0.1:{port}\"\n",
|
||||
"wait_for_server(url)"
|
||||
"wait_for_server(url, process=server_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -404,7 +404,7 @@
|
||||
"source": [
|
||||
"### OpenAI-compatible API usage\n",
|
||||
"\n",
|
||||
"You can use LoRA adapters via the OpenAI-compatible APIs by specifying the adapter in the `model` field using the `base-model:adapter-name` syntax (for example, `qwen/qwen2.5-0.5b-instruct:adapter_a`). For more details and examples, see the “Using LoRA Adapters” section in the OpenAI API documentation: [openai_api_completions](../basic_usage/openai_api_completions).\n"
|
||||
"You can use LoRA adapters via the OpenAI-compatible APIs by specifying the adapter in the `model` field using the `base-model:adapter-name` syntax (for example, `qwen/qwen2.5-0.5b-instruct:adapter_a`). For more details and examples, see the “Using LoRA Adapters” section in the OpenAI API documentation: [openai_api_completions.ipynb](../basic_usage/openai_api_completions.ipynb).\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -439,7 +439,7 @@
|
||||
" --max-lora-rank 256 \\\n",
|
||||
" --lora-target-modules all \\\n",
|
||||
" --lora-paths \\\n",
|
||||
" {\"lora_name\":\"lora0\",\"lora_path\":\"Nutanix/Meta-Llama-3.1-8B-Instruct_lora_4_alpha_16\",\"pinned\":true} \\\n",
|
||||
" {\"lora_name\":\"lora0\",\"lora_path\":\"Nutanix/Meta-Llama-3.1-8B-Instruct_SFT_lora_4_alpha_16_humaneval_raw_json\",\"pinned\":true} \\\n",
|
||||
" {\"lora_name\":\"lora1\",\"lora_path\":\"algoprog/fact-generation-llama-3.1-8b-instruct-lora\"} \\\n",
|
||||
" lora2=philschmid/code-llama-3-1-8b-text-to-sql-lora\n",
|
||||
" --log-level warning\n",
|
||||
@@ -447,7 +447,7 @@
|
||||
"\n",
|
||||
"\n",
|
||||
"url = f\"http://127.0.0.1:{port}\"\n",
|
||||
"wait_for_server(url)"
|
||||
"wait_for_server(url, process=server_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -581,7 +581,7 @@
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"lora0 = \"Nutanix/Meta-Llama-3.1-8B-Instruct_lora_4_alpha_16\"\n",
|
||||
"lora0 = \"Nutanix/Meta-Llama-3.1-8B-Instruct_SFT_lora_4_alpha_16_humaneval_raw_json\"\n",
|
||||
"lora1 = \"algoprog/fact-generation-llama-3.1-8b-instruct-lora\"\n",
|
||||
"lora2 = \"philschmid/code-llama-3-1-8b-text-to-sql-lora\"\n",
|
||||
"\n",
|
||||
@@ -591,7 +591,7 @@
|
||||
" --model-path meta-llama/Meta-Llama-3.1-8B-Instruct \\\n",
|
||||
" --enable-lora \\\n",
|
||||
" --enable-lora-overlap-loading \\\n",
|
||||
" --lora-paths lora0=Nutanix/Meta-Llama-3.1-8B-Instruct_lora_4_alpha_16 \\\n",
|
||||
" --lora-paths lora0=Nutanix/Meta-Llama-3.1-8B-Instruct_SFT_lora_4_alpha_16_humaneval_raw_json \\\n",
|
||||
" lora1=algoprog/fact-generation-llama-3.1-8b-instruct-lora \\\n",
|
||||
" lora2=philschmid/code-llama-3-1-8b-text-to-sql-lora \\\n",
|
||||
" --max-lora-rank 256 \\\n",
|
||||
@@ -600,7 +600,7 @@
|
||||
" \"\"\")\n",
|
||||
"\n",
|
||||
"url = f\"http://127.0.0.1:{port}\"\n",
|
||||
"wait_for_server(url)"
|
||||
"wait_for_server(url, process=server_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
|
||||
@@ -43,7 +43,7 @@ From client side, the user needs to provide a list of strings as input batch, an
|
||||
|
||||
**Note:** SGLang supports LoRA adapters through two APIs:
|
||||
|
||||
1. **OpenAI-Compatible API** (`/v1/chat/completions`, `/v1/completions`): Use the `model:adapter-name` syntax. See [OpenAI API with LoRA](../basic_usage/openai_api_completions.ipynb#Using-LoRA-Adapters) for examples.
|
||||
1. **OpenAI-Compatible API** (`/v1/chat/completions`, `/v1/completions`): Use the `model:adapter-name` syntax. See [OpenAI API with LoRA](../basic_usage/openai_api_completions#using-lora-adapters) for examples.
|
||||
|
||||
2. **Native API** (`/generate`): Pass `lora_path` in the request body (shown below).
|
||||
|
||||
@@ -108,7 +108,7 @@ server_process, port = launch_server_cmd(
|
||||
python3 -m sglang.launch_server --model-path meta-llama/Meta-Llama-3.1-8B-Instruct \
|
||||
--enable-lora \
|
||||
--lora-paths lora0=algoprog/fact-generation-llama-3.1-8b-instruct-lora \
|
||||
lora1=Nutanix/Meta-Llama-3.1-8B-Instruct_lora_4_alpha_16 \
|
||||
lora1=Nutanix/Meta-Llama-3.1-8B-Instruct_SFT_lora_4_alpha_16_humaneval_raw_json \
|
||||
--max-loras-per-batch 2 \
|
||||
--log-level warning \
|
||||
"""
|
||||
@@ -152,7 +152,7 @@ When using dynamic LoRA loading, it's recommended to explicitly specify both `--
|
||||
|
||||
|
||||
```python Example
|
||||
lora0 = "Nutanix/Meta-Llama-3.1-8B-Instruct_lora_4_alpha_16" # rank - 4, target modules - q_proj, k_proj, v_proj, o_proj, gate_proj
|
||||
lora0 = "Nutanix/Meta-Llama-3.1-8B-Instruct_SFT_lora_4_alpha_16_humaneval_raw_json" # rank - 4, target modules - q_proj, k_proj, v_proj, o_proj, gate_proj
|
||||
lora1 = "algoprog/fact-generation-llama-3.1-8b-instruct-lora" # rank - 64, target modules - q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
|
||||
lora0_new = "philschmid/code-llama-3-1-8b-text-to-sql-lora" # rank - 256, target modules - q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
|
||||
|
||||
@@ -317,7 +317,7 @@ server_process, port = launch_server_cmd(
|
||||
--max-lora-rank 256 \
|
||||
--lora-target-modules all \
|
||||
--lora-paths \
|
||||
{"lora_name":"lora0","lora_path":"Nutanix/Meta-Llama-3.1-8B-Instruct_lora_4_alpha_16","pinned":true} \
|
||||
{"lora_name":"lora0","lora_path":"Nutanix/Meta-Llama-3.1-8B-Instruct_SFT_lora_4_alpha_16_humaneval_raw_json","pinned":true} \
|
||||
{"lora_name":"lora1","lora_path":"algoprog/fact-generation-llama-3.1-8b-instruct-lora"} \
|
||||
lora2=philschmid/code-llama-3-1-8b-text-to-sql-lora
|
||||
--log-level warning
|
||||
@@ -418,7 +418,7 @@ By using the `--enable-lora-overlap-loading` server argument, the SGLang engine
|
||||
|
||||
|
||||
```python Example
|
||||
lora0 = "Nutanix/Meta-Llama-3.1-8B-Instruct_lora_4_alpha_16"
|
||||
lora0 = "Nutanix/Meta-Llama-3.1-8B-Instruct_SFT_lora_4_alpha_16_humaneval_raw_json"
|
||||
lora1 = "algoprog/fact-generation-llama-3.1-8b-instruct-lora"
|
||||
lora2 = "philschmid/code-llama-3-1-8b-text-to-sql-lora"
|
||||
|
||||
@@ -429,7 +429,7 @@ server_process, port = launch_server_cmd(
|
||||
--model-path meta-llama/Meta-Llama-3.1-8B-Instruct \
|
||||
--enable-lora \
|
||||
--enable-lora-overlap-loading \
|
||||
--lora-paths lora0=Nutanix/Meta-Llama-3.1-8B-Instruct_lora_4_alpha_16 \
|
||||
--lora-paths lora0=Nutanix/Meta-Llama-3.1-8B-Instruct_SFT_lora_4_alpha_16_humaneval_raw_json \
|
||||
lora1=algoprog/fact-generation-llama-3.1-8b-instruct-lora \
|
||||
lora2=philschmid/code-llama-3-1-8b-text-to-sql-lora \
|
||||
--max-lora-rank 256 \
|
||||
|
||||
@@ -26,7 +26,7 @@ When you need to profile prefill or decode workers in PD disaggregation mode, pl
|
||||
|
||||
## Router Integration
|
||||
|
||||
For deploying PD disaggregation at scale with load balancing and fault tolerance, SGLang provides a router. The router can distribute requests between prefill and decode instances using various routing policies. For detailed information on setting up routing with PD disaggregation, including configuration options and deployment patterns, see the [SGLang Model Gateway (former Router)](../advanced_features/sgl_model_gateway.md#prefill-decode-disaggregation).
|
||||
For deploying PD disaggregation at scale with load balancing and fault tolerance, SGLang provides a router. The router can distribute requests between prefill and decode instances using various routing policies. For detailed information on setting up routing with PD disaggregation, including configuration options and deployment patterns, see the [SGLang Model Gateway (former Router)](./sgl_model_gateway#prefill-decode-disaggregation).
|
||||
|
||||
|
||||
## Mooncake
|
||||
@@ -132,11 +132,13 @@ PD Disaggregation with Mooncake supports the following environment variables for
|
||||
#### NVLink Transport Configuration
|
||||
To enable NVLink transport for KV cache transfers with the mooncake backend (recommended for NVL72 deployments), set the following environment variables. Note that auxiliary data transfer will still use TCP as a temporary workaround.
|
||||
|
||||
```bash
|
||||
export SGLANG_MOONCAKE_CUSTOM_MEM_POOL=True
|
||||
```bash Command
|
||||
export SGLANG_MOONCAKE_CUSTOM_MEM_POOL=NVLINK
|
||||
export MC_FORCE_MNNVL=True
|
||||
```
|
||||
|
||||
The `SGLANG_MOONCAKE_CUSTOM_MEM_POOL` environment variable enables the custom memory pool. Supported values are `NVLINK` (or `True`), `BAREX`, and `INTRA_NODE_NVLINK`.
|
||||
|
||||
#### Prefill Server Configuration
|
||||
<table style={{width: "100%", borderCollapse: "collapse", tableLayout: "fixed"}}>
|
||||
<colgroup>
|
||||
@@ -155,11 +157,11 @@ export MC_FORCE_MNNVL=True
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>**`SGLANG_DISAGGREGATION_THREAD_POOL_SIZE`**</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Controls the total number of worker threads for KVCache transfer operations per TP rank</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>A dynamic value calculated by `int(0.75 * os.cpu_count()) // 8)`, which is limited to be larger than 4 and less than 12 to ensure efficiency and prevent thread race conditions</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>A dynamic value calculated by <code>int(0.75 * os.cpu_count()) // 8)</code>, which is limited to be larger than 4 and less than 12 to ensure efficiency and prevent thread race conditions</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>**`SGLANG_DISAGGREGATION_QUEUE_SIZE`**</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Sets the number of parallel transfer queues. KVCache transfer requests from multiple decode instances will be sharded into these queues so that they can share the threads and the transfer bandwidth at the same time. If it is set to `1`, then we transfer requests one by one according to fcfs strategy</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Sets the number of parallel transfer queues. KVCache transfer requests from multiple decode instances will be sharded into these queues so that they can share the threads and the transfer bandwidth at the same time. If it is set to <code>1</code>, then we transfer requests one by one according to fcfs strategy</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>`4`</td>
|
||||
</tr>
|
||||
<tr>
|
||||
@@ -167,7 +169,12 @@ export MC_FORCE_MNNVL=True
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Timeout (seconds) for receiving destination KV indices during request initialization</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>`300`</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong><code>SGLANG_DISAGGREGATION_BOOTSTRAP_ENTRY_CLEANUP_INTERVAL</code></strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Interval (seconds) between cleanups of bootstrap entries</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}><code>120</code></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
If a greater mean TTFT is acceptable, you can `export SGLANG_DISAGGREGATION_BOOTSTRAP_TIMEOUT=600` (10 minutes) to relax the timeout condition.
|
||||
@@ -209,6 +216,84 @@ Please be aware that this setting will cause prefill instances to take a longer
|
||||
If a greater mean TTFT is acceptable, you can `export SGLANG_DISAGGREGATION_WAITING_TIMEOUT=600` (10 minutes) to relax the timeout condition.
|
||||
|
||||
|
||||
## Heterogeneous TP with GPU Staging Buffer
|
||||
|
||||
When prefill and decode use different tensor parallelism (TP) sizes (e.g., prefill TP=4, decode DP attention with TP=1), the KV cache memory layout differs between the two sides. The **GPU staging buffer** solves this by gathering KV head slices into a contiguous buffer on the prefill side, performing bulk RDMA transfer, then scattering into the correct KV cache pages on the decode side. This provides **2–5x throughput improvement** over the default per-token slice approach at high concurrency and matches homogeneous TP baselines within ~5%.
|
||||
|
||||
Enable the staging buffer when prefill and decode use **different TP sizes** with the **Mooncake** transfer backend. When both sides use the same TP size, staging is automatically bypassed even if enabled.
|
||||
|
||||
> **Note:** The staging buffer is designed for non-MLA models (e.g. GQA, MHA). MLA models (e.g. DeepSeek-V2/V3) should not enable this flag.
|
||||
|
||||
### Environment Variables
|
||||
|
||||
<table style={{width: "100%", borderCollapse: "collapse", tableLayout: "fixed"}}>
|
||||
<colgroup>
|
||||
<col style={{width: "30%"}} />
|
||||
<col style={{width: "50%"}} />
|
||||
<col style={{width: "20%"}} />
|
||||
</colgroup>
|
||||
<thead>
|
||||
<tr style={{borderBottom: "2px solid #d55816"}}>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.02)"}}>Variable</th>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.05)"}}>Description</th>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.02)"}}>Default</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong><code>SGLANG_DISAGG_STAGING_BUFFER</code></strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Enable GPU staging buffer for heterogeneous TP KV transfer</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}><code>False</code></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong><code>SGLANG_DISAGG_STAGING_BUFFER_SIZE_MB</code></strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Prefill-side per-worker staging buffer size in MB</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}><code>64</code></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong><code>SGLANG_DISAGG_STAGING_POOL_SIZE_MB</code></strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Decode-side ring buffer pool total size in MB</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}><code>4096</code></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
### Usage Example
|
||||
|
||||
```bash Command
|
||||
# Set staging buffer environment variables on BOTH prefill and decode
|
||||
export SGLANG_DISAGG_STAGING_BUFFER=1
|
||||
export SGLANG_DISAGG_STAGING_BUFFER_SIZE_MB=64
|
||||
export SGLANG_DISAGG_STAGING_POOL_SIZE_MB=4096
|
||||
|
||||
# Prefill with TP=4
|
||||
python -m sglang.launch_server \
|
||||
--model-path $MODEL_PATH \
|
||||
--disaggregation-mode prefill \
|
||||
--port 30000 \
|
||||
--tp 4 \
|
||||
--trust-remote-code \
|
||||
--disaggregation-ib-device mlx5_1,mlx5_2
|
||||
|
||||
# Decode with TP=1 (or DP attention with effective attention TP=1)
|
||||
python -m sglang.launch_server \
|
||||
--model-path $MODEL_PATH \
|
||||
--disaggregation-mode decode \
|
||||
--port 30001 \
|
||||
--tp 4 \
|
||||
--dp 4 \
|
||||
--enable-dp-attention \
|
||||
--trust-remote-code \
|
||||
--disaggregation-ib-device mlx5_3,mlx5_4
|
||||
|
||||
# Router
|
||||
python -m sglang_router.launch_router \
|
||||
--pd-disaggregation \
|
||||
--prefill http://127.0.0.1:30000 \
|
||||
--decode http://127.0.0.1:30001 \
|
||||
--host 0.0.0.0 --port 8000
|
||||
```
|
||||
|
||||
## NIXL
|
||||
### Requirements
|
||||
|
||||
@@ -343,8 +428,8 @@ python -m sglang.launch_server \
|
||||
|
||||
Use ascend backend with [memfabric_hybrid](https://gitcode.com/Ascend/memfabric_hybrid) and ASCEND_MF_STORE_URL being set
|
||||
|
||||
```bash
|
||||
pip install memfabric-hybrid==1.0.5
|
||||
```bash Command
|
||||
pip install memfabric-hybrid==1.0.0
|
||||
export ASCEND_MF_STORE_URL="tcp://xxx.xx.xxx.xxx:xxxx"
|
||||
```
|
||||
Use mooncake backend, more details can be found in mooncake section.
|
||||
|
||||
@@ -20,11 +20,276 @@ or [NeuralMagic](https://huggingface.co/collections/neuralmagic) collections on
|
||||
popular quality validated quantized models. Quantized models must be validated via benchmarks post-quantization
|
||||
to guard against abnormal quantization loss regressions.
|
||||
|
||||
## Platform Compatibility
|
||||
|
||||
The following table summarizes quantization method support across NVIDIA and AMD GPUs, Ascend NPUs.
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Method</th>
|
||||
<th>NVIDIA GPUs</th>
|
||||
<th>AMD GPUs (MI300X/MI325X/MI350X)</th>
|
||||
<th>Ascend NPUs (A2/A3)</th>
|
||||
<th>Notes</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td><code>fp8</code></td>
|
||||
<td>Yes</td>
|
||||
<td>Yes</td>
|
||||
<td>WIP</td>
|
||||
<td>Aiter or Triton backend on AMD</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>mxfp4</code></td>
|
||||
<td>Yes</td>
|
||||
<td>Yes</td>
|
||||
<td>WIP</td>
|
||||
<td>Requires CDNA3/CDNA4 with MXFP support; uses Aiter</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>blockwise_int8</code></td>
|
||||
<td>Yes</td>
|
||||
<td>Yes</td>
|
||||
<td>No</td>
|
||||
<td>Triton-based, works on both platforms</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>w8a8_int8</code></td>
|
||||
<td>Yes</td>
|
||||
<td>Yes</td>
|
||||
<td>No</td>
|
||||
<td></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>w8a8_fp8</code></td>
|
||||
<td>Yes</td>
|
||||
<td>Yes</td>
|
||||
<td>No</td>
|
||||
<td>Aiter or Triton FP8 on AMD</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>awq</code></td>
|
||||
<td>Yes</td>
|
||||
<td>Yes</td>
|
||||
<td>Yes</td>
|
||||
<td>Uses Triton dequantize on AMD (vs. optimized CUDA kernels on NVIDIA). Uses CANN kernels on Ascend</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>gptq</code></td>
|
||||
<td>Yes</td>
|
||||
<td>Yes</td>
|
||||
<td>Yes</td>
|
||||
<td>Uses Triton or vLLM kernels on AMD. Uses CANN kernels on Ascend</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>compressed-tensors</code></td>
|
||||
<td>Yes</td>
|
||||
<td>Yes</td>
|
||||
<td>Partial</td>
|
||||
<td>Aiter paths for FP8/MoE on AMD. Uses CANN kernels on Ascend, <code>FP8</code> not supported yet</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>quark</code></td>
|
||||
<td>Yes</td>
|
||||
<td>Yes</td>
|
||||
<td>No</td>
|
||||
<td>AMD Quark quantization; Aiter GEMM paths on AMD</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>auto-round</code></td>
|
||||
<td>Yes</td>
|
||||
<td>Yes</td>
|
||||
<td>Partial</td>
|
||||
<td>Platform-agnostic (Intel auto-round). Uses CANN kernels on Ascend</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>quark_int4fp8_moe</code></td>
|
||||
<td>No</td>
|
||||
<td>Yes</td>
|
||||
<td>No</td>
|
||||
<td>AMD-only; online INT4-to-FP8 MoE quantization (CDNA3/CDNA4)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>awq_marlin</code></td>
|
||||
<td>Yes</td>
|
||||
<td>No</td>
|
||||
<td>No</td>
|
||||
<td>Marlin kernels are CUDA-only</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>gptq_marlin</code></td>
|
||||
<td>Yes</td>
|
||||
<td>No</td>
|
||||
<td>No</td>
|
||||
<td>Marlin kernels are CUDA-only</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>gguf</code></td>
|
||||
<td>Yes</td>
|
||||
<td>No</td>
|
||||
<td>WIP</td>
|
||||
<td>CUDA-only kernels in sgl-kernel</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>modelopt</code> / <code>modelopt_fp8</code></td>
|
||||
<td>Yes (Hopper/SM90+)</td>
|
||||
<td>No</td>
|
||||
<td>No</td>
|
||||
<td><a href="https://github.com/NVIDIA/Model-Optimizer">NVIDIA ModelOpt</a>; requires NVIDIA hardware</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>modelopt_fp4</code></td>
|
||||
<td>Yes (Blackwell/SM100+)</td>
|
||||
<td>No</td>
|
||||
<td>No</td>
|
||||
<td><a href="https://github.com/NVIDIA/Model-Optimizer">NVIDIA ModelOpt</a>; native FP4 on Blackwell (B200, GB200)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>petit_nvfp4</code></td>
|
||||
<td>No</td>
|
||||
<td>Yes (MI250/MI300X/MI325X)</td>
|
||||
<td>No</td>
|
||||
<td>Enables NVFP4 on ROCm via <a href="https://github.com/causalflow-ai/petit-kernel">Petit</a>; use <code>modelopt_fp4</code> on NVIDIA Blackwell. Auto-selected when loading NVFP4 models on AMD. See <a href="https://lmsys.org/blog/2025-09-21-petit-amdgpu/">LMSYS blog</a> and <a href="https://rocm.blogs.amd.com/artificial-intelligence/fp4-mixed-precision/README.html">AMD ROCm blog</a>.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>bitsandbytes</code></td>
|
||||
<td>Yes</td>
|
||||
<td>Experimental</td>
|
||||
<td>No</td>
|
||||
<td>Depends on bitsandbytes ROCm support</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>torchao</code> (<code>int4wo</code>, etc.)</td>
|
||||
<td>Yes</td>
|
||||
<td>Partial</td>
|
||||
<td>No</td>
|
||||
<td><code>int4wo</code> not supported on AMD; other methods may work</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>modelslim</code></td>
|
||||
<td>No</td>
|
||||
<td>No</td>
|
||||
<td>Yes</td>
|
||||
<td>Ascend quantization; Uses CANN kernels</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
On AMD, several of these methods use [Aiter](https://github.com/ROCm/aiter) for acceleration -- set `SGLANG_USE_AITER=1` where noted. See [AMD GPU setup](../hardware-platforms/amd_gpu) for installation and configuration details.
|
||||
|
||||
On Ascend, various layers quantization configurations are supported, see [Ascend NPU quantization](../hardware-platforms/ascend-npus/ascend_npu_quantization) for details.
|
||||
|
||||
## GEMM Backends for FP4/FP8 Quantization
|
||||
|
||||
<Note>
|
||||
Backend selection is supported only for **blockwise FP8** and **NVFP4** GEMM. When running FP8 or FP4 quantized models, you can select the GEMM backend via `--fp8-gemm-backend` and `--fp4-gemm-backend`.
|
||||
</Note>
|
||||
|
||||
### `--fp8-gemm-backend` (Blockwise FP8 GEMM)
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Backend</th>
|
||||
<th>Hardware</th>
|
||||
<th>Description</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td><code>auto</code></td>
|
||||
<td>All</td>
|
||||
<td>Auto-selects based on hardware</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>deep_gemm</code></td>
|
||||
<td>SM90, SM100</td>
|
||||
<td>JIT-compiled; enabled when DeepGEMM is installed</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>flashinfer_trtllm</code></td>
|
||||
<td>SM100</td>
|
||||
<td>FlashInfer TensorRT-LLM backend; optimal for low-latency</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>flashinfer_cutlass</code></td>
|
||||
<td>SM100/120</td>
|
||||
<td>FlashInfer CUTLASS groupwise FP8 GEMM</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>flashinfer_deepgemm</code></td>
|
||||
<td>SM90</td>
|
||||
<td>Uses swapAB optimization for small M dimensions in decoding</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>cutlass</code></td>
|
||||
<td>SM90, SM100/120</td>
|
||||
<td>sgl-kernel CUTLASS</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>triton</code></td>
|
||||
<td>All</td>
|
||||
<td>Fallback; widely compatible</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>aiter</code></td>
|
||||
<td>ROCm</td>
|
||||
<td>AMD AITER backend</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
**`auto` selection order:** 1) DeepGEMM (SM90/SM100, installed); 2) FlashInfer TRTLLM (SM100, FlashInfer available); 3) CUTLASS (SM90/SM100/120); 4) AITER (AMD); 5) Triton. **Exception:** SM120 always resolves to Triton.
|
||||
|
||||
### `--fp4-gemm-backend` (NVFP4 GEMM)
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Backend</th>
|
||||
<th>Hardware</th>
|
||||
<th>Description</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td><code>auto</code></td>
|
||||
<td>SM100/120</td>
|
||||
<td>Auto-selects: <code>flashinfer_cudnn</code> on SM120; <code>flashinfer_cutlass</code> on SM100</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>cutlass</code></td>
|
||||
<td>SM100/120</td>
|
||||
<td>SGLang CUTLASS kernel</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>flashinfer_cutlass</code></td>
|
||||
<td>SM100/120</td>
|
||||
<td>FlashInfer CUTLASS backend</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>flashinfer_cudnn</code></td>
|
||||
<td>SM100/120 (CUDA 13+, cuDNN 9.15+)</td>
|
||||
<td>FlashInfer cuDNN backend; used on SM120 for performance</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>flashinfer_trtllm</code></td>
|
||||
<td>SM100</td>
|
||||
<td>FlashInfer TensorRT-LLM backend</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
When FlashInfer is unavailable for NVFP4, the SGLang CUTLASS kernel is used as an automatic fallback.
|
||||
|
||||
## Offline Quantization
|
||||
|
||||
To load already quantized models, simply load the model weights and config. **Again, if the model has been quantized offline,
|
||||
there's no need to add `--quantization` argument when starting the engine. The quantization method will be parsed from the
|
||||
downloaded Hugging Face config. For example, DeepSeek V3/R1 models are already in FP8, so do not add redundant parameters.**
|
||||
downloaded Hugging Face or msModelSlim config. For example, DeepSeek V3/R1 models are already in FP8, so do not add redundant parameters.**
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.launch_server \
|
||||
@@ -194,23 +459,85 @@ python3 -m sglang.launch_server \
|
||||
|
||||
#### Using [NVIDIA ModelOpt](https://github.com/NVIDIA/Model-Optimizer)
|
||||
|
||||
NVIDIA Model Optimizer (ModelOpt) provides advanced quantization techniques optimized for NVIDIA hardware. SGLang includes a streamlined workflow for quantizing models with ModelOpt and automatically exporting them for deployment.
|
||||
NVIDIA Model Optimizer (ModelOpt) provides advanced quantization techniques optimized for NVIDIA hardware.
|
||||
|
||||
**Offline vs. Online Quantization:**
|
||||
|
||||
SGLang supports two modes for ModelOpt.
|
||||
|
||||
* **Offline Quantization (pre-quantized):**
|
||||
* **Usage:** Download a pre-quantized model from Hugging Face or run `hf_ptq.py` once to create a new quantized checkpoint. Then load this quantized checkpoint.
|
||||
* **Pros:** Fast server startup, quantization can be validated before deployment, efficient resource usage.
|
||||
* **Cons:** Requires an extra preparation step.
|
||||
|
||||
* **Online Quantization (quant and serve):**
|
||||
* **Usage:** Load a standard BF16/FP16 model and add a flag. The engine applies quantization *on startup*.
|
||||
* **Pros:** Convenient (no new checkpoint needed).
|
||||
* **Cons:** **High startup time**, increases VRAM usage during initialization (risk of OOM).
|
||||
|
||||
The following sections guide you through using the Offline path: loading pre-quantized models or creating your own checkpoints.
|
||||
|
||||
##### Using Pre-Quantized Checkpoints
|
||||
|
||||
If a model is already quantized (e.g., from Hugging Face), you can load it directly.
|
||||
|
||||
* **FP8 Models:**
|
||||
Use `--quantization modelopt_fp8`.
|
||||
```bash Command
|
||||
python3 -m sglang.launch_server \
|
||||
--model-path nvidia/Llama-3.1-8B-Instruct-FP8 \
|
||||
--quantization modelopt_fp8 \
|
||||
--port 30000
|
||||
```
|
||||
|
||||
* **FP4 Models:**
|
||||
Use `--quantization modelopt_fp4`.
|
||||
```bash Command
|
||||
python3 -m sglang.launch_server \
|
||||
--model-path nvidia/Llama-3.3-70B-Instruct-NVFP4 \
|
||||
--quantization modelopt_fp4 \
|
||||
--port 30000
|
||||
```
|
||||
|
||||
##### Creating Your Own Quantized Checkpoints
|
||||
|
||||
If a pre-quantized checkpoint is not available for your model, you can create one using NVIDIA Model Optimizer's `hf_ptq.py` script.
|
||||
|
||||
**Why quantize?**
|
||||
- Reduce VRAM usage
|
||||
- Higher throughput and lower latency
|
||||
- More flexible deployment (on smaller GPUs)
|
||||
|
||||
**What can be quantized?**
|
||||
- The entire model
|
||||
- MLP layers only
|
||||
- KV cache
|
||||
|
||||
**Key options in `hf_ptq.py`:**
|
||||
|
||||
`--qformat`: Quantization formats `fp8`, `nvfp4`, `nvfp4_mlp_only`
|
||||
|
||||
`--kv_cache_qformat`: KV cache quantization format (default: `fp8`)
|
||||
|
||||
**Note:** The default `kv_cache_qformat` may not be optimal for all use cases. Consider setting this explicitly.
|
||||
|
||||
**Hardware requirements:** Hopper and higher are recommended. Insufficient GPU memory may cause weight offloading, resulting in extremely long quantization time.
|
||||
|
||||
For detailed usage and supported model architectures, see [NVIDIA Model Optimizer LLM PTQ](https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/llm_ptq).
|
||||
|
||||
SGLang includes a streamlined workflow for quantizing models with ModelOpt and automatically exporting them for deployment.
|
||||
|
||||
##### Installation
|
||||
|
||||
First, install ModelOpt. You can either install it directly or as an optional SGLang dependency:
|
||||
First, install ModelOpt:
|
||||
|
||||
```bash Command
|
||||
# Option 1: Install ModelOpt directly
|
||||
pip install nvidia-modelopt
|
||||
|
||||
# Option 2: Install SGLang with ModelOpt support (recommended)
|
||||
pip install sglang[modelopt]
|
||||
```
|
||||
|
||||
##### Quantization and Export Workflow
|
||||
|
||||
SGLang provides an example script that demonstrates the complete ModelOpt quantization and export workflow:
|
||||
SGLang provides an example script that demonstrates the complete ModelOpt quantization and export workflow. Run from the SGLang repository root (see [modelopt_quantize_and_export.py](https://github.com/sgl-project/sglang/blob/main/examples/usage/modelopt_quantize_and_export.py)):
|
||||
|
||||
```bash Command
|
||||
# Quantize and export a model using ModelOpt FP8 quantization
|
||||
@@ -219,7 +546,7 @@ python examples/usage/modelopt_quantize_and_export.py quantize \
|
||||
--export-dir ./quantized_tinyllama_fp8 \
|
||||
--quantization-method modelopt_fp8
|
||||
|
||||
# For FP4 quantization
|
||||
# For FP4 quantization (requires Blackwell GPU)
|
||||
python examples/usage/modelopt_quantize_and_export.py quantize \
|
||||
--model-path TinyLlama/TinyLlama-1.1B-Chat-v1.0 \
|
||||
--export-dir ./quantized_tinyllama_fp4 \
|
||||
@@ -275,25 +602,39 @@ python -m sglang.launch_server \
|
||||
--port 30000 --host 0.0.0.0
|
||||
```
|
||||
|
||||
Or using the Python API:
|
||||
Or using the Python API (use the same path as `modelopt_export_path` from the quantize step):
|
||||
|
||||
```python Example
|
||||
import sglang as sgl
|
||||
|
||||
# Deploy exported ModelOpt quantized model
|
||||
llm = sgl.Engine(
|
||||
model_path="./quantized_tinyllama_fp8",
|
||||
quantization="modelopt"
|
||||
)
|
||||
def main():
|
||||
# Deploy exported ModelOpt quantized model
|
||||
# Path must match modelopt_export_path from quantize step (e.g., ./exported_model)
|
||||
llm = sgl.Engine(
|
||||
model_path="./exported_model",
|
||||
quantization="modelopt",
|
||||
)
|
||||
|
||||
# Run inference
|
||||
prompts = ["Hello, how are you?", "What is the capital of France?"]
|
||||
sampling_params = {"temperature": 0.8, "top_p": 0.95, "max_new_tokens": 100}
|
||||
outputs = llm.generate(prompts, sampling_params)
|
||||
# Run inference
|
||||
prompts = [
|
||||
"Hello, how are you?",
|
||||
"What is the capital of France?",
|
||||
]
|
||||
sampling_params = {
|
||||
"temperature": 0.8,
|
||||
"top_p": 0.95,
|
||||
"max_new_tokens": 100,
|
||||
}
|
||||
|
||||
outputs = llm.generate(prompts, sampling_params)
|
||||
|
||||
for i, output in enumerate(outputs):
|
||||
print(f"Prompt: {prompts[i]}")
|
||||
print(f"Output: {output['text']}")
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
|
||||
for i, output in enumerate(outputs):
|
||||
print(f"Prompt: {prompts[i]}")
|
||||
print(f"Output: {output.outputs[0].text}")
|
||||
```
|
||||
|
||||
##### Advanced Features
|
||||
@@ -311,7 +652,7 @@ python examples/usage/modelopt_quantize_and_export.py quantize \
|
||||
# The checkpoint can be reused for future quantization runs and skip calibration
|
||||
```
|
||||
|
||||
**Export-only Workflow**: If you have a pre-existing fake quantized ModelOpt checkpoint, you can export it directly:
|
||||
**Export-only Workflow**: If you have a pre-existing fake quantized ModelOpt checkpoint, you can export it directly. See [LoadConfig](https://github.com/sgl-project/sglang/blob/main/python/sglang/srt/configs/load_config.py) for the full API:
|
||||
|
||||
```python Example
|
||||
from sglang.srt.configs.device_config import DeviceConfig
|
||||
@@ -330,7 +671,7 @@ load_config = LoadConfig(
|
||||
modelopt_export_path="./exported_model",
|
||||
)
|
||||
|
||||
# Load and export the model
|
||||
# Load and export the model (DeviceConfig defaults to device="cuda")
|
||||
model_loader = get_model_loader(load_config, model_config)
|
||||
model_loader.load_model(model_config=model_config, device_config=DeviceConfig())
|
||||
```
|
||||
@@ -343,6 +684,74 @@ model_loader.load_model(model_config=model_config, device_config=DeviceConfig())
|
||||
- **Calibration-based**: Uses calibration datasets for optimal quantization quality
|
||||
- **Production Ready**: Enterprise-grade quantization with NVIDIA support
|
||||
|
||||
#### Using [ModelSlim](https://gitcode.com/Ascend/msmodelslim)
|
||||
MindStudio-ModelSlim (msModelSlim) is a model offline quantization compression tool launched by MindStudio and optimized for Ascend hardware.
|
||||
|
||||
- **Installation**
|
||||
|
||||
```bash Command
|
||||
# Clone repo and install msmodelslim:
|
||||
git clone https://gitcode.com/Ascend/msmodelslim.git
|
||||
cd msmodelslim
|
||||
bash install.sh
|
||||
```
|
||||
|
||||
- **LLM quantization**
|
||||
|
||||
Download the original floating-point weights of the large model. Taking Qwen3-32B as an example, you can go to [Qwen3-32B](https://huggingface.co/Qwen/Qwen3-32B) to obtain the original model weights. Then install other dependencies (related to the model, refer to the huggingface model card).
|
||||
> Note: You can find pre-quantized validated models on [modelscope/Eco-Tech](https://modelscope.cn/models/Eco-Tech).
|
||||
|
||||
_Traditional quantification methods require the preparation of calibration data files (```.jsonl``` formats) for calibration in the quantification process._
|
||||
```bash Command
|
||||
Qwen3-32B/ # floating-point model downloaded from official HF (or modelscope) repo
|
||||
msmodelslim/ # msmodelslim repo
|
||||
|----- lab_calib # calibration date folder (put your dataset here in ```.jsonl``` format or use pre-prepared ones)
|
||||
|----- some file (such as laos_calib.jsonl)
|
||||
|----- lab_practice # best practice folder with configs for quantization
|
||||
|----- model folder (such as qwen3_5_moe folder) # folder with quantization configs
|
||||
|----- quant_config (such as qwen3_5_moe_w8a8.yaml) # quantization config
|
||||
|----- another folders
|
||||
output_folder/ # generated by below command
|
||||
|----- quant_model_weights-00001-of-0001.safetensors # quantized weights
|
||||
|----- quant_model_description.json # file with description of the quantization methods for each layer (```W4A4_DYNAMIC```, etc.)
|
||||
|----- another files (such as config.json, tokenizer.json, etc.)
|
||||
```
|
||||
Run quantization using one-click quantization (recommended):
|
||||
```bash Command
|
||||
msmodelslim quant \
|
||||
--model_path ${MODEL_PATH} \
|
||||
--save_path ${SAVE_PATH} \
|
||||
--device npu:0,1 \
|
||||
--model_type Qwen3-32B \
|
||||
--quant_type w8a8 \
|
||||
--trust_remote_code True
|
||||
```
|
||||
|
||||
- **Usage Example**
|
||||
```bash Command
|
||||
python3 -m sglang.launch_server \
|
||||
--model-path $PWD/Qwen3-32B-w8a8 \
|
||||
--port 30000 --host 0.0.0.0
|
||||
```
|
||||
|
||||
- **Available Quantization Methods**:
|
||||
- [x] ```W4A4_DYNAMIC``` linear with online quantization of activations
|
||||
- [x] ```W8A8``` linear with offline quantization of activations
|
||||
- [x] ```W8A8_DYNAMIC``` linear with online quantization of activations
|
||||
- [x] ```W4A4_DYNAMIC``` MOE with online quantization of activations
|
||||
- [x] ```W4A8_DYNAMIC``` MOE with online quantization of activations
|
||||
- [x] ```W8A8_DYNAMIC``` MOE with online quantization of activations
|
||||
- [ ] ```W4A8``` linear TBD
|
||||
- [ ] ```W4A16``` linear TBD
|
||||
- [ ] ```W48A16``` linear TBD
|
||||
- [ ] ```W4A16``` MoE in progress
|
||||
- [ ] ```W8A16``` MoE in progress
|
||||
- [ ] ```KV Cache``` in progress
|
||||
- [ ] ```Attention``` in progress
|
||||
|
||||
|
||||
For more detailed examples of quantization of models, as well as information about their support, see the [examples](https://gitcode.com/Ascend/msmodelslim/blob/master/example/README.md) section in ModelSLim repo.
|
||||
|
||||
## Online Quantization
|
||||
|
||||
To enable online quantization, you can simply specify `--quantization` in the command line. For example, you can launch the server with the following command to enable `FP8` quantization for model `meta-llama/Meta-Llama-3.1-8B-Instruct`:
|
||||
@@ -381,7 +790,7 @@ python3 -m sglang.launch_server \
|
||||
|
||||
### `quark_int4fp8_moe` online quantization method
|
||||
|
||||
SGLang running on AMD GPUs (CDNA3 or CDNA4 architecture) supports the quantization method `--quantization quark_int4fp8_moe`, that will replace [MoE layers](https://github.com/sgl-project/sglang/blob/main/python/sglang/srt/layers/moe/fused_moe_triton/layer.py) originally in high precision (bfloat16, float16 or float32) to use weights dynamically quantized to int4, that are upcasted to float8 during inference to run compute in float8 precision with activations dynamically quantized on the fly to float8.
|
||||
SGLang running on AMD GPUs (CDNA3 or CDNA4 architecture) supports the quantization method `--quantization quark_int4fp8_moe`, that will replace [MoE layers](https://github.com/sgl-project/sglang/blob/v0.4.8/python/sglang/srt/layers/moe/fused_moe_triton/layer.py#L271) originally in high precision (bfloat16, float16 or float32) to use weights dynamically quantized to int4, that are upcasted to float8 during inference to run compute in float8 precision with activations dynamically quantized on the fly to float8.
|
||||
|
||||
Other layers (e.g. projections in the attention layers) have their weights quantized online to float8 directly.
|
||||
|
||||
@@ -390,6 +799,9 @@ Other layers (e.g. projections in the attention layers) have their weights quant
|
||||
- [GPTQModel](https://github.com/ModelCloud/GPTQModel)
|
||||
- [LLM Compressor](https://github.com/vllm-project/llm-compressor/)
|
||||
- [NVIDIA Model Optimizer (ModelOpt)](https://github.com/NVIDIA/Model-Optimizer)
|
||||
- [NVIDIA Model Optimizer LLM PTQ](https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/llm_ptq)
|
||||
- [Petit: NVFP4 on ROCm](https://github.com/causalflow-ai/petit-kernel) — [LMSYS blog](https://lmsys.org/blog/2025-09-21-petit-amdgpu/), [AMD ROCm blog](https://rocm.blogs.amd.com/artificial-intelligence/fp4-mixed-precision/README.html)
|
||||
- [Torchao: PyTorch Architecture Optimization](https://github.com/pytorch/ao)
|
||||
- [vLLM Quantization](https://docs.vllm.ai/en/latest/quantization/)
|
||||
- [auto-round](https://github.com/intel/auto-round)
|
||||
- [ModelSlim](https://gitcode.com/Ascend/msmodelslim)
|
||||
|
||||
@@ -5,7 +5,7 @@ metatags:
|
||||
---
|
||||
R-Fork (Tensor Remote Fork) is a novel weight loading methodology that leverages efficient inter-node GPU-to-GPU data transfer path to load tensors from a running SGLang instance to a new instance with zero-copy. It can significantly optimize the SGLang instance boot-up time by reducing model weights loading from several minutes to mere seconds.
|
||||
|
||||
To learn more details about R-Fork, please check **[R-Fork blog](https://lmsys.org/blog/2025-12-10-rfork/)**
|
||||
To learn more details about R-Fork, please check **<a href="https://lmsys.org/blog/2025-12-10-rfork/"> R-Fork blog </a>**
|
||||
|
||||
## Usage
|
||||
|
||||
@@ -27,25 +27,29 @@ To learn more details about R-Fork, please check **[R-Fork blog](https://lmsys.o
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>remote-instance-weight-loader-backend</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>`nccl` or `transfer_engine`, default value is `nccl`</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><code>nccl</code>, <code>transfer_engine</code>, or <code>modelexpress</code>. Default is <code>nccl</code>.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>remote-instance-weight-loader-seed-instance-ip</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>IP address of the seed instance who will provide the model weight</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>IP address of the seed instance who will provide the model weight. Used by <code>nccl</code> and <code>transfer_engine</code> backends.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>remote-instance-weight-loader-seed-instance-service-port</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>the port that the seed instance's HTTP server is listening on</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>the port that the seed instance's HTTP server is listening on. Used by <code>nccl</code> and <code>transfer_engine</code> backends.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>remote-instance-weight-loader-send-weights-group-ports</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>the list of available ports on the seed instance that will be used to build NCCL communication groups between seed and client instance. This argument is only needed by `nccl` backend.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>the list of available ports on the seed instance that will be used to build NCCL communication groups between seed and client instance. Only needed by <code>nccl</code> backend.</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>remote-instance-weight-loader-start-seed-via-transfer-engine</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>set to start seed service that supports TransferEngine as backend. It is needed for seed instances when using `transfer_engine` as backend.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>set to start seed service that supports TransferEngine as backend. Needed for seed instances when using <code>transfer_engine</code> as backend.</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>modelexpress-config</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>JSON config for <code>modelexpress</code> backend. Keys: <code>"url"</code> (required, gRPC host:port of ModelExpress server), <code>"model_name"</code> (optional, defaults to <code>--model-path</code>), <code>"source"</code> (optional bool, <code>true</code> for seed mode).</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
### NCCL as backend
|
||||
@@ -80,3 +84,25 @@ python -m sglang.launch_server [args] \
|
||||
--remote-instance-weight-loader-seed-instance-service-port [seed_instance_service_port] \
|
||||
--remote-instance-weight-loader-backend transfer_engine
|
||||
```
|
||||
|
||||
### ModelExpress as backend
|
||||
|
||||
[ModelExpress](https://github.com/ai-dynamo/modelexpress) is a coordination service that manages P2P weight transfer metadata. It removes the need for direct seed IP/port configuration by providing a centralized registry that seeds publish to and clients discover from. Under the hood it uses TransferEngine (Mooncake) for the actual RDMA data transfer.
|
||||
|
||||
A running ModelExpress server is required. See the [ModelExpress documentation](https://github.com/ai-dynamo/modelexpress) for setup instructions.
|
||||
|
||||
seed instance:
|
||||
```bash Command
|
||||
python -m sglang.launch_server [args] \
|
||||
--modelexpress-config '{"url": "[modelexpress_grpc_host:port]", "model_name": "[model_name]", "source": true}'
|
||||
```
|
||||
|
||||
client instance:
|
||||
```bash Command
|
||||
python -m sglang.launch_server [args] \
|
||||
--load-format remote_instance \
|
||||
--remote-instance-weight-loader-backend modelexpress \
|
||||
--modelexpress-config '{"url": "[modelexpress_grpc_host:port]", "model_name": "[model_name]"}'
|
||||
```
|
||||
|
||||
The seed publishes its TransferEngine session ID and tensor layout to ModelExpress. The client queries ModelExpress to discover the seed, then pulls weights directly via RDMA. This enables dynamic seed discovery without hardcoding IPs, and supports multiple models through a single ModelExpress instance.
|
||||
|
||||
@@ -70,7 +70,7 @@
|
||||
" \"python3 -m sglang.launch_server --model-path deepseek-ai/DeepSeek-R1-Distill-Qwen-7B --host 0.0.0.0 --reasoning-parser deepseek-r1 --log-level warning\"\n",
|
||||
")\n",
|
||||
"\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")"
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=server_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -105,7 +105,7 @@ Enable memory saver support when launching the server:
|
||||
|
||||
## Open-To-Use Refit Functionality
|
||||
|
||||
After training completes each step, rollout engines must be refit with new weights. SGLang supports three refit strategies so you can match your infrastructure style (co-located vs disaggregated) and scaling needs. Each strategy maps to a concrete API with clear request schemas. For a deeper dive into SGLang's weight update utilities, see [RL System Deep Thinking: Weight Update Mechanisms](https://github.com/zhaochenyang20/Awesome-ML-SYS-Tutorial/blob/main/rlhf/sys-design/readme-1-EN).
|
||||
After training completes each step, rollout engines must be refit with new weights. SGLang supports three refit strategies so you can match your infrastructure style (co-located vs disaggregated) and scaling needs. Each strategy maps to a concrete API with clear request schemas. For a deeper dive into SGLang's weight update utilities, see [RL System Deep Thinking: Weight Update Mechanisms](https://github.com/zhaochenyang20/Awesome-ML-SYS-Tutorial/blob/main/rlhf/sys-design/readme-1-EN.md).
|
||||
|
||||
**How to choose:**
|
||||
|
||||
@@ -248,6 +248,86 @@ This path trades some I/O overhead for simplicity and flexibility. It integrates
|
||||
|
||||
**Python Engine API:** `engine.update_weights_from_disk(model_path, load_format=None)`
|
||||
|
||||
**Diffusion engine (SGLang-Diffusion):** The diffusion engine exposes the same `POST /update_weights_from_disk` endpoint with the following behavior:
|
||||
|
||||
- **All-or-nothing with rollback:** if any module fails to load, all previously updated modules are rolled back to the original weights by reloading from the original model path. No partial updates are left behind. If rollback itself fails, the exception propagates so the caller knows the model is in an inconsistent state.
|
||||
- **Offload-aware:** when layerwise offload (`--dit-layerwise-offload`) is enabled, the diffusion offload manager replaces GPU parameters with small `torch.empty((1,))` placeholders while real weights live in consolidated pinned CPU buffers. A naive `param.data.copy_()` would fail with a shape mismatch. Instead, the updater dynamically detects active offload managers and writes new weights directly into their CPU buffers, bypassing the placeholders entirely. For any layer that happens to be prefetched on GPU at update time, the live GPU tensor is also updated so the change takes effect immediately. This requires no extra GPU memory and does not disturb the offload state.
|
||||
- **DTensor-aware:** parameters distributed via `torch.distributed.tensor` (tensor parallelism) are updated through `distribute_tensor` so that each shard is correctly placed on the right device mesh.
|
||||
|
||||
**Request body:**
|
||||
|
||||
<table style={{width: "100%", borderCollapse: "collapse", tableLayout: "fixed"}}>
|
||||
<colgroup>
|
||||
<col style={{width: "25%"}} />
|
||||
<col style={{width: "25%"}} />
|
||||
<col style={{width: "25%"}} />
|
||||
<col style={{width: "25%"}} />
|
||||
</colgroup>
|
||||
<thead>
|
||||
<tr style={{borderBottom: "2px solid #d55816"}}>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.02)"}}>Field</th>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.05)"}}>Description</th>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.02)"}}>Defaults</th>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.05)"}}>Options</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><code>model_path</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>The model path with the new weights.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>Required</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Type: str</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><code>flush_cache</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Flush TeaCache state after update.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}><code>True</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Type: bool</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><code>target_modules</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>List of module names to update (e.g. <code>["transformer"]</code>). If omitted, all <code>nn.Module</code> components are updated.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}><code>None</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Type: list[str]</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
**Response body:**
|
||||
|
||||
<table style={{width: "100%", borderCollapse: "collapse", tableLayout: "fixed"}}>
|
||||
<colgroup>
|
||||
<col style={{width: "25%"}} />
|
||||
<col style={{width: "25%"}} />
|
||||
<col style={{width: "25%"}} />
|
||||
<col style={{width: "25%"}} />
|
||||
</colgroup>
|
||||
<thead>
|
||||
<tr style={{borderBottom: "2px solid #d55816"}}>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.02)"}}>Field</th>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.05)"}}>Description</th>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.02)"}}>Defaults</th>
|
||||
<th style={{textAlign: "left", padding: "10px 12px", fontWeight: 700, whiteSpace: "nowrap", backgroundColor: "rgba(255,255,255,0.05)"}}>Options</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><code>success</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Whether the update succeeded.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>-</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Type: bool</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><code>message</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Status / error message.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>-</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Type: str</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
> **Note:** The diffusion engine (SGLang-Diffusion) does not currently support hot refit (updating weights while inference is in progress). The diffusion scheduler processes one request at a time and completes the entire inference before handling the next request, so weight updates and inference never run concurrently.
|
||||
|
||||
### Update Weights from Tensor
|
||||
|
||||
**When to use:**
|
||||
@@ -280,33 +360,33 @@ This strategy requires the training process and rollout engine to share access t
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>`serialized_named_tensors`</td>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><code>serialized_named_tensors</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Per-TP serialized tensor payloads.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>Required</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Type: list[str</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Type: list[str|bytes]</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>`load_format`</td>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><code>load_format</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Optional load format selector.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>`None`</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>`None`, `direct`, `flattened_bucket`, or a custom loader path string</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}><code>None</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}><code>None</code>, <code>direct</code>, <code>flattened_bucket</code>, or a custom loader path string</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>`flush_cache`</td>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><code>flush_cache</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Flush KV cache after update.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>`True`</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}><code>True</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Type: bool</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>`abort_all_requests`</td>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><code>abort_all_requests</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Abort all running requests before update.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>`False`</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}><code>False</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Type: bool</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>`weight_version`</td>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><code>weight_version</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Optional version label tracked by the server.</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>`None`</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}><code>None</code></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Type: str</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
@@ -551,7 +631,7 @@ SGLang exposes explicit pause/resume APIs so you can pause slow requests and con
|
||||
|
||||
In many RL stacks, rollout and training are implemented with different kernels or batching behavior. Even when weights are identical, token probabilities can drift, silently breaking the on-policy assumption. This is the training–inference mismatch problem.
|
||||
|
||||
SGLang supports a deterministic inference mode that reduces non-determinism across batch shapes. This mitigates variance introduced by runtime batching and kernel selection. To further achieve true on-policy training, you need to modify the training engine to use the same deterministic kernels. For implementation details, see these miles examples: [True On-Policy](https://github.com/radixark/miles/tree/main/examples/true_on_policy) and [True On-Policy for VLM](https://github.com/radixark/miles/tree/main/examples/true_on_policy_vlm). For additional context, see the blog post [Let Speed Be With Stability: All-In-One Solution to Training-Inference Mismatch with Miles](https://github.com/zhaochenyang20/Awesome-ML-SYS-Tutorial/blob/main/rlhf/slime/mismatch/blog-en).
|
||||
SGLang supports a deterministic inference mode that reduces non-determinism across batch shapes. This mitigates variance introduced by runtime batching and kernel selection. To further achieve true on-policy training, you need to modify the training engine to use the same deterministic kernels. For implementation details, see these miles examples: [True On-Policy](https://github.com/radixark/miles/tree/main/examples/true_on_policy) and [True On-Policy for VLM](https://github.com/radixark/miles/tree/main/examples/true_on_policy_vlm). For additional context, see the blog post [Let Speed Be With Stability: All-In-One Solution to Training-Inference Mismatch with Miles](https://github.com/zhaochenyang20/Awesome-ML-SYS-Tutorial/blob/main/rlhf/slime/mismatch/blog-en.md).
|
||||
|
||||
**Server flag:**
|
||||
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -19,7 +19,7 @@
|
||||
"- [Outlines](https://github.com/dottxt-ai/outlines): Supports JSON schema and regular expression constraints.\n",
|
||||
"- [Llguidance](https://github.com/guidance-ai/llguidance): Supports JSON schema, regular expression, and EBNF constraints.\n",
|
||||
"\n",
|
||||
"We suggest using XGrammar for its better performance and utility. XGrammar currently uses the [GGML BNF format](https://github.com/ggerganov/llama.cpp/blob/master/grammars/README). For more details, see [XGrammar technical overview](https://blog.mlc.ai/2024/11/22/achieving-efficient-flexible-portable-structured-generation-with-xgrammar).\n",
|
||||
"We suggest using XGrammar for its better performance and utility. XGrammar currently uses the [GGML BNF format](https://github.com/ggerganov/llama.cpp/blob/master/grammars/README.md). For more details, see [XGrammar technical overview](https://blog.mlc.ai/2024/11/22/achieving-efficient-flexible-portable-structured-generation-with-xgrammar).\n",
|
||||
"\n",
|
||||
"To use Outlines, simply add `--grammar-backend outlines` when launching the server.\n",
|
||||
"To use llguidance, add `--grammar-backend llguidance` when launching the server.\n",
|
||||
@@ -54,7 +54,7 @@
|
||||
" \"python -m sglang.launch_server --model-path meta-llama/Meta-Llama-3.1-8B-Instruct --host 0.0.0.0 --log-level warning\"\n",
|
||||
")\n",
|
||||
"\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=server_process)\n",
|
||||
"client = openai.Client(base_url=f\"http://127.0.0.1:{port}/v1\", api_key=\"None\")"
|
||||
]
|
||||
},
|
||||
@@ -356,8 +356,7 @@
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# Support for XGrammar latest structural tag format\n",
|
||||
"# https://xgrammar.mlc.ai/docs/tutorials/structural_tag.html\n",
|
||||
"\n",
|
||||
"# <https://xgrammar.mlc.ai/docs/tutorials/structural_tag.html>\n",
|
||||
"response = client.chat.completions.create(\n",
|
||||
" model=\"meta-llama/Meta-Llama-3.1-8B-Instruct\",\n",
|
||||
" messages=messages,\n",
|
||||
@@ -645,8 +644,7 @@
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# Support for XGrammar latest structural tag format\n",
|
||||
"# https://xgrammar.mlc.ai/docs/tutorials/structural_tag.html\n",
|
||||
"\n",
|
||||
"# <https://xgrammar.mlc.ai/docs/tutorials/structural_tag.html>\n",
|
||||
"payload = {\n",
|
||||
" \"text\": text,\n",
|
||||
" \"sampling_params\": {\n",
|
||||
@@ -925,8 +923,7 @@
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"# Support for XGrammar latest structural tag format\n",
|
||||
"# https://xgrammar.mlc.ai/docs/tutorials/structural_tag.html\n",
|
||||
"\n",
|
||||
"# <https://xgrammar.mlc.ai/docs/tutorials/structural_tag.html>\n",
|
||||
"sampling_params = {\n",
|
||||
" \"temperature\": 0.8,\n",
|
||||
" \"top_p\": 0.95,\n",
|
||||
|
||||
@@ -11,7 +11,7 @@ SGLang supports three grammar backends:
|
||||
- [Outlines](https://github.com/dottxt-ai/outlines): Supports JSON schema and regular expression constraints.
|
||||
- [Llguidance](https://github.com/guidance-ai/llguidance): Supports JSON schema, regular expression, and EBNF constraints.
|
||||
|
||||
We suggest using XGrammar for its better performance and utility. XGrammar currently uses the [GGML BNF format](https://github.com/ggerganov/llama.cpp/blob/master/grammars/README). For more details, see [XGrammar technical overview](https://blog.mlc.ai/2024/11/22/achieving-efficient-flexible-portable-structured-generation-with-xgrammar).
|
||||
We suggest using XGrammar for its better performance and utility. XGrammar currently uses the [GGML BNF format](https://github.com/ggerganov/llama.cpp/blob/master/grammars/README.md). For more details, see [XGrammar technical overview](https://blog.mlc.ai/2024/11/22/achieving-efficient-flexible-portable-structured-generation-with-xgrammar).
|
||||
|
||||
To use Outlines, simply add `--grammar-backend outlines` when launching the server.
|
||||
To use llguidance, add `--grammar-backend llguidance` when launching the server.
|
||||
@@ -247,9 +247,9 @@ If a you choose to call a function ONLY reply in the following format:
|
||||
where
|
||||
start_tag => `<function`
|
||||
parameters => a JSON dict with the function argument name as key and function argument value as value.
|
||||
end_tag => `</function>`
|
||||
end_tag => `</function>`
|
||||
Here is an example,
|
||||
<function=example_function_name>{{"example_name": "example_value"}}</function>
|
||||
<function=example_function_name>{{"example_name": "example_value"}}</function>
|
||||
Reminder:
|
||||
- Function calls MUST follow the specified format
|
||||
- Required parameters MUST be specified
|
||||
@@ -274,17 +274,17 @@ response = client.chat.completions.create(
|
||||
"type": "structural_tag",
|
||||
"structures": [
|
||||
{
|
||||
"begin": "<function=get_current_weather>",
|
||||
"begin": "<function=get_current_weather>",
|
||||
"schema": schema_get_current_weather,
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
{
|
||||
"begin": "<function=get_current_date>",
|
||||
"begin": "<function=get_current_date>",
|
||||
"schema": schema_get_current_date,
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
],
|
||||
"triggers": ["<function="],
|
||||
"triggers": ["<function="],
|
||||
},
|
||||
)
|
||||
|
||||
@@ -303,23 +303,23 @@ response = client.chat.completions.create(
|
||||
"type": "structural_tag",
|
||||
"format": {
|
||||
"type": "triggered_tags",
|
||||
"triggers": ["<function="],
|
||||
"triggers": ["<function="],
|
||||
"tags": [
|
||||
{
|
||||
"begin": "<function=get_current_weather>",
|
||||
"begin": "<function=get_current_weather>",
|
||||
"content": {
|
||||
"type": "json_schema",
|
||||
"json_schema": schema_get_current_weather,
|
||||
},
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
{
|
||||
"begin": "<function=get_current_date>",
|
||||
"begin": "<function=get_current_date>",
|
||||
"content": {
|
||||
"type": "json_schema",
|
||||
"json_schema": schema_get_current_date,
|
||||
},
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
],
|
||||
"at_least_one": False,
|
||||
@@ -506,17 +506,17 @@ payload = {
|
||||
"type": "structural_tag",
|
||||
"structures": [
|
||||
{
|
||||
"begin": "<function=get_current_weather>",
|
||||
"begin": "<function=get_current_weather>",
|
||||
"schema": schema_get_current_weather,
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
{
|
||||
"begin": "<function=get_current_date>",
|
||||
"begin": "<function=get_current_date>",
|
||||
"schema": schema_get_current_date,
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
],
|
||||
"triggers": ["<function="],
|
||||
"triggers": ["<function="],
|
||||
}
|
||||
)
|
||||
},
|
||||
@@ -541,23 +541,23 @@ payload = {
|
||||
"type": "structural_tag",
|
||||
"format": {
|
||||
"type": "triggered_tags",
|
||||
"triggers": ["<function="],
|
||||
"triggers": ["<function="],
|
||||
"tags": [
|
||||
{
|
||||
"begin": "<function=get_current_weather>",
|
||||
"begin": "<function=get_current_weather>",
|
||||
"content": {
|
||||
"type": "json_schema",
|
||||
"json_schema": schema_get_current_weather,
|
||||
},
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
{
|
||||
"begin": "<function=get_current_date>",
|
||||
"begin": "<function=get_current_date>",
|
||||
"content": {
|
||||
"type": "json_schema",
|
||||
"json_schema": schema_get_current_date,
|
||||
},
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
],
|
||||
"at_least_one": False,
|
||||
@@ -727,17 +727,17 @@ sampling_params = {
|
||||
"type": "structural_tag",
|
||||
"structures": [
|
||||
{
|
||||
"begin": "<function=get_current_weather>",
|
||||
"begin": "<function=get_current_weather>",
|
||||
"schema": schema_get_current_weather,
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
{
|
||||
"begin": "<function=get_current_date>",
|
||||
"begin": "<function=get_current_date>",
|
||||
"schema": schema_get_current_date,
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
],
|
||||
"triggers": ["<function="],
|
||||
"triggers": ["<function="],
|
||||
}
|
||||
),
|
||||
}
|
||||
@@ -763,23 +763,23 @@ sampling_params = {
|
||||
"type": "structural_tag",
|
||||
"format": {
|
||||
"type": "triggered_tags",
|
||||
"triggers": ["<function="],
|
||||
"triggers": ["<function="],
|
||||
"tags": [
|
||||
{
|
||||
"begin": "<function=get_current_weather>",
|
||||
"begin": "<function=get_current_weather>",
|
||||
"content": {
|
||||
"type": "json_schema",
|
||||
"json_schema": schema_get_current_weather,
|
||||
},
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
{
|
||||
"begin": "<function=get_current_date>",
|
||||
"begin": "<function=get_current_date>",
|
||||
"content": {
|
||||
"type": "json_schema",
|
||||
"json_schema": schema_get_current_date,
|
||||
},
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
],
|
||||
"at_least_one": False,
|
||||
|
||||
@@ -50,7 +50,7 @@
|
||||
" \"python -m sglang.launch_server --model-path deepseek-ai/DeepSeek-R1-Distill-Qwen-7B --host 0.0.0.0 --reasoning-parser deepseek-r1 --log-level warning\"\n",
|
||||
")\n",
|
||||
"\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=server_process)\n",
|
||||
"client = openai.Client(base_url=f\"http://127.0.0.1:{port}/v1\", api_key=\"None\")"
|
||||
]
|
||||
},
|
||||
|
||||
@@ -3,17 +3,17 @@ title: "Structured Outputs For Reasoning Models"
|
||||
metatags:
|
||||
description: "SGLang structured outputs for reasoning models: free-form thinking with constrained final output for DeepSeek R1, QwQ models."
|
||||
---
|
||||
When working with reasoning models that use special tokens like `<think>...</think>` to denote reasoning sections, you might want to allow free-form text within these sections while still enforcing grammar constraints on the rest of the output.
|
||||
When working with reasoning models that use special tokens like `<think>...</think>` to denote reasoning sections, you might want to allow free-form text within these sections while still enforcing grammar constraints on the rest of the output.
|
||||
|
||||
SGLang provides a feature to disable grammar restrictions within reasoning sections. This is particularly useful for models that need to perform complex reasoning steps before providing a structured output.
|
||||
|
||||
To enable this feature, use the `--reasoning-parser` flag which decide the think_end_token, such as `</think>`, when launching the server. You can also specify the reasoning parser using the `--reasoning-parser` flag.
|
||||
To enable this feature, use the `--reasoning-parser` flag which decide the think_end_token, such as `</think>`, when launching the server. You can also specify the reasoning parser using the `--reasoning-parser` flag.
|
||||
|
||||
## Supported Models
|
||||
|
||||
Currently, SGLang supports the following reasoning models:
|
||||
- [DeepSeek R1 series](https://huggingface.co/collections/deepseek-ai/deepseek-r1-678e1e131c0169c0bc89728d): The reasoning content is wrapped with `<think>` and `</think>` tags.
|
||||
- [QwQ](https://huggingface.co/Qwen/QwQ-32B): The reasoning content is wrapped with `<think>` and `</think>` tags.
|
||||
- [DeepSeek R1 series](https://huggingface.co/collections/deepseek-ai/deepseek-r1-678e1e131c0169c0bc89728d): The reasoning content is wrapped with `<think>` and `</think>` tags.
|
||||
- [QwQ](https://huggingface.co/Qwen/QwQ-32B): The reasoning content is wrapped with `<think>` and `</think>` tags.
|
||||
|
||||
|
||||
## Usage
|
||||
@@ -252,9 +252,9 @@ If a you choose to call a function ONLY reply in the following format:
|
||||
where
|
||||
start_tag => `<function`
|
||||
parameters => a JSON dict with the function argument name as key and function argument value as value.
|
||||
end_tag => `</function>`
|
||||
end_tag => `</function>`
|
||||
Here is an example,
|
||||
<function=example_function_name>{{"example_name": "example_value"}}</function>
|
||||
<function=example_function_name>{{"example_name": "example_value"}}</function>
|
||||
Reminder:
|
||||
- Function calls MUST follow the specified format
|
||||
- Required parameters MUST be specified
|
||||
@@ -280,17 +280,17 @@ response = client.chat.completions.create(
|
||||
"max_new_tokens": 2048,
|
||||
"structures": [
|
||||
{
|
||||
"begin": "<function=get_current_weather>",
|
||||
"begin": "<function=get_current_weather>",
|
||||
"schema": schema_get_current_weather,
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
{
|
||||
"begin": "<function=get_current_date>",
|
||||
"begin": "<function=get_current_date>",
|
||||
"schema": schema_get_current_date,
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
],
|
||||
"triggers": ["<function="],
|
||||
"triggers": ["<function="],
|
||||
},
|
||||
)
|
||||
|
||||
@@ -351,8 +351,8 @@ response = requests.post(
|
||||
print(response.json())
|
||||
|
||||
|
||||
reasoing_content = response.json()["text"].split("</think>")[0]
|
||||
content = response.json()["text"].split("</think>")[1]
|
||||
reasoing_content = response.json()["text"].split("</think>")[0]
|
||||
content = response.json()["text"].split("</think>")[1]
|
||||
print_highlight(f"reasoing_content: {reasoing_content}\n\ncontent: {content}")
|
||||
```
|
||||
|
||||
@@ -460,17 +460,17 @@ payload = {
|
||||
"type": "structural_tag",
|
||||
"structures": [
|
||||
{
|
||||
"begin": "<function=get_current_weather>",
|
||||
"begin": "<function=get_current_weather>",
|
||||
"schema": schema_get_current_weather,
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
{
|
||||
"begin": "<function=get_current_date>",
|
||||
"begin": "<function=get_current_date>",
|
||||
"schema": schema_get_current_date,
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
],
|
||||
"triggers": ["<function="],
|
||||
"triggers": ["<function="],
|
||||
}
|
||||
),
|
||||
},
|
||||
@@ -634,17 +634,17 @@ sampling_params = {
|
||||
"type": "structural_tag",
|
||||
"structures": [
|
||||
{
|
||||
"begin": "<function=get_current_weather>",
|
||||
"begin": "<function=get_current_weather>",
|
||||
"schema": schema_get_current_weather,
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
{
|
||||
"begin": "<function=get_current_date>",
|
||||
"begin": "<function=get_current_date>",
|
||||
"schema": schema_get_current_date,
|
||||
"end": "</function>",
|
||||
"end": "</function>",
|
||||
},
|
||||
],
|
||||
"triggers": ["<function="],
|
||||
"triggers": ["<function="],
|
||||
}
|
||||
),
|
||||
}
|
||||
|
||||
@@ -60,7 +60,7 @@
|
||||
"server_process, port = launch_server_cmd(\n",
|
||||
" \"python3 -m sglang.launch_server --model-path Qwen/Qwen2.5-7B-Instruct --tool-call-parser qwen25 --host 0.0.0.0 --log-level warning\" # qwen25\n",
|
||||
")\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")"
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=server_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -550,7 +550,9 @@
|
||||
"server_process_tool_choice, port_tool_choice = launch_server_cmd(\n",
|
||||
" \"python3 -m sglang.launch_server --model-path Qwen/Qwen2.5-7B-Instruct --tool-call-parser qwen25 --host 0.0.0.0 --log-level warning\"\n",
|
||||
")\n",
|
||||
"wait_for_server(f\"http://localhost:{port_tool_choice}\")\n",
|
||||
"wait_for_server(\n",
|
||||
" f\"http://localhost:{port_tool_choice}\", process=server_process_tool_choice\n",
|
||||
")\n",
|
||||
"\n",
|
||||
"# Initialize client for tool choice examples\n",
|
||||
"client_tool_choice = OpenAI(\n",
|
||||
@@ -695,7 +697,7 @@
|
||||
"server_process, port = launch_server_cmd(\n",
|
||||
" \" python3 -m sglang.launch_server --model-path meta-llama/Llama-3.2-1B-Instruct --tool-call-parser pythonic --tp 1 --log-level warning\" # llama-3.2-1b-instruct\n",
|
||||
")\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=server_process)\n",
|
||||
"\n",
|
||||
"tools = [\n",
|
||||
" {\n",
|
||||
|
||||
@@ -64,8 +64,11 @@
|
||||
"\n",
|
||||
"nest_asyncio.apply()\n",
|
||||
"\n",
|
||||
"import sglang.test.doc_patch # noqa: F401\n",
|
||||
"\n",
|
||||
"model_path = \"Qwen/Qwen2.5-VL-3B-Instruct\"\n",
|
||||
"chat_template = \"qwen2-vl\""
|
||||
"chat_template = \"qwen2-vl\"\n",
|
||||
"example_image_url = \"https://raw.githubusercontent.com/sgl-project/sglang/main/examples/assets/example_image.png\""
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -81,13 +84,7 @@
|
||||
"\n",
|
||||
"from sglang.srt.parser.conversation import chat_templates\n",
|
||||
"\n",
|
||||
"image = Image.open(\n",
|
||||
" BytesIO(\n",
|
||||
" requests.get(\n",
|
||||
" \"https://github.com/sgl-project/sglang/blob/main/examples/assets/example_image.png?raw=true\"\n",
|
||||
" ).content\n",
|
||||
" )\n",
|
||||
")\n",
|
||||
"image = Image.open(BytesIO(requests.get(example_image_url).content))\n",
|
||||
"\n",
|
||||
"conv = chat_templates[chat_template].copy()\n",
|
||||
"conv.append_message(conv.roles[0], f\"What's shown here: {conv.image_token}?\")\n",
|
||||
@@ -185,9 +182,8 @@
|
||||
"from transformers import Qwen2_5_VLForConditionalGeneration\n",
|
||||
"\n",
|
||||
"processor = AutoProcessor.from_pretrained(model_path, use_fast=True)\n",
|
||||
"vision = (\n",
|
||||
" Qwen2_5_VLForConditionalGeneration.from_pretrained(model_path).eval().visual.cuda()\n",
|
||||
")"
|
||||
"model = Qwen2_5_VLForConditionalGeneration.from_pretrained(model_path).eval()\n",
|
||||
"vision = model.model.visual.cuda()"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -206,6 +202,7 @@
|
||||
"precomputed_embeddings = vision(\n",
|
||||
" processor_output[\"pixel_values\"].cuda(), processor_output[\"image_grid_thw\"].cuda()\n",
|
||||
")\n",
|
||||
"precomputed_embeddings = precomputed_embeddings.pooler_output\n",
|
||||
"\n",
|
||||
"multi_modal_item = dict(\n",
|
||||
" processor_output,\n",
|
||||
@@ -238,13 +235,7 @@
|
||||
"from sglang.srt.parser.conversation import chat_templates\n",
|
||||
"\n",
|
||||
"# Download the same example image\n",
|
||||
"image = Image.open(\n",
|
||||
" BytesIO(\n",
|
||||
" requests.get(\n",
|
||||
" \"https://github.com/sgl-project/sglang/blob/main/examples/assets/example_image.png?raw=true\"\n",
|
||||
" ).content\n",
|
||||
" )\n",
|
||||
")\n",
|
||||
"image = Image.open(BytesIO(requests.get(example_image_url).content))\n",
|
||||
"\n",
|
||||
"conv = chat_templates[chat_template].copy()\n",
|
||||
"conv.append_message(conv.roles[0], f\"What's shown here: {conv.image_token}?\")\n",
|
||||
|
||||
@@ -9,7 +9,6 @@ This tutorial demonstrates how to use SGLang's **offline Engine API** to query V
|
||||
2. **Processor Output**: Use HuggingFace processor for data preprocessing.
|
||||
3. **Precomputed Embeddings**: Pre-calculate image features to improve inference efficiency.
|
||||
|
||||
|
||||
## Understanding the Three Input Formats
|
||||
|
||||
SGLang supports three ways to pass visual data, each optimized for different scenarios:
|
||||
@@ -35,21 +34,20 @@ SGLang supports three ways to pass visual data, each optimized for different sce
|
||||
|
||||
The examples below demonstrate all three approaches with both Qwen2.5-VL and Llama 4 models.
|
||||
|
||||
|
||||
## Querying Qwen2.5-VL Model
|
||||
|
||||
|
||||
|
||||
```python Example
|
||||
import nest_asyncio
|
||||
|
||||
nest_asyncio.apply()
|
||||
|
||||
import sglang.test.doc_patch # noqa: F401
|
||||
|
||||
model_path = "Qwen/Qwen2.5-VL-3B-Instruct"
|
||||
chat_template = "qwen2-vl"
|
||||
example_image_url = "https://raw.githubusercontent.com/sgl-project/sglang/main/examples/assets/example_image.png"
|
||||
```
|
||||
|
||||
|
||||
```python Example
|
||||
from io import BytesIO
|
||||
import requests
|
||||
@@ -57,13 +55,7 @@ from PIL import Image
|
||||
|
||||
from sglang.srt.parser.conversation import chat_templates
|
||||
|
||||
image = Image.open(
|
||||
BytesIO(
|
||||
requests.get(
|
||||
"https://github.com/sgl-project/sglang/blob/main/examples/assets/example_image.png?raw=true"
|
||||
).content
|
||||
)
|
||||
)
|
||||
image = Image.open(BytesIO(requests.get(example_image_url).content))
|
||||
|
||||
conv = chat_templates[chat_template].copy()
|
||||
conv.append_message(conv.roles[0], f"What's shown here: {conv.image_token}?")
|
||||
@@ -78,16 +70,12 @@ image
|
||||
|
||||
### Basic Offline Engine API Call
|
||||
|
||||
|
||||
|
||||
```python Example
|
||||
from sglang import Engine
|
||||
|
||||
|
||||
llm = Engine(model_path=model_path, chat_template=chat_template, log_level="warning")
|
||||
```
|
||||
|
||||
|
||||
```python Example
|
||||
out = llm.generate(prompt=conv.get_prompt(), image_data=[image])
|
||||
print("Model response:")
|
||||
@@ -98,8 +86,6 @@ print(out["text"])
|
||||
|
||||
Using a HuggingFace processor to preprocess text and images, and passing the `processor_output` directly into `Engine.generate`.
|
||||
|
||||
|
||||
|
||||
```python Example
|
||||
from transformers import AutoProcessor
|
||||
|
||||
@@ -120,19 +106,15 @@ print(out["text"])
|
||||
|
||||
You can pre-calculate image features to avoid repeated visual encoding processes.
|
||||
|
||||
|
||||
|
||||
```python Example
|
||||
from transformers import AutoProcessor
|
||||
from transformers import Qwen2_5_VLForConditionalGeneration
|
||||
|
||||
processor = AutoProcessor.from_pretrained(model_path, use_fast=True)
|
||||
vision = (
|
||||
Qwen2_5_VLForConditionalGeneration.from_pretrained(model_path).eval().visual.cuda()
|
||||
)
|
||||
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(model_path).eval()
|
||||
vision = model.model.visual.cuda()
|
||||
```
|
||||
|
||||
|
||||
```python Example
|
||||
processor_output = processor(
|
||||
images=[image], text=conv.get_prompt(), return_tensors="pt"
|
||||
@@ -143,6 +125,7 @@ input_ids = processor_output["input_ids"][0].detach().cpu().tolist()
|
||||
precomputed_embeddings = vision(
|
||||
processor_output["pixel_values"].cuda(), processor_output["image_grid_thw"].cuda()
|
||||
)
|
||||
precomputed_embeddings = precomputed_embeddings.pooler_output
|
||||
|
||||
multi_modal_item = dict(
|
||||
processor_output,
|
||||
@@ -170,13 +153,7 @@ from PIL import Image
|
||||
from sglang.srt.parser.conversation import chat_templates
|
||||
|
||||
# Download the same example image
|
||||
image = Image.open(
|
||||
BytesIO(
|
||||
requests.get(
|
||||
"https://github.com/sgl-project/sglang/blob/main/examples/assets/example_image.png?raw=true"
|
||||
).content
|
||||
)
|
||||
)
|
||||
image = Image.open(BytesIO(requests.get(example_image_url).content))
|
||||
|
||||
conv = chat_templates[chat_template].copy()
|
||||
conv.append_message(conv.roles[0], f"What's shown here: {conv.image_token}?")
|
||||
@@ -190,7 +167,6 @@ print(f"Image size: {image.size}")
|
||||
image
|
||||
```
|
||||
|
||||
|
||||
### Llama 4 Basic Call
|
||||
|
||||
Llama 4 requires more computational resources, so it's configured with multi-GPU parallelism (tp_size=4) and larger context length.
|
||||
@@ -209,7 +185,6 @@ print("Llama 4 response:")
|
||||
print(out["text"])
|
||||
```
|
||||
|
||||
|
||||
### Call with Processor Output
|
||||
|
||||
Using HuggingFace processor to preprocess data can reduce computational overhead during inference.
|
||||
@@ -230,7 +205,6 @@ print("Response using processor output:")
|
||||
print(out)
|
||||
```
|
||||
|
||||
|
||||
### Call with Precomputed Embeddings
|
||||
|
||||
```python Example
|
||||
|
||||
@@ -0,0 +1,58 @@
|
||||
---
|
||||
title: "DeepSeek OCR (OCR-1 / OCR-2)"
|
||||
metatags:
|
||||
description: "DeepSeek OCR models are multimodal (image + text) models for OCR and document understanding."
|
||||
---
|
||||
|
||||
DeepSeek OCR models are multimodal (image + text) models for OCR and document understanding.
|
||||
|
||||
## Launch server
|
||||
|
||||
```shell
|
||||
python -m sglang.launch_server \
|
||||
--model-path deepseek-ai/DeepSeek-OCR-2 \
|
||||
--trust-remote-code \
|
||||
--host 0.0.0.0 \
|
||||
--port 30000
|
||||
```
|
||||
|
||||
> You can replace `deepseek-ai/DeepSeek-OCR-2` with `deepseek-ai/DeepSeek-OCR`.
|
||||
|
||||
## Prompt examples
|
||||
|
||||
Recommended prompts from the model card:
|
||||
|
||||
```
|
||||
<image>
|
||||
<|grounding|>Convert the document to markdown.
|
||||
```
|
||||
|
||||
```
|
||||
<image>
|
||||
Free OCR.
|
||||
```
|
||||
|
||||
## OpenAI-compatible request example
|
||||
|
||||
```python
|
||||
import requests
|
||||
|
||||
url = "http://localhost:30000/v1/chat/completions"
|
||||
|
||||
data = {
|
||||
"model": "deepseek-ai/DeepSeek-OCR-2",
|
||||
"messages": [
|
||||
{
|
||||
"role": "user",
|
||||
"content": [
|
||||
{"type": "text", "text": "<image>\n<|grounding|>Convert the document to markdown."},
|
||||
{"type": "image_url", "image_url": {"url": "https://example.com/your_image.jpg"}},
|
||||
],
|
||||
}
|
||||
],
|
||||
"max_tokens": 512,
|
||||
}
|
||||
|
||||
response = requests.post(url, json=data)
|
||||
print(response.text)
|
||||
```
|
||||
@@ -25,86 +25,125 @@ To run DeepSeek V3.1/V3/R1 models, the recommended settings are as follows:
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}} rowSpan={5}>**Full precision [FP8](https://huggingface.co/deepseek-ai/DeepSeek-R1-0528)** *(recommended)*</td>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}} rowSpan={5}><strong>Full precision <a href="https://huggingface.co/deepseek-ai/DeepSeek-R1-0528">FP8</a></strong><br>*(recommended)*</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>8 x H200</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>8 x B200</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>8 x MI300X</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>2 x 8 x H100/800/20</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Xeon 6980P CPU</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}} rowSpan={4}>**Full precision ([BF16](https://huggingface.co/unsloth/DeepSeek-R1-0528-BF16))** (upcast from original FP8)</td>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}} rowSpan={4}><strong>Full precision (<a href="https://huggingface.co/unsloth/DeepSeek-R1-0528-BF16">BF16</a>)</strong> (upcast from original FP8)</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>2 x 8 x H200</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>2 x 8 x MI300X</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>4 x 8 x H100/800/20</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>4 x 8 x A100/A800</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}} rowSpan={4}>**Quantized weights ([INT8](https://huggingface.co/meituan/DeepSeek-R1-Channel-INT8))**</td>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}} rowSpan={4}><strong>Quantized weights (<a href="https://huggingface.co/meituan/DeepSeek-R1-Channel-INT8">INT8</a>)</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>16 x A100/800</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>32 x L40S</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>Xeon 6980P CPU</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>4 x Atlas 800I A3</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>**Quantized weights ([W4A8](https://huggingface.co/novita/Deepseek-R1-0528-W4AFP8))**</td>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong>Quantized weights (<a href="https://huggingface.co/novita/Deepseek-R1-0528-W4AFP8">W4A8</a>)</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>8 x H20/100, 4 x H200</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}} rowSpan={2}>**Quantized weights ([AWQ](https://huggingface.co/QuixiAI/DeepSeek-R1-0528-AWQ))**</td>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}} rowSpan={2}><strong>Quantized weights (<a href="https://huggingface.co/QuixiAI/DeepSeek-R1-0528-AWQ">AWQ</a>)</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>8 x H100/800/20</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>8 x A100/A800</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>**Quantized weights ([MXFP4](https://huggingface.co/amd/DeepSeek-R1-MXFP4-Preview))**</td>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong>Quantized weights (<a href="https://huggingface.co/amd/DeepSeek-R1-MXFP4-Preview">MXFP4</a>)</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>8, 4 x MI355X/350X</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>**Quantized weights ([NVFP4](https://huggingface.co/nvidia/DeepSeek-R1-0528-NVFP4-v2))**</td>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}><strong>Quantized weights (<a href="https://huggingface.co/nvidia/DeepSeek-R1-0528-NVFP4-v2">NVFP4</a>)</strong></td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>8, 4 x B200</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
<style>
|
||||
.md-typeset__table {
|
||||
width: 100%;
|
||||
}
|
||||
|
||||
<Callout icon="key" color="#FFC107" iconType="regular">
|
||||
.md-typeset__table table {
|
||||
border-collapse: collapse;
|
||||
margin: 1em 0;
|
||||
border: 2px solid var(--md-typeset-table-color);
|
||||
table-layout: fixed;
|
||||
}
|
||||
|
||||
.md-typeset__table th {
|
||||
border: 1px solid var(--md-typeset-table-color);
|
||||
border-bottom: 2px solid var(--md-typeset-table-color);
|
||||
background-color: var(--md-default-bg-color--lighter);
|
||||
padding: 12px;
|
||||
}
|
||||
|
||||
.md-typeset__table td {
|
||||
border: 1px solid var(--md-typeset-table-color);
|
||||
padding: 12px;
|
||||
}
|
||||
|
||||
.md-typeset__table tr:nth-child(2n) {
|
||||
background-color: var(--md-default-bg-color--lightest);
|
||||
}
|
||||
</style>
|
||||
|
||||
<Warning>
|
||||
The official DeepSeek V3 is already in FP8 format, so you should not run it with any quantization arguments like `--quantization fp8`.
|
||||
</Callout>
|
||||
</Warning>
|
||||
|
||||
Detailed commands for reference:
|
||||
|
||||
- [8 x H200](https://github.com/sgl-project/sglang/tree/main/benchmark/deepseek_v3#using-docker-recommended)
|
||||
- [4 x B200, 8 x B200](https://github.com/sgl-project/sglang/tree/main/benchmark/deepseek_v3#example-serving-with-one-b200-node)
|
||||
- [8 x MI300X](../hardware-platforms/amd-gpus#running-deepseek-v3)
|
||||
- [2 x 8 x H200](https://github.com/sgl-project/sglang/tree/main/benchmark/deepseek_v3#example-serving-with-two-h208-nodes)
|
||||
- [8 x MI300X](../hardware-platforms/amd_gpu#running-deepseek-v3)
|
||||
- [2 x 8 x H200](https://github.com/sgl-project/sglang/tree/main/benchmark/deepseek_v3#example-serving-with-two-h2008-nodes-and-docker)
|
||||
- [4 x 8 x A100](https://github.com/sgl-project/sglang/tree/main/benchmark/deepseek_v3#example-serving-with-four-a1008-nodes)
|
||||
- [8 x A100 (AWQ)](https://github.com/sgl-project/sglang/tree/main/benchmark/deepseek_v3#example-serving-with-8-a100a800-with-awq-quantization)
|
||||
- [16 x A100 (INT8)](https://github.com/sgl-project/sglang/tree/main/benchmark/deepseek_v3#example-serving-with-16-a100a800-with-int8-quantization)
|
||||
- [32 x L40S (INT8)](https://github.com/sgl-project/sglang/tree/main/benchmark/deepseek_v3#example-serving-with-32-l40s-with-int8-quantization)
|
||||
- [Xeon 6980P CPU](../hardware-platforms/cpu-server#example-running-deepseek-v31-terminus)
|
||||
- [4 x Atlas 800I A3 (int8)](../hardware-platforms/ascend-npus/DeepSeek-Examples#running-deepseek-with-pd-disaggregation-on-4-x-atlas-800i-a3)
|
||||
- [Xeon 6980P CPU](../hardware-platforms/cpu_server#example-running-deepseek-r1)
|
||||
- [4 x Atlas 800I A3 (int8)](../hardware-platforms/ascend-npus/ascend_npu_deepseek_example#running-deepseek-with-pd-disaggregation-on-4-x-atlas-800i-a3)
|
||||
|
||||
### Download Weights
|
||||
If you encounter errors when starting the server, ensure the weights have finished downloading. It's recommended to download them beforehand or restart multiple times until all weights are downloaded. Please refer to [DeepSeek V3](https://huggingface.co/deepseek-ai/DeepSeek-V3-Base#61-inference-with-deepseek-infer-demo-example-only) official guide to download the weights.
|
||||
@@ -116,7 +155,7 @@ Please refer to [the example](https://github.com/sgl-project/sglang/tree/main/be
|
||||
|
||||
- [Deploying DeepSeek on GB200 NVL72 with PD and Large Scale EP](https://lmsys.org/blog/2025-06-16-gb200-part-1/) ([Part I](https://lmsys.org/blog/2025-06-16-gb200-part-1/), [Part II](https://lmsys.org/blog/2025-09-25-gb200-part-2/)) - Comprehensive guide on GB200 optimizations.
|
||||
|
||||
- [Deploying DeepSeek with PD Disaggregation and Large-Scale Expert Parallelism on 96 H100 GPUs](https://lmsys.org/blog/2025-05-05-deepseek-pd-ep/) - Guide on PD disaggregation and large-scale EP.
|
||||
- [Deploying DeepSeek with PD Disaggregation and Large-Scale Expert Parallelism on 96 H100 GPUs](https://lmsys.org/blog/2025-05-05-large-scale-ep/) - Guide on PD disaggregation and large-scale EP.
|
||||
|
||||
- [Serving with two H20*8 nodes](https://github.com/sgl-project/sglang/tree/main/benchmark/deepseek_v3#example-serving-with-two-h208-nodes).
|
||||
|
||||
@@ -144,9 +183,9 @@ Please refer to [the example](https://github.com/sgl-project/sglang/tree/main/be
|
||||
|
||||
Overall, with these optimizations, we have achieved up to **7x** acceleration in output throughput compared to the previous version.
|
||||
|
||||
<Frame>
|
||||
<img src="https://lmsys.org/images/blog/sglang_v0_3/deepseek_mla.svg" alt="Multi-head Latent Attention for DeepSeek Series Models"/>
|
||||
</Frame>
|
||||
<p align="center">
|
||||
<img src="https://lmsys.org/images/blog/sglang_v0_3/deepseek_mla.svg" alt="Multi-head Latent Attention for DeepSeek Series Models" />
|
||||
</p>
|
||||
|
||||
**Usage**: MLA optimization is enabled by default.
|
||||
|
||||
@@ -156,15 +195,15 @@ Overall, with these optimizations, we have achieved up to **7x** acceleration in
|
||||
|
||||
**Description**: This optimization involves data parallelism (DP) for the MLA attention mechanism of DeepSeek Series Models, which allows for a significant reduction in the KV cache size, enabling larger batch sizes. Each DP worker independently handles different types of batches (prefill, decode, idle), which are then synchronized before and after processing through the Mixture-of-Experts (MoE) layer. If you do not use DP attention, KV cache will be duplicated among all TP ranks.
|
||||
|
||||
<Frame>
|
||||
<img src="https://lmsys.org/images/blog/sglang_v0_4/dp_attention.svg" alt="Data Parallelism Attention for DeepSeek Series Models"/>
|
||||
</Frame>
|
||||
<p align="center">
|
||||
<img src="https://lmsys.org/images/blog/sglang_v0_4/dp_attention.svg" alt="Data Parallelism Attention for DeepSeek Series Models" />
|
||||
</p>
|
||||
|
||||
With data parallelism attention enabled, we have achieved up to **1.9x** decoding throughput improvement compared to the previous version.
|
||||
|
||||
<Frame>
|
||||
<img src="https://lmsys.org/images/blog/sglang_v0_4/deepseek_coder_v2.svg" alt="Data Parallelism Attention Performance Comparison"/>
|
||||
</Frame>
|
||||
<p align="center">
|
||||
<img src="https://lmsys.org/images/blog/sglang_v0_4/deepseek_coder_v2.svg" alt="Data Parallelism Attention Performance Comparison" />
|
||||
</p>
|
||||
|
||||
**Usage**:
|
||||
- Append `--enable-dp-attention --tp 8 --dp 8` to the server arguments when using 8 H200 GPUs. This optimization improves peak throughput in high batch size scenarios where the server is limited by KV cache capacity.
|
||||
@@ -180,7 +219,7 @@ Data parallelism attention is not recommended for low-latency, small-batch use c
|
||||
|
||||
**Description**: For users with limited memory on a single node, SGLang supports serving DeepSeek Series Models, including DeepSeek V3, across multiple nodes using tensor parallelism. This approach partitions the model parameters across multiple GPUs or nodes to handle models that are too large for one node's memory.
|
||||
|
||||
**Usage**: Check [here](https://github.com/sgl-project/sglang/tree/main/benchmark/deepseek_v3#example-serving-with-2-h208) for usage examples.
|
||||
**Usage**: Check [here](https://github.com/sgl-project/sglang/tree/main/benchmark/deepseek_v3#example-serving-with-two-h2008-nodes-and-docker) for usage examples.
|
||||
|
||||
### Block-wise FP8
|
||||
|
||||
@@ -215,7 +254,7 @@ python3 -m sglang.launch_server \
|
||||
--tp 8
|
||||
```
|
||||
- The default configuration for DeepSeek models is `--speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4`. The best configuration for `--speculative-num-steps`, `--speculative-eagle-topk` and `--speculative-num-draft-tokens` can be searched with [bench_speculative.py](https://github.com/sgl-project/sglang/blob/main/scripts/playground/bench_speculative.py) script for given batch size. The minimum configuration is `--speculative-num-steps 1 --speculative-eagle-topk 1 --speculative-num-draft-tokens 2`, which can achieve speedup for larger batch sizes.
|
||||
- Most MLA attention backends fully support MTP usage. See [MLA Backends](../advanced_features/attention_backend.md#mla-backends) for details.
|
||||
- Most MLA attention backends fully support MTP usage. See [MLA Backends](../advanced_features/attention_backend#mla-backends) for details.
|
||||
|
||||
<Note>
|
||||
To enable DeepSeek MTP for large batch sizes (>48), you need to adjust some parameters (Reference [this discussion](https://github.com/sgl-project/sglang/issues/4543#issuecomment-2737413756)):
|
||||
@@ -250,10 +289,10 @@ python3 -m sglang.launch_server \
|
||||
|
||||
Sample Request:
|
||||
|
||||
```text Output
|
||||
```
|
||||
curl "http://127.0.0.1:30000/v1/chat/completions" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"temperature": 0, "max_tokens": 100, "model": "deepseek-ai/DeepSeek-V3-0324", "tools": [{"type": "function", "function": {"name": "query_weather", "description": "Get weather of an city, the user should supply a city first", "parameters": {"type": "object", "properties": {"city": {"type": "string", "description": "The city, e.g. Beijing"}}, "required": ["city"]}}}], "messages": [{"role": "user", "content": "Hows the weather like in Qingdao today"}]}'
|
||||
-d '{"temperature": 0, "max_tokens": 100, "model": "deepseek-ai/DeepSeek-V3-0324", "tools": [{"type": "function", "function": {"name": "query_weather", "description": "Get weather of a city, the user should supply a city first", "parameters": {"type": "object", "properties": {"city": {"type": "string", "description": "The city, e.g. Beijing"}}, "required": ["city"]}}}], "messages": [{"role": "user", "content": "How'\''s the weather like in Qingdao today"}]}'
|
||||
```
|
||||
|
||||
Expected Response
|
||||
@@ -263,10 +302,10 @@ Expected Response
|
||||
|
||||
```
|
||||
Sample Streaming Request:
|
||||
```text Output
|
||||
```
|
||||
curl "http://127.0.0.1:30000/v1/chat/completions" \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"temperature": 0, "max_tokens": 100, "model": "deepseek-ai/DeepSeek-V3-0324","stream":true,"tools": [{"type": "function", "function": {"name": "query_weather", "description": "Get weather of an city, the user should supply a city first", "parameters": {"type": "object", "properties": {"city": {"type": "string", "description": "The city, e.g. Beijing"}}, "required": ["city"]}}}], "messages": [{"role": "user", "content": "Hows the weather like in Qingdao today"}]}'
|
||||
-d '{"temperature": 0, "max_tokens": 100, "model": "deepseek-ai/DeepSeek-V3-0324","stream":true,"tools": [{"type": "function", "function": {"name": "query_weather", "description": "Get weather of a city, the user should supply a city first", "parameters": {"type": "object", "properties": {"city": {"type": "string", "description": "The city, e.g. Beijing"}}, "required": ["city"]}}}], "messages": [{"role": "user", "content": "How'\''s the weather like in Qingdao today"}]}'
|
||||
```
|
||||
Expected Streamed Chunks (simplified for clarity):
|
||||
```text Output
|
||||
@@ -284,10 +323,11 @@ The client needs to concatenate all arguments fragments to reconstruct the compl
|
||||
```text Output
|
||||
{"city": "Qingdao"}
|
||||
```
|
||||
<Callout icon="key" color="#FFC107" iconType="regular">
|
||||
|
||||
<Warning>
|
||||
1. Use a lower `"temperature"` value for better results.
|
||||
2. To receive more consistent tool call results, it is recommended to use `--chat-template examples/chat_template/tool_chat_template_deepseekv3.jinja`. It provides an improved unified prompt.
|
||||
</Callout>
|
||||
</Warning>
|
||||
|
||||
|
||||
### Thinking Budget for DeepSeek R1
|
||||
@@ -302,7 +342,6 @@ python3 -m sglang.launch_server --model deepseek-ai/DeepSeek-R1 --tp 8 --port 30
|
||||
|
||||
Sample Request:
|
||||
|
||||
<CodeGroup>
|
||||
```python Sample Request
|
||||
import openai
|
||||
from rich.pretty import pprint
|
||||
@@ -328,7 +367,6 @@ response = client.chat.completions.create(
|
||||
)
|
||||
pprint(response)
|
||||
```
|
||||
</CodeGroup>
|
||||
|
||||
## FAQ
|
||||
|
||||
|
||||
@@ -1,13 +1,12 @@
|
||||
---
|
||||
title: "DeepSeek V3.2 Usage"
|
||||
title: "DeepSeek V3.2/GLM-5 Usage"
|
||||
metatags:
|
||||
description: "Deploy DeepSeek V3.2 with SGLang: DeepSeek Sparse Attention (DSA), long-context optimization, MTP speculative decoding, function calling. Supports H200, B200, MI300X, MI350."
|
||||
description: "Deploy DeepSeek V3.2/GLM-5 with SGLang: DeepSeek Sparse Attention (DSA), long-context optimization, MTP speculative decoding, function calling. Supports H200, B200, MI300X, MI350."
|
||||
---
|
||||
DeepSeek-V3.2 model family equips DeepSeek-V3.1-Terminus with DeepSeek Sparse Attention (DSA) through continued training. With DSA, a fine-grained sparse attention mechanism powered by a lightning indexer, DeepSeek-V3.2 achieves efficiency improvements in long-context scenarios.
|
||||
|
||||
For reporting issues or tracking upcoming features, please refer to this [Roadmap](https://github.com/sgl-project/sglang/issues/11060).
|
||||
|
||||
Note: This document is originally written for the usage of [DeepSeek-V3.2-Exp](https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Exp) model. The usage of [DeepSeek-V3.2](https://huggingface.co/deepseek-ai/DeepSeek-V3.2) or [DeepSeek-V3.2-Speciale](https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Speciale) is the same as DeepSeek-V3.2-Exp except for the tool call parser.
|
||||
Note: This document is originally written for the usage of [DeepSeek-V3.2-Exp](https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Exp) model. The usage of [DeepSeek-V3.2](https://huggingface.co/deepseek-ai/DeepSeek-V3.2) or [DeepSeek-V3.2-Speciale](https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Speciale) is the same as DeepSeek-V3.2-Exp except for the tool call parser. [GLM-5](https://huggingface.co/zai-org/GLM-5) model also applies DSA (DeepSeek Sparse Attention) structure, so it can share most of the usage here, except for the reasoning parser and tool call parser.
|
||||
|
||||
|
||||
## Installation
|
||||
@@ -41,7 +40,8 @@ cd sglang
|
||||
pip3 install pip --upgrade
|
||||
pip3 install -e "python"
|
||||
```
|
||||
## Launch DeepSeek V3.2 with SGLang
|
||||
|
||||
## Launch DeepSeek V3.2/GLM-5 with SGLang
|
||||
|
||||
To serve [DeepSeek-V3.2-Exp](https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Exp) on 8xH200/B200 GPUs:
|
||||
|
||||
@@ -59,19 +59,25 @@ python -m sglang.launch_server --model deepseek-ai/DeepSeek-V3.2-Exp --tp 8
|
||||
python3 -m sglang.launch_server --model deepseek-ai/DeepSeek-V3.2-Exp --tp 8 --nsa-prefill-backend tilelang --nsa-decode-backend tilelang
|
||||
```
|
||||
|
||||
To serve GLM-5, just replace the `--model` argument with `zai-org/GLM-5-FP8`.
|
||||
|
||||
### Configuration Tips
|
||||
- **DP Attention (Recommended)**: For DeepSeek V3.2 model, the kernels are customized for the use case of `dp_size=8`, so DP attention (`--dp 8 --enable-dp-attention`) is the recommended configuration for better stability and performance. All test cases use this configuration by default.
|
||||
- **Pure TP Mode**: Launching with pure TP (without `--dp` and `--enable-dp-attention`) is also supported. Note that this mode has not been fully validated in PD disaggregation scenarios.
|
||||
- **Short-sequence MHA prefill (adaptive)**: For short prefill sequences (default threshold: **2048 tokens**), the NSA backend uses standard MHA automatically (no extra flags). On H200 (SM90) this path uses the FlashAttention variable-length kernel; on B200 (SM100) it uses TRT-LLM ragged MHA. MHA uses `MHA_ONE_SHOT` for best performance. `MHA_ONE_SHOT` computes multi-head attention over all tokens (both cached prefix and newly extended tokens) in a single kernel invocation, avoiding the overhead of chunked KV cache processing. This achieves optimal throughput for short sequences where total sequence length fits within the chunk capacity limit.
|
||||
- **DP Attention**: To enable [DP Attention](../advanced_features/dp_dpa_smg_guide), please include `--enable-dp-attention --dp <dp-size>` in command. DP Attention is better for large concurrency scenarios.
|
||||
- **TP Attention**: Launching with TP attention is also supported. TP attention is better for low latency scenarios.
|
||||
- **Short-sequence MHA prefill (adaptive)**: For short prefill sequences (default threshold: **2048 tokens**), the NSA backend uses standard MHA automatically (no extra flags). On H200 (SM90) this path uses the FlashAttention variable-length kernel; on B200 (SM100) it uses TRT-LLM ragged MHA. MHA uses `MHA_ONE_SHOT` for best performance, which computes multi-head attention over all tokens (both cached prefix and newly extended tokens) in a single kernel invocation, avoiding the overhead of chunked KV cache processing. This achieves optimal throughput for short sequences where total sequence length fits within the chunk capacity limit.
|
||||
- **MHA prefill threshold relaxation**: To apply MHA attention to requests longer than 2048 tokens, please set the flag `SGLANG_NSA_PREFILL_DENSE_ATTN_KV_LEN_THRESHOLD` to a value larger than 2048. As threshold grows larger, the prefill performance can be improved, but at the cost of potential accuracy drop.
|
||||
- **Choices of Attention Kernels**: The attention backend is automatically set to `nsa` attention backend for DeepSeek V3.2 model. In this backend, different kernels for sparse prefilling/decoding are implemented, which can be specified by `--nsa-prefill-backend` and `--nsa-decode-backend` server arguments. The choices of nsa prefill/decode attention kernels include:
|
||||
- `flashmla_sparse`: `flash_mla_sparse_fwd` kernel from `flash_mla` library. Can run on both Hopper and Blackwell GPUs. It requires bf16 q, kv inputs.
|
||||
- `flashmla_kv`: `flash_mla_with_kvcache` kernel from `flash_mla` library. Can run on both Hopper and Blackwell GPUs. It requires bf16 q, fp8 k_cache inputs.
|
||||
- `flashmla_auto`: enables automatic selection of either `flashmla_sparse` or `flashmla_kv` kernel for prefill based on KV cache dtype, hardware, and heuristics. With BF16 KV cache, `flashmla_sparse` is always used on both Hopper and Blackwell. With FP8 KV cache: On Hopper (SM90), it unconditionally uses `flashmla_kv`; On Blackwell (SM100), it uses `flashmla_sparse` when `total_kv_tokens < total_q_tokens * 512`, otherwise falls back to `flashmla_kv`. The heuristics may need to be tuned if the performance of either kernel changes significantly.
|
||||
- `fa3`: `flash_attn_with_kvcache` kernel from `flash_attn` library. Can only run on Hopper GPUs. It requires bf16 q, kv inputs.
|
||||
- `tilelang`: `tilelang` implementation that can run on GPU, HPU and NPU.
|
||||
- `aiter`: Aiter kernel on AMD HPUs. Can only be used as decode kernel.
|
||||
- On the basis of performance benchmarks, the default configuration on H200 and B200 are set as follows :
|
||||
- H200: `flashmla_sparse` prefill attention (short-seq prefill uses MHA via FlashAttention varlen), `fa3` decode attention, `bf16` kv cache dtype.
|
||||
- B200: `flashmla_auto` prefill attention (short-seq prefill uses MHA via TRT-LLM ragged), `flashmla_kv` decode attention, `fp8_e4m3` kv cache dtype. `flashmla_auto` enables automatic selection of either `flashmla_sparse` or `flashmla_kv` kernel for prefill based on KV cache dtype, hardware, and heuristics. When FP8 KV cache is enabled and `total_kv_tokens < total_q_tokens * 512`, it uses the `flashmla_sparse` kernel; otherwise, it falls back to the `flashmla_kv` kernel. The heuristics may need to be tuned if the performance of either the `flashmla_sparse` or `flashmla_kv` kernel changes significantly.
|
||||
- `trtllm`: `trtllm-mla` sparse kernel from flashinfer library. Only run on blackwell GPUs. It requires q,k,v to be uniformly bf16 or fp8_e4m3 format.
|
||||
- On the basis of performance benchmarks, the default configuration of DSA kernels on Hopper and Blackwell are set as follows :
|
||||
- Bfloat 16 kv cache: On Hopper, `flashmla_sparse` prefill attention, `fa3` decode attention; On Blackwell, `flashmla_sparse` prefill attention, `trtllm` decode attention
|
||||
- Float8_e4m3fn KV cache: On Hopper, `flashmla_kv` prefill attention, `flashmla_kv` decode attention; On Blackwell, `trtllm` prefill attention and `trtllm` decode attention.
|
||||
- **Index Cache**: Introduce in [this paper](https://arxiv.org/abs/2603.12201), IndexCache improves speed by reusing the result of indexer across different layers, only at cost of negligible accuracy loss. For **GLM-5** model, we recommend appending `--json-model-override-args '{"index_topk_pattern": "FFSFSSSFSSFFFSSSFFFSFSSSSSSFFSFFSFFSSFFFFFFSFFFFFSFFSSSSSSFSFFFSFSSSFSFFSFFSSS"}'` to command for better tradeoff between speedup and performance.
|
||||
|
||||
## Multi-token Prediction
|
||||
SGLang implements Multi-Token Prediction (MTP) for DeepSeek V3.2 based on [EAGLE speculative decoding](../advanced_features/speculative_decoding#EAGLE-Decoding). With this optimization, the decoding speed can be improved significantly on small batch sizes. Please look at [this PR](https://github.com/sgl-project/sglang/pull/11652) for more information.
|
||||
@@ -90,7 +96,7 @@ python -m sglang.launch_server --model deepseek-ai/DeepSeek-V3.2-Exp --tp 8 --sp
|
||||
- The default value of `--max-running-requests` is set to `48` for MTP. For larger batch sizes, this value should be increased beyond the default value.
|
||||
|
||||
<Tip>
|
||||
To enable the experimental overlap scheduler for EAGLE speculative decoding, set the environment variable `SGLANG_ENABLE_SPEC_V2=1`. This can improve performance by enabling overlap scheduling between draft and verification stages.
|
||||
To enable overlap scheduler for EAGLE speculative decoding, we recommend setting the environment variable `SGLANG_ENABLE_SPEC_V2=1`. This can improve performance by enabling overlap scheduling between draft and verification stages.
|
||||
</Tip>
|
||||
|
||||
|
||||
@@ -98,10 +104,7 @@ To enable the experimental overlap scheduler for EAGLE speculative decoding, set
|
||||
The usage of function calling and reasoning parser is the same as DeepSeek V3.1. Please refer to [Reasoning Parser](../advanced_features/separate_reasoning) and [Tool Parser](../advanced_features/tool_parser) documents.
|
||||
|
||||
To launch `DeepSeek-V3.2-Exp` with function calling and reasoning parser:
|
||||
<Note>
|
||||
It is recommended to specify the chat-template, ensuring that you are within the sglang's root directory.
|
||||
</Note>
|
||||
|
||||
> Note: It is recommended to specify the chat-template, ensuring that you are within the sglang's root directory.
|
||||
```bash Command
|
||||
python3 -m sglang.launch_server \
|
||||
--model-path deepseek-ai/DeepSeek-V3.2-Exp \
|
||||
@@ -122,7 +125,7 @@ python3 -m sglang.launch_server \
|
||||
--reasoning-parser deepseek-v3
|
||||
```
|
||||
|
||||
`DeepSeek-V3.2-Speciale` doesn't support tool calling, so can only be launched with reasoning parser:
|
||||
`DeepSeek-V3.2-Speciale` does not support tool calling, so it can only be launched with the reasoning parser:
|
||||
```bash Command
|
||||
python3 -m sglang.launch_server \
|
||||
--model-path deepseek-ai/DeepSeek-V3.2-Speciale \
|
||||
@@ -131,6 +134,23 @@ python3 -m sglang.launch_server \
|
||||
--reasoning-parser deepseek-v3
|
||||
```
|
||||
|
||||
To launch `GLM-5` with function calling and reasoning parser:
|
||||
```bash Command
|
||||
python -m sglang.launch_server \
|
||||
--model zai-org/GLM-5-FP8 \
|
||||
--tp-size 8 --dp-size 8 --enable-dp-attention \
|
||||
--tool-call-parser glm47 \
|
||||
--reasoning-parser glm45 \
|
||||
```
|
||||
|
||||
## NVFP4 Checkpoint
|
||||
|
||||
To launch deepseek v3.2 [NVFP4 checkpoint](https://huggingface.co/nvidia/DeepSeek-V3.2-NVFP4) on Blackwell devices, the user needs to specify the quantization method as `modelopt_fp4`, and moe runner backend as one of `flashinfer_trtllm`(recommended), `flashinfer_cutlass` and `flashinfer_cutedsl`. Any other usage (parallelism, reasoning parser, ...) is the same as FP8 checkpoint.
|
||||
|
||||
An example launching command can be:
|
||||
```bash Command
|
||||
python -m sglang.launch_server --model nvidia/DeepSeek-V3.2-NVFP4 --tp 4 --quantization modelopt_fp4 --moe-runner-backend flashinfer_trtllm --tool-call-parser deepseekv32 --reasoning-parser deepseek-v3
|
||||
```
|
||||
|
||||
## PD Disaggregation
|
||||
|
||||
@@ -174,7 +194,7 @@ python -m sglang_router.launch_router --pd-disaggregation \
|
||||
--port 8000 \
|
||||
```
|
||||
|
||||
If you need more advanced deployment methods or production-ready deployment methods, such as RBG or LWS-based deployment, please refer to [references/multi_node_deployment/rbg_pd/deepseekv32_pd](../references/multi_node_deployment/rbg_pd/deepseekv32_pd). Additionally, you can also find startup commands for DeepEP-based EP parallelism in the aforementioned documentation.
|
||||
If you need more advanced deployment methods or production-ready deployment methods, such as RBG or LWS-based deployment, please refer to [references/multi_node_deployment/rbg_pd/deepseekv32_pd.md](../references/multi_node_deployment/rbg_pd/deepseekv32_pd). Additionally, you can also find startup commands for DeepEP-based EP parallelism in the aforementioned documentation.
|
||||
|
||||
|
||||
## Benchmarking Results
|
||||
@@ -215,7 +235,7 @@ Repeat: 8, mean: 0.797
|
||||
Scores: ['0.808', '0.798', '0.808', '0.798', '0.783', '0.788', '0.803', '0.793']
|
||||
```
|
||||
|
||||
For Deepseek V3.2, Deepseek recommends setting the sampling parameters to temperature = 1.0, top_p = 0.95:
|
||||
For DeepSeek V3.2, DeepSeek recommends setting the sampling parameters to temperature = 1.0, top_p = 0.95:
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.test.run_eval --port 30000 --eval-name gpqa --num-examples 198 --max-tokens 128000 --repeat 8 --top-p 0.95 --temperature 1.0 --thinking-mode deepseek-v3
|
||||
@@ -223,13 +243,13 @@ python3 -m sglang.test.run_eval --port 30000 --eval-name gpqa --num-examples 198
|
||||
Repeat: 8, mean: 0.840
|
||||
Scores: ['0.848', '0.808', '0.848', '0.838', '0.879', '0.813', '0.838', '0.848']
|
||||
```
|
||||
which matches the official score, 0.824, as reported in the [Deepseek-V3.2 technical report](https://huggingface.co/deepseek-ai/DeepSeek-V3.2/blob/main/assets/paper.pdf).
|
||||
which matches the official score, 0.824, as reported in the [DeepSeek-V3.2 technical report](https://huggingface.co/deepseek-ai/DeepSeek-V3.2/blob/main/assets/paper.pdf).
|
||||
|
||||
### Accuracy Test with `aime 2025`
|
||||
|
||||
Prepare the environment by installing NeMo-Skills in the docker or your own virtual environment:
|
||||
|
||||
```text Output
|
||||
```
|
||||
pip install git+https://github.com/NVIDIA/NeMo-Skills.git --ignore-installed blinker
|
||||
```
|
||||
|
||||
@@ -272,7 +292,7 @@ ns eval \
|
||||
|
||||
Test results (8*B200):
|
||||
|
||||
DeepSeek-V3.2-Exp:
|
||||
DeepSeek-V3.2-Exp:
|
||||
|
||||
<table style={{width: "100%", borderCollapse: "collapse", tableLayout: "fixed"}}>
|
||||
<colgroup>
|
||||
@@ -420,23 +440,19 @@ DeepSeek-V3.2-Speciale:
|
||||
</table>
|
||||
|
||||
|
||||
|
||||
## DSA long sequence context parallel optimization(experimental)
|
||||
|
||||
**Note: This feature is only verified on Hopper machines**
|
||||
|
||||
For context parallel in DeepSeek V3.2 model, we provide two different modes of splitting tokens, which can be controlled with argument `--nsa-prefill-cp-mode`.
|
||||
|
||||
### In sequence splitting (default setting)
|
||||
### In sequence splitting
|
||||
|
||||
The first mode can be enabled by `--nsa-prefill-cp-mode in-seq-split`. This mode implements context parallel for DSA by splitting the sequence uniformly between context parallel ranks. At attention stage, each cp rank computes the indexer results of sharded sequence, and collects the whole kv cache through all gather operator.
|
||||
The first mode can be enabled by `--nsa-prefill-cp-mode in-seq-split`. This mode implements context parallel for DSA by splitting the sequence uniformly between context parallel ranks. At attention stage, each cp rank computes the indexer results of sharded sequence, and collects the whole kv cache through all gather operator. Add `attn_cp_size` for communication group for context parallel.
|
||||
|
||||
The communication group for context parallel reuses the one for attention tp, thus `cp_size` equals `atten_tp_size = tp_size / dp_size`.
|
||||
|
||||
Note that in sequence splitting mode has the following restrictions:
|
||||
Note that the in-sequence splitting mode has the following restrictions:
|
||||
- The batch size is restricted to 1 for prefill batches
|
||||
- Multi-node/PD disaggregation is still not supported
|
||||
- `moe_dense_tp_size=1`, `kv_cache_dtype = "bf16"`, `moe_a2a_backend = "deepep"`
|
||||
- `moe_dense_tp_size=1`, `moe_a2a_backend = "deepep"`
|
||||
- To ensure `cp_size > 1`, the passed in `tp_size` must be larger than `dp_size`
|
||||
|
||||
For more details, please refer to PR https://github.com/sgl-project/sglang/pull/12065.
|
||||
@@ -444,21 +460,21 @@ For more details, please refer to PR https://github.com/sgl-project/sglang/pull/
|
||||
Example:
|
||||
```bash Command
|
||||
# In-seq splitting mode launched with EP + DP
|
||||
python -m sglang.launch_server --model deepseek-ai/DeepSeek-V3.2-Exp --tp 8 --ep 8 --dp 2 --enable-dp-attention --enable-nsa-prefill-context-parallel --nsa-prefill-cp-mode in-seq-split --max-running-requests 32
|
||||
python -m sglang.launch_server --model deepseek-ai/DeepSeek-V3.2-Exp --tp 8 --ep 8 --dp 2 --enable-dp-attention --enable-nsa-prefill-context-parallel --attn-cp-size 4 --nsa-prefill-cp-mode in-seq-split --max-running-requests 32
|
||||
```
|
||||
|
||||
### Round robin splitting
|
||||
### Round robin splitting (default setting)
|
||||
|
||||
This mode can be enabled by specifying the parameter `--nsa-prefill-cp-mode round-robin-split`, which distributes tokens across ranks based on `token_idx % cp_size`.
|
||||
|
||||
In this scenario, compared with the aforementioned method, it additionally supports the fused MoE backend (the fused MoE backend may deliver better performance than DeepEP in single-machine scenarios), FP8 KV-cache, and multi-batch prefill inference. But it cannot be enabled with dp attention together.
|
||||
In this scenario, compared to the in-sequence splitting method, it additionally supports the fused MoE backend (the fused MoE backend may deliver better performance than DeepEP in single-machine scenarios), FP8 KV-cache, and multi-batch prefill inference. However, it cannot be enabled with DP attention together.
|
||||
|
||||
For more details, please refer to PR https://github.com/sgl-project/sglang/pull/13959.
|
||||
|
||||
Example usage:
|
||||
```bash Command
|
||||
# Launch with FusedMoe + CP8
|
||||
python -m sglang.launch_server --model deepseek-ai/DeepSeek-V3.2-Exp --tp 8 --enable-nsa-prefill-context-parallel --nsa-prefill-cp-mode round-robin-split --max-running-requests 32
|
||||
python -m sglang.launch_server --model deepseek-ai/DeepSeek-V3.2-Exp --tp 8 --enable-nsa-prefill-context-parallel --attn-cp-size 8 --nsa-prefill-cp-mode round-robin-split --max-running-requests 32
|
||||
```
|
||||
### Pipeline Parallel + Context Parallel (PP + CP)
|
||||
|
||||
@@ -482,6 +498,7 @@ python3 -m sglang.launch_server \
|
||||
--tp 8 --pp-size 2 \
|
||||
--dp-size 1 --moe-dense-tp-size 1 \
|
||||
--enable-nsa-prefill-context-parallel \
|
||||
--attn-cp-size 8 \
|
||||
--nsa-prefill-cp-mode round-robin-split \
|
||||
--trust-remote-code \
|
||||
--disable-radix-cache \
|
||||
@@ -505,6 +522,7 @@ python3 -m sglang.launch_server \
|
||||
--tp 8 --pp-size 2 \
|
||||
--dp-size 1 --moe-dense-tp-size 1 \
|
||||
--enable-nsa-prefill-context-parallel \
|
||||
--attn-cp-size 8 \
|
||||
--nsa-prefill-cp-mode round-robin-split \
|
||||
--trust-remote-code \
|
||||
--disable-radix-cache \
|
||||
@@ -532,6 +550,7 @@ python -m sglang.launch_server \
|
||||
--tp 8 --pp-size 2 \
|
||||
--dp-size 1 --moe-dense-tp-size 1 \
|
||||
--enable-nsa-prefill-context-parallel \
|
||||
--attn-cp-size 8 \
|
||||
--nsa-prefill-cp-mode round-robin-split \
|
||||
--disaggregation-ib-device mlx5_bond_0,mlx5_bond_1,mlx5_bond_2,mlx5_bond_3 \
|
||||
--trust-remote-code \
|
||||
@@ -557,6 +576,7 @@ python -m sglang.launch_server \
|
||||
--tp 8 --pp-size 2 \
|
||||
--dp-size 1 --moe-dense-tp-size 1 \
|
||||
--enable-nsa-prefill-context-parallel \
|
||||
--attn-cp-size 8 \
|
||||
--nsa-prefill-cp-mode round-robin-split \
|
||||
--disaggregation-ib-device mlx5_bond_0,mlx5_bond_1,mlx5_bond_2,mlx5_bond_3 \
|
||||
--trust-remote-code \
|
||||
@@ -573,3 +593,9 @@ python -m sglang.launch_server \
|
||||
```
|
||||
|
||||
For the Decode nodes, it is recommended to use the **EP mode**.
|
||||
|
||||
## HiSparse: Hierarchical Sparse Attention for DSA (experimental)
|
||||
|
||||
HiSparse reduces per-request GPU memory during decode by keeping only a small "hot" KV buffer on GPU while storing complete KV data in CPU pinned memory. A CUDA kernel dynamically swaps in the top-k most relevant KV entries from host memory on each decode step. This enables significantly higher decode concurrency for long-context DSA models.
|
||||
|
||||
HiSparse currently requires PD disaggregation mode and is enabled on the decode instance only. For detailed design, configuration, and deployment instructions, see the [HiSparse Guide](../advanced_features/hisparse_guide).
|
||||
|
||||
@@ -3,7 +3,6 @@ title: "Launch GLM-4.5 / GLM-4.6 / GLM-4.7 with SGLang"
|
||||
metatags:
|
||||
description: "Deploy GLM-4.5/4.6/4.7 models with SGLang: FP8 inference, EAGLE speculative decoding, function calling support. Optimized for H100/H200 GPUs."
|
||||
---
|
||||
|
||||
## Launch GLM-4.5 / GLM-4.6 / GLM-4.7 with SGLang
|
||||
|
||||
To serve GLM-4.5 / GLM-4.6 FP8 models on 8xH100/H200 GPUs:
|
||||
|
||||
@@ -136,4 +136,4 @@ python -m sglang.launch_server \
|
||||
|
||||
In SGLang, we can implement thinking budget with `CustomLogitProcessor`.
|
||||
|
||||
Launch a server with `--enable-custom-logit-processor` flag on. and using `Glm4MoeThinkingBudgetLogitProcessor` in the request likes `GLM-4.6` example in [glm45](./glm45).
|
||||
Launch a server with the `--enable-custom-logit-processor` flag. Then, use `Glm4MoeThinkingBudgetLogitProcessor` in the request, similar to the `GLM-4.6` example in [glm45.md](./glm45).
|
||||
|
||||
@@ -1,16 +1,17 @@
|
||||
---
|
||||
title: "MiniMax M2.1/M2 Usage"
|
||||
title: "MiniMax M2.5/M2.1/M2 Usage"
|
||||
metatags:
|
||||
description: "Deploy MiniMax M2.1/M2 with SGLang: 230B MoE model (10B active), up to 3M context, optimized for coding and agentic tasks, tool use support."
|
||||
description: "Deploy MiniMax M2.5/M2.1/M2 with SGLang: 230B MoE model (10B active), up to 3M context, optimized for coding and agentic tasks, tool use support."
|
||||
---
|
||||
[MiniMax-M2.1](https://huggingface.co/MiniMaxAI/MiniMax-M2.1) and [MiniMax-M2](https://huggingface.co/MiniMaxAI/MiniMax-M2) are advanced large language models created by [MiniMax](https://www.minimax.io/).
|
||||
[MiniMax-M2.5](https://huggingface.co/MiniMaxAI/MiniMax-M2.5), [MiniMax-M2.1](https://huggingface.co/MiniMaxAI/MiniMax-M2.1), and [MiniMax-M2](https://huggingface.co/MiniMaxAI/MiniMax-M2) are advanced large language models created by [MiniMax](https://www.minimax.io/).
|
||||
|
||||
MiniMax-M2 series redefines efficiency for agents. It's a compact, fast, and cost-effective MoE model (230 billion total parameters with 10 billion active parameters) built for elite performance in coding and agentic tasks, all while maintaining powerful general intelligence. With just 10 billion activated parameters, MiniMax-M2 provides the sophisticated, end-to-end tool use performance expected from today's leading models, but in a streamlined form factor that makes deployment and scaling easier than ever.
|
||||
The MiniMax-M2 series redefines efficiency for agents. These compact, fast, and cost-effective MoE models (230 billion total parameters with 10 billion active parameters) are built for elite performance in coding and agentic tasks, all while maintaining powerful general intelligence. With just 10 billion activated parameters, the MiniMax-M2 series provides sophisticated, end-to-end tool use performance expected from today's leading models, but in a streamlined form factor that makes deployment and scaling easier than ever.
|
||||
|
||||
## Supported Models
|
||||
|
||||
This guide applies to the following models. You only need to update the model name during deployment. The following examples use **MiniMax-M2**:
|
||||
|
||||
- [MiniMaxAI/MiniMax-M2.5](https://huggingface.co/MiniMaxAI/MiniMax-M2.5)
|
||||
- [MiniMaxAI/MiniMax-M2.1](https://huggingface.co/MiniMaxAI/MiniMax-M2.1)
|
||||
- [MiniMaxAI/MiniMax-M2](https://huggingface.co/MiniMaxAI/MiniMax-M2)
|
||||
|
||||
@@ -52,6 +53,24 @@ python -m sglang.launch_server \
|
||||
--mem-fraction-static 0.85
|
||||
```
|
||||
|
||||
### AMD GPUs (MI300X/MI325X/MI355X)
|
||||
|
||||
8-GPU deployment command:
|
||||
|
||||
```bash Command
|
||||
SGLANG_USE_AITER=1 python -m sglang.launch_server \
|
||||
--model-path MiniMaxAI/MiniMax-M2.5 \
|
||||
--tp-size 8 \
|
||||
--ep-size 8 \
|
||||
--attention-backend aiter \
|
||||
--tool-call-parser minimax-m2 \
|
||||
--reasoning-parser minimax-append-think \
|
||||
--host 0.0.0.0 \
|
||||
--trust-remote-code \
|
||||
--port 8000 \
|
||||
--mem-fraction-static 0.85
|
||||
```
|
||||
|
||||
## Testing Deployment
|
||||
|
||||
After startup, you can test the SGLang OpenAI-compatible API with the following command:
|
||||
|
||||
@@ -10,7 +10,7 @@
|
||||
"\n",
|
||||
"- `/generate` (text generation model)\n",
|
||||
"- `/get_model_info`\n",
|
||||
"- `/get_server_info`\n",
|
||||
"- `/server_info`\n",
|
||||
"- `/health`\n",
|
||||
"- `/health_generate`\n",
|
||||
"- `/flush_cache`\n",
|
||||
@@ -49,7 +49,7 @@
|
||||
" \"python3 -m sglang.launch_server --model-path qwen/qwen2.5-0.5b-instruct --host 0.0.0.0 --log-level warning\"\n",
|
||||
")\n",
|
||||
"\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")"
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=server_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -57,7 +57,7 @@
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Generate (text generation model)\n",
|
||||
"Generate completions. This is similar to the `/v1/completions` in OpenAI API. Detailed parameters can be found in the [sampling parameters](sampling_params)."
|
||||
"Generate completions. This is similar to the `/v1/completions` in OpenAI API. Detailed parameters can be found in the [sampling parameters](sampling_params.md)."
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -140,7 +140,7 @@
|
||||
"metadata": {},
|
||||
"outputs": [],
|
||||
"source": [
|
||||
"url = f\"http://localhost:{port}/get_server_info\"\n",
|
||||
"url = f\"http://localhost:{port}/server_info\"\n",
|
||||
"\n",
|
||||
"response = requests.get(url)\n",
|
||||
"print_highlight(response.text)"
|
||||
@@ -185,7 +185,15 @@
|
||||
"source": [
|
||||
"## Flush Cache\n",
|
||||
"\n",
|
||||
"Flush the radix cache. It will be automatically triggered when the model weights are updated by the `/update_weights` API."
|
||||
"Flush the radix cache. It will be automatically triggered when the model weights are updated by the `/update_weights` API.\n",
|
||||
"\n",
|
||||
"Parameters:\n",
|
||||
"- `timeout` (query, float, default `0`, unit: seconds): Wait time for idle state before flushing. `0` means fail fast if not idle. When HiCache async operations are in-flight, a non-zero timeout allows the server to wait until idle before flushing, avoiding unnecessary 400 errors.\n",
|
||||
"\n",
|
||||
"```bash\n",
|
||||
"# With timeout (wait up to 30s for idle state)\n",
|
||||
"curl -s -X POST \"http://127.0.0.1:30000/flush_cache?timeout=30\"\n",
|
||||
"```"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -265,7 +273,7 @@
|
||||
"source": [
|
||||
"## Encode (embedding model)\n",
|
||||
"\n",
|
||||
"Encode text into embeddings. Note that this API is only available for [embedding models](openai_api_embeddings) and will raise an error for generation models.\n",
|
||||
"Encode text into embeddings. Note that this API is only available for [embedding models](openai_api_embeddings.ipynb) and will raise an error for generation models.\n",
|
||||
"Therefore, we launch a new server to server an embedding model."
|
||||
]
|
||||
},
|
||||
@@ -280,7 +288,7 @@
|
||||
" --host 0.0.0.0 --is-embedding --log-level warning\n",
|
||||
"\"\"\")\n",
|
||||
"\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")"
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=embedding_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -327,7 +335,7 @@
|
||||
" --host 0.0.0.0 --disable-radix-cache --chunked-prefill-size -1 --attention-backend triton --is-embedding --log-level warning\n",
|
||||
"\"\"\")\n",
|
||||
"\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")"
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=reranker_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -393,7 +401,7 @@
|
||||
" --host 0.0.0.0 --log-level warning\n",
|
||||
"\"\"\")\n",
|
||||
"\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")"
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=score_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -454,7 +462,7 @@
|
||||
"python3 -m sglang.launch_server --model-path Skywork/Skywork-Reward-Llama-3.1-8B-v0.2 --host 0.0.0.0 --is-embedding --log-level warning\n",
|
||||
"\"\"\")\n",
|
||||
"\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")"
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=reward_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -518,7 +526,7 @@
|
||||
" \"python3 -m sglang.launch_server --model-path Qwen/Qwen1.5-MoE-A2.7B --host 0.0.0.0 --expert-distribution-recorder-mode stat --log-level warning\"\n",
|
||||
")\n",
|
||||
"\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")"
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=expert_record_server_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -571,7 +579,7 @@
|
||||
"python3 -m sglang.launch_server --model-path qwen/qwen2.5-0.5b-instruct\n",
|
||||
"\"\"\")\n",
|
||||
"\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")"
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=tokenizer_free_server_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
|
||||
@@ -7,7 +7,7 @@ Apart from the OpenAI compatible APIs, the SGLang Runtime also provides its nati
|
||||
|
||||
- `/generate` (text generation model)
|
||||
- `/get_model_info`
|
||||
- `/get_server_info`
|
||||
- `/server_info`
|
||||
- `/health`
|
||||
- `/health_generate`
|
||||
- `/flush_cache`
|
||||
@@ -35,7 +35,7 @@ server_process, port = launch_server_cmd(
|
||||
"python3 -m sglang.launch_server --model-path qwen/qwen2.5-0.5b-instruct --host 0.0.0.0 --log-level warning"
|
||||
)
|
||||
|
||||
wait_for_server(f"http://localhost:{port}")
|
||||
wait_for_server(f"http://localhost:{port}", process=server_process)
|
||||
```
|
||||
|
||||
## Generate (text generation model)
|
||||
@@ -96,7 +96,7 @@ Gets the server information including CLI arguments, token limits, and memory po
|
||||
- `get_max_total_num_tokens`
|
||||
|
||||
```python Example
|
||||
url = f"http://localhost:{port}/get_server_info"
|
||||
url = f"http://localhost:{port}/server_info"
|
||||
|
||||
response = requests.get(url)
|
||||
print_highlight(response.text)
|
||||
@@ -124,6 +124,14 @@ print_highlight(response.text)
|
||||
|
||||
Flush the radix cache. It will be automatically triggered when the model weights are updated by the `/update_weights` API.
|
||||
|
||||
Parameters:
|
||||
- `timeout` (query, float, default `0`, unit: seconds): Wait time for idle state before flushing. `0` means fail fast if not idle. When HiCache async operations are in-flight, a non-zero timeout allows the server to wait until idle before flushing, avoiding unnecessary 400 errors.
|
||||
|
||||
```bash Command
|
||||
# With timeout (wait up to 30s for idle state)
|
||||
curl -s -X POST "http://127.0.0.1:30000/flush_cache?timeout=30"
|
||||
```
|
||||
|
||||
```python Example
|
||||
url = f"http://localhost:{port}/flush_cache"
|
||||
|
||||
@@ -176,14 +184,12 @@ Encode text into embeddings. Note that this API is only available for [embedding
|
||||
Therefore, we launch a new server to server an embedding model.
|
||||
|
||||
```python Example
|
||||
embedding_process, port = launch_server_cmd(
|
||||
"""
|
||||
embedding_process, port = launch_server_cmd("""
|
||||
python3 -m sglang.launch_server --model-path Alibaba-NLP/gte-Qwen2-1.5B-instruct \
|
||||
--host 0.0.0.0 --is-embedding --log-level warning
|
||||
"""
|
||||
)
|
||||
""")
|
||||
|
||||
wait_for_server(f"http://localhost:{port}")
|
||||
wait_for_server(f"http://localhost:{port}", process=embedding_process)
|
||||
```
|
||||
|
||||
```python Example
|
||||
@@ -205,14 +211,12 @@ terminate_process(embedding_process)
|
||||
Rerank a list of documents given a query using a cross-encoder model. Note that this API is only available for cross encoder model like [BAAI/bge-reranker-v2-m3](https://huggingface.co/BAAI/bge-reranker-v2-m3) with `attention-backend` `triton` and `torch_native`.
|
||||
|
||||
```python Example
|
||||
reranker_process, port = launch_server_cmd(
|
||||
"""
|
||||
reranker_process, port = launch_server_cmd("""
|
||||
python3 -m sglang.launch_server --model-path BAAI/bge-reranker-v2-m3 \
|
||||
--host 0.0.0.0 --disable-radix-cache --chunked-prefill-size -1 --attention-backend triton --is-embedding --log-level warning
|
||||
"""
|
||||
)
|
||||
""")
|
||||
|
||||
wait_for_server(f"http://localhost:{port}")
|
||||
wait_for_server(f"http://localhost:{port}", process=reranker_process)
|
||||
```
|
||||
|
||||
```python Example
|
||||
@@ -253,14 +257,12 @@ Parameters:
|
||||
The response contains `scores` - a list of probability lists, one per item, each in the order of `label_token_ids`.
|
||||
|
||||
```python Example
|
||||
score_process, port = launch_server_cmd(
|
||||
"""
|
||||
score_process, port = launch_server_cmd("""
|
||||
python3 -m sglang.launch_server --model-path qwen/qwen2.5-0.5b-instruct \
|
||||
--host 0.0.0.0 --log-level warning
|
||||
"""
|
||||
)
|
||||
""")
|
||||
|
||||
wait_for_server(f"http://localhost:{port}")
|
||||
wait_for_server(f"http://localhost:{port}", process=score_process)
|
||||
```
|
||||
|
||||
```python Example
|
||||
@@ -297,13 +299,11 @@ SGLang Runtime also supports reward models. Here we use a reward model to classi
|
||||
# Note that SGLang now treats embedding models and reward models as the same type of models.
|
||||
# This will be updated in the future.
|
||||
|
||||
reward_process, port = launch_server_cmd(
|
||||
"""
|
||||
reward_process, port = launch_server_cmd("""
|
||||
python3 -m sglang.launch_server --model-path Skywork/Skywork-Reward-Llama-3.1-8B-v0.2 --host 0.0.0.0 --is-embedding --log-level warning
|
||||
"""
|
||||
)
|
||||
""")
|
||||
|
||||
wait_for_server(f"http://localhost:{port}")
|
||||
wait_for_server(f"http://localhost:{port}", process=reward_process)
|
||||
```
|
||||
|
||||
```python Example
|
||||
@@ -347,7 +347,7 @@ expert_record_server_process, port = launch_server_cmd(
|
||||
"python3 -m sglang.launch_server --model-path Qwen/Qwen1.5-MoE-A2.7B --host 0.0.0.0 --expert-distribution-recorder-mode stat --log-level warning"
|
||||
)
|
||||
|
||||
wait_for_server(f"http://localhost:{port}")
|
||||
wait_for_server(f"http://localhost:{port}", process=expert_record_server_process)
|
||||
```
|
||||
|
||||
```python Example
|
||||
@@ -376,13 +376,11 @@ terminate_process(expert_record_server_process)
|
||||
This example demonstrates how to use the /tokenize and /detokenize endpoints together. We first tokenize a string, then detokenize the resulting IDs to reconstruct the original text. This workflow is useful when you need to handle tokenization externally but still leverage the server for detokenization.
|
||||
|
||||
```python Example
|
||||
tokenizer_free_server_process, port = launch_server_cmd(
|
||||
"""
|
||||
tokenizer_free_server_process, port = launch_server_cmd("""
|
||||
python3 -m sglang.launch_server --model-path qwen/qwen2.5-0.5b-instruct
|
||||
"""
|
||||
)
|
||||
""")
|
||||
|
||||
wait_for_server(f"http://localhost:{port}")
|
||||
wait_for_server(f"http://localhost:{port}", process=tokenizer_free_server_process)
|
||||
```
|
||||
|
||||
```python Example
|
||||
|
||||
@@ -66,7 +66,7 @@
|
||||
"import asyncio\n",
|
||||
"\n",
|
||||
"import sglang as sgl\n",
|
||||
"import sglang.test.doc_patch\n",
|
||||
"import sglang.test.doc_patch # noqa: F401\n",
|
||||
"from sglang.utils import async_stream_and_merge, stream_and_merge\n",
|
||||
"\n",
|
||||
"llm = sgl.Engine(model_path=\"qwen/qwen2.5-0.5b-instruct\")"
|
||||
|
||||
@@ -14,7 +14,7 @@
|
||||
"- `chat/completions`\n",
|
||||
"- `completions`\n",
|
||||
"\n",
|
||||
"Check out other tutorials to learn about [vision APIs](openai_api_vision) for vision-language models and [embedding APIs](openai_api_embeddings) for embedding models."
|
||||
"Check out other tutorials to learn about [vision APIs](openai_api_vision.ipynb) for vision-language models and [embedding APIs](openai_api_embeddings.ipynb) for embedding models."
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -39,7 +39,7 @@
|
||||
" \"python3 -m sglang.launch_server --model-path qwen/qwen2.5-0.5b-instruct --host 0.0.0.0 --log-level warning\"\n",
|
||||
")\n",
|
||||
"\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=server_process)\n",
|
||||
"print(f\"Server started on http://localhost:{port}\")"
|
||||
]
|
||||
},
|
||||
@@ -477,7 +477,7 @@
|
||||
"source": [
|
||||
"## Structured Outputs (JSON, Regex, EBNF)\n",
|
||||
"\n",
|
||||
"For OpenAI compatible structured outputs API, refer to [Structured Outputs](../advanced_features/structured_outputs) for more details.\n"
|
||||
"For OpenAI compatible structured outputs API, refer to [Structured Outputs](../advanced_features/structured_outputs.ipynb) for more details.\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -496,7 +496,7 @@
|
||||
" --lora-paths adapter_a=/path/to/adapter_a adapter_b=/path/to/adapter_b\n",
|
||||
"```\n",
|
||||
"\n",
|
||||
"For more details on LoRA serving configuration, see the [LoRA documentation](../advanced_features/lora).\n",
|
||||
"For more details on LoRA serving configuration, see the [LoRA documentation](../advanced_features/lora.ipynb).\n",
|
||||
"\n",
|
||||
"**API Call:**\n",
|
||||
"\n",
|
||||
|
||||
@@ -9,7 +9,7 @@
|
||||
"SGLang provides OpenAI-compatible APIs to enable a smooth transition from OpenAI services to self-hosted local models.\n",
|
||||
"A complete reference for the API is available in the [OpenAI API Reference](https://platform.openai.com/docs/guides/embeddings).\n",
|
||||
"\n",
|
||||
"This tutorial covers the embedding APIs for embedding models. For a list of the supported models see the [corresponding overview page](../supported_models/embedding_models)\n"
|
||||
"This tutorial covers the embedding APIs for embedding models. For a list of the supported models see the [corresponding overview page](../supported_models/retrieval_ranking/embedding_models.md)\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -35,7 +35,7 @@
|
||||
" --host 0.0.0.0 --is-embedding --log-level warning\n",
|
||||
"\"\"\")\n",
|
||||
"\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")"
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=embedding_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -171,7 +171,7 @@
|
||||
"metadata": {},
|
||||
"source": [
|
||||
"## Multi-Modal Embedding Model\n",
|
||||
"Please refer to [Multi-Modal Embedding Model](../supported_models/embedding_models)"
|
||||
"Please refer to [Multi-Modal Embedding Model](../supported_models/retrieval_ranking/embedding_models.md)"
|
||||
]
|
||||
}
|
||||
],
|
||||
|
||||
@@ -10,7 +10,7 @@
|
||||
"A complete reference for the API is available in the [OpenAI API Reference](https://platform.openai.com/docs/guides/vision).\n",
|
||||
"This tutorial covers the vision APIs for vision language models.\n",
|
||||
"\n",
|
||||
"SGLang supports various vision language models such as Llama 3.2, LLaVA-OneVision, Qwen2.5-VL, Gemma3 and [more](../supported_models/multimodal_language_models).\n",
|
||||
"SGLang supports various vision language models such as Llama 3.2, LLaVA-OneVision, Qwen2.5-VL, Gemma3 and [more](../supported_models/text_generation/multimodal_language_models.md).\n",
|
||||
"\n",
|
||||
"As an alternative to the OpenAI API, you can also use the [SGLang offline engine](https://github.com/sgl-project/sglang/blob/main/examples/runtime/engine/offline_batch_inference_vlm.py)."
|
||||
]
|
||||
@@ -33,11 +33,16 @@
|
||||
"from sglang.test.doc_patch import launch_server_cmd\n",
|
||||
"from sglang.utils import wait_for_server, print_highlight, terminate_process\n",
|
||||
"\n",
|
||||
"example_image_url = \"https://raw.githubusercontent.com/sgl-project/sglang/main/examples/assets/example_image.png\"\n",
|
||||
"logo_image_url = (\n",
|
||||
" \"https://raw.githubusercontent.com/sgl-project/sglang/main/assets/logo.png\"\n",
|
||||
")\n",
|
||||
"\n",
|
||||
"vision_process, port = launch_server_cmd(\"\"\"\n",
|
||||
"python3 -m sglang.launch_server --model-path Qwen/Qwen2.5-VL-7B-Instruct --log-level warning\n",
|
||||
"\"\"\")\n",
|
||||
"\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")"
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=vision_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -73,7 +78,7 @@
|
||||
" {{\n",
|
||||
" \"type\": \"image_url\",\n",
|
||||
" \"image_url\": {{\n",
|
||||
" \"url\": \"https://github.com/sgl-project/sglang/blob/main/examples/assets/example_image.png?raw=true\"\n",
|
||||
" \"url\": \"{example_image_url}\"\n",
|
||||
" }}\n",
|
||||
" }}\n",
|
||||
" ]\n",
|
||||
@@ -117,9 +122,7 @@
|
||||
" {\"type\": \"text\", \"text\": \"What’s in this image?\"},\n",
|
||||
" {\n",
|
||||
" \"type\": \"image_url\",\n",
|
||||
" \"image_url\": {\n",
|
||||
" \"url\": \"https://github.com/sgl-project/sglang/blob/main/examples/assets/example_image.png?raw=true\"\n",
|
||||
" },\n",
|
||||
" \"image_url\": {\"url\": example_image_url},\n",
|
||||
" },\n",
|
||||
" ],\n",
|
||||
" }\n",
|
||||
@@ -160,9 +163,7 @@
|
||||
" },\n",
|
||||
" {\n",
|
||||
" \"type\": \"image_url\",\n",
|
||||
" \"image_url\": {\n",
|
||||
" \"url\": \"https://github.com/sgl-project/sglang/blob/main/examples/assets/example_image.png?raw=true\"\n",
|
||||
" },\n",
|
||||
" \"image_url\": {\"url\": example_image_url},\n",
|
||||
" },\n",
|
||||
" ],\n",
|
||||
" }\n",
|
||||
@@ -201,13 +202,13 @@
|
||||
" {\n",
|
||||
" \"type\": \"image_url\",\n",
|
||||
" \"image_url\": {\n",
|
||||
" \"url\": \"https://github.com/sgl-project/sglang/blob/main/examples/assets/example_image.png?raw=true\",\n",
|
||||
" \"url\": example_image_url,\n",
|
||||
" },\n",
|
||||
" },\n",
|
||||
" {\n",
|
||||
" \"type\": \"image_url\",\n",
|
||||
" \"image_url\": {\n",
|
||||
" \"url\": \"https://raw.githubusercontent.com/sgl-project/sglang/main/assets/logo.png\",\n",
|
||||
" \"url\": logo_image_url,\n",
|
||||
" },\n",
|
||||
" },\n",
|
||||
" {\n",
|
||||
|
||||
@@ -7,36 +7,34 @@ SGLang provides OpenAI-compatible APIs to enable a smooth transition from OpenAI
|
||||
A complete reference for the API is available in the [OpenAI API Reference](https://platform.openai.com/docs/guides/vision).
|
||||
This tutorial covers the vision APIs for vision language models.
|
||||
|
||||
SGLang supports various vision language models such as Llama 3.2, LLaVA-OneVision, Qwen2.5-VL, Gemma3 and [more](../supported-models).
|
||||
SGLang supports various vision language models such as Llama 3.2, LLaVA-OneVision, Qwen2.5-VL, Gemma3 and [more](../supported-models/multimodal_language_models).
|
||||
|
||||
As an alternative to the OpenAI API, you can also use the [SGLang offline engine](https://github.com/sgl-project/sglang/blob/main/examples/runtime/engine/offline_batch_inference_vlm.py).
|
||||
|
||||
|
||||
## Launch A Server
|
||||
|
||||
Launch the server in your terminal and wait for it to initialize.
|
||||
|
||||
|
||||
|
||||
```python Example
|
||||
from sglang.test.doc_patch import launch_server_cmd
|
||||
from sglang.utils import wait_for_server, print_highlight, terminate_process
|
||||
|
||||
vision_process, port = launch_server_cmd(
|
||||
"""
|
||||
python3 -m sglang.launch_server --model-path Qwen/Qwen2.5-VL-7B-Instruct --log-level warning
|
||||
"""
|
||||
example_image_url = "https://raw.githubusercontent.com/sgl-project/sglang/main/examples/assets/example_image.png"
|
||||
logo_image_url = (
|
||||
"https://raw.githubusercontent.com/sgl-project/sglang/main/assets/logo.png"
|
||||
)
|
||||
|
||||
wait_for_server(f"http://localhost:{port}")
|
||||
vision_process, port = launch_server_cmd("""
|
||||
python3 -m sglang.launch_server --model-path Qwen/Qwen2.5-VL-7B-Instruct --log-level warning
|
||||
""")
|
||||
|
||||
wait_for_server(f"http://localhost:{port}", process=vision_process)
|
||||
```
|
||||
|
||||
## Using cURL
|
||||
|
||||
Once the server is up, you can send test requests using curl or requests.
|
||||
|
||||
|
||||
|
||||
```python Example
|
||||
import subprocess
|
||||
|
||||
@@ -56,7 +54,7 @@ curl -s http://localhost:{port}/v1/chat/completions \\
|
||||
{{
|
||||
"type": "image_url",
|
||||
"image_url": {{
|
||||
"url": "https://github.com/sgl-project/sglang/blob/main/examples/assets/example_image.png?raw=true"
|
||||
"url": "{example_image_url}"
|
||||
}}
|
||||
}}
|
||||
]
|
||||
@@ -76,8 +74,6 @@ print_highlight(response)
|
||||
|
||||
## Using Python Requests
|
||||
|
||||
|
||||
|
||||
```python Example
|
||||
import requests
|
||||
|
||||
@@ -92,9 +88,7 @@ data = {
|
||||
{"type": "text", "text": "What’s in this image?"},
|
||||
{
|
||||
"type": "image_url",
|
||||
"image_url": {
|
||||
"url": "https://github.com/sgl-project/sglang/blob/main/examples/assets/example_image.png?raw=true"
|
||||
},
|
||||
"image_url": {"url": example_image_url},
|
||||
},
|
||||
],
|
||||
}
|
||||
@@ -108,8 +102,6 @@ print_highlight(response.text)
|
||||
|
||||
## Using OpenAI Python Client
|
||||
|
||||
|
||||
|
||||
```python Example
|
||||
from openai import OpenAI
|
||||
|
||||
@@ -127,9 +119,7 @@ response = client.chat.completions.create(
|
||||
},
|
||||
{
|
||||
"type": "image_url",
|
||||
"image_url": {
|
||||
"url": "https://github.com/sgl-project/sglang/blob/main/examples/assets/example_image.png?raw=true"
|
||||
},
|
||||
"image_url": {"url": example_image_url},
|
||||
},
|
||||
],
|
||||
}
|
||||
@@ -144,8 +134,6 @@ print_highlight(response.choices[0].message.content)
|
||||
|
||||
The server also supports multiple images and interleaved text and images if the model supports it.
|
||||
|
||||
|
||||
|
||||
```python Example
|
||||
from openai import OpenAI
|
||||
|
||||
@@ -160,13 +148,13 @@ response = client.chat.completions.create(
|
||||
{
|
||||
"type": "image_url",
|
||||
"image_url": {
|
||||
"url": "https://github.com/sgl-project/sglang/blob/main/examples/assets/example_image.png?raw=true",
|
||||
"url": example_image_url,
|
||||
},
|
||||
},
|
||||
{
|
||||
"type": "image_url",
|
||||
"image_url": {
|
||||
"url": "https://raw.githubusercontent.com/sgl-project/sglang/main/assets/logo.png",
|
||||
"url": logo_image_url,
|
||||
},
|
||||
},
|
||||
{
|
||||
@@ -183,7 +171,6 @@ response = client.chat.completions.create(
|
||||
print_highlight(response.choices[0].message.content)
|
||||
```
|
||||
|
||||
|
||||
```python Example
|
||||
terminate_process(vision_process)
|
||||
```
|
||||
|
||||
@@ -2,13 +2,16 @@
|
||||
title: "Popular Model Usage (DeepSeek, GPT-OSS, GLM, Llama, MiniMax, Qwen, and more)"
|
||||
description: "Documentation for Popular Model Usage (DeepSeek, GPT-OSS, GLM, Llama, MiniMax, Qwen, and more)"
|
||||
---
|
||||
For more usage examples and recipes, visit the [SGLang Cookbook](https://cookbook.sglang.io/).
|
||||
|
||||
- [Deepseek V3](./deepseek_v3)
|
||||
- [Deepseek V32](./deepseek_v32)
|
||||
- [Glm45](./glm45)
|
||||
- [Glmv](./glmv)
|
||||
- [Gpt Oss](./gpt_oss)
|
||||
- [Kimi K2 5](./kimi_k2_5)
|
||||
- [Minimax M2](./minimax_m2)
|
||||
- [Qwen3](./qwen3)
|
||||
- [Qwen3 5](./qwen3_5)
|
||||
- [Qwen3 Vl](./qwen3_vl)
|
||||
- [Deepseek Ocr](./deepseek_ocr)
|
||||
- [Llama4](./llama4)
|
||||
|
||||
@@ -0,0 +1,80 @@
|
||||
---
|
||||
title: "Qwen 3.5 Usage"
|
||||
metatags:
|
||||
description: "Qwen 3.5 is Alibaba's latest generation LLM featuring a hybrid attention architecture, advanced MoE with shared experts, and native multimodal capabilities."
|
||||
---
|
||||
|
||||
Qwen 3.5 is Alibaba's latest generation LLM featuring a hybrid attention architecture, advanced MoE with shared experts, and native multimodal capabilities.
|
||||
|
||||
Key architecture features:
|
||||
- **Hybrid Attention**: Gated Delta Networks (linear, O(n) complexity) combined with full attention every 4th layer for high associative recall
|
||||
- **MoE with Shared Experts**: Top-8 active out of 64 routed experts plus a dedicated shared expert for universal features
|
||||
- **Multimodal**: DeepStack Vision Transformer with Conv3d for native image and video understanding
|
||||
|
||||
## Launch Qwen 3.5 with SGLang
|
||||
|
||||
### Dense Model
|
||||
|
||||
To serve `Qwen/Qwen3.5-397B-A17B` on 8 GPUs:
|
||||
|
||||
```bash
|
||||
python3 -m sglang.launch_server \
|
||||
--model-path Qwen/Qwen3.5-397B-A17B \
|
||||
--tp 8 \
|
||||
--trust-remote-code
|
||||
```
|
||||
|
||||
### AMD GPU (MI300X / MI325X / MI35X)
|
||||
|
||||
On AMD Instinct GPUs, use the `triton` attention backend. Both the full attention layers and the Gated Delta Net (linear attention) layers use Triton-based kernels on ROCm:
|
||||
|
||||
```bash
|
||||
SGLANG_USE_AITER=1 python3 -m sglang.launch_server \
|
||||
--model-path Qwen/Qwen3.5-397B-A17B \
|
||||
--tp 8 \
|
||||
--attention-backend triton \
|
||||
--trust-remote-code
|
||||
```
|
||||
|
||||
<Tip>
|
||||
Set `SGLANG_USE_AITER=1` to enable AMD's optimized aiter kernels for MoE and GEMM operations.
|
||||
</Tip>
|
||||
|
||||
### Configuration Tips
|
||||
|
||||
- `--attention-backend`: Use `triton` on AMD GPUs for Qwen 3.5. The hybrid attention architecture (Gated Delta Networks + full attention) works best with the Triton backend on ROCm. The linear attention (GDN) layers always use Triton kernels internally via the `GDNAttnBackend`.
|
||||
- `--watchdog-timeout`: Increase to `1200` or higher for this large model, as weight loading takes significant time.
|
||||
- `--model-loader-extra-config '{"enable_multithread_load": true}'`: Enables parallel weight loading for faster startup.
|
||||
|
||||
### Reasoning and Tool Calling
|
||||
|
||||
Qwen 3.5 supports reasoning and tool calling via the Qwen3 parsers:
|
||||
|
||||
```bash
|
||||
python3 -m sglang.launch_server \
|
||||
--model-path Qwen/Qwen3.5-397B-A17B \
|
||||
--tp 8 \
|
||||
--trust-remote-code \
|
||||
--reasoning-parser qwen3 \
|
||||
--tool-call-parser qwen3_coder
|
||||
```
|
||||
|
||||
## Accuracy Evaluation
|
||||
|
||||
You can evaluate the model accuracy using `lm-eval`:
|
||||
|
||||
```bash
|
||||
pip install lm-eval[api]
|
||||
|
||||
lm_eval --model local-completions \
|
||||
--model_args '{"base_url": "http://localhost:8000/v1/completions", "model": "Qwen/Qwen3.5-397B-A17B", "num_concurrent": 256, "max_retries": 10, "max_gen_toks": 2048}' \
|
||||
--tasks gsm8k \
|
||||
--batch_size auto \
|
||||
--num_fewshot 5 \
|
||||
--trust_remote_code
|
||||
```
|
||||
|
||||
## Additional Resources
|
||||
|
||||
- [AMD Day 0 Support for Qwen 3.5 on AMD Instinct GPUs](https://www.amd.com/en/developer/resources/technical-articles/2026/day-0-support-for-qwen-3-5-on-amd-instinct-gpus.html)
|
||||
- [HuggingFace Model Card](https://huggingface.co/Qwen/Qwen3.5-397B-A17B)
|
||||
@@ -7,9 +7,9 @@
|
||||
"# Sending Requests\n",
|
||||
"This notebook provides a quick-start guide to use SGLang in chat completions after installation. Once your server is running, API documentation is available at `http://localhost:30000/docs` (Swagger UI), `http://localhost:30000/redoc` (ReDoc), or `http://localhost:30000/openapi.json` (OpenAPI spec, useful for AI agents). Replace `30000` with your port if using a different one.\n",
|
||||
"\n",
|
||||
"- For Vision Language Models, see [OpenAI APIs - Vision](openai_api_vision).\n",
|
||||
"- For Embedding Models, see [OpenAI APIs - Embedding](openai_api_embeddings) and [Encode (embedding model)](native_api#encode-embedding-model).\n",
|
||||
"- For Reward Models, see [Classify (reward model)](native_api#classify-reward-model)."
|
||||
"- For Vision Language Models, see [OpenAI APIs - Vision](openai_api_vision.ipynb).\n",
|
||||
"- For Embedding Models, see [OpenAI APIs - Embedding](openai_api_embeddings.ipynb) and [Encode (embedding model)](native_api.html#Encode-(embedding-model)).\n",
|
||||
"- For Reward Models, see [Classify (reward model)](native_api.html#Classify-(reward-model))."
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -36,7 +36,7 @@
|
||||
" --host 0.0.0.0 --log-level warning\n",
|
||||
"\"\"\")\n",
|
||||
"\n",
|
||||
"wait_for_server(f\"http://localhost:{port}\")"
|
||||
"wait_for_server(f\"http://localhost:{port}\", process=server_process)"
|
||||
]
|
||||
},
|
||||
{
|
||||
@@ -158,7 +158,7 @@
|
||||
"source": [
|
||||
"## Using Native Generation APIs\n",
|
||||
"\n",
|
||||
"You can also use the native `/generate` endpoint with requests, which provides more flexibility. An API reference is available at [Sampling Parameters](sampling_params)."
|
||||
"You can also use the native `/generate` endpoint with requests, which provides more flexibility. An API reference is available at [Sampling Parameters](sampling_params.md)."
|
||||
]
|
||||
},
|
||||
{
|
||||
|
||||
@@ -142,7 +142,7 @@ python3 -m sglang.bench_serving \
|
||||
|
||||
- `--output-file FILE.jsonl`: append JSONL results to file; auto-named if unspecified
|
||||
- `--output-details`: include per-request arrays (generated texts, errors, ttfts, itls, input/output lens)
|
||||
- `--extra-request-body '{"top_p":0.9,"temperature":0.6}'`: merged into payload (sampling params, etc.)
|
||||
- `--extra-request-body '{"top_p":0.9,"temperature":0.6}'`: merged into payload (sampling params, etc.)
|
||||
- `--disable-ignore-eos`: pass through EOS behavior (varies by backend)
|
||||
- `--warmup-requests N`: run warmup requests with short output first (default 1)
|
||||
- `--flush-cache`: call `/flush_cache` (sglang) before main run
|
||||
@@ -335,7 +335,7 @@ python3 -m sglang.bench_serving \
|
||||
python3 -m sglang.bench_serving \
|
||||
--backend sglang \
|
||||
--host 127.0.0.1 --port 30000 \
|
||||
--model mode-name \
|
||||
--model model-name \
|
||||
--dataset-name mooncake \
|
||||
--mooncake-slowdown-factor 1.0 \
|
||||
--mooncake-num-rounds 1000 \
|
||||
@@ -344,6 +344,41 @@ python3 -m sglang.bench_serving \
|
||||
--random-output-len 256
|
||||
```
|
||||
|
||||
10) Fake decode stress testing (PD disaggregation, decode-only):
|
||||
|
||||
When benchmarking pure decode performance in a PD disaggregation setup, you can bypass the prefill node entirely by using `--fake-prefill`. This requires the decode server to be started with `--disaggregation-transfer-backend fake`:
|
||||
|
||||
```bash Command
|
||||
# Step 1: Start a decode-only server with fake transfer backend
|
||||
python -m sglang.launch_server \
|
||||
--model-path meta-llama/Llama-3.1-8B-Instruct \
|
||||
--disaggregation-mode decode \
|
||||
--disaggregation-transfer-backend fake \
|
||||
--port 30001
|
||||
|
||||
# Step 2: Run bench_serving with --fake-prefill
|
||||
python3 -m sglang.bench_serving \
|
||||
--backend sglang \
|
||||
--host 127.0.0.1 --port 30001 \
|
||||
--model meta-llama/Llama-3.1-8B-Instruct \
|
||||
--dataset-name random \
|
||||
--num-prompts 500 \
|
||||
--random-input-len 1024 --random-output-len 256 \
|
||||
--fake-prefill
|
||||
```
|
||||
|
||||
Similarly, `bench_one_batch_server` also supports `--fake-prefill`:
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.bench_one_batch_server \
|
||||
--base-url http://127.0.0.1:30001 \
|
||||
--model-path meta-llama/Llama-3.1-8B-Instruct \
|
||||
--batch-size 32 --input-len 1024 --output-len 256 \
|
||||
--fake-prefill
|
||||
```
|
||||
|
||||
The `--fake-prefill` flag automatically injects special sentinel values into each request, telling the decode server to skip real KV transfer and generate fake KV data locally.
|
||||
|
||||
### Troubleshooting
|
||||
|
||||
- All requests failed: verify `--backend`, server URL/port, `--model`, and authentication. Check warmup errors printed by the script.
|
||||
@@ -355,4 +390,4 @@ python3 -m sglang.bench_serving \
|
||||
### Notes
|
||||
|
||||
- The script raises the file descriptor soft limit (`RLIMIT_NOFILE`) to help with many concurrent connections.
|
||||
- For sglang, `/get_server_info` is queried post-run to report speculative decoding accept length when available.
|
||||
- For sglang, `/server_info` is queried post-run to report speculative decoding accept length when available.
|
||||
|
||||
@@ -5,28 +5,69 @@ metatags:
|
||||
---
|
||||
## Benchmark
|
||||
|
||||
- Benchmark the latency of running a single static batch without a server. The arguments are the same as for `launch_server.py`.
|
||||
Note that this is a simplified test script without a dynamic batching server, so it may run out of memory for a batch size that a real server can handle. A real server truncates the prefill into several batches, while this simplified script does not.
|
||||
- Without a server (do not need to launch a server)
|
||||
```bash Command
|
||||
python -m sglang.bench_one_batch --model-path meta-llama/Meta-Llama-3.1-8B-Instruct --batch 32 --input-len 256 --output-len 32
|
||||
```
|
||||
- With a server (please use `sglang.launch_server` to launch a server first and run the following command.)
|
||||
```bash Command
|
||||
python -m sglang.bench_one_batch_server --base-url http://127.0.0.1:30000 --model-path meta-llama/Meta-Llama-3.1-8B-Instruct --batch-size 32 --input-len 256 --output-len 32
|
||||
```
|
||||
SGLang provides four benchmark tools that operate at different levels of the stack. The table below summarizes their key differences:
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Tool</th>
|
||||
<th>HTTP Server</th>
|
||||
<th>Scheduler</th>
|
||||
<th>Use Case</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td><code>bench_serving</code></td>
|
||||
<td>Yes (async HTTP client to a running server)</td>
|
||||
<td>Yes (indirectly, via server)</td>
|
||||
<td>Realistic online serving benchmarks with latency metrics (TTFT, TPOT, ITL)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>bench_one_batch_server</code></td>
|
||||
<td>Yes (sends HTTP requests to a running server)</td>
|
||||
<td>Yes (indirectly, via server)</td>
|
||||
<td>End-to-end single-batch latency including HTTP and scheduler overhead</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>bench_offline_throughput</code></td>
|
||||
<td>No</td>
|
||||
<td>Yes (directly uses <code>Engine</code> in-process)</td>
|
||||
<td>Maximum throughput measurement without HTTP overhead</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>bench_one_batch</code></td>
|
||||
<td>No</td>
|
||||
<td>No (directly calls <code>ModelRunner</code>)</td>
|
||||
<td>Kernel-level latency profiling of a single static batch</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
- Benchmark offline processing. This script will start an offline engine and run the benchmark.
|
||||
Use `bench_serving` by default unless there are specific needs.
|
||||
|
||||
**`bench_serving`** is an async HTTP load-testing client that sends requests at controlled rates with configurable concurrency to a running server. It measures realistic online serving metrics including time-to-first-token (TTFT), time-per-output-token (TPOT), inter-token latency (ITL), and throughput. Use `num-prompts >= 5 * max-concurrency` to measure steady-state performance. Launch a server with `sglang.launch_server` first.
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.bench_serving --backend sglang --max-concurrency 16 --num-prompts 80 --random-input-len 256 --random-output-len 32 --dataset-name random
|
||||
```
|
||||
|
||||
**`bench_one_batch_server`** sends a single batch as one HTTP request to a running server. Due to only having a single batch, the server is never in a steady-state and metrics will be biased. Launch a server with `sglang.launch_server` first.
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.bench_one_batch_server --base-url http://127.0.0.1:30000 --model-path meta-llama/Meta-Llama-3.1-8B-Instruct --batch-size 32 --input-len 256 --output-len 32
|
||||
```
|
||||
|
||||
**`bench_offline_throughput`** directly instantiates the `Engine` object in-process (no HTTP server) and submits all requests at once via `engine.generate()`. The engine's scheduler handles batching and execution. This measures maximum achievable throughput without any network overhead.
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.bench_offline_throughput --model-path meta-llama/Meta-Llama-3.1-8B-Instruct --num-prompts 10
|
||||
```
|
||||
|
||||
- Benchmark online serving. Please use `sglang.launch_server` to launch a server first and run the following command.
|
||||
**`bench_one_batch`** is the lowest-level tool. It directly instantiates a `ModelRunner` and calls `extend()` / `decode()` on a fixed static batch, bypassing the scheduler entirely. The prefill and decode phases are run separately, making profiling easier but rendering the metrics unrealistic. Because there is no dynamic batching, it may run out of memory for batch sizes that a real server can handle (a real server chunks prefill into smaller batches). This is best suited for profiling individual kernel performance.
|
||||
|
||||
```bash Command
|
||||
python3 -m sglang.bench_serving --backend sglang --num-prompt 10
|
||||
python3 -m sglang.bench_one_batch --model-path meta-llama/Meta-Llama-3.1-8B-Instruct --batch-size 32 --input-len 256 --output-len 32
|
||||
```
|
||||
|
||||
## Profile with PyTorch Profiler
|
||||
@@ -46,7 +87,10 @@ python -m sglang.launch_server --model-path meta-llama/Llama-3.1-8B-Instruct
|
||||
python -m sglang.bench_serving --backend sglang --model meta-llama/Llama-3.1-8B-Instruct --num-prompts 10 --sharegpt-output-len 100 --profile
|
||||
```
|
||||
|
||||
Please make sure that the `SGLANG_TORCH_PROFILER_DIR` should be set at both server and client side, otherwise the trace file cannot be generated correctly . A secure way will be setting `SGLANG_TORCH_PROFILER_DIR` in the `.*rc` file of shell (e.g. `~/.bashrc` for bash shells).
|
||||
For `bench_serving --profile`, the output directory is selected on the client side from `--profile-output-dir` or `SGLANG_TORCH_PROFILER_DIR` (fallback: `/tmp`), then sent in the `/start_profile` request.
|
||||
If you call `/start_profile` directly and do not provide `output_dir`, the server uses its own `SGLANG_TORCH_PROFILER_DIR` (fallback: `/tmp`).
|
||||
|
||||
Setting `SGLANG_TORCH_PROFILER_DIR` on both server and client is still recommended to avoid confusion about where traces are written.
|
||||
|
||||
For more details, please refer to [Bench Serving Guide](./bench_serving).
|
||||
|
||||
@@ -147,7 +191,7 @@ curl -X POST http://127.0.0.1:30000/start_profile \
|
||||
**Parameters:**
|
||||
|
||||
- `output_dir` (optional): Directory where profile traces will be saved. If not specified, uses `SGLANG_TORCH_PROFILER_DIR` environment variable, or `/tmp` as the default
|
||||
- `num_steps` (optional): Number of steps to profile. If not specified, profiling continues until manually stopped with `/end_profile`
|
||||
- `num_steps` (optional): Number of steps to profile. If not specified, profiling continues until manually stopped with `/stop_profile`
|
||||
- `start_step` (optional): Step number at which to start profiling (inclusive). Useful for skipping warmup iterations
|
||||
- `activities` (optional): List of activities to profile, e.g., `["CPU", "GPU"]`. Default is `["CPU", "GPU"]`
|
||||
- `merge_profiles` (optional): Whether to merge distributed traces. Default is `false`
|
||||
@@ -171,17 +215,17 @@ curl -X POST http://127.0.0.1:30000/start_profile \
|
||||
**Continuous profiling (manual stop):**
|
||||
|
||||
```bash Command
|
||||
# Start profiling without num_steps - must manually stop with /end_profile
|
||||
# Start profiling without num_steps - must manually stop with /stop_profile
|
||||
curl -X POST http://127.0.0.1:30000/start_profile
|
||||
```
|
||||
|
||||
#### Using `/end_profile` endpoint
|
||||
#### Using `/stop_profile` endpoint
|
||||
|
||||
The `/end_profile` endpoint stops an ongoing profiling session and saves the trace file.
|
||||
The `/stop_profile` endpoint stops an ongoing profiling session and saves the trace file.
|
||||
|
||||
```bash Command
|
||||
# Stop profiling and save traces
|
||||
curl -X POST http://127.0.0.1:30000/end_profile
|
||||
curl -X POST http://127.0.0.1:30000/stop_profile
|
||||
```
|
||||
|
||||
This is only needed when you start profiling without specifying `num_steps`. If `num_steps` is specified, profiling will automatically stop after that many steps.
|
||||
@@ -204,7 +248,7 @@ curl -X POST http://127.0.0.1:30000/start_profile \
|
||||
python -m sglang.bench_serving --backend sglang --num-prompts 100
|
||||
|
||||
# Terminal 2: Stop profiling when done
|
||||
curl -X POST http://127.0.0.1:30000/end_profile
|
||||
curl -X POST http://127.0.0.1:30000/stop_profile
|
||||
```
|
||||
|
||||
### Profiler Trace Merger for Distributed Traces
|
||||
@@ -246,8 +290,8 @@ python -m sglang.profiler \
|
||||
#### Output Files
|
||||
|
||||
The profile merger generates:
|
||||
- Individual rank trace files: `{profile_id}-TP-{tp}-DP-{dp}-PP-{pp}-EP-{ep}.trace.json.gz`
|
||||
- Merged trace file: `merged-{profile_id}.trace.json.gz`
|
||||
- Individual rank trace files: `{profile_id}-TP-{tp}-DP-{dp}-PP-{pp}-EP-{ep}.trace.json.gz`
|
||||
- Merged trace file: `merged-{profile_id}.trace.json.gz`
|
||||
|
||||
### Possible PyTorch bugs
|
||||
If in any cases you encounter the following error (for example, using qwen 2.5 VL):
|
||||
@@ -398,10 +442,10 @@ This method allows you to control exactly when profiling starts/stops via HTTP A
|
||||
|
||||
```bash Command
|
||||
# Terminal 2: Only needed if num_steps was not specified
|
||||
curl -X POST http://127.0.0.1:30000/end_profile
|
||||
curl -X POST http://127.0.0.1:30000/stop_profile
|
||||
```
|
||||
|
||||
The `--capture-range=cudaProfilerApi` option tells Nsight Systems to only capture data between `cudaProfilerStart()` and `cudaProfilerStop()` calls (triggered by `/start_profile` and `/end_profile`), reducing overhead and file size. The `start_step` parameter skips the first 3 steps to avoid capturing warmup overhead.
|
||||
The `--capture-range=cudaProfilerApi` option tells Nsight Systems to only capture data between `cudaProfilerStart()` and `cudaProfilerStop()` calls (triggered by `/start_profile` and `/stop_profile`), reducing overhead and file size. The `start_step` parameter skips the first 3 steps to avoid capturing warmup overhead.
|
||||
|
||||
**Method 2: Simpler approach without `/start_profile` API**
|
||||
|
||||
|
||||
@@ -4,7 +4,7 @@ mode: wide
|
||||
metatags:
|
||||
description: "SGLang contribution guide: source install, pre-commit, unit tests, CI triggers, code style, sgl-kernel updates."
|
||||
---
|
||||
Welcome to **SGLang**! We appreciate your interest in contributing. This guide provides a concise overview of how to set up your environment, run tests, build documentation, and open a Pull Request (PR). Whether you're fixing a small bug or developing a major feature, we encourage following these steps for a smooth contribution process.
|
||||
Welcome to **SGLang**! We appreciate your interest in contributing. This guide provides a concise overview of how to set up your environment, run tests, build documentation, and open a Pull Request (PR). Whether you’re fixing a small bug or developing a major feature, we encourage following these steps for a smooth contribution process.
|
||||
|
||||
## Install SGLang from Source
|
||||
|
||||
@@ -18,7 +18,7 @@ git clone https://github.com/<your_user_name>/sglang.git
|
||||
|
||||
### Build from source
|
||||
|
||||
Refer to [Install SGLang from Source](../get-started/installation).
|
||||
Refer to [Install SGLang from Source](../get-started/install#method-2-from-source).
|
||||
|
||||
## Format code with pre-commit
|
||||
|
||||
@@ -32,17 +32,50 @@ pre-commit run --all-files
|
||||
|
||||
- **`pre-commit run --all-files`** manually runs all configured checks, applying fixes if possible. If it fails the first time, re-run it to ensure lint errors are fully resolved. Make sure your code passes all checks **before** creating a Pull Request.
|
||||
- **Do not commit** directly to the `main` branch. Always create a new branch (e.g., `feature/my-new-feature`), push your changes, and open a PR from that branch.
|
||||
- Link checking with lychee is **enforced in CI**. By default, it is not blocking local commits.
|
||||
- To run local link checks manually, use: `pre-commit run --hook-stage manual lychee --all-files`.
|
||||
|
||||
## Run and add unit tests
|
||||
|
||||
If you add a new feature or fix a bug, please add corresponding unit tests to ensure coverage and prevent regression.
|
||||
SGLang uses Python's built-in [unittest](https://docs.python.org/3/library/unittest.html) framework.
|
||||
For detailed instructions on running tests and integrating them into CI, refer to [test/README](https://github.com/sgl-project/sglang/tree/main/test/README).
|
||||
|
||||
### Unit tests (no server required)
|
||||
|
||||
Unit tests live under [`test/registered/unit/`](https://github.com/sgl-project/sglang/tree/main/test/registered/unit), organized to mirror the `python/sglang/srt/` source tree. These tests validate component logic **without** launching a server or loading real model weights.
|
||||
SGLang uses Python's built-in [unittest](https://docs.python.org/3/library/unittest.html) framework with [pytest](https://docs.pytest.org/) as the test runner.
|
||||
|
||||
**When to add a unit test:** If you modify a file under `python/sglang/srt/`, check whether a corresponding test exists in `test/registered/unit/` and add coverage for your changes. For example:
|
||||
|
||||
```
|
||||
srt/mem_cache/radix_cache.py → unit/mem_cache/test_radix_cache.py
|
||||
srt/sampling/sampling_params.py → unit/sampling/test_sampling_params.py
|
||||
```
|
||||
|
||||
**Run unit tests locally:**
|
||||
|
||||
```bash Command
|
||||
pytest test/registered/unit/ -v # all unit tests
|
||||
pytest test/registered/unit/mem_cache/ -v # one module
|
||||
```
|
||||
|
||||
**Run with coverage:**
|
||||
|
||||
```bash Command
|
||||
pytest test/registered/unit/ --cov --cov-config=.coveragerc -v
|
||||
```
|
||||
|
||||
For conventions on CI registration, test structure, and examples, see [`test/registered/unit/README.md`](https://github.com/sgl-project/sglang/tree/main/test/registered/unit/README.md).
|
||||
|
||||
### E2E tests (server required)
|
||||
|
||||
For tests that require launching a server, refer to [`test/registered/README.md`](https://github.com/sgl-project/sglang/tree/main/test/registered/README.md) for guidance on where to place your test.
|
||||
|
||||
For detailed instructions on running tests and integrating them into CI, refer to [test/README.md](https://github.com/sgl-project/sglang/tree/main/test/README.md).
|
||||
|
||||
## Write documentations
|
||||
|
||||
We recommend new contributors start from writing documentation, which helps you quickly understand SGLang codebase.
|
||||
For more details, please refer to [docs/README](https://github.com/sgl-project/sglang/blob/main/docs/README.md).
|
||||
For more details, please refer to [docs/README.md](https://github.com/sgl-project/sglang/tree/main/docs/README.md).
|
||||
|
||||
## Test the accuracy
|
||||
If your code changes the model output, please run the accuracy tests. A quick sanity check is the few-shot GSM8K.
|
||||
@@ -56,19 +89,19 @@ python3 -m sglang.test.few_shot_gsm8k --num-questions 200
|
||||
```
|
||||
|
||||
Please note that the above script is primarily a sanity check, not a rigorous accuracy or speed test.
|
||||
This test can have significant variance (1%-5%) in accuracy due to batching and the non-deterministic nature of the inference engine.
|
||||
This test can have significant variance (1%–5%) in accuracy due to batching and the non-deterministic nature of the inference engine.
|
||||
Also, do not rely on the "Latency/Output throughput" from this script, as it is not a proper speed test.
|
||||
|
||||
GSM8K is too easy for state-of-the-art models nowadays. Please try your own more challenging accuracy tests.
|
||||
You can find additional accuracy eval examples in:
|
||||
- [test_eval_accuracy_large.py](https://github.com/sgl-project/sglang/blob/main/test/srt/test_eval_accuracy_large.py)
|
||||
- [test_gpt_oss_1gpu.py](https://github.com/sgl-project/sglang/blob/main/test/srt/test_gpt_oss_1gpu.py)
|
||||
- [test_eval_accuracy_large.py](https://github.com/sgl-project/sglang/blob/main/test/registered/eval/test_eval_accuracy_large.py)
|
||||
- [test_gpt_oss_1gpu.py](https://github.com/sgl-project/sglang/blob/main/test/registered/core/test_gpt_oss_1gpu.py)
|
||||
|
||||
## Benchmark the speed
|
||||
Refer to [Benchmark and Profiling](../developer_guide/benchmark_and_profiling).
|
||||
Refer to [Benchmark and Profiling](./benchmark_and_profiling).
|
||||
|
||||
## Requesting a review for merge
|
||||
You can follow the pull request merge process described in [MAINTAINER](https://github.com/sgl-project/sglang/blob/main/.github/MAINTAINER).
|
||||
You can follow the pull request merge process described in [MAINTAINER.md](https://github.com/sgl-project/sglang/blob/main/.github/MAINTAINER.md).
|
||||
You will need to work with the Merge Oncall, Codeowner, and other reviewers to get their approvals.
|
||||
Then your PR can be merged.
|
||||
|
||||
@@ -77,6 +110,8 @@ Then your PR can be merged.
|
||||
We have a lot of open PRs but limited CI machines, so only top and trusted contributors have permission to trigger CI tests.
|
||||
Users with permission are listed in the [CI_PERMISSIONS.json](https://github.com/sgl-project/sglang/blob/main/.github/CI_PERMISSIONS.json)
|
||||
|
||||
**PR authors** can always use `/rerun-failed-ci` on their own PRs, even if they are not listed in `CI_PERMISSIONS.json`.
|
||||
|
||||
For CI to run on a pull request, it must have the "run-ci" label. Authorized users can add the label or rerun failed tests by commenting on the PR with one of these commands:
|
||||
|
||||
- `/tag-run-ci-label`: Adds the "run-ci" label. Every future commit will trigger CI.
|
||||
@@ -84,18 +119,17 @@ For CI to run on a pull request, it must have the "run-ci" label. Authorized use
|
||||
- `/tag-and-rerun-ci`: A single command that performs both `/tag-run-ci-label` and `/rerun-failed-ci`.
|
||||
- `/rerun-stage <stage-name>`: Reruns a specific test stage without waiting for its dependencies. This is useful when you want to quickly validate a fix for a specific test failure instead of waiting ~30 minutes for preceding stages to complete.
|
||||
|
||||
If you have permission, the [Slash Command Handler](https://github.com/sgl-project/sglang/actions/workflows/slash-command-handler.yml) will run your command and react with a +1 to your comment. It may take up to a few minutes for the reaction to appear. Here's a usage [example](https://github.com/sgl-project/sglang/pull/14253#issuecomment-3599509302).
|
||||
If you have permission, the [Slash Command Handler](https://github.com/sgl-project/sglang/actions/workflows/slash-command-handler.yml) will run your command and react with a 👍 to your comment. It may take up to a few minutes for the reaction to appear. Here’s a usage [example](https://github.com/sgl-project/sglang/pull/14253#issuecomment-3599509302).
|
||||
|
||||
To avoid spamming a PR with too many `/rerun-failed-ci` comments, you can also trigger the command by editing an existing comment and adding any suffix (e.g., `/rerun-failed-ci try again`).
|
||||
|
||||
Example of rerunning a single test stage: `/rerun-stage unit-test-backend-4-gpu`.
|
||||
|
||||
If you don't have permission, please ask maintainers to trigger CI for you.
|
||||
If you don’t have permission and you’re not the PR author, please ask maintainers to trigger CI for you.
|
||||
|
||||
### CI rate limits
|
||||
|
||||
Due to CI scheduling and limited resources, higher-priority PRs may preempt running jobs. In such cases, you may need to rerun the tests.
|
||||
|
||||
We apply CI rate limits to prevent abuse and ensure fair usage of our CI resources.
|
||||
|
||||
Each CI workflow has a default limit defined in its workflow configuration file. For example, in [pr-gate.yml](https://github.com/sgl-project/sglang/blob/main/.github/workflows/pr-gate.yml), the default cooldown period is 120 minutes, and each workflow can override it via the `cool-down-minutes` input parameter:
|
||||
@@ -107,8 +141,7 @@ cool-down-minutes:
|
||||
default: 120
|
||||
```
|
||||
|
||||
Users listed in [CI_PERMISSIONS.json](https://github.com/sgl-project/sglang/blob/main/.github/CI_PERMISSIONS.json) may have a per-user cooldown interval. In practice, we use the minimum of the workflow's default window and the user-specific interval.
|
||||
|
||||
Users listed in [CI_PERMISSIONS.json](https://github.com/sgl-project/sglang/blob/main/.github/CI_PERMISSIONS.json) may have a per-user cooldown interval. In practice, we use the minimum of the workflow’s default window and the user-specific interval.
|
||||
|
||||
## Code style guidance
|
||||
- Avoid code duplication. If the same code snippet (more than five lines) appears multiple times, extract it into a shared function.
|
||||
@@ -121,28 +154,34 @@ Users listed in [CI_PERMISSIONS.json](https://github.com/sgl-project/sglang/blob
|
||||
- If a single test file run longer than 500 seconds, split it into multiple smaller files (e.g., `test_eagle_infer_a.py`, `test_eagle_infer_b.py`).
|
||||
- If a single job in a github workflow runs longer than 30 mins, split it into smaller jobs/steps.
|
||||
- Reuse server launches in your unit tests to make tests run faster.
|
||||
- Never use `pickle.loads()`, `pickle.load()`, or `recv_pyobj()` to deserialize untrusted or network-received data. Python's [pickle module is not secure](https://docs.python.org/3/library/pickle.html) — it can execute arbitrary code during deserialization. Use safe serialization formats such as [msgpack](https://github.com/jcrist/msgspec) or JSON instead.
|
||||
- When supporting new hardware or features, follow these guidelines:
|
||||
- Do not drastically change existing code.
|
||||
- Always prefer new files to introduce specific components for your new hardware (e.g., `allocator_ascend.py`).
|
||||
- If you write multiple if/else blocks for new features, ensure the common path (e.g., NVIDIA hardware or the existing code path) is the first branch.
|
||||
|
||||
## How to update sgl-kernel
|
||||
Since sglang and sgl-kernel are separate Python packages, our current GitHub CI infrastructure does not support updating a kernel and using it immediately within the same pull request (PR).
|
||||
To add a new kernel or modify an existing one in the sgl-kernel package, you must use multiple PRs.
|
||||
Since sglang and the `sglang-kernel` (prior `sgl-kernel`) distribution are separate Python packages, our current GitHub CI infrastructure does not support updating a kernel and using it immediately within the same pull request (PR).
|
||||
To add a new kernel or modify an existing one in the `sgl-kernel/` source tree, you must use multiple PRs.
|
||||
|
||||
Follow these steps:
|
||||
|
||||
1. Submit a PR to update the sgl-kernel source code without using it in sglang python package (e.g., [#8884](https://github.com/sgl-project/sglang/pull/8884/files)).
|
||||
2. Bump the version of sgl-kernel (e.g., [#9220](https://github.com/sgl-project/sglang/pull/9220/files)).
|
||||
- Once merged, this will trigger an automatic release of the sgl-kernel wheel to PyPI.
|
||||
2. Bump the version of the kernel package (e.g., [#9220](https://github.com/sgl-project/sglang/pull/9220/files)).
|
||||
- Once merged, this will trigger an automatic release of the `sglang-kernel` wheel to PyPI.
|
||||
- If not urgent, you can wait for other people to release the wheel. A new version will typically be released within one week.
|
||||
3. Apply the changes:
|
||||
- Update the sgl-kernel version in `sglang/python/pyproject.toml` to use the modified kernels.
|
||||
- Update the `sglang-kernel` version in `sglang/python/pyproject.toml` to use the modified kernels.
|
||||
- Update the related caller code in the sglang to use the new kernel.
|
||||
|
||||
## Tips for newcomers
|
||||
|
||||
If you want to contribute but don't have a specific idea in mind, pick issues labeled ["good first issue" or "help wanted"](https://github.com/sgl-project/sglang/issues?q=is%3Aissue+label%3A%22good+first+issue%22%2C%22help+wanted%22). These tasks typically have lower complexity and provide an excellent introduction to the codebase. Also check out this [code walk-through](https://github.com/zhaochenyang20/Awesome-ML-SYS-Tutorial/tree/main/sglang/code-walk-through) for a deeper look into SGLang's workflow.
|
||||
If you want to contribute but don’t have a specific idea in mind, pick issues labeled [“good first issue” or “help wanted”](https://github.com/sgl-project/sglang/issues?q=is%3Aissue+label%3A%22good+first+issue%22%2C%22help+wanted%22). These tasks typically have lower complexity and provide an excellent introduction to the codebase.
|
||||
|
||||
Also check out the following materials as startup guide:
|
||||
- [Mini-SGLang](https://github.com/sgl-project/mini-sglang) for a quick overview on the structure of sglang.
|
||||
- [Code Walk-through](https://github.com/zhaochenyang20/Awesome-ML-SYS-Tutorial/tree/main/sglang/code-walk-through) for a deeper look into SGLang’s workflow.
|
||||
- [GTC-2026 Training Lab](https://drive.google.com/file/d/1mwOZEtipNLJzrflCTodj34KhuOZEoEw5/view?usp=drive_link) for hands-on practices of how to do optimization, benchmarking, or profiling on a launched SGLang instance.
|
||||
|
||||
If you have any questions or want to start a discussion, please feel free to ask in our [Slack channel](https://slack.sglang.io).
|
||||
|
||||
|
||||
@@ -7,7 +7,7 @@ metatags:
|
||||
## Setup VSCode on a Remote Host
|
||||
(Optional - you can skip this step if you plan to run sglang dev container locally)
|
||||
|
||||
1. In the remote host, download `code` from [Https://code.visualstudio.com/docs/?dv=linux64cli](https://code.visualstudio.com/download) and run `code tunnel` in a shell.
|
||||
1. In the remote host, download `code` from [VSCode](https://code.visualstudio.com/download) and run `code tunnel` in a shell.
|
||||
|
||||
Example
|
||||
```bash Command
|
||||
|
||||
+425
-266
@@ -1,266 +1,425 @@
|
||||
---
|
||||
title: "Development Guide for JIT Kernels"
|
||||
sidebarTitle: "JIT Kernels"
|
||||
metatags:
|
||||
description: "SGLang JIT kernel development: clangd setup, TensorMatcher, LaunchKernel, add_constant example walkthrough."
|
||||
---
|
||||
## Environment Setup
|
||||
|
||||
We strongly recommend using `clangd` as the language server for JIT kernel development.
|
||||
For Ubuntu/Debian, you can download clangd from [apt.llvm.org](https://apt.llvm.org/).
|
||||
If you are using VS Code, we recommend installing the `clangd` extension for better IDE integration.
|
||||
|
||||
All JIT-related files are located in `python/sglang/jit_kernel`.
|
||||
Unlike `sgl-kernel`, which compiles CUDA/C++ binaries ahead of time (AOT), just-in-time (JIT) kernels are compiled at runtime.
|
||||
Consequently, a static `compile_commands.json` cannot be generated.
|
||||
To enable code completion with `clangd`, run `python -m sglang.jit_kernel` to generate a `.clangd` configuration file in your current directory.
|
||||
After generating the file, restart the clangd language server. It should now recognize all JIT kernel files.
|
||||
|
||||
## Code Structure
|
||||
|
||||
### C++ Implementation
|
||||
|
||||
C++ source code is located in `python/sglang/jit_kernel/csrc`.
|
||||
Reusable functions should be placed in `python/sglang/jit_kernel/include`.
|
||||
|
||||
We use [tvm-ffi](https://github.com/apache/tvm-ffi) for efficient foreign language bindings.
|
||||
Refer to the [documentation](https://tvm.apache.org/ffi/) for advanced usage, such as exporting C++ objects.
|
||||
Typically, `tvm::ffi::TensorView` is sufficient for passing PyTorch Tensors from Python.
|
||||
|
||||
### Python Interface
|
||||
|
||||
Python interfaces are defined in `python/sglang/jit_kernel`.
|
||||
The `load_jit` utility function in `python/sglang/jit_kernel/utils.py` loads and returns the compiled module.
|
||||
To export a C++ function (e.g., `cpp_func`), pass `cuda_wrappers=[("func", "cpp_func")]` to `load_jit`.
|
||||
The function can then be called in Python as `module.func`.
|
||||
|
||||
### C++ Utilities
|
||||
|
||||
The following C++ utilities are available:
|
||||
|
||||
#### Integer Range
|
||||
|
||||
Similar to PyTorch, we provide an `irange` function to represent an integer range.
|
||||
|
||||
```C++ Example
|
||||
#include <sgl_kernel/utils.h>
|
||||
|
||||
void test() {
|
||||
for (auto i : host::irange(100)) { // [0, 100)
|
||||
// do something
|
||||
}
|
||||
for (auto i : host::irange(0, 100)) { // [0, 100)
|
||||
// do something
|
||||
}
|
||||
}
|
||||
|
||||
```
|
||||
|
||||
#### Runtime Checking
|
||||
|
||||
`RuntimeCheck` validates conditions at runtime. It accepts optional arguments for error reporting.
|
||||
If the check fails, these arguments are output to aid debugging.
|
||||
`RuntimeDeviceCheck` verifies the status of the last kernel launch.
|
||||
|
||||
```C++ Example
|
||||
#include <sgl_kernel/utils.h>
|
||||
#include <sgl_kernel/utils.cuh>
|
||||
|
||||
void test() {
|
||||
host::RuntimeCheck(1 + 1 == 2, 1 + 1, " != ", 2);
|
||||
host::RuntimeDeviceCheck();
|
||||
// check the provided `cudaError_t`
|
||||
host::RuntimeDeviceCheck(cudaGetLastError());
|
||||
}
|
||||
|
||||
```
|
||||
|
||||
#### Tensor Checking
|
||||
|
||||
`TensorMatcher` provides a readable way to validate and extract tensor shape information.
|
||||
|
||||
```cpp Example
|
||||
#include <sgl_kernel/tensor.h>
|
||||
|
||||
void test(const tvm::ffi::TensorView k_cache, const tvm::ffi::TensorView v_cache) {
|
||||
using namespace host;
|
||||
|
||||
auto D = SymbolicSize{"D"}; // cache dimension
|
||||
auto N = SymbolicSize{"N"}; // kvcache stride
|
||||
auto dtype = SymbolicDType{};
|
||||
auto device = SymbolicDevice{};
|
||||
|
||||
TensorMatcher({-1, D}) //
|
||||
.with_strides({N, 1})
|
||||
.with_dtype<int32_t, int64_t>(dtype)
|
||||
.with_device<kDLCUDA, kDLCPU>(device)
|
||||
.verify(k_cache)
|
||||
.verify(v_cache);
|
||||
}
|
||||
```
|
||||
|
||||
Configure the `TensorMatcher` with expected stride, dtype, and device properties before verification.
|
||||
- If `with_strides` is omitted, the tensor is expected to be contiguous.
|
||||
- Template arguments in `with_dtype` restrict the allowed data types.
|
||||
- Template arguments in `with_device` restrict the allowed devices.
|
||||
- Values passed to `with_xxx` methods enforce equality checks.
|
||||
- Passing `-1` for size or stride allows matching any value.
|
||||
|
||||
A `Symbolic` variable must resolve to the same value across all verifications.
|
||||
Use `.unwrap()` to retrieve the matched value after verification.
|
||||
|
||||
<Note>
|
||||
`TensorMatcher` is a temporary expression and should not be stored in a variable.
|
||||
</Note>
|
||||
|
||||
<Tip>
|
||||
Add `//` at the end of the `TensorMatcher` chain to enforce proper indentation.
|
||||
</Tip>
|
||||
|
||||
#### Kernel Launching
|
||||
|
||||
`LaunchKernel::resolve_device` retrieves the current `cudaStream` from PyTorch.
|
||||
Kernels can also be launched directly using `LaunchKernel`.
|
||||
|
||||
```cpp Example
|
||||
#include <sgl_kernel/utils.cuh>
|
||||
|
||||
#include <dlpack/dlpack.h>
|
||||
|
||||
__global__ void kernel() {}
|
||||
|
||||
void test() {
|
||||
const auto num_blocks = 1;
|
||||
const auto num_threads = 32;
|
||||
const auto dynamic_smem = 0;
|
||||
|
||||
DLDevice dev; // suppose this is initialized properly
|
||||
host::LaunchKernel(num_blocks, num_threads, dev)(kernel);
|
||||
|
||||
cudaStream_t stream = host::LaunchKernel::resolve_device(dev);
|
||||
host::LaunchKernel(num_blocks, num_threads, stream, dynamic_smem)(kernel);
|
||||
}
|
||||
|
||||
```
|
||||
|
||||
## Add new kernels
|
||||
|
||||
This section walks through a complete, end-to-end example of adding a new JIT kernel to the system.
|
||||
We use a simple add_constant kernel as a running example, which adds a constant integer value to every element of an input tensor.
|
||||
|
||||
Conceptually, the Python interface looks like this:
|
||||
|
||||
```python Example
|
||||
def add_constant(src: torch.Tensor, c: int):
|
||||
return src + c
|
||||
```
|
||||
|
||||
### STEP 1: Write the C++ kernel
|
||||
|
||||
Write your CUDA kernel in [jit_kernel/csrc/add_constant.cuh](https://github.com/sgl-project/sglang/blob/main/python/sglang/jit_kernel/csrc/add_constant.cuh). For demonstration purposes, we pass the constant value as a template parameter.
|
||||
|
||||
```cpp Example
|
||||
#include <sgl_kernel/tensor.h> // For TensorMatcher, SymbolicSize, SymbolicDevice
|
||||
#include <sgl_kernel/utils.cuh> // For LaunchKernel
|
||||
#include <sgl_kernel/utils.h> // For div_ceil, RuntimeCheck
|
||||
|
||||
#include <dlpack/dlpack.h>
|
||||
#include <tvm/ffi/container/tensor.h>
|
||||
|
||||
#include <cstddef>
|
||||
#include <cstdint>
|
||||
|
||||
namespace {
|
||||
|
||||
template <int32_t kConstant>
|
||||
__global__ void add_constant_kernel(int32_t* dst, const int32_t* src, size_t length) {
|
||||
size_t idx = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (idx < length) {
|
||||
dst[idx] = src[idx] + kConstant;
|
||||
}
|
||||
}
|
||||
|
||||
constexpr size_t kBlockSize = 256;
|
||||
|
||||
// You can also use struct with static method as an alternative
|
||||
template <int32_t kConstant>
|
||||
void add_constant(tvm::ffi::TensorView dst, tvm::ffi::TensorView src) {
|
||||
using namespace host;
|
||||
|
||||
// 1. Validate input tensors
|
||||
SymbolicSize N = {"num_elements"};
|
||||
SymbolicDevice device_;
|
||||
TensorMatcher({N}) // 1D tensor, must be contiguous
|
||||
.with_dtype<int32_t>() // must be int32
|
||||
.with_device<kDLCUDA>(device_) // must be on CUDA device
|
||||
.verify(dst) // check tensor dst
|
||||
.verify(src); // check tensor src
|
||||
|
||||
// 2. Extract required parameters, prepare for kernel launch
|
||||
const size_t num_elements = N.unwrap();
|
||||
const size_t grid_size = div_ceil(num_elements, kBlockSize);
|
||||
const DLDevice device = device_.unwrap();
|
||||
// some extra runtime checks using host::RuntimeCheck
|
||||
RuntimeCheck(num_elements > 0, "We only support non-empty tensors, got num_elements = ", num_elements);
|
||||
|
||||
// 3. Launch the kernel. Error code will be automatically checked.
|
||||
LaunchKernel(grid_size, kBlockSize, device /*, dynamic_smem*/)(
|
||||
// kernel function
|
||||
add_constant_kernel<kConstant>,
|
||||
// kernel arguments
|
||||
static_cast<int32_t*>(dst.data_ptr()),
|
||||
static_cast<int32_t*>(src.data_ptr()),
|
||||
num_elements);
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
```
|
||||
|
||||
### STEP 2: Create Python Interfaces
|
||||
|
||||
Next, expose the kernel through a Python wrapper.
|
||||
Create a new file at [jit_kernel/add_constant.py](https://github.com/sgl-project/sglang/blob/main/python/sglang/jit_kernel/add_constant.py) and expose the needed interfaces.
|
||||
|
||||
```python Example
|
||||
from __future__ import annotations
|
||||
|
||||
import functools
|
||||
from typing import TYPE_CHECKING
|
||||
|
||||
import torch
|
||||
|
||||
from sglang.jit_kernel.utils import load_jit, make_cpp_args
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from tvm_ffi.module import Module
|
||||
|
||||
|
||||
@functools.cache
|
||||
def _jit_add_constant_module(constant: int) -> Module:
|
||||
args = make_cpp_args(constant) # pass all the template argument
|
||||
return load_jit(
|
||||
"add_constant",
|
||||
*args,
|
||||
cuda_files=["add_constant.cuh"],
|
||||
cuda_wrappers=[("add_constant", f"add_constant<{args}>")],
|
||||
)
|
||||
|
||||
|
||||
def add_constant(src: torch.Tensor, constant: int) -> torch.Tensor:
|
||||
dst = torch.empty_like(src)
|
||||
module = _jit_add_constant_module(constant)
|
||||
module.add_constant(dst, src)
|
||||
return dst
|
||||
|
||||
```
|
||||
|
||||
### STEP 3: Use your kernel
|
||||
|
||||
Finally, import and use the kernel like a regular Python function:
|
||||
|
||||
```python Example
|
||||
from sglang.jit_kernel.add_constant import add_constant
|
||||
```
|
||||
|
||||
For a complete, runnable example, refer to [test_add_constant.py](https://github.com/sgl-project/sglang/blob/main/python/sglang/jit_kernel/test_add_constant.py).
|
||||
---
|
||||
title: "Development Guide for JIT Kernels"
|
||||
sidebarTitle: "JIT Kernels"
|
||||
metatags:
|
||||
description: "SGLang JIT kernel development: clangd setup, TensorMatcher, LaunchKernel, add_constant example walkthrough."
|
||||
---
|
||||
## Environment Setup
|
||||
|
||||
We strongly recommend using `clangd` as the language server for JIT kernel development.
|
||||
For Ubuntu/Debian, you can download clangd from [apt.llvm.org](https://apt.llvm.org/).
|
||||
If you are using VS Code, we recommend installing the `clangd` extension for better IDE integration.
|
||||
|
||||
All JIT-related files are located in `python/sglang/jit_kernel`.
|
||||
Unlike `sgl-kernel`, which compiles CUDA/C++ binaries ahead of time (AOT), just-in-time (JIT) kernels are compiled at runtime.
|
||||
Consequently, a static `compile_commands.json` cannot be generated.
|
||||
To enable code completion with `clangd`, run `python -m sglang.jit_kernel` to generate a `.clangd` configuration file in your current directory.
|
||||
After generating the file, restart the clangd language server. It should now recognize all JIT kernel files.
|
||||
|
||||
## Code Structure
|
||||
|
||||
### C++ Implementation
|
||||
|
||||
C++ source code is located in `python/sglang/jit_kernel/csrc`.
|
||||
Reusable functions should be placed in `python/sglang/jit_kernel/include`.
|
||||
|
||||
We use [tvm-ffi](https://github.com/apache/tvm-ffi) for efficient foreign language bindings.
|
||||
Refer to the [documentation](https://tvm.apache.org/ffi/) for advanced usage, such as exporting C++ objects.
|
||||
Typically, `tvm::ffi::TensorView` is sufficient for passing PyTorch Tensors from Python.
|
||||
|
||||
### Python Interface
|
||||
|
||||
Python interfaces are defined in `python/sglang/jit_kernel`.
|
||||
The `load_jit` utility function in `python/sglang/jit_kernel/utils.py` loads and returns the compiled module.
|
||||
To export a C++ function (e.g., `cpp_func`), pass `cuda_wrappers=[("func", "cpp_func")]` to `load_jit`.
|
||||
The function can then be called in Python as `module.func`.
|
||||
|
||||
For caching compiled modules, prefer `sglang.jit_kernel.utils.cache_once` over `functools.lru_cache`.
|
||||
`functools.lru_cache` is not compatible with `torch.compile`.
|
||||
|
||||
### C++ Utilities
|
||||
|
||||
The following C++ utilities are available:
|
||||
|
||||
#### Integer Range
|
||||
|
||||
Similar to PyTorch, we provide an `irange` function to represent an integer range.
|
||||
|
||||
```C++ Example
|
||||
#include <sgl_kernel/utils.h>
|
||||
|
||||
void test() {
|
||||
for (auto i : host::irange(100)) { // [0, 100)
|
||||
// do something
|
||||
}
|
||||
for (auto i : host::irange(0, 100)) { // [0, 100)
|
||||
// do something
|
||||
}
|
||||
}
|
||||
|
||||
```
|
||||
|
||||
#### Runtime Checking
|
||||
|
||||
`RuntimeCheck` validates conditions at runtime. It accepts optional arguments for error reporting.
|
||||
If the check fails, these arguments are output to aid debugging.
|
||||
`RuntimeDeviceCheck` verifies the status of the last kernel launch.
|
||||
|
||||
```C++ Example
|
||||
#include <sgl_kernel/utils.h>
|
||||
#include <sgl_kernel/utils.cuh>
|
||||
|
||||
void test() {
|
||||
host::RuntimeCheck(1 + 1 == 2, 1 + 1, " != ", 2);
|
||||
host::RuntimeDeviceCheck();
|
||||
// check the provided `cudaError_t`
|
||||
host::RuntimeDeviceCheck(cudaGetLastError());
|
||||
}
|
||||
|
||||
```
|
||||
|
||||
#### Tensor Checking
|
||||
|
||||
`TensorMatcher` provides a readable way to validate and extract tensor shape information.
|
||||
|
||||
```cpp Example
|
||||
#include <sgl_kernel/tensor.h>
|
||||
|
||||
void test(const tvm::ffi::TensorView k_cache, const tvm::ffi::TensorView v_cache) {
|
||||
using namespace host;
|
||||
|
||||
auto D = SymbolicSize{"D"}; // cache dimension
|
||||
auto N = SymbolicSize{"N"}; // kvcache stride
|
||||
auto dtype = SymbolicDType{};
|
||||
auto device = SymbolicDevice{};
|
||||
|
||||
TensorMatcher({-1, D}) //
|
||||
.with_strides({N, 1})
|
||||
.with_dtype<int32_t, int64_t>(dtype)
|
||||
.with_device<kDLCUDA, kDLCPU>(device)
|
||||
.verify(k_cache)
|
||||
.verify(v_cache);
|
||||
}
|
||||
```
|
||||
|
||||
Configure the `TensorMatcher` with expected stride, dtype, and device properties before verification.
|
||||
- If `with_strides` is omitted, the tensor is expected to be contiguous.
|
||||
- Template arguments in `with_dtype` restrict the allowed data types.
|
||||
- Template arguments in `with_device` restrict the allowed devices.
|
||||
- Values passed to `with_xxx` methods enforce equality checks.
|
||||
- Passing `-1` for size or stride allows matching any value.
|
||||
|
||||
A `Symbolic` variable must resolve to the same value across all verifications.
|
||||
Use `.unwrap()` to retrieve the matched value after verification.
|
||||
|
||||
> Note: `TensorMatcher` is a temporary expression and should not be stored in a variable.
|
||||
|
||||
> Tip: Add `//` at the end of the `TensorMatcher` chain to enforce proper indentation.
|
||||
|
||||
#### Kernel Launching
|
||||
|
||||
`LaunchKernel::resolve_device` retrieves the current `cudaStream` from PyTorch.
|
||||
Kernels can also be launched directly using `LaunchKernel`.
|
||||
|
||||
```cpp Example
|
||||
#include <sgl_kernel/utils.cuh>
|
||||
|
||||
#include <dlpack/dlpack.h>
|
||||
|
||||
__global__ void kernel() {}
|
||||
|
||||
void test() {
|
||||
const auto num_blocks = 1;
|
||||
const auto num_threads = 32;
|
||||
const auto dynamic_smem = 0;
|
||||
|
||||
DLDevice dev; // suppose this is initialized properly
|
||||
host::LaunchKernel(num_blocks, num_threads, dev)(kernel);
|
||||
|
||||
cudaStream_t stream = host::LaunchKernel::resolve_device(dev);
|
||||
host::LaunchKernel(num_blocks, num_threads, stream, dynamic_smem)(kernel);
|
||||
}
|
||||
|
||||
```
|
||||
|
||||
## Add new kernels
|
||||
|
||||
This section walks through a complete, end-to-end example of adding a new JIT kernel to the system.
|
||||
We use a simple add_constant kernel as a running example, which adds a constant integer value to every element of an input tensor.
|
||||
|
||||
Conceptually, the Python interface looks like this:
|
||||
|
||||
```python Example
|
||||
def add_constant(src: torch.Tensor, c: int):
|
||||
return src + c
|
||||
```
|
||||
|
||||
### STEP 1: Write the C++ kernel
|
||||
|
||||
Write your CUDA kernel in [jit_kernel/csrc/add_constant.cuh](https://github.com/sgl-project/sglang/blob/main/python/sglang/jit_kernel/csrc/add_constant.cuh). For demonstration purposes, we pass the constant value as a template parameter.
|
||||
|
||||
```cpp Example
|
||||
#include <sgl_kernel/tensor.h> // For TensorMatcher, SymbolicSize, SymbolicDevice
|
||||
#include <sgl_kernel/utils.cuh> // For LaunchKernel
|
||||
#include <sgl_kernel/utils.h> // For div_ceil, RuntimeCheck
|
||||
|
||||
#include <dlpack/dlpack.h>
|
||||
#include <tvm/ffi/container/tensor.h>
|
||||
|
||||
#include <cstddef>
|
||||
#include <cstdint>
|
||||
|
||||
namespace {
|
||||
|
||||
template <int32_t kConstant>
|
||||
__global__ void add_constant_kernel(int32_t* dst, const int32_t* src, size_t length) {
|
||||
size_t idx = blockIdx.x * blockDim.x + threadIdx.x;
|
||||
if (idx < length) {
|
||||
dst[idx] = src[idx] + kConstant;
|
||||
}
|
||||
}
|
||||
|
||||
constexpr size_t kBlockSize = 256;
|
||||
|
||||
// You can also use struct with static method as an alternative
|
||||
template <int32_t kConstant>
|
||||
void add_constant(tvm::ffi::TensorView dst, tvm::ffi::TensorView src) {
|
||||
using namespace host;
|
||||
|
||||
// 1. Validate input tensors
|
||||
SymbolicSize N = {"num_elements"};
|
||||
SymbolicDevice device_;
|
||||
TensorMatcher({N}) // 1D tensor, must be contiguous
|
||||
.with_dtype<int32_t>() // must be int32
|
||||
.with_device<kDLCUDA>(device_) // must be on CUDA device
|
||||
.verify(dst) // check tensor dst
|
||||
.verify(src); // check tensor src
|
||||
|
||||
// 2. Extract required parameters, prepare for kernel launch
|
||||
const size_t num_elements = N.unwrap();
|
||||
const size_t grid_size = div_ceil(num_elements, kBlockSize);
|
||||
const DLDevice device = device_.unwrap();
|
||||
// some extra runtime checks using host::RuntimeCheck
|
||||
RuntimeCheck(num_elements > 0, "We only support non-empty tensors, got num_elements = ", num_elements);
|
||||
|
||||
// 3. Launch the kernel. Error code will be automatically checked.
|
||||
LaunchKernel(grid_size, kBlockSize, device /*, dynamic_smem*/)(
|
||||
// kernel function
|
||||
add_constant_kernel<kConstant>,
|
||||
// kernel arguments
|
||||
static_cast<int32_t*>(dst.data_ptr()),
|
||||
static_cast<int32_t*>(src.data_ptr()),
|
||||
num_elements);
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
```
|
||||
|
||||
### STEP 2: Create Python Interfaces
|
||||
|
||||
Next, expose the kernel through a Python wrapper.
|
||||
Create a new file at [jit_kernel/add_constant.py](https://github.com/sgl-project/sglang/blob/main/python/sglang/jit_kernel/add_constant.py) and expose the needed interfaces.
|
||||
|
||||
```python Example
|
||||
from __future__ import annotations
|
||||
from typing import TYPE_CHECKING
|
||||
|
||||
import torch
|
||||
|
||||
from sglang.jit_kernel.utils import cache_once, load_jit, make_cpp_args
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from tvm_ffi.module import Module
|
||||
|
||||
|
||||
@cache_once
|
||||
def _jit_add_constant_module(constant: int) -> Module:
|
||||
args = make_cpp_args(constant) # pass all the template argument
|
||||
return load_jit(
|
||||
"add_constant",
|
||||
*args,
|
||||
cuda_files=["add_constant.cuh"],
|
||||
cuda_wrappers=[("add_constant", f"add_constant<{args}>")],
|
||||
)
|
||||
|
||||
|
||||
def add_constant(src: torch.Tensor, constant: int) -> torch.Tensor:
|
||||
if not src.is_cuda:
|
||||
raise RuntimeError("src must be a CUDA tensor")
|
||||
if src.dtype != torch.int32:
|
||||
raise RuntimeError(f"Unsupported dtype {src.dtype}. Supported: int32")
|
||||
dst = torch.empty_like(src)
|
||||
module = _jit_add_constant_module(constant)
|
||||
module.add_constant(dst, src)
|
||||
return dst
|
||||
|
||||
```
|
||||
|
||||
Keep the Python wrapper thin, but still validate the basic invariants such as device and dtype before dispatch. In the current JIT/FFI path, invalid tensors are not always rejected safely before launch.
|
||||
|
||||
### STEP 3: Use your kernel
|
||||
|
||||
Finally, import and use the kernel like a regular Python function:
|
||||
|
||||
```python Example
|
||||
from sglang.jit_kernel.add_constant import add_constant
|
||||
```
|
||||
|
||||
For a complete, runnable example, refer to [test_add_constant.py](https://github.com/sgl-project/sglang/blob/main/python/sglang/jit_kernel/tests/test_add_constant.py).
|
||||
|
||||
## C++ Include Library Reference
|
||||
|
||||
The JIT kernel framework provides a set of reusable C++ headers in
|
||||
`python/sglang/jit_kernel/include/sgl_kernel/`. Each header is designed
|
||||
to be lightweight and self-contained. Below is a summary of each header
|
||||
and its key APIs.
|
||||
|
||||
### Core Utilities
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Header</th>
|
||||
<th>Namespace</th>
|
||||
<th>Purpose</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td><code>utils.h</code></td>
|
||||
<td><code>host</code></td>
|
||||
<td>Host-side essentials: <code>RuntimeCheck</code>, <code>Panic</code>, <code>div_ceil</code>, <code>irange</code></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>utils.cuh</code></td>
|
||||
<td><code>device</code> / <code>host</code></td>
|
||||
<td>Type aliases (<code>fp16_t</code>, <code>bf16_t</code>, ...), <code>SGL_DEVICE</code> macro, PDL helpers, <code>LaunchKernel</code>, <code>RuntimeDeviceCheck</code></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>source_location.h</code></td>
|
||||
<td>(global)</td>
|
||||
<td>Portable <code>std::source_location</code> wrapper for error reporting</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>runtime.cuh</code></td>
|
||||
<td><code>host::runtime</code></td>
|
||||
<td>CUDA runtime queries: <code>get_blocks_per_sm</code>, <code>get_sm_count</code>, <code>get_cc_major</code>, <code>get_runtime_version</code>, <code>get_available_dynamic_smem_per_block</code></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
### Tensor Validation
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Header</th>
|
||||
<th>Namespace</th>
|
||||
<th>Purpose</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td><code>tensor.h</code></td>
|
||||
<td><code>host</code></td>
|
||||
<td><code>TensorMatcher</code>, <code>SymbolicSize</code>, <code>SymbolicDType</code>, <code>SymbolicDevice</code></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
### Math & Type System
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Header</th>
|
||||
<th>Namespace</th>
|
||||
<th>Purpose</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td><code>math.cuh</code></td>
|
||||
<td><code>device::math</code></td>
|
||||
<td><code>max</code>, <code>min</code>, <code>abs</code>, <code>sqrt</code>, <code>rsqrt</code>, <code>exp</code>, <code>sin</code>, <code>cos</code>, constants</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>type.cuh</code></td>
|
||||
<td>(global) / <code>device</code></td>
|
||||
<td><code>dtype_trait<T></code>, <code>packed_t<T></code>, <code>device::cast<To>(from)</code></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
### Memory Access
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Header</th>
|
||||
<th>Namespace</th>
|
||||
<th>Purpose</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td><code>vec.cuh</code></td>
|
||||
<td><code>device</code></td>
|
||||
<td><code>AlignedVector<T, N></code> - vectorized load/store (up to 128-bit; 256-bit requires Blackwell GPUs)</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>tile.cuh</code></td>
|
||||
<td><code>device::tile</code></td>
|
||||
<td><code>Memory<T></code> - cooperative tiled memory I/O (thread/warp/CTA)</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
### Parallel Primitives
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Header</th>
|
||||
<th>Namespace</th>
|
||||
<th>Purpose</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td><code>warp.cuh</code></td>
|
||||
<td><code>device::warp</code></td>
|
||||
<td><code>reduce_sum</code>, <code>reduce_max</code> via <code>__shfl_xor_sync</code></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>cta.cuh</code></td>
|
||||
<td><code>device::cta</code></td>
|
||||
<td><code>reduce_max</code> across warps via shared memory</td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>atomic.cuh</code></td>
|
||||
<td><code>device::atomic</code></td>
|
||||
<td><code>max</code> - atomic float max (CUDA + ROCm fallback)</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
### Reusable Kernel Templates
|
||||
|
||||
<table>
|
||||
<thead>
|
||||
<tr>
|
||||
<th>Header</th>
|
||||
<th>Namespace</th>
|
||||
<th>Purpose</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr>
|
||||
<td><code>impl/norm.cuh</code></td>
|
||||
<td><code>host::norm</code> / <code>device::norm</code></td>
|
||||
<td>RMSNorm building blocks (warp & CTA paths, <code>StorageType</code>)</td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
@@ -5,7 +5,7 @@ description: Contributing to SGLang — development setup, benchmarking, and eva
|
||||
|
||||
- [Contribution Guide](./contribution_guide)
|
||||
- [Development Guide (Docker)](./development_guide_using_docker)
|
||||
- [JIT Kernels](./JIT_kernels)
|
||||
- [JIT Kernels](./development_jit_kernel_guide)
|
||||
- [Benchmark and Profiling](./benchmark_and_profiling)
|
||||
- [Bench Serving](./bench_serving)
|
||||
- [Evaluating New Models](./evaluating_new_models)
|
||||
|
||||
@@ -10,14 +10,14 @@ metatags:
|
||||
**You can mount a folder for the shared huggingface model weights cache. **
|
||||
The command below uses `/tmp/huggingface` as an example.
|
||||
|
||||
```text Output
|
||||
```
|
||||
docker pull nvidia/cuda:12.9.1-devel-ubuntu22.04
|
||||
# Nvidia
|
||||
docker run --shm-size 128g -it -v /tmp/huggingface:/hf_home --gpus all nvidia/cuda:12.9.1-devel-ubuntu22.04 /bin/bash
|
||||
# AMD
|
||||
docker run --rm --device=/dev/kfd --device=/dev/dri --group-add video --shm-size 128g -it -v /tmp/huggingface:/hf_home lmsysorg/sglang:v0.5.0rc1-rocm630 /bin/bash
|
||||
docker run --rm --device=/dev/kfd --device=/dev/dri --group-add video --shm-size 128g -it -v /tmp/huggingface:/hf_home lmsysorg/sglang:v0.5.8-rocm700-mi30x /bin/bash
|
||||
# AMD just the last 2 GPUs
|
||||
docker run --rm --device=/dev/kfd --device=/dev/dri/renderD176 --device=/dev/dri/renderD184 --group-add video --shm-size 128g -it -v /tmp/huggingface:/hf_home lmsysorg/sglang:v0.5.0rc1-rocm630 /bin/bash
|
||||
docker run --rm --device=/dev/kfd --device=/dev/dri/renderD176 --device=/dev/dri/renderD184 --group-add video --shm-size 128g -it -v /tmp/huggingface:/hf_home lmsysorg/sglang:v0.5.8-rocm700-mi30x /bin/bash
|
||||
```
|
||||
|
||||
### Step 2: Configure the runner by `config.sh`
|
||||
@@ -30,11 +30,11 @@ pip install --upgrade pip
|
||||
export RUNNER_ALLOW_RUNASROOT=1
|
||||
```
|
||||
|
||||
Then follow https://github.com/sgl-project/sglang/settings/actions/runners/new?arch=x64&os=linux to run `config.sh`
|
||||
Then follow https://docs.github.com/en/actions/hosting-your-own-runners/managing-self-hosted-runners/adding-self-hosted-runners to run `config.sh`
|
||||
|
||||
**Notes**
|
||||
- Do not need to specify the runner group
|
||||
- Give it a name (e.g., `test-sgl-gpu-0`) and some labels (e.g., `1-gpu-runner`). The labels can be edited later in Github Settings.
|
||||
- Give it a name (e.g., `test-sgl-gpu-0`) and some labels (e.g., `1-gpu-h100`). The labels can be edited later in Github Settings.
|
||||
- Do not need to change the work folder.
|
||||
|
||||
### Step 3: Run the runner by `run.sh`
|
||||
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user