[NPU] [DOC] Upgrade recommendeded sglang version on Ascend NPU (#33976)
This commit is contained in:
@@ -26,7 +26,7 @@ First, launch the SGLang environment using the provided container image:
|
||||
<Tab title="Atlas 800I A3">
|
||||
|
||||
```shell Command
|
||||
export IMAGE=quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
export IMAGE=quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
docker run -it --rm --privileged --network=host --ipc=host --shm-size=16g \
|
||||
--device=/dev/davinci0 --device=/dev/davinci1 --device=/dev/davinci2 --device=/dev/davinci3 \
|
||||
@@ -49,7 +49,7 @@ docker run -it --rm --privileged --network=host --ipc=host --shm-size=16g \
|
||||
<Tab title="Atlas 800I A2">
|
||||
|
||||
```shell Command
|
||||
export IMAGE=quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
export IMAGE=quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
docker run -it --rm --privileged --network=host --ipc=host --shm-size=16g \
|
||||
--device=/dev/davinci0 --device=/dev/davinci1 --device=/dev/davinci2 --device=/dev/davinci3 \
|
||||
|
||||
@@ -297,7 +297,7 @@ Replace model weights, e.g., switch Qwen/Qwen3.5-27B-FP8 to Eco-Tech/Qwen3.5-27B
|
||||
|
||||
Docker images for Ascend NPU are available in two types:
|
||||
|
||||
- **Stable release** — validated version with a specific tag, e.g., `quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3`.
|
||||
- **Stable release** — validated version with a specific tag, e.g., `quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16`.
|
||||
Recommended for production deployments.
|
||||
- **Daily build** — includes the latest development changes, e.g., `quay.io/ascend/sglang:main-cann8.5.0-a3`. Use this
|
||||
if you need the latest features or bug fixes that have not yet been included in a stable release.
|
||||
|
||||
@@ -206,7 +206,7 @@ Ensure sufficient disk space before pulling images. Each Docker image requires a
|
||||
1. Download SGLang image
|
||||
|
||||
<Tip>
|
||||
We publish both **stable releases** and **daily builds**. Choose a stable release tag (e.g., `v0.5.13.post1-cann9.0.0-a3`) if you prefer a validated version, or a daily build tag (e.g., `main-cann9.0.0-a3`) if you need the latest development changes.
|
||||
We publish both **stable releases** and **daily builds**. Choose a stable release tag (e.g., `cann9.0.0-a3-v0.5.16`) if you prefer a validated version, or a daily build tag (e.g., `main-cann9.0.0-a3`) if you need the latest development changes.
|
||||
</Tip>
|
||||
|
||||
<Tabs>
|
||||
@@ -214,7 +214,7 @@ We publish both **stable releases** and **daily builds**. Choose a stable releas
|
||||
|
||||
```bash Command
|
||||
# Stable release
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
# Daily build
|
||||
docker pull quay.io/ascend/sglang:main-cann9.0.0-a3
|
||||
@@ -225,7 +225,7 @@ docker pull quay.io/ascend/sglang:main-cann9.0.0-a3
|
||||
|
||||
```bash Command
|
||||
# Stable release
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
# Daily build
|
||||
docker pull quay.io/ascend/sglang:main-cann9.0.0-910b
|
||||
|
||||
@@ -40,7 +40,7 @@ at least **30GB** of free space. If you need to download model weights, check th
|
||||
</Warning>
|
||||
|
||||
<Tip>
|
||||
We publish both **stable releases** and **daily builds**. Choose a stable release tag (e.g., `v0.5.14-cann9.0.0-a3`)
|
||||
We publish both **stable releases** and **daily builds**. Choose a stable release tag (e.g., `cann9.0.0-a3-v0.5.16`)
|
||||
if you prefer a validated version, or a daily build tag (e.g., `main-cann9.0.0-a3`) if you need the latest development changes.
|
||||
</Tip>
|
||||
|
||||
@@ -52,7 +52,7 @@ adding `--volume /path/to/model:/path/to/model` to the `docker run` command belo
|
||||
|
||||
```shell Command
|
||||
# Choose one (uncomment the line you want):
|
||||
export IMAGE=quay.io/ascend/sglang:v0.5.14-cann9.0.0-a3 # Stable release
|
||||
export IMAGE=quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16 # Stable release
|
||||
# export IMAGE=quay.io/ascend/sglang:main-cann9.0.0-a3 # Daily build
|
||||
|
||||
docker run -it --rm --privileged --network=host --ipc=host --shm-size=16g \
|
||||
@@ -77,7 +77,7 @@ docker run -it --rm --privileged --network=host --ipc=host --shm-size=16g \
|
||||
|
||||
```shell Command
|
||||
# Choose one (uncomment the line you want):
|
||||
export IMAGE=quay.io/ascend/sglang:v0.5.14-cann9.0.0-910b # Stable release
|
||||
export IMAGE=quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16 # Stable release
|
||||
# export IMAGE=quay.io/ascend/sglang:main-cann9.0.0-910b # Daily build
|
||||
|
||||
docker run -it --rm --privileged --network=host --ipc=host --shm-size=16g \
|
||||
|
||||
+1
-1
@@ -9,7 +9,7 @@ This page focuses on optimal configuration and benchmark results for Qwen3.6-35B
|
||||
|
||||
On A3 each card has 2 dies, so `--tp-size` is twice the card count; see [Ascend NPU Reference](/docs/hardware-platforms/ascend-npus/reference/glossary#hardware) for details.
|
||||
|
||||
Use image **SGLang `>= v0.5.14`** for these NEXTN configurations. Without `--dataset-path`, `bench_serving --dataset-name random` downloads ShareGPT from Hugging Face; in offline environments, pass a local dataset path (for example a ShareGPT JSON file).
|
||||
Use image **SGLang `>= v0.5.16`** for these NEXTN configurations. Without `--dataset-path`, `bench_serving --dataset-name random` downloads ShareGPT from Hugging Face; in offline environments, pass a local dataset path (for example a ShareGPT JSON file).
|
||||
</Note>
|
||||
|
||||
### Low Latency
|
||||
|
||||
@@ -14,9 +14,9 @@ reinforcement learning-based training.
|
||||
This document demonstrates the deployment of DeepSeek-R1 on Ascend NPUs using SGLang, including single-node PD mixed
|
||||
mode, multi-node PD disaggregation mode, feature configuration, and performance optimization.
|
||||
|
||||
This document is validated and written based on **SGLang v0.5.13**. The current model (DeepSeek-R1) is fully supported in
|
||||
This document is validated and written based on **SGLang v0.5.16**. The current model (DeepSeek-R1) is fully supported in
|
||||
this version. To use the latest features (e.g., PD disaggregation, speculative decoding), it is recommended to use
|
||||
v0.5.13 or a later version.
|
||||
v0.5.16 or a later version.
|
||||
|
||||
## Supported features
|
||||
|
||||
@@ -97,7 +97,7 @@ For details, see [Docker image versions](/docs/hardware-platforms/ascend-npus/fa
|
||||
<Tab title="Atlas 800I A3">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -125,14 +125,14 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
<Tab title="Atlas 800I A2">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -152,7 +152,7 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
+6
-6
@@ -15,9 +15,9 @@ improvements in long-context scenarios.
|
||||
This document demonstrates the deployment of DeepSeek-V3.2 on Ascend NPUs using SGLang, including multi-node PD
|
||||
disaggregation mode, feature configuration, and performance optimization.
|
||||
|
||||
This document is validated and written based on **SGLang v0.5.13**. The current model (DeepSeek-V3.2) is fully supported in
|
||||
This document is validated and written based on **SGLang v0.5.16**. The current model (DeepSeek-V3.2) is fully supported in
|
||||
this version. To use the latest features (e.g., PD disaggregation, speculative decoding, DSA context parallel), it is
|
||||
recommended to use v0.5.13 or a later version.
|
||||
recommended to use v0.5.16 or a later version.
|
||||
|
||||
## Supported features
|
||||
|
||||
@@ -92,7 +92,7 @@ For details, see [Docker image versions](/docs/hardware-platforms/ascend-npus/fa
|
||||
<Tab title="Atlas 800I A3">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -120,14 +120,14 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
<Tab title="Atlas 800I A2">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -147,7 +147,7 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
@@ -16,9 +16,9 @@ mode (direct response), with a native context window of approximately 200k token
|
||||
This document demonstrates the deployment of GLM-5.1 on Ascend NPUs using SGLang, including multi-node PD mixed mode,
|
||||
multi-node PD disaggregation mode, feature configuration, and performance optimization.
|
||||
|
||||
This document is validated and written based on **SGLang v0.5.13**. The current model (GLM-5.1) is fully supported in
|
||||
This document is validated and written based on **SGLang v0.5.16**. The current model (GLM-5.1) is fully supported in
|
||||
this version. To use the latest features (e.g., speculative decoding, multi-node deployment), it is recommended to use
|
||||
v0.5.13 or a later version.
|
||||
v0.5.16 or a later version.
|
||||
|
||||
## Supported features
|
||||
|
||||
@@ -96,7 +96,7 @@ For details, see [Docker image versions](/docs/hardware-platforms/ascend-npus/fa
|
||||
<Tab title="Atlas 800I A3">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -124,14 +124,14 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
<Tab title="Atlas 800I A2">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -151,7 +151,7 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
@@ -13,10 +13,6 @@ Attention (DSA) and multi-token prediction (MTP), and supports high-throughput i
|
||||
This document demonstrates the deployment of GLM-5.2 on Ascend NPUs using SGLang, including single-node deployment,
|
||||
multi-node deployment, prefill-decode disaggregation, feature configuration, and performance optimization.
|
||||
|
||||
> GLM-5.2 currently uses a daily build image because 0Day support was released before the related code was merged into
|
||||
> the SGLang main branch. The deployment commands and image tags in this document will be updated to a stable release
|
||||
> once the support lands in a stable SGLang release.
|
||||
|
||||
## Supported features
|
||||
|
||||
| Feature | Example usage |
|
||||
@@ -85,16 +81,11 @@ Ensure sufficient disk space before pulling images. The Docker image requires at
|
||||
The dependencies required for the NPU runtime environment have been integrated into a Docker image and uploaded to the
|
||||
online platform. You can directly pull it.
|
||||
|
||||
<Note>
|
||||
The GLM-5.2 images below use daily build tags because 0Day support was released before the related code was merged into
|
||||
the main branch. These tags will be switched to stable release images after the support lands in a stable release.
|
||||
</Note>
|
||||
|
||||
<Tabs>
|
||||
<Tab title="Atlas 800I A3">
|
||||
|
||||
```bash Command
|
||||
docker pull swr.cn-southwest-2.myhuaweicloud.com/base_image/dockerhub/lmsysorg/sglang:cann9.0.0-a3-glm5.2-20260615
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -122,14 +113,14 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
swr.cn-southwest-2.myhuaweicloud.com/base_image/dockerhub/lmsysorg/sglang:cann9.0.0-a3-glm5.2-20260615
|
||||
quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
<Tab title="Atlas 800I A2">
|
||||
|
||||
```bash Command
|
||||
docker pull swr.cn-southwest-2.myhuaweicloud.com/base_image/dockerhub/lmsysorg/sglang:cann9.0.0-910b-glm5.2-20260615
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -157,7 +148,7 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
swr.cn-southwest-2.myhuaweicloud.com/base_image/dockerhub/lmsysorg/sglang:cann9.0.0-910b-glm5.2-20260615
|
||||
quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
@@ -503,10 +494,10 @@ For performance data and benchmark commands, see [Performance Testing on Ascend
|
||||
|
||||
### Best practice configuration reference
|
||||
|
||||
GLM-5.2 has no standalone best practice page yet, as 0Day support is still on daily build images and tuning parameters
|
||||
are not finalized. Instead, the [Online service deployment](#online-service-deployment) section above provides
|
||||
ready-to-use scripts for single-node, multi-node, and PD disaggregation, each embedding the recommended feature
|
||||
combinations and tuning parameters (e.g., DeepEP mode, speculative decoding, overlap schedule).
|
||||
GLM-5.2 has no standalone best practice page yet, as tuning parameters are not finalized. Instead, the
|
||||
[Online service deployment](#online-service-deployment) section above provides ready-to-use scripts for single-node,
|
||||
multi-node, and PD disaggregation, each embedding the recommended feature combinations and tuning parameters (e.g.,
|
||||
DeepEP mode, speculative decoding, overlap schedule).
|
||||
|
||||
For the full catalog of optimization features and their parameter and compatibility details,
|
||||
see [Optimization on Ascend NPU](/docs/hardware-platforms/ascend-npus/optimization/parameter_tuning).
|
||||
|
||||
@@ -16,9 +16,9 @@ This document demonstrates the deployment of Kimi-K2.6 on Ascend NPUs using SGLa
|
||||
mode, multi-node PD mixed mode, multi-node PD disaggregation mode, feature configuration, and performance
|
||||
optimization.
|
||||
|
||||
This document is validated and written based on **SGLang v0.5.13**. The current model (Kimi-K2.6) is fully supported in
|
||||
This document is validated and written based on **SGLang v0.5.16**. The current model (Kimi-K2.6) is fully supported in
|
||||
this version. To use the latest features (e.g., speculative decoding, multimodal), it is recommended to use
|
||||
v0.5.13 or a later version.
|
||||
v0.5.16 or a later version.
|
||||
|
||||
## Supported features
|
||||
|
||||
@@ -98,7 +98,7 @@ For details, see [Docker image versions](/docs/hardware-platforms/ascend-npus/fa
|
||||
<Tab title="Atlas 800I A3">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -126,14 +126,14 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
<Tab title="Atlas 800I A2">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -153,7 +153,7 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
+6
-6
@@ -13,9 +13,9 @@ latency scenarios through PD disaggregation deployment.
|
||||
This document demonstrates the deployment of MiMo-V2-Flash on Ascend NPUs using SGLang, including multi-node PD
|
||||
disaggregation mode, feature configuration, and performance optimization.
|
||||
|
||||
This document is validated and written based on **SGLang v0.5.13**. The current model (MiMo-V2-Flash) is fully supported in
|
||||
This document is validated and written based on **SGLang v0.5.16**. The current model (MiMo-V2-Flash) is fully supported in
|
||||
this version. To use the latest features (e.g., PD disaggregation, speculative decoding), it is recommended to use
|
||||
v0.5.13 or a later version.
|
||||
v0.5.16 or a later version.
|
||||
|
||||
## Supported features
|
||||
|
||||
@@ -90,7 +90,7 @@ For details, see [Docker image versions](/docs/hardware-platforms/ascend-npus/fa
|
||||
<Tab title="Atlas 800I A3">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -118,14 +118,14 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
<Tab title="Atlas 800I A2">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -145,7 +145,7 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
+6
-6
@@ -15,9 +15,9 @@ long-context tasks.
|
||||
This document demonstrates the deployment of MiniMax-M2.5 on Ascend NPUs using SGLang, including single-node PD mixed
|
||||
mode, feature configuration, and performance optimization.
|
||||
|
||||
This document is validated and written based on **SGLang v0.5.13**. The current model (MiniMax-M2.5) is fully supported in
|
||||
This document is validated and written based on **SGLang v0.5.16**. The current model (MiniMax-M2.5) is fully supported in
|
||||
this version. To use the latest features (e.g., speculative decoding), it is recommended to use
|
||||
v0.5.13 or a later version.
|
||||
v0.5.16 or a later version.
|
||||
|
||||
## Supported features
|
||||
|
||||
@@ -90,7 +90,7 @@ For details, see [Docker image versions](/docs/hardware-platforms/ascend-npus/fa
|
||||
<Tab title="Atlas 800I A3">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -118,14 +118,14 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
<Tab title="Atlas 800I A2">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -145,7 +145,7 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
+6
-6
@@ -16,9 +16,9 @@ This document demonstrates the deployment of Qwen3-235B-A22B on Ascend NPUs usin
|
||||
mode, multi-node PD disaggregation mode, 256k long-sequence inference, Prefill Context Parallel, feature configuration,
|
||||
and performance optimization.
|
||||
|
||||
This document is validated and written based on **SGLang v0.5.13**. The current model (Qwen3-235B-A22B) is fully supported in
|
||||
This document is validated and written based on **SGLang v0.5.16**. The current model (Qwen3-235B-A22B) is fully supported in
|
||||
this version. To use the latest features (e.g., PD disaggregation, speculative decoding), it is recommended to use
|
||||
v0.5.13 or a later version.
|
||||
v0.5.16 or a later version.
|
||||
|
||||
## Supported features
|
||||
|
||||
@@ -98,7 +98,7 @@ For details, see [Docker image versions](/docs/hardware-platforms/ascend-npus/fa
|
||||
<Tab title="Atlas 800I A3">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -126,14 +126,14 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
<Tab title="Atlas 800I A2">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -153,7 +153,7 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
+6
-6
@@ -15,9 +15,9 @@ and thinking/reasoning-enhanced editions.
|
||||
This document demonstrates the deployment of Qwen3-30B-A3B on Ascend NPUs using SGLang, including single-node PD mixed
|
||||
mode, feature configuration, and performance optimization.
|
||||
|
||||
This document is validated and written based on **SGLang v0.5.13**. The current model (Qwen3-30B-A3B) is fully supported in
|
||||
This document is validated and written based on **SGLang v0.5.16**. The current model (Qwen3-30B-A3B) is fully supported in
|
||||
this version. To use the latest features (e.g., speculative decoding), it is recommended to use
|
||||
v0.5.13 or a later version.
|
||||
v0.5.16 or a later version.
|
||||
|
||||
## Supported features
|
||||
|
||||
@@ -89,7 +89,7 @@ For details, see [Docker image versions](/docs/hardware-platforms/ascend-npus/fa
|
||||
<Tab title="Atlas 800I A3">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -117,14 +117,14 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
<Tab title="Atlas 800I A2">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -144,7 +144,7 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
@@ -14,8 +14,8 @@ decoding for accelerated inference and supports both standard and thinking/reaso
|
||||
This document demonstrates the deployment of Qwen3-32B on Ascend NPUs using SGLang, including single-node PD mixed mode,
|
||||
feature configuration, and performance optimization.
|
||||
|
||||
This document is validated and written based on **SGLang v0.5.13**. The current model (Qwen3-32B) is fully supported in
|
||||
this version. To use the latest features (e.g., speculative decoding), it is recommended to use v0.5.13 or a later
|
||||
This document is validated and written based on **SGLang v0.5.16**. The current model (Qwen3-32B) is fully supported in
|
||||
this version. To use the latest features (e.g., speculative decoding), it is recommended to use v0.5.16 or a later
|
||||
version.
|
||||
|
||||
## Supported features
|
||||
@@ -89,7 +89,7 @@ For details, see [Docker image versions](/docs/hardware-platforms/ascend-npus/fa
|
||||
<Tab title="Atlas 800I A3">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -117,14 +117,14 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
<Tab title="Atlas 800I A2">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -144,7 +144,7 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
+6
-6
@@ -16,9 +16,9 @@ prediction (MTP) for speculative decoding.
|
||||
This document demonstrates the deployment of Qwen3.5-397B-A17B on Ascend NPUs using SGLang, including single-node PD
|
||||
mixed mode, feature configuration, and performance optimization.
|
||||
|
||||
This document is validated and written based on **SGLang v0.5.13**. The current model (Qwen3.5-397B-A17B) is fully
|
||||
This document is validated and written based on **SGLang v0.5.16**. The current model (Qwen3.5-397B-A17B) is fully
|
||||
supported in this version. To use the latest features (e.g., speculative decoding, multimodal), it is recommended to
|
||||
use v0.5.13 or a later version.
|
||||
use v0.5.16 or a later version.
|
||||
|
||||
## Supported features
|
||||
|
||||
@@ -91,7 +91,7 @@ For details, see [Docker image versions](/docs/hardware-platforms/ascend-npus/fa
|
||||
<Tab title="Atlas 800I A3">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -119,14 +119,14 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
<Tab title="Atlas 800I A2">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -146,7 +146,7 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
@@ -15,9 +15,9 @@ reasoning, text comprehension, and tool usage.
|
||||
This document demonstrates the deployment of Qwen3.6-27B on Ascend NPUs using SGLang, including single-node PD mixed
|
||||
mode, feature configuration, and performance optimization.
|
||||
|
||||
This document is validated and written based on **SGLang v0.5.13**. The current model (Qwen3.6-27B) is fully supported
|
||||
This document is validated and written based on **SGLang v0.5.16**. The current model (Qwen3.6-27B) is fully supported
|
||||
in this version. To use the latest features (e.g., speculative decoding, multimodal), it is recommended to use
|
||||
v0.5.13 or a later version.
|
||||
v0.5.16 or a later version.
|
||||
|
||||
## Supported features
|
||||
|
||||
@@ -87,7 +87,7 @@ For details, see [Docker image versions](/docs/hardware-platforms/ascend-npus/fa
|
||||
<Tab title="Atlas 800I A3">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -115,14 +115,14 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
<Tab title="Atlas 800I A2">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -142,7 +142,7 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
+7
-8
@@ -16,10 +16,9 @@ comprehension, and tool usage.
|
||||
This document demonstrates the deployment of Qwen3.6-35B-A3B on Ascend NPUs using SGLang, including single-node PD mixed
|
||||
mode, feature configuration, and performance optimization.
|
||||
|
||||
This document is validated and written based on **SGLang v0.5.14**. The current model (Qwen3.6-35B-A3B) is fully supported
|
||||
in this version. To use the latest features (e.g., speculative decoding / NEXTN, multimodal), use
|
||||
**v0.5.14 or a later version**. Ascend images based on `v0.5.13(.post1)` can fail when NEXTN is enabled
|
||||
(`NPUMHATokenToKVPool` missing `data_strides`).
|
||||
This document is validated and written based on **SGLang v0.5.16**. The current model (Qwen3.6-35B-A3B) is fully supported
|
||||
in this version. To use the latest features (e.g., speculative decoding / NEXTN, multimodal), use **v0.5.16 or a later
|
||||
version**.
|
||||
|
||||
## Supported features
|
||||
|
||||
@@ -88,7 +87,7 @@ For details, see [Docker image versions](/docs/hardware-platforms/ascend-npus/fa
|
||||
<Tab title="Atlas 800I A3">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.14-cann9.0.0-a3
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -116,14 +115,14 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.14-cann9.0.0-a3
|
||||
quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
<Tab title="Atlas 800I A2">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.14-cann9.0.0-910b
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -143,7 +142,7 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.14-cann9.0.0-910b
|
||||
quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
@@ -15,8 +15,8 @@ editions.
|
||||
This document demonstrates the deployment of Qwen3-8B on Ascend NPUs using SGLang, including single-node PD mixed mode,
|
||||
feature configuration, and performance optimization.
|
||||
|
||||
This document is validated and written based on **SGLang v0.5.13**. The current model (Qwen3-8B) is fully supported in
|
||||
this version. To use the latest features (e.g., speculative decoding), it is recommended to use v0.5.13 or a later
|
||||
This document is validated and written based on **SGLang v0.5.16**. The current model (Qwen3-8B) is fully supported in
|
||||
this version. To use the latest features (e.g., speculative decoding), it is recommended to use v0.5.16 or a later
|
||||
version.
|
||||
|
||||
## Supported features
|
||||
@@ -90,7 +90,7 @@ For details, see [Docker image versions](/docs/hardware-platforms/ascend-npus/fa
|
||||
<Tab title="Atlas 800I A3">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -118,14 +118,14 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
<Tab title="Atlas 800I A2">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -145,7 +145,7 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
+6
-6
@@ -15,9 +15,9 @@ following, reasoning, and text generation tasks.
|
||||
This document demonstrates the deployment of Qwen3-Next-80B-A3B-Instruct on Ascend NPUs using SGLang, including
|
||||
single-node PD mixed mode, feature configuration, and performance optimization.
|
||||
|
||||
This document is validated and written based on **SGLang v0.5.13**. The current model (Qwen3-Next-80B-A3B-Instruct) is
|
||||
This document is validated and written based on **SGLang v0.5.16**. The current model (Qwen3-Next-80B-A3B-Instruct) is
|
||||
fully supported in this version. To use the latest features (e.g., speculative decoding), it is recommended to use
|
||||
v0.5.13 or a later version.
|
||||
v0.5.16 or a later version.
|
||||
|
||||
## Supported features
|
||||
|
||||
@@ -92,7 +92,7 @@ For details, see [Docker image versions](/docs/hardware-platforms/ascend-npus/fa
|
||||
<Tab title="Atlas 800I A3">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -120,14 +120,14 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-a3
|
||||
quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
<Tab title="Atlas 800I A2">
|
||||
|
||||
```bash Command
|
||||
docker pull quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
|
||||
docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--privileged=true --net=host \
|
||||
@@ -147,7 +147,7 @@ docker run -itd --shm-size=16g --name ${NAME} \
|
||||
--device=/dev/davinci_manager:/dev/davinci_manager \
|
||||
--device=/dev/hisi_hdc:/dev/hisi_hdc \
|
||||
--entrypoint=bash \
|
||||
quay.io/ascend/sglang:v0.5.13.post1-cann9.0.0-910b
|
||||
quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16
|
||||
```
|
||||
|
||||
</Tab>
|
||||
|
||||
@@ -88,6 +88,11 @@ This document provides a list of commonly used environment variables and aims to
|
||||
</td>
|
||||
<td><code>0</code></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td><code>DEEPEP_HCCL_BUFFSIZE</code></td>
|
||||
<td>Configures the HCCL buffer size (in MB) for the process groups used by <br/> DeepEP MoE All-to-All communication (the default group and <code>moe</code>-named <br/> groups), allowing the MoE communication buffer to be tuned independently from <br/> non-MoE groups. If unset, falls back to <code>HCCL_BUFFSIZE</code>, then to <code>200</code>. <br/> A larger buffer reduces All-to-All latency at the cost of more HBM usage.</td>
|
||||
<td><code>200</code></td>
|
||||
</tr>
|
||||
</tbody>
|
||||
</table>
|
||||
|
||||
|
||||
@@ -21,7 +21,7 @@ SGLang supports the following Ascend inference hardware:
|
||||
<Note>
|
||||
Throughout these docs, **A2** and **A3** are used as shorthand for the hardware above.
|
||||
Docker image tags use `910b` for A2 and `a3` for A3.
|
||||
For example, `v0.5.13.post1-cann9.0.0-910b` and `v0.5.13.post1-cann9.0.0-a3`.
|
||||
For example, `cann9.0.0-910b-v0.5.16` and `cann9.0.0-a3-v0.5.16`.
|
||||
|
||||
On A3, each card has 2 dies, giving 16 devices vs 8 on A2.
|
||||
Benchmark pages use "Cards" to refer to physical cards, so `Cards: 4` on A3 means `--tp-size 8`.
|
||||
|
||||
@@ -1376,7 +1376,7 @@ click [Server Arguments](../../../advanced_features/server_arguments).
|
||||
<tr>
|
||||
<td style={{padding: "9px 12px", fontWeight: 500, backgroundColor: "rgba(255,255,255,0.02)"}}>`--speculative-algorithm`</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>`None`</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>`EAGLE3`,<br/> `NEXTN`</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.02)"}}>`EAGLE`,<br/> `EAGLE3`,<br/> `NEXTN`</td>
|
||||
<td style={{padding: "9px 12px", backgroundColor: "rgba(255,255,255,0.05)"}}>A2, A3</td>
|
||||
</tr>
|
||||
<tr>
|
||||
|
||||
Reference in New Issue
Block a user