[NPU] chore: bump basic software version to 8.3.rc2 (#14614)
This commit is contained in:
@@ -18,7 +18,7 @@ conda activate sglang_npu
|
||||
|
||||
#### CANN
|
||||
|
||||
Prior to start work with SGLang on Ascend you need to install CANN Toolkit, Kernels operator package and NNAL version 8.3.RC1 or higher, check the [installation guide](https://www.hiascend.com/document/detail/zh/CANNCommunityEdition/83RC1/softwareinst/instg/instg_0008.html?Mode=PmIns&InstallType=local&OS=openEuler&Software=cannToolKit)
|
||||
Prior to start work with SGLang on Ascend you need to install CANN Toolkit, Kernels operator package and NNAL version 8.3.RC2 or higher, check the [installation guide](https://www.hiascend.com/document/detail/zh/CANNCommunityEdition/83RC1/softwareinst/instg/instg_0008.html?Mode=PmIns&InstallType=local&OS=openEuler&Software=cannToolKit)
|
||||
|
||||
#### MemFabric Adaptor
|
||||
|
||||
|
||||
@@ -30,7 +30,7 @@ python3 -m sglang.launch_server \
|
||||
--trust-remote-code \
|
||||
--attention-backend ascend \
|
||||
--device npu \
|
||||
--quantization w8a8_int8 \
|
||||
--quantization modelslim \
|
||||
--watchdog-timeout 9000 \
|
||||
--host 127.0.0.1 \
|
||||
--port 6688 \
|
||||
@@ -89,7 +89,7 @@ python -m sglang.launch_server \
|
||||
--mem-fraction-static 0.6 \
|
||||
--attention-backend ascend \
|
||||
--device npu \
|
||||
--quantization w8a8_int8 \
|
||||
--quantization modelslim \
|
||||
--disaggregation-transfer-backend ascend \
|
||||
--max-running-requests 8 \
|
||||
--context-length 8192 \
|
||||
@@ -145,7 +145,7 @@ python -m sglang.launch_server \
|
||||
--max-running-requests 352 \
|
||||
--attention-backend ascend \
|
||||
--device npu \
|
||||
--quantization w8a8_int8 \
|
||||
--quantization modelslim \
|
||||
--moe-a2a-backend deepep \
|
||||
--enable-dp-attention \
|
||||
--deepep-mode low_latency \
|
||||
@@ -214,7 +214,7 @@ do
|
||||
--mem-fraction-static 0.81 \
|
||||
--attention-backend ascend \
|
||||
--device npu \
|
||||
--quantization w8a8_int8 \
|
||||
--quantization modelslim \
|
||||
--disaggregation-transfer-backend ascend \
|
||||
--max-running-requests 8 \
|
||||
--context-length 8192 \
|
||||
@@ -275,7 +275,7 @@ do
|
||||
--max-running-requests 832 \
|
||||
--attention-backend ascend \
|
||||
--device npu \
|
||||
--quantization w8a8_int8 \
|
||||
--quantization modelslim \
|
||||
--moe-a2a-backend deepep \
|
||||
--enable-dp-attention \
|
||||
--deepep-mode low_latency \
|
||||
|
||||
@@ -7,7 +7,7 @@ MindSpore is a high-performance AI framework optimized for Ascend NPUs. This doc
|
||||
## Requirements
|
||||
|
||||
MindSpore currently only supports Ascend NPU devices. Users need to first install Ascend CANN software packages.
|
||||
The CANN software packages can be downloaded from the [Ascend Official Website](https://www.hiascend.com). The recommended version is 8.3.RC1.
|
||||
The CANN software packages can be downloaded from the [Ascend Official Website](https://www.hiascend.com). The recommended version is 8.3.RC2.
|
||||
|
||||
## Supported Models
|
||||
|
||||
|
||||
Reference in New Issue
Block a user