[Intel CPU/XPU] SGL doc updates (#23547)
merge this one as doc change only.
This commit is contained in:
@@ -222,12 +222,12 @@ Notes:
|
||||
Example command to launch SGLang serving:
|
||||
|
||||
```bash Launch Server
|
||||
python -m sglang.launch_server \
|
||||
--model <MODEL_ID_OR_PATH> \
|
||||
--trust-remote-code \
|
||||
--disable-overlap-schedule \
|
||||
--device cpu \
|
||||
--host 0.0.0.0 \
|
||||
sglang serve \
|
||||
--model-path <MODEL_ID_OR_PATH> \
|
||||
--trust-remote-code \
|
||||
--disable-overlap-schedule \
|
||||
--device cpu \
|
||||
--host 0.0.0.0 \
|
||||
--tp 6
|
||||
```
|
||||
|
||||
@@ -320,64 +320,64 @@ or MoE (Mixture of Experts) models with fewer than 10B activated parameters, can
|
||||
An example command to launch service of W8A8_INT8 DeepSeek-V3.1-Terminus on a Xeon® 6980P server:
|
||||
|
||||
```bash W8A8_INT8
|
||||
python -m sglang.launch_server \
|
||||
--model IntervitensInc/DeepSeek-V3.1-Terminus-Channel-int8 \
|
||||
--trust-remote-code \
|
||||
--disable-overlap-schedule \
|
||||
--device cpu \
|
||||
--quantization w8a8_int8 \
|
||||
--host 0.0.0.0 \
|
||||
--enable-torch-compile \
|
||||
--torch-compile-max-bs 4 \
|
||||
sglang serve \
|
||||
--model-path IntervitensInc/DeepSeek-V3.1-Terminus-Channel-int8 \
|
||||
--trust-remote-code \
|
||||
--disable-overlap-schedule \
|
||||
--device cpu \
|
||||
--quantization w8a8_int8 \
|
||||
--enable-torch-compile \
|
||||
--torch-compile-max-bs 4 \
|
||||
--host 0.0.0.0 \
|
||||
--tp 6
|
||||
```
|
||||
|
||||
Similarly, an example command to launch service of FP8 DeepSeek-V3.1-Terminus would be:
|
||||
|
||||
```bash FP8
|
||||
python -m sglang.launch_server \
|
||||
--model deepseek-ai/DeepSeek-V3.1-Terminus \
|
||||
--trust-remote-code \
|
||||
--disable-overlap-schedule \
|
||||
--device cpu \
|
||||
--host 0.0.0.0 \
|
||||
--enable-torch-compile \
|
||||
--torch-compile-max-bs 4 \
|
||||
sglang serve \
|
||||
--model-path deepseek-ai/DeepSeek-V3.1-Terminus \
|
||||
--trust-remote-code \
|
||||
--disable-overlap-schedule \
|
||||
--device cpu \
|
||||
--enable-torch-compile \
|
||||
--torch-compile-max-bs 4 \
|
||||
--host 0.0.0.0 \
|
||||
--tp 6
|
||||
```
|
||||
|
||||
Note: Please set `--torch-compile-max-bs` to the maximum desired batch size for your deployment,
|
||||
which can be up to 16. The value `4` in the examples is illustrative.
|
||||
Note: Please set `--torch-compile-max-bs` to the maximum desired batch size for your deployment.
|
||||
The value `4` in the examples is illustrative.
|
||||
|
||||
### Example: Running Llama-3.2-3B
|
||||
|
||||
An example command to launch service of Llama-3.2-3B with BF16 precision:
|
||||
|
||||
```bash BF16
|
||||
python -m sglang.launch_server \
|
||||
--model meta-llama/Llama-3.2-3B-Instruct \
|
||||
--trust-remote-code \
|
||||
--disable-overlap-schedule \
|
||||
--device cpu \
|
||||
--host 0.0.0.0 \
|
||||
--enable-torch-compile \
|
||||
--torch-compile-max-bs 16 \
|
||||
--tp 2
|
||||
sglang serve \
|
||||
--model-path meta-llama/Llama-3.2-3B-Instruct \
|
||||
--trust-remote-code \
|
||||
--disable-overlap-schedule \
|
||||
--device cpu \
|
||||
--enable-torch-compile \
|
||||
--torch-compile-max-bs 16 \
|
||||
--host 0.0.0.0 \
|
||||
--tp 3
|
||||
```
|
||||
|
||||
The example command to launch service of W8A8_INT8 version of Llama-3.2-3B:
|
||||
|
||||
```bash W8A8_INT8
|
||||
python -m sglang.launch_server \
|
||||
--model RedHatAI/Llama-3.2-3B-quantized.w8a8 \
|
||||
--trust-remote-code \
|
||||
--disable-overlap-schedule \
|
||||
--device cpu \
|
||||
--quantization w8a8_int8 \
|
||||
--host 0.0.0.0 \
|
||||
--enable-torch-compile \
|
||||
--torch-compile-max-bs 16 \
|
||||
--tp 2
|
||||
sglang serve \
|
||||
--model-path RedHatAI/Llama-3.2-3B-quantized.w8a8 \
|
||||
--trust-remote-code \
|
||||
--disable-overlap-schedule \
|
||||
--device cpu \
|
||||
--quantization w8a8_int8 \
|
||||
--enable-torch-compile \
|
||||
--torch-compile-max-bs 16 \
|
||||
--host 0.0.0.0 \
|
||||
--tp 3
|
||||
```
|
||||
|
||||
Note: The `--torch-compile-max-bs` and `--tp` settings are examples that should be adjusted for your setup.
|
||||
|
||||
@@ -71,15 +71,40 @@ pip install -v . --extra-index-url https://download.pytorch.org/whl/xpu
|
||||
|
||||
### Install Using Docker
|
||||
|
||||
The docker for XPU is under active development. Please stay tuned.
|
||||
[The SGLang XPU Dockerfile](https://github.com/sgl-project/sglang/blob/main/docker/xpu.Dockerfile) is provided to facilitate the installation.
|
||||
Replace `<secret>` below with your [HuggingFace access token](https://huggingface.co/docs/hub/en/security-tokens).
|
||||
|
||||
```bash Command
|
||||
# Clone the SGLang repository
|
||||
git clone https://github.com/sgl-project/sglang.git
|
||||
cd sglang/docker
|
||||
|
||||
# Build the docker image
|
||||
docker build -t sglang-xpu:latest -f xpu.Dockerfile .
|
||||
|
||||
# Initiate a docker container
|
||||
docker run \
|
||||
-it \
|
||||
--privileged \
|
||||
--ipc=host \
|
||||
--network=host \
|
||||
--group-add $(getent group video | cut -d: -f3) \
|
||||
--device /dev/dri \
|
||||
-v /dev/dri/by-path:/dev/dri/by-path \
|
||||
-v /dev/shm:/dev/shm \
|
||||
-v ~/.cache/huggingface:/root/.cache/huggingface \
|
||||
-p 30000:30000 \
|
||||
-e "HF_TOKEN=<secret>" \
|
||||
sglang-xpu:latest /bin/bash
|
||||
```
|
||||
|
||||
## Launch of the Serving Engine
|
||||
|
||||
Example command to launch SGLang serving:
|
||||
|
||||
```bash
|
||||
python -m sglang.launch_server \
|
||||
--model <MODEL_ID_OR_PATH> \
|
||||
sglang serve \
|
||||
--model-path <MODEL_ID_OR_PATH> \
|
||||
--trust-remote-code \
|
||||
--disable-overlap-schedule \
|
||||
--device xpu \
|
||||
|
||||
Reference in New Issue
Block a user