From f148eb6e6ea51eff7ef8725a77c71818e3fe6fe3 Mon Sep 17 00:00:00 2001 From: Jinyan Yi Date: Tue, 11 Aug 2026 21:20:30 +0800 Subject: [PATCH] Add Hunyuan3 On Ascend Doc (#30223) --- docs/docs.json | 1 + .../model-deployment/tutorials/hy3.mdx | 388 ++++++++++++++++++ 2 files changed, 389 insertions(+) create mode 100644 docs/docs/hardware-platforms/ascend-npus/model-deployment/tutorials/hy3.mdx diff --git a/docs/docs.json b/docs/docs.json index 5b91d773f..c828898d9 100644 --- a/docs/docs.json +++ b/docs/docs.json @@ -1106,6 +1106,7 @@ "docs/hardware-platforms/ascend-npus/model-deployment/tutorials/deepseek_v3_2", "docs/hardware-platforms/ascend-npus/model-deployment/tutorials/glm_5_1", "docs/hardware-platforms/ascend-npus/model-deployment/tutorials/glm_5_2", + "docs/hardware-platforms/ascend-npus/model-deployment/tutorials/hy3", "docs/hardware-platforms/ascend-npus/model-deployment/tutorials/kimi_k2_6", "docs/hardware-platforms/ascend-npus/model-deployment/tutorials/minimax_m2_5", "docs/hardware-platforms/ascend-npus/model-deployment/tutorials/mimo_v2_flash", diff --git a/docs/docs/hardware-platforms/ascend-npus/model-deployment/tutorials/hy3.mdx b/docs/docs/hardware-platforms/ascend-npus/model-deployment/tutorials/hy3.mdx new file mode 100644 index 000000000..0ad0b420a --- /dev/null +++ b/docs/docs/hardware-platforms/ascend-npus/model-deployment/tutorials/hy3.mdx @@ -0,0 +1,388 @@ +--- +title: "Hy3" +metatags: + description: "Deploy Hy3 model with SGLang on Ascend NPUs, including single-node deployment and speculative decoding." +--- + +## Introduction + +Hy3 is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer +parameters, developed by the Tencent Hy Team. It features 192 experts with top-8 activation per token, GQA attention +(64 heads, 8 KV heads, head dim 128), and a 256K context length. The model supports built-in reasoning via +`reasoning_effort`, tool calling, and multi-token prediction (MTP) for speculative decoding. Hy3 significantly +outperforms similar-size models and rivals flagship open-source models with 2–5× the parameters, with notable +improvements in agentic workflows, long-context tasks, and production reliability. + +This document demonstrates the deployment of Hy3 on Ascend NPUs using SGLang, including single-node (Atlas 800I A3) and +multi-node (Atlas 800I A2) PD mixed mode and speculative decoding. + +## Supported features + +| Feature | Example usage | +|--------------------------|----------------------------------------------------------------------------------------------------------------------------------------| +| Tensor Parallelism | `--tp-size 16` | +| NPU Graph | enabled by default; disable with `--disable-cuda-graph`;
control range via `--cuda-graph-bs`; e.g. `--cuda-graph-bs 4 8 16 20 24 28 32 36 40` | +| Speculative Decoding | `--speculative-algorithm EAGLE \`
`--speculative-num-steps 2 \`
`--speculative-eagle-topk 1 \`
`--speculative-num-draft-tokens 3` | +| Overlap Schedule | `export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1` | +| Reasoning Mode | `extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}}` — deep chain-of-thought;
`"reasoning_effort": "low"` — short thinking chain;
`"reasoning_effort": "no_think"` — direct response, no thinking | +| Tool Calling | `--tool-call-parser auto --reasoning-parser auto` | + + +The values in the **Example usage** column are for illustration only. Adjust them according to your hardware, deployment +mode, and workload. For parameter details, see +[Feature descriptions](/docs/hardware-platforms/ascend-npus/optimization/parameter_tuning#feature-descriptions). + + +For feature compatibility and conflict information between features, +see [Feature Compatibility](/docs/hardware-platforms/ascend-npus/optimization/parameter_tuning#feature-compatibility). + +## Prerequisites + +### Environment + +Before following this tutorial, complete the environment setup in the documents below: + +- [Ascend NPU Quickstart](/docs/hardware-platforms/ascend-npus/getting-started/quick_start) — the fastest way to get started. + It walks you through launching the official container image, starting the SGLang server, and sending a test request. + Recommended if you are new to SGLang on Ascend. +- [SGLang Installation with NPU Support](/docs/hardware-platforms/ascend-npus/getting-started/installation) — the full installation guide. + It covers the component version mapping (CANN, TorchNPU, Triton, kernels, etc.), building from source or from a + Dockerfile, and recommended system settings (CPU power scheme, NUMA, swap). Use it when you need to install or customize + the environment instead of using the official image. + +### Model weights + + +Hy3 has 295B total parameters. If you need to download model weights, check the model size before downloading to reserve enough space. + + +- [Hy3](https://www.modelscope.cn/models/Tencent-Hunyuan/Hy3) (BF16, 597.60GB) + + +The BF16 variant (597.60GB) can be deployed on 16 × 64GB of device memory (`--tp-size 16`), which corresponds to one full +Atlas 800I A3 node (8 cards × 2 dies) or two Atlas 800I A2 nodes (8 cards × 1 die each). + + +It is recommended to download the model weights to a shared directory accessible from within the container. + +For the hardware specifications (memory per die, dies per card, and the difference between A2 and A3), +see [Ascend NPU Reference — Hardware](/docs/hardware-platforms/ascend-npus/reference/glossary#hardware). + +## Installation + + +Ensure sufficient disk space before pulling images. The Docker image requires at least **30GB** of free space. + + +The dependencies required for the NPU runtime environment have been integrated into a Docker image and uploaded to the +online platform. You can directly pull it. + + +Both **stable releases** and **daily builds** are available. The following command is based on the stable release tag. +For details, see +[Docker image versions](/docs/hardware-platforms/ascend-npus/faq#8-docker-image-versions-stable-release-vs-daily-build). + + + + + +```bash Command +docker pull quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16 + +docker run -itd --shm-size=64g --name ${NAME} \ +--privileged=true --net=host \ +-v /var/queue_schedule:/var/queue_schedule \ +-v /etc/ascend_install.info:/etc/ascend_install.info \ +-v /usr/local/sbin:/usr/local/sbin \ +-v /usr/local/Ascend/driver:/usr/local/Ascend/driver \ +-v /usr/local/Ascend/firmware:/usr/local/Ascend/firmware \ +--device=/dev/davinci0:/dev/davinci0 \ +--device=/dev/davinci1:/dev/davinci1 \ +--device=/dev/davinci2:/dev/davinci2 \ +--device=/dev/davinci3:/dev/davinci3 \ +--device=/dev/davinci4:/dev/davinci4 \ +--device=/dev/davinci5:/dev/davinci5 \ +--device=/dev/davinci6:/dev/davinci6 \ +--device=/dev/davinci7:/dev/davinci7 \ +--device=/dev/davinci8:/dev/davinci8 \ +--device=/dev/davinci9:/dev/davinci9 \ +--device=/dev/davinci10:/dev/davinci10 \ +--device=/dev/davinci11:/dev/davinci11 \ +--device=/dev/davinci12:/dev/davinci12 \ +--device=/dev/davinci13:/dev/davinci13 \ +--device=/dev/davinci14:/dev/davinci14 \ +--device=/dev/davinci15:/dev/davinci15 \ +--device=/dev/davinci_manager:/dev/davinci_manager \ +--device=/dev/hisi_hdc:/dev/hisi_hdc \ +--entrypoint=bash \ +quay.io/ascend/sglang:cann9.0.0-a3-v0.5.16 +``` + + + + +```bash Command +docker pull quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16 + +docker run -itd --shm-size=64g --name ${NAME} \ +--privileged=true --net=host \ +-v /var/queue_schedule:/var/queue_schedule \ +-v /etc/ascend_install.info:/etc/ascend_install.info \ +-v /usr/local/sbin:/usr/local/sbin \ +-v /usr/local/Ascend/driver:/usr/local/Ascend/driver \ +-v /usr/local/Ascend/firmware:/usr/local/Ascend/firmware \ +--device=/dev/davinci0:/dev/davinci0 \ +--device=/dev/davinci1:/dev/davinci1 \ +--device=/dev/davinci2:/dev/davinci2 \ +--device=/dev/davinci3:/dev/davinci3 \ +--device=/dev/davinci4:/dev/davinci4 \ +--device=/dev/davinci5:/dev/davinci5 \ +--device=/dev/davinci6:/dev/davinci6 \ +--device=/dev/davinci7:/dev/davinci7 \ +--device=/dev/davinci_manager:/dev/davinci_manager \ +--device=/dev/hisi_hdc:/dev/hisi_hdc \ +--entrypoint=bash \ +quay.io/ascend/sglang:cann9.0.0-910b-v0.5.16 +``` + + + + + +- If the model weights have already been downloaded to a shared directory, use `-v` to mount the model path into the + container, for example: `-v /path/to/models:/models`. +- Replace `${NAME}` with your own container name or remove `--name` to use default name. +- `--shm-size=64g` is recommended for 16-NPU deployments with large batch sizes. + + +## Online service deployment + +### Single-node online deployment + +Single-node deployment completes both prefill and decode within the same node (PD mixed mode). Hy3 with 295B +parameters requires all 16 logical NPUs on an Atlas 800I A3 server (8 physical NPUs × 2 logical cores each, 64 GB HBM +per logical NPU). + +Set the following environment variables before launching the server: + +```bash +# Performance tuning +export SGLANG_SET_CPU_AFFINITY=1 +echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor +sysctl -w vm.swappiness=0 +sysctl -w kernel.numa_balancing=0 + +# CANN environment (adjust paths to match your installation) +source /usr/local/Ascend/ascend-toolkit/set_env.sh +source /usr/local/Ascend/nnal/atb/set_env.sh + +# Ascend NPU settings +export ASCEND_USE_FIA=1 +export STREAMS_PER_DEVICE=32 +export HCCL_BUFFSIZE=3000 +export HCCL_OP_EXPANSION_MODE=AIV +export HCCL_SOCKET_IFNAME=lo +export GLOO_SOCKET_IFNAME=lo + +# SGLang settings +export SGLANG_ENABLE_SPEC_V2=1 +export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1 +export DEEP_NORMAL_MODE_USE_INT8_QUANT=1 +``` + +Launch the server: + +```bash +MODEL_PATH="/path/to/Hy3" + +python3 -m sglang.launch_server \ + --model-path ${MODEL_PATH} \ + --attention-backend ascend \ + --reasoning-parser auto \ + --tool-call-parser auto \ + --device npu \ + --tp-size 16 \ + --host 0.0.0.0 \ + --port 9999 \ + --mem-fraction-static 0.84 \ + --dtype bfloat16 \ + --base-gpu-id 0 \ + --prefill-max-requests 40 \ + --max-running-requests 40 \ + --cuda-graph-bs 4 8 16 20 24 28 32 36 40 \ + --speculative-algorithm EAGLE \ + --speculative-num-steps 2 \ + --speculative-eagle-topk 1 \ + --speculative-num-draft-tokens 3 +``` + +### Multi-node PD mixed deployment on Atlas 800I A2 + +Each Atlas 800I A2 node provides 8 NPUs (Ascend 910B, 64 GB HBM per NPU). Hy3 requires all 16 logical NPUs across +**2 × Atlas 800I A2** nodes in PD mixed mode (prefill and decode are colocated on the same 2-node cluster). + +On each node, set the following environment variables before launching the server: + +```bash +# ============================================================ +# Before running, update the following variables: +# HCCL_SOCKET_IFNAME: network interface name for HCCL (use `ifconfig` to find) +# GLOO_SOCKET_IFNAME: network interface name for Gloo +# ============================================================ + +# Performance tuning +export SGLANG_SET_CPU_AFFINITY=1 +echo performance | tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor +sysctl -w vm.swappiness=0 +sysctl -w kernel.numa_balancing=0 + +# CANN environment (adjust paths to match your installation) +source /usr/local/Ascend/ascend-toolkit/set_env.sh +source /usr/local/Ascend/nnal/atb/set_env.sh + +# Ascend NPU settings +export ASCEND_USE_FIA=1 +export STREAMS_PER_DEVICE=32 +export HCCL_BUFFSIZE=3000 +export HCCL_OP_EXPANSION_MODE=AIV +export HCCL_SOCKET_IFNAME= +export GLOO_SOCKET_IFNAME= + +# SGLang settings +export SGLANG_ENABLE_SPEC_V2=1 +export SGLANG_ENABLE_OVERLAP_PLAN_STREAM=1 +export DEEP_NORMAL_MODE_USE_INT8_QUANT=1 +``` + +Launch the server on both nodes by running the same script on each node. `NODE_IPS[0]` is the master node and must be +reachable from the other node: + +```bash +# ============================================================ +# Before running, update the following variables: +# MODEL_PATH: path to the model weights directory (shared) +# NODE_IPS: IP addresses of the two nodes, e.g. ('10.0.0.1' '10.0.0.2') +# ============================================================ + +MODEL_PATH="/path/to/Hy3" +NODE_IPS=('' '') + +for i in "${!NODE_IPS[@]}"; do + if [[ "$(hostname -I | awk '{print $1}')" == "${NODE_IPS[$i]}" ]]; then + python3 -m sglang.launch_server \ + --model-path ${MODEL_PATH} \ + --attention-backend ascend \ + --reasoning-parser auto \ + --tool-call-parser auto \ + --device npu \ + --tp-size 16 \ + --nnodes 2 \ + --dist-init-addr ${NODE_IPS[0]}:5000 \ + --node-rank $i \ + --host 0.0.0.0 \ + --port 9999 \ + --mem-fraction-static 0.84 \ + --dtype bfloat16 \ + --base-gpu-id 0 \ + --prefill-max-requests 40 \ + --max-running-requests 40 \ + --cuda-graph-bs 4 8 16 20 24 28 32 36 40 \ + --speculative-algorithm EAGLE \ + --speculative-num-steps 2 \ + --speculative-eagle-topk 1 \ + --speculative-num-draft-tokens 3 + break + fi +done +``` + + +- `HCCL_SOCKET_IFNAME` and `GLOO_SOCKET_IFNAME` must be set to the cluster network interface on both nodes. Do **not** use + `lo` for multi-node deployments. +- Model weights must be accessible from both nodes, e.g. mounted from a shared directory. +- Send requests only to the master node (`NODE_IPS[0]:9999`), not to the worker node. + + +## Functional verification + +After the service is started, you can invoke the model by sending a prompt: + +```shell +# ============================================================ +# Before running, update the following variables: +# HOST: the server host address (e.g., localhost) +# PORT: the server port number (e.g., 9999) +# ============================================================ + +curl http://${HOST}:${PORT}/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Hy3", + "messages": [{"role": "user", "content": "What is the capital of France?"}], + "max_tokens": 64, + "temperature": 0.9, + "extra_body": {"chat_template_kwargs": {"reasoning_effort": "no_think"}} + }' +``` + +Expected result: an HTTP 200 response with `"Paris"` in the content field. + +To enable chain-of-thought reasoning for math or coding tasks, set `reasoning_effort` to `"high"`: + +```shell +curl http://${HOST}:${PORT}/v1/chat/completions \ + -H "Content-Type: application/json" \ + -d '{ + "model": "Hy3", + "messages": [{"role": "user", "content": "Solve: if 3x + 7 = 22, what is x?"}], + "max_tokens": 8192, + "temperature": 0.9, + "extra_body": {"chat_template_kwargs": {"reasoning_effort": "high"}} + }' +``` + +The response will contain a `reasoning_content` field with the thinking process and a `content` field with the final +answer. + +Once the server prints `The server is fired up and ready to roll!` in the logs, it is ready to accept requests. For more testing examples (Health Check, Generate, Chat Completions, and port usage guidance), see [Testing the Service](/docs/hardware-platforms/ascend-npus/getting-started/installation#testing-the-service). + +## Accuracy evaluation + +For accuracy evaluation methods and datasets, see [Accuracy Evaluation on Ascend NPU](/docs/hardware-platforms/ascend-npus/evaluation/accuracy_evaluation). + +## Performance + +For performance data and benchmark commands, see [Performance Testing on Ascend NPU](/docs/hardware-platforms/ascend-npus/evaluation/performance_testing). + +## Best practices + +### Best practice configuration reference + +Hy3 has no standalone best practice page yet. The [Single-node online deployment](#single-node-online-deployment) and +[Multi-node PD mixed deployment on Atlas 800I A2](#multi-node-pd-mixed-deployment-on-atlas-800i-a2) sections above provide +the ready-to-use scripts for the supported PD mixed cases. + +For the full catalog of optimization features and their parameter and compatibility details, +see [Optimization on Ascend NPU](/docs/hardware-platforms/ascend-npus/optimization/parameter_tuning). + +## Performance tuning + +For the full list of supported features, see [Supported features](#supported-features). For detailed optimization +guidance, see [Optimization on Ascend NPU](/docs/hardware-platforms/ascend-npus/optimization/parameter_tuning). + +### Reasoning mode selection + +Hy3 supports three reasoning modes via `chat_template_kwargs`: + +| `reasoning_effort` | Behavior | Recommended for | +|--------------------|-----------------------------------|------------------------------------| +| `no_think` | No chain-of-thought, direct reply | Simple Q&A, low-latency tasks | +| `low` | Short thinking chain | Moderate reasoning tasks | +| `high` | Full chain-of-thought | Math, coding, complex reasoning | + +Use `no_think` for interactive or latency-sensitive applications to reduce output token count significantly. + +## FAQ + +For common environment, installation, and general parameter issues, please refer to the [Ascend NPU FAQ](/docs/hardware-platforms/ascend-npus/faq). +This section only covers model-specific issues.