  
|
33ecf4bcd8
|
Add pr tests (#31952)
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
Co-authored-by: Cherry_ming <136634645@qq.com>
|
2026-08-01 15:03:21 +08:00 |
|
amote-i
|
0d6bef6b6d
|
[NPU] [DOC] renew triton-ascend installation guide location (#32986)
|
2026-07-31 14:46:04 +08:00 |
|
amote-i
|
c039e1a7ee
|
[NPU][DOC] Restructure ascend-npus docs into layered navigation (#32857)
|
2026-07-31 09:49:46 +08:00 |
|
amote-i
|
20d5b91e5b
|
[NPU] [DOC] update feature name to follow the code changement (#32749)
|
2026-07-30 09:54:08 +08:00 |
|
Xiaoyu Zhang
|
c32c4ef79c
|
[Kernel] Move sgl-kernel under sglang.kernels.aot (#32648)
|
2026-07-29 17:25:00 +08:00 |
|
 
|
f05c92fb6d
|
✨ [llm][npu][quant] Add W8A8 MXFP8 quantization for Qwen3 MoE on Ascend NPU (#30768)
Co-authored-by: Артем Савкин <58187114+OrangeRedeng@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-29 10:39:36 +03:00 |
|
amote-i
|
cb12a1547b
|
[NPU] [DOC] update supported features on ascend npu (#32647)
|
2026-07-29 09:05:21 +08:00 |
|
 sglang-botandsglang-bot
|
1ef2b85ef7
|
chore: bump docs install version to 0.5.16 (#32347)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-07-24 17:19:39 -07:00 |
|
 Polisetty V R K Jyothendra VarmaandMa Mingfei
|
319055c191
|
[Intel GPU] Add XPU Platform support (#31949)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-07-24 12:49:29 +08:00 |
|
Jinyan Yi
|
b98a577fbe
|
Doc/update ascend quickstart (#32205)
|
2026-07-23 20:25:20 +08:00 |
|
axx-ty911
|
7dcf3f7cbf
|
[Doc] Rename A5 product name (#31998)
|
2026-07-22 15:09:05 +08:00 |
|
Rahul Vijayaraghavan
|
fa0ced195e
|
[XPU] Enable breakable prefill CUDA graph on XPU (#30273)
|
2026-07-21 09:09:40 +08:00 |
|
amote-i
|
97e0647bdc
|
[NPU] [DOC] Fix issues about npu docs found by aidd (#31302)
|
2026-07-20 16:31:27 +08:00 |
|
 Brayden ZhongandBrayden Zhong
|
7fc3fb9657
|
Remove deprecated Mamba flags from doc, wrong FP8 GEMM docstrings and change Nemotron image to 0.5.15 (#31094)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-07-17 14:34:29 -07:00 |
|
Junlin Wu
|
bbd2a3fe4a
|
✨ [llm][npu][quant] Add W4A4 MXFP4 quantization support for Qwen3 Dense on Ascend NPU (#23795)
|
2026-07-17 09:06:30 +03:00 |
|
amote-i
|
6275114548
|
[NPU] [DOC] Update model names supported on Ascend NPU (#31316)
|
2026-07-16 14:45:59 +08:00 |
|
 Артем Савкинandronnie_zheng
|
8ed82afcc8
|
[MoE Refactor] [NPU] Refactor Ascend MoE implementation to reduce code duplication and align with community design (#25663)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-15 14:59:42 +03:00 |
|
 
|
bdc9848c25
|
[Doc]Standardize the names of PyTorch NPU-related software throughout the documentation by replacing them all with TorchNPU. (#29886)
Signed-off-by: a60124901 <anyuxin4@h-partners.com>
Signed-off-by: axx-ty911 <anyuxin4@h-partners.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-14 21:56:08 +03:00 |
|
zijiexia
|
50ed4c011f
|
Remove legacy Sphinx docs/ and finish the Mintlify cutover (#28964)
|
2026-07-13 15:06:08 -07:00 |
|
Liangsheng Yin
|
e2728ac504
|
[Spec] Remove dead padded_static_len and stale SGLANG_ENABLE_SPEC_V2 references (#30998)
|
2026-07-13 15:30:31 -05:00 |
|
 sglang-botandsglang-bot
|
805385414e
|
chore: bump docs install version to 0.5.15 (#31058)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-07-13 18:58:48 +00:00 |
|
 
|
11a82af5f8
|
[Platform] Route pin memory availability through current_platform (#28113)
Co-authored-by: N3u0ns <N3u0ns@users.noreply.github.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-07-13 11:37:59 -07:00 |
|
amote-i
|
be9791071a
|
[NPU] [DOC] Fix Ascend NPU docs issues found by AIDD (#31036)
|
2026-07-13 23:45:35 +08:00 |
|
amote-i
|
f391c71758
|
[NPU] [DOC] --pp-size can not be used witgh --tp-size (#31039)
|
2026-07-13 22:35:20 +08:00 |
|
xutizhou
|
eb31b5310c
|
Support Waterfill with MegaMoE backend (#27350)
|
2026-07-13 03:56:46 -07:00 |
|
amote-i
|
2225817424
|
[NPU] [DOC] Optimize and fix docs issues on Ascend NPU (#30767)
|
2026-07-13 15:35:02 +08:00 |
|
loading66
|
79f096d43b
|
[DOCS][NPU]update npu support features and models (#30843)
|
2026-07-11 15:21:29 +08:00 |
|
amote-i
|
4153410477
|
[NPU] [DOC] remove unsupported models from Ascend NPU models list (#30647)
|
2026-07-09 20:01:00 +08:00 |
|
amote-i
|
69ddbf9ef6
|
[NPU] [DOC] fix model name error on Ascend NPU (#30577)
|
2026-07-09 11:28:29 +08:00 |
|
amote-i
|
042228a195
|
[NPU] [DOC] Remove unsupported options of features on Ascend NPU (#30504)
|
2026-07-08 17:35:06 +08:00 |
|
amote-i
|
cfd3fdc54f
|
[NPU] [DOC] Update features and mainstream models on ascend npu (#30370)
|
2026-07-07 19:03:43 +08:00 |
|
qinsir5522
|
5e9032c527
|
[NPU]Modify LoRA heading in ascend_npu_support_features.mdx to specify Qwen model limitations. (#30358)
|
2026-07-07 16:01:51 +08:00 |
|
ZeyuanChen2000
|
2d9f0b3317
|
[NPU] [DOC] Update arguments detail to NPU support features page (#30328)
|
2026-07-07 14:09:46 +08:00 |
|
loading66
|
998acf7df8
|
[DOCS][NPU]update npu support features (#30324)
|
2026-07-07 11:37:19 +08:00 |
|
 Junlin Wuandronnie_zheng
|
3abdbab9bb
|
✨ [llm][npu][quant] Add W4A8 MXFP quantization support for Qwen3 Dense on Ascend NPU (#23650)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-06 19:23:26 +03:00 |
|
Cao E
|
9df16b5ba9
|
[XPU] Remove redundant xpu graph backend and make xpu graph opt-in by default (#29911)
|
2026-07-03 15:58:56 +08:00 |
|
amote-i
|
2fe7182e75
|
[DOC] [NPU] update supported features on ascend npu (#30011)
|
2026-07-03 15:36:41 +08:00 |
|
ming_wang
|
d8a4f7a7aa
|
add mimo-v2-flash model tutorial (#29932)
|
2026-07-03 11:02:24 +08:00 |
|
Peng Xingchen
|
a6bc7fef90
|
glm5.2 on ascend doc (new version) (#29828)
|
2026-07-03 10:19:44 +08:00 |
|
amote-i
|
70a813493f
|
[NPU] [DOC] add missing DEEP_NORMAL_MODE_USE_INT8_QUANT for w8a8+deepep scenarios (#29937)
|
2026-07-03 10:18:22 +08:00 |
|
Cao E
|
926140d789
|
[XPU] Enable XPU graph support (decode full-graph + prefill tc_piecewise) (#29053)
|
2026-07-02 13:24:35 +08:00 |
|
Zaili Wang
|
cb06c4e6ce
|
[CPU] Fix model failures on Xeon (#29497)
|
2026-07-02 13:20:18 +08:00 |
|
qinsir5522
|
a7390b17f8
|
[NPU]Modify --lora-backend & --moe-runner-backend description. (#29793)
|
2026-07-01 11:15:46 +08:00 |
|
 
|
97fc4dfd73
|
[Doc]Checking and modifying Markdown formatting issues and link validity (#28586)
Signed-off-by: a60124901 <anyuxin4@h-partners.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-06-30 20:27:38 +03:00 |
|
amote-i
|
489017b3d6
|
[NPU] [DOC] Update deterministic inference feature support status to A2, A3 (#29632)
|
2026-06-29 19:17:06 +08:00 |
|
jianzhao-xu
|
2260e612f6
|
[NPU] update best practicce docs from testcase (#29492)
|
2026-06-29 11:30:27 +08:00 |
|
Liangsheng Yin
|
909123ddb8
|
[misc] Use --cuda-graph-max-bs-decode in tests, examples, and docs (#29591)
|
2026-06-28 18:38:28 -07:00 |
|
amote-i
|
a3c5e286f6
|
[NPU] [DOC] Fix and update Ascend NPU docs (#29501)
|
2026-06-27 18:23:11 +08:00 |
|
jianzhao-xu
|
cc294829aa
|
[NPU] fix best practicce docs (#29303)
|
2026-06-26 11:17:50 +08:00 |
|
amote-i
|
10ff3c1dcb
|
[NPU] [DOC] Add environment prerequisites to model tutorials (#29293)
|
2026-06-26 09:51:06 +08:00 |
|