amote-i
|
4153410477
|
[NPU] [DOC] remove unsupported models from Ascend NPU models list (#30647)
|
2026-07-09 20:01:00 +08:00 |
|
amote-i
|
69ddbf9ef6
|
[NPU] [DOC] fix model name error on Ascend NPU (#30577)
|
2026-07-09 11:28:29 +08:00 |
|
amote-i
|
042228a195
|
[NPU] [DOC] Remove unsupported options of features on Ascend NPU (#30504)
|
2026-07-08 17:35:06 +08:00 |
|
amote-i
|
cfd3fdc54f
|
[NPU] [DOC] Update features and mainstream models on ascend npu (#30370)
|
2026-07-07 19:03:43 +08:00 |
|
qinsir5522
|
5e9032c527
|
[NPU]Modify LoRA heading in ascend_npu_support_features.mdx to specify Qwen model limitations. (#30358)
|
2026-07-07 16:01:51 +08:00 |
|
ZeyuanChen2000
|
2d9f0b3317
|
[NPU] [DOC] Update arguments detail to NPU support features page (#30328)
|
2026-07-07 14:09:46 +08:00 |
|
loading66
|
998acf7df8
|
[DOCS][NPU]update npu support features (#30324)
|
2026-07-07 11:37:19 +08:00 |
|
 Junlin Wuandronnie_zheng
|
3abdbab9bb
|
✨ [llm][npu][quant] Add W4A8 MXFP quantization support for Qwen3 Dense on Ascend NPU (#23650)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-06 19:23:26 +03:00 |
|
Cao E
|
9df16b5ba9
|
[XPU] Remove redundant xpu graph backend and make xpu graph opt-in by default (#29911)
|
2026-07-03 15:58:56 +08:00 |
|
amote-i
|
2fe7182e75
|
[DOC] [NPU] update supported features on ascend npu (#30011)
|
2026-07-03 15:36:41 +08:00 |
|
ming_wang
|
d8a4f7a7aa
|
add mimo-v2-flash model tutorial (#29932)
|
2026-07-03 11:02:24 +08:00 |
|
Peng Xingchen
|
a6bc7fef90
|
glm5.2 on ascend doc (new version) (#29828)
|
2026-07-03 10:19:44 +08:00 |
|
amote-i
|
70a813493f
|
[NPU] [DOC] add missing DEEP_NORMAL_MODE_USE_INT8_QUANT for w8a8+deepep scenarios (#29937)
|
2026-07-03 10:18:22 +08:00 |
|
Cao E
|
926140d789
|
[XPU] Enable XPU graph support (decode full-graph + prefill tc_piecewise) (#29053)
|
2026-07-02 13:24:35 +08:00 |
|
Zaili Wang
|
cb06c4e6ce
|
[CPU] Fix model failures on Xeon (#29497)
|
2026-07-02 13:20:18 +08:00 |
|
qinsir5522
|
a7390b17f8
|
[NPU]Modify --lora-backend & --moe-runner-backend description. (#29793)
|
2026-07-01 11:15:46 +08:00 |
|
 
|
97fc4dfd73
|
[Doc]Checking and modifying Markdown formatting issues and link validity (#28586)
Signed-off-by: a60124901 <anyuxin4@h-partners.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-06-30 20:27:38 +03:00 |
|
amote-i
|
489017b3d6
|
[NPU] [DOC] Update deterministic inference feature support status to A2, A3 (#29632)
|
2026-06-29 19:17:06 +08:00 |
|
jianzhao-xu
|
2260e612f6
|
[NPU] update best practicce docs from testcase (#29492)
|
2026-06-29 11:30:27 +08:00 |
|
Liangsheng Yin
|
909123ddb8
|
[misc] Use --cuda-graph-max-bs-decode in tests, examples, and docs (#29591)
|
2026-06-28 18:38:28 -07:00 |
|
amote-i
|
a3c5e286f6
|
[NPU] [DOC] Fix and update Ascend NPU docs (#29501)
|
2026-06-27 18:23:11 +08:00 |
|
jianzhao-xu
|
cc294829aa
|
[NPU] fix best practicce docs (#29303)
|
2026-06-26 11:17:50 +08:00 |
|
amote-i
|
10ff3c1dcb
|
[NPU] [DOC] Add environment prerequisites to model tutorials (#29293)
|
2026-06-26 09:51:06 +08:00 |
|
amote-i
|
73d976e375
|
[NPU] [DOC] Fix TOC of Ascend NPU Docs (#29129)
|
2026-06-24 16:47:10 +08:00 |
|
amote-i
|
84338df6f0
|
[NPU] [DOC] Update contribution guide of Ascend NPU (#28909)
|
2026-06-23 10:00:12 +08:00 |
|
jianzhao-xu
|
4e1d25117b
|
[NPU] update best practice docs from testcase (#28621)
|
2026-06-22 16:58:48 +08:00 |
|
amote-i
|
93553a67a3
|
[NPU] [DOC] Create deployment tutorials for mainstream models on Ascend NPU (#27893)
|
2026-06-22 16:21:13 +08:00 |
|
amote-i
|
5deca2d39f
|
[DOC] [NPU] Update features on Ascend NPU (#28643)
|
2026-06-22 09:50:51 +08:00 |
|
cctry
|
6d4ca9bc54
|
Cap SWA pool sizing with chunk cache (#28755)
|
2026-06-21 01:06:59 -07:00 |
|
zijiexia
|
74e2e48c82
|
Introduce CpuDeviceMixin and CpuSRTPlatform (#26385)
|
2026-06-17 17:41:17 -07:00 |
|
chenxu214
|
71b090a8e7
|
[Ascend]GLM 5.2 deployment (#28433)
|
2026-06-17 09:55:50 +08:00 |
|
amote-i
|
74155d32bd
|
[NPU] [DOC] Update Ascend NPU docs: HDK 25.5.2, Triton 3.2.1.dev20260530 (#28311)
|
2026-06-17 09:38:31 +08:00 |
|
 Junlin Wuandronnie_zheng
|
2a8ea70059
|
✨ [llm][npu][quant] Add W8A8 MXFP8 quantization support for Qwen3 Dense on Ascend NPU (#22352)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-06-16 09:45:18 +03:00 |
|
qinsir5522
|
81166f382d
|
Fix inaccuracies and add NPU constraints in ascend_npu_profiling.mdx. (#28283)
|
2026-06-15 21:12:18 +08:00 |
|
jianzhao-xu
|
f8d1d397b6
|
[NPU] fix ascend_docs (#28279)
|
2026-06-15 21:11:55 +08:00 |
|
McZyWu
|
8bdb007e58
|
[NPU] Docs op performance optimize (#28277)
|
2026-06-15 20:28:18 +08:00 |
|
loading66
|
f768344b1a
|
[DOCS][NPU]Supplementary Notes (#28295)
|
2026-06-15 20:06:05 +08:00 |
|
longxin9715
|
7bd1a9d163
|
Update documentation for Ascend NPU Guide (#28284)
|
2026-06-15 20:02:50 +08:00 |
|
amote-i
|
edd5eff519
|
[NPU] [DOC] fix issues in ascend_npu_support_new_models (#28296)
|
2026-06-15 19:58:33 +08:00 |
|
Chang Min Bark
|
93b402580c
|
feat: add decode clear steps env var (#28160)
|
2026-06-13 17:15:42 -07:00 |
|
jianzhao-xu
|
29128f31fd
|
[NPU] Best Practice Docs Splitting (#27663)
|
2026-06-13 22:16:20 +08:00 |
|
amote-i
|
c26669e37b
|
[NPU] [DOC] Update server arguments to NPU support features page (#28083)
|
2026-06-12 17:16:52 -07:00 |
|
ming_wang
|
f308abc052
|
Revise the mimo-v2-flash best practice (#28016)
|
2026-06-12 17:12:33 +08:00 |
|
Liangsheng Yin
|
c0480a88be
|
[Spec] Retire Spec V1 (#27964)
|
2026-06-11 16:15:15 -07:00 |
|
ming_wang
|
ce0ff154a5
|
add mimo best practice (#27665)
|
2026-06-11 10:15:31 +08:00 |
|
amote-i
|
fc1fee528c
|
[NPU] [DOC] replace <code> with backticks and remove obsolete params (#27677)
|
2026-06-11 09:33:31 +08:00 |
|
 
|
bcd9c5a903
|
update pytorch-xpu to 2.12 (#27133)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: MingxuZh <109504044+MingxuZh@users.noreply.github.com>
|
2026-06-10 09:25:30 +08:00 |
|
 
|
52a5c01eba
|
[plugin] enable OOT platforms to provide custom quant configs (#25347)
Signed-off-by: Devashish Lal <devcode@fb.com>
Co-authored-by: Devashish Lal <devcode@fb.com>
Co-authored-by: Devashish Lal <laldevashish@gmail.com>
|
2026-06-07 12:48:55 +08:00 |
|
Артем Савкин
|
9f28512737
|
[NPU] Update documentation for software version upgrades (#26731)
|
2026-06-06 15:06:13 +03:00 |
|
McZyWu
|
d8487bad06
|
Update best practice for qwen3-next-80b-a3b-instruct (#27353)
|
2026-06-05 17:02:38 +08:00 |
|