Commit Graph
13 Commits
Author SHA1 Message Date
chenxu214 532470bcca [NPU] add new fusion operator DispatchFFNCombine (#20245) 2026-03-18 15:22:04 +08:00
chenxu214 b912d7ae19 [OPT]Skip the first delayer to maximize the BS of the decoding. (#19836) 2026-03-06 08:53:19 +08:00
chenxu214 88cfa6c11d [NPU]Releasing redundant memory of w13_weight and nz when the ascend_fuseep feature is enabled (#19813) 2026-03-04 19:26:29 +08:00
chenxu214andsglang-npu-bot 5f07ff9271 Added the prefill delayer policy: The prefill deplay range is expanded. (#17456)
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
2026-02-28 08:56:49 +08:00
chenxu214 fd5a45d5cf Update ascend_npu_support.rst (#18868) 2026-02-16 01:41:38 +08:00
chenxu214 f2d72866e9 Create ascend_npu_qwen3_5_examples.md (#18864) 2026-02-16 01:15:20 +08:00
chenxu214 4e162d4b1b change npu.dockerfile (#18835) 2026-02-15 20:43:15 +08:00
chenxu214 1edc69be08 [Ascend]Support qwen3.5 (#18544)
This PR affects only the NPU. If any issues arise, please contact iforgetmyname.
2026-02-12 15:22:47 +08:00
chenxu214 444b9521e4 [Bugfix]Repeated add modelslim quant_config and bugfix with "enable-piecewise-cuda-graph" on NPU (#17511) 2026-01-26 09:51:07 +08:00
5d299c25c0 [NPU] bugfix with Kimi-k2 and bge-reranker-v2 model (#17478)
Co-authored-by: amote-i <49533125+amote-i@users.noreply.github.com>
Co-authored-by: cy <chenyang08056032@163.com>
2026-01-22 22:02:05 +08:00
chenxu214 a4dc432587 Change naming for graph mode on multiplatform (#17469) 2026-01-22 19:25:06 +08:00
chenxu214 53dca74f47 Bugfix: EagleDraftWorker has not attribute "eagle_use_aux_hidden_state" (#16480) 2026-01-12 20:14:35 +08:00
chenxu214 7dd679cbb9 [NPU][Bugfix] Fix qwen3 error when enable-dp-lm-head (#16115) 2026-01-08 15:15:43 +08:00