9 Commits
Author SHA1 Message Date
AndyLi429andAndyLi429 d903351a66 [NPU][bugfix] update low latency quantization input and update MXFP8 tests (#38831)
Co-authored-by: AndyLi429 <AndyLi429@noreply.gitcode.com>
2026-09-20 09:55:30 +08:00
e970453b43 [NPU]Refactor weight processing and add NPUSwigluLimit activation (#38420)
Co-authored-by: AndyLi429 <AndyLi429@noreply.gitcode.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
2026-09-17 16:40:32 +08:00
62a4a6ea0e [NPU] Add NPU arch35 support and enhance DSV4 processing in DeepSeek-V4 (#37373)
Co-authored-by: AndyLi429 <AndyLi429@noreply.gitcode.com>
Co-authored-by: Kailong Lu <kelonlu@163.com>
Co-authored-by: cx <chengxin65@huawei.com>
Co-authored-by: ranjiewen <ranjiewen@huawei.com>
Co-authored-by: HEX1A0A <1a0ahex@gmail.com>
Co-authored-by: vstone-w <374330057@qq.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
Co-authored-by: ClownBin <chaobin1993@126.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
2026-09-07 21:08:06 +08:00
AndyLi429 62d7929b3a [NPU][Bugfix] Accept in_capture in Ascend replay metadata (#29598) 2026-06-29 14:37:45 +08:00
AndyLi429 72cac88022 [NPU] perf: precompute mamba conv-state track indices once per batch (#29105) 2026-06-25 14:04:56 +08:00
AndyLi429 cd6efcb947 [NPU][Bugfix] fix MTP accuracy regression on Qwen3 hybrid models (#27202) 2026-06-09 15:47:23 +08:00
AndyLi429 fe4b29d391 [Bugfix] Fix Ascend NPU CP attention for batch size > 1 (#26705) 2026-05-30 15:07:39 +08:00
AndyLi429 3640116397 [NPU]Bugfix:Set default values for npu_wrapper_preprocess parameters (#25130) 2026-05-14 14:21:11 +08:00
AndyLi429 4c1eefca4f [NPU] ascend backend support qwen3 moe attention cp (#21685) 2026-04-29 19:25:17 +08:00