gjsheu
|
02aeed5387
|
[NPU] DFlash Speculative Decoding Support NPU (#23122)
|
2026-05-30 15:13:59 +08:00 |
|
gjsheu
|
d9d719b270
|
[npu] [bugfix] Add contiguous operation during quantized weight loading. (#26309)
|
2026-05-27 19:55:58 +08:00 |
|
  ![gemini-code-assist[bot]](/assets/img/avatar_default.png)  
|
36c93fc6fb
|
[NPU] [Diffusion] Use fused operator to improve Wan model E2E performance. (#24028)
Co-authored-by: gengjinsong <gengjinsong@huawei.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: gengjinsong <904939979@qq.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-05-11 07:17:31 +03:00 |
|
 gjsheuandgengjinsong
|
e708ea6d94
|
[diffusion] fix: restore cache-dit support for LTX2 (#23235)
Co-authored-by: gengjinsong <gengjinsong@huawei.com>
|
2026-04-25 18:10:43 +08:00 |
|
 gjsheuandgengjinsong
|
d9e96153de
|
[NPU] Support Hybrid KV Cache for Ascend backend (#18032)
Co-authored-by: gengjinsong <gengjinsong@huawei.com>
|
2026-03-26 11:27:36 +08:00 |
|