cen121212
|
f8f03910f2
|
【NPU】Support EAGLE when PP enabled in prefill nodes (#32207)
|
2026-09-08 15:45:36 +08:00 |
|
cen121212
|
b78d3999b5
|
【NPU】fix decode MTP + eagle shape error (#32791)
|
2026-07-30 21:34:16 +08:00 |
|
cen121212
|
0ffed946f2
|
[NPU] Add extra topk_weights input in deepep ll dispatch (#29480)
|
2026-07-09 09:23:29 +08:00 |
|
cen121212
|
9305d10099
|
[NPU] adapt_fused_rope_qk_mqa_optimize (#28872)
|
2026-06-25 08:55:52 +08:00 |
|
 
|
b421e60eed
|
【NPU】【bugfix】fix server error when mtp unquant (#26389)
Co-authored-by: cen121212 <luochen23@huawei.com>
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
|
2026-05-30 15:01:19 +03:00 |
|
cen121212
|
461bc8af49
|
[NPU][Doc] Update GLM-5 docs, enabling deepep by default (#23708)
|
2026-05-08 11:12:35 +08:00 |
|
cen121212
|
b1e1fe8eee
|
【NPU】【bugfix】accuracy fix when enable both nsa cp and prefixcache (#23268)
|
2026-04-28 09:08:28 +08:00 |
|
cen121212
|
ba6d54d0f0
|
[NPU] GLM-5 optimize with fused kernels (#18617)
|
2026-03-30 22:48:15 +08:00 |
|
cen121212
|
fc543df289
|
[NPU] qwen3_vl encoder support graph
|
2026-03-09 10:13:35 +08:00 |
|
cen121212
|
0c2993eed0
|
Optimize Qwen3-VL video memory usage (#16366)
|
2026-01-22 09:10:08 +08:00 |
|
cen121212
|
25b48564c3
|
[NPU][Bugfix] fix Qwen3-VL-30B-A3B-Instruct accuracy loss (#15597)
|
2025-12-31 15:57:38 +08:00 |
|