5 Commits
Author SHA1 Message Date
wenxuewuhdandronnie_zheng 36afd442c7 [DLLM] vectorized joint/low-confidence decoding and skip redundant attn init (#21094)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-07-30 09:13:14 +03:00
wenxuewuhd c0ed009f5b [NPU] Fix LLaDA2 MoE OOM after the FRACTAL_NZ cast, re-enabling the NZ speedup (#31772) 2026-07-21 16:09:34 +08:00
e856eae921 use sgl_kernel_npu rmsrope accelerate llada2 (#27127)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-07-20 22:46:17 +03:00
wenxuewuhdandronnie_zheng a9270250c3 Dllm radix cache npu (#27144)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-06-25 06:44:03 +03:00
11b76d24dc [NPU] [DLLM]DLLM LLaDA2.x graph mode support with NPU speedup modifications (#18485)
Co-authored-by: Zhang-Xiaoxue <xiaoxuezhang17@outlook.com>
Co-authored-by: dawncc <dawn.cc022@gmail.com>
Co-authored-by: lixinqi7 <li_xinqi7@163.com>
Co-authored-by: rangejay <rangejay1st@163.com>
2026-03-09 22:41:05 +08:00