14 Commits
Author SHA1 Message Date
3a770da756 [Unified Tree] Support Branching-Point Caching for the SWA Component (#34565)
Co-authored-by: alphabetc1 <2508695655@qq.com>
Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com>
2026-09-05 12:29:16 +08:00
Jincong Chen 3c533acec6 [Hicache][2/2]Support Mamba branching in Unified Radix Cache with HiCache (#33639) 2026-08-10 17:16:42 +08:00
7c4b22fae5 [Hicache][1/2]Support Mamba branching in Unified Radix Cache with HiCache (#31181)
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-07-25 19:44:40 +08:00
Jincong Chen e27d4fb70f [Perf][Qwen3.5] Add case 512 to topkGatingSoftmaxKernelLauncher, (#25775) 2026-05-25 16:08:21 +08:00
Jincong Chen 2bac219d0c [Perf] Precompute gemma_weight to avoid redundant add on every forward (#22673) 2026-04-17 23:37:41 +08:00
Jincong Chen 6760c790bd [bugfix] avoid attention padding tokens computation in pcg (#17706) 2026-04-14 16:08:23 +08:00
Jincong Chen 0668a7f51a [Perf] Remove two operations in gdn_backend extend verify path (#22444) 2026-04-10 17:53:57 +08:00
Jincong Chen 03e4f2858d [Perf]Remove H2D for Qwen3.5 SpecV2 (#20864) 2026-03-31 11:54:58 +08:00
Jincong Chenandgemini-code-assist[bot] c77d7c629e [Bugfix] Fix MTP prefill cuda graph logging (#20279)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-17 16:36:52 -07:00
Jincong Chen 165aff38e1 Add CI permission for Chen-0210 (#18494) 2026-02-12 09:33:35 +08:00
Jincong Chen a72f4f839c Tiny fix for fp8 moe backend flashinfer_trtllm naming (#18243) 2026-02-04 19:58:04 +08:00
Jincong Chen 72f790bf6f [BUGFIX] Skip Mamba Cache Slot 0 to Avoid Using Dummy Cache (#17404) 2026-01-22 23:08:38 +08:00
Jincong Chen 350fbbf4dc fix ds3.2 nsa backend prefill TBO (#14901) 2025-12-21 13:16:46 -08:00
Chen1022 3c7886ec4c Fix attention backend logic for Qwen3-Next on SM100 (#14560) 2025-12-06 22:03:34 -08:00