Commit Graph
17516 Commits
Author SHA1 Message Date
Yuwei AnandClaude Fable 5 07d84ebd6d [2/N][Mixed] Mixed chunk prefill with spec enabled (#36933)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-31 10:48:07 -07:00
9cf157c252 [Radix Cache] Add Rust TreeCore backend with shared parity tests (#32710)
Co-authored-by: alphabetc1 <2508695655@qq.com>
Co-authored-by: ispobock <ispobaoke@gmail.com>
2026-09-01 00:26:20 +08:00
Xiaoyu Zhang 52e1c24744 [Diffusion] Fuse FLUX.2 token concatenation and NVFP4 quantization (#37141) 2026-08-31 21:33:02 +08:00
Xiaoyu ZhangandSong Bian d60d658f5f [Kernel] Add GB300 Triton MoE configs for GLM-4.5 FP8 (#37159)
Co-authored-by: Song Bian <biansonghz@gmail.com>
2026-08-31 21:31:04 +08:00
Xiaoyu Zhang 771e613d96 [Diffusion] Fuse Qwen-Image final adaptive LayerNorm (#37144) 2026-08-31 18:22:20 +08:00
Xiaoyu ZhangandCursor 1ed9bfac2c [Diffusion] Fuse Qwen-Image FP8 norm and activation quantization (#37156)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-31 18:19:42 +08:00
pllimax a874b83c2c test: re-enable DSV4-Flash W8A8 8p nightly perf cases (#37214) 2026-08-31 17:20:46 +08:00
pllimax 63b2adbeac [NPU] Fix evalscope accuracy parsing and add glm5_1 aime26 request timeout (#36459) 2026-08-31 17:19:03 +08:00
YAMY a53718e88c fix(ci): update attention backend test fixtures (#37219) 2026-08-31 02:15:17 -07:00
YC Yen-Ching Tseng 0674be736c [AMD] build gfx1250 release image from main (#37225) 2026-08-31 17:04:04 +08:00
6580d5cd9a weight cache: key daemon paths by GPU UUID (#36101)
Co-authored-by: siyu <liusy58@linux.alibaba.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
2026-08-31 01:44:45 -07:00
+1 3865efc9f7 [AMD] support gfx1250 on ROCM 10 (#36871)
Co-authored-by: HAI <hixiao@gmail.com>
Co-authored-by: Kao <akao@amd.com>
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
Co-authored-by: sogalin_codegen <39478626+sogalin@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
2026-08-31 01:19:11 -07:00
Yuan Luoandluoyuan.luo 712a720c8a [KDA] Fused-accept state advance for FlashInfer KDA MTP verify (#33722)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-08-31 16:11:49 +08:00
Cheng Wan f61bb7b40a [unified-memory] Drop the vacated 'dense' qualifier and the restating comments (#37170) 2026-08-31 00:54:13 -07:00
8bb776dc48 feat(unified-memory): read unified pool from attention backends fa3/flashinfer/trtllm_mha/flashmla (#34613)
Co-authored-by: Caihua Li <caihua.li@bytedance.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
2026-08-30 23:58:24 -07:00
29578d5578 refactor(unified-memory): translate the KV write location once, at ForwardBatch construction (#35245)
Co-authored-by: Caihua Li <caihua.li@bytedance.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
2026-08-30 23:52:14 -07:00
Ziang Li 4f997a432a test: re-enable FlashInfer per-token NVFP4 coverage (#36985) 2026-08-30 23:44:09 -07:00
Lianmin ZhengandMing Yang 3ed3326631 Decouple speculative draft capacity from runtime state (#36897)
Co-authored-by: Ming Yang <minos.future@gmail.com>
2026-08-30 23:31:25 -07:00
Mick 2cb3f32b03 [diffusion] chore: preserve exact component identity during loading (#36875) 2026-08-31 14:17:57 +08:00
Alex NailsandClaude Fable 5 5e679b0cad [CI] Speed up lint: cache pre-commit envs + mint, drop redundant work (#37203)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-30 23:16:31 -07:00
YAMYandgithub-actions[bot] b77cac06a9 [PP] Support prefill CUDA graph proxy tensors (#36248)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-08-30 23:11:45 -07:00
Zhanghengand晟海 5d92e60783 [Unified Cache Linker][3/N]: Add backend-independent linker core (#37151)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
2026-08-31 14:07:39 +08:00
ashwini rathi 10b67aa7a1 xpu: move prefill-only model tests to the nightly-xpu-1-gpu grid (#36814) 2026-08-31 14:03:48 +08:00
Mick 62c470697e [diffusion] chore: enforce component attention backend application (#36907) 2026-08-31 14:00:02 +08:00
Mick 28690f5aa5 [diffusion] chore: detect quantized transformer replacements (#36916) 2026-08-31 13:55:50 +08:00
Liangsheng Yin 3a6ed55999 [Fix] Shut hicache test servers down gracefully before SIGKILL (#37194) 2026-08-30 22:45:40 -07:00
Zaili WangandAlex Nails afbca97f74 add suffix for xpu kernel upload space (#36422)
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
2026-08-30 22:41:22 -07:00
Liangsheng Yin 5d12ad4fd7 [mem_cache] Move mamba state and retraction_backup into ReqKvInfo (#37164) 2026-08-30 22:03:01 -07:00
Aurick Qiao 9a9e167179 [Bugfix] Fix full prefill CUDA graph padding and EAGLE capture (#35588) 2026-08-30 21:30:24 -07:00
YAMY 2ea6d17eab fix(staging): make empty staging rings reusable (#37166) 2026-08-30 20:59:57 -07:00
YC Yen-Ching TsengandAlex Nails 5972211977 [AMD] Fix the QuickReduce bf16 cast failing to build for CDNA (#37132)
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
2026-08-30 20:57:29 -07:00
Shuwen Wang 62f86ce470 [CI] Fix unreachable FakeReq field initialization (#37182) 2026-08-30 20:51:40 -07:00
Mick 881cbfe54c [diffusion] feat: add exact component precision overrides (#36991) 2026-08-31 11:12:07 +08:00
Pavan Sivaram GirijalaandYanbingJiang 4dc7dc8518 [Fix] Transformers-fallback (GPT-NeoX) + KV pool config (DeepSeek-VL2) (#35244)
Co-authored-by: YanbingJiang <yanbing.jiang@intel.com>
2026-08-31 11:05:50 +08:00
Piotr MazurekandKhoa Pham 046454404a [Spec] Add LFM2 and LFM2-MoE DSpark speculative decoding support (#31041)
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
2026-08-30 20:03:59 -07:00
jambow0320andShangming Cai 7700602278 [PD] Align defensive protocol behavior across Mooncake, NIXL, and Mori (#35281)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-08-31 10:57:56 +08:00
Polisetty V R K Jyothendra Varma df75ec5f77 [Intel GPU] Add rust support to XPU docker images (#35877) 2026-08-31 10:52:57 +08:00
Cui LilyandMa Mingfei 3139ceaeec [XPU] Use SYCL kernels for topk_transform on XPU (#33318)
Signed-off-by: Cui, Lily <lily.cui@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-08-31 10:33:19 +08:00
YanbingJiang a9d5ca723a [CPU] Fix weight missing issue in fused_input_proj_cpu for GPTQ INT4 for Qwen 3.5 (#35805) 2026-08-31 10:14:41 +08:00
Mohammad Miadh Angkad 4f761e8649 [Deps] Bump FlashInfer to 0.6.18 (#36954) 2026-08-30 19:02:39 -07:00
Ali Ihsan Nergiz 0da6a66856 [MLX] Fix startup crash when reporting preloaded weights (#37035) 2026-08-30 17:36:13 -07:00
Xiaoyu Zhang bb5e619860 [diffusion] perf: absorb Qwen-Image output projection biases (#37116) 2026-08-31 08:25:37 +08:00
Liangsheng Yin 4bb8de34cc [mem_cache] Share one ReqKvInfo between a streaming session slot and its request (#37108) 2026-08-30 16:36:48 -07:00
4bea51d885 feat(unified-memory): dense KV views for uniform-row MHA/SWA models (#34602)
Co-authored-by: Caihua Li <caihua.li@bytedance.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
2026-08-30 15:10:12 -07:00
Liangsheng Yin 007ef5e23a [mem_cache] Move req_pool_idx into ReqKvInfo (#37094) 2026-08-30 14:46:21 -07:00
Yuzhen ZhouandJiajun Li 8a87079dbb Fix stale GLM MoE routing after runtime weight updates (#35883)
Co-authored-by: Jiajun Li <jiajun.li@radixark.ai>
2026-08-30 14:13:34 -07:00
Xiaoyu Zhang 5ab97c4f44 [Diffusion] Cache Qwen-Image modulation across serial CFG branches (#37090) 2026-08-31 01:40:29 +08:00
Xiaoyu Zhang 8c28cdd116 [Diffusion][Kernel] Fuse Wan2.2 NVFP4 bias + GELU on Blackwell (#37075) 2026-08-31 01:37:35 +08:00
YAMY 9a03bc2dc3 [CI] Fix stale GPU capability test patches (#37148) 2026-08-30 10:12:32 -07:00
Zhanghengand晟海 6a9366f036 [Unified Cache Linker][2/N]: Add device pool assembly for external linkers (#37098)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
2026-08-30 23:54:00 +08:00