Commit Graph
9058 Commits
Author SHA1 Message Date
Kevin Flansburg 52f2fe456a fix(disagg): correct DSA/SWA state-page transfer mismatch in PD disaggregation (#27004) 2026-06-03 14:33:41 +08:00
Mick dae86f51f5 [diffusion] chore: polish realtime webui waiting state (#27068) 2026-06-03 14:29:26 +08:00
Khoa Pham b5560ffc36 Fix flashinfer autotune oom glm51 (#24195) 2026-06-02 23:28:57 -07:00
Cheng Wan 202e618898 Revert "Fix hybrid linear attention misrouting plain-RadixAttention linear layers to the full backend (Ring-2.5-1T)" (#27116) 2026-06-02 23:27:57 -07:00
Jimmy Shong 0ef39784ef [Bugfix] Gate DP-attention even-token padding to CP-enabled configs (#26911) 2026-06-03 02:06:52 -04:00
Duyi-Wang ab7c4ab6bb [AMD] Fix correctness for AITER MLA backend with --page-size > 1 (#25556) 2026-06-02 23:01:19 -07:00
Shaun Kotek b8d7351a74 Feat/add w4a16 moe support to nemotron (#25655) 2026-06-02 22:42:26 -07:00
gaopengff aa510bda45 Support specific pass of bias_grouped_topk for xpu (#26349) 2026-06-03 13:13:48 +08:00
f4e7a98fe5 [HiCache] feat: support draft offload for mooncake (#24984)
Co-authored-by: huangtingwei9988 <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: stmatengss <11641725+stmatengss@users.noreply.github.com>
2026-06-02 21:42:04 -07:00
CrazyCoder c3aaafc5f2 [Bugfix] Clean up failed NIXL sender state (#27011) 2026-06-03 12:15:15 +08:00
zhaoshangandShangming Cai 3e681d7fff Add per-rank staggered weight loading for improved TP I/O concurrency (#26937)
Signed-off-by: zhaoshang <zhaoshangsjtu@linux.alibaba.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-06-03 11:25:21 +08:00
zijiexiaandYihao Wang 1ebc7438ac model: support Command A plus (#26106)
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
2026-06-03 11:23:04 +08:00
Mick 71a747cf15 [diffusion] fix: fix lingbot realtime consistency gt pin (#27080) 2026-06-03 10:29:37 +08:00
Chi McIsaac 3d54056391 [diffusion] optimize: optimize Cosmos3 i2v latent prep (#27084) 2026-06-03 10:23:37 +08:00
Mick 3715a07e87 [diffusion] improve: clamp wanvae decode output in place (#27086) 2026-06-03 10:16:10 +08:00
Mick f0e18be0ba [diffusion] improve: use conv2d width padding in wanvae (#27081) 2026-06-03 10:14:50 +08:00
Brayden Zhongandb8zhong 13852d3f31 Support NextN = 2/4 in DSV32 (#24870)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
2026-06-02 19:06:06 -07:00
MingxuZh b678448b8a ci(xeon): merge 2 partitions into 1 job to reduce runner contention (#26904) 2026-06-03 09:46:27 +08:00
gaopengff eda21f6839 Add fused_rope and for xpu (#25773) 2026-06-03 09:41:42 +08:00
Mick 83bc776612 [diffusion] optimize: preserve dtype in wanvae nearest upsample (#27077) 2026-06-03 08:32:28 +08:00
Qiaolin Yu c55548ba11 [perf] Replicate embed_tokens to drop the post-embed all-reduce (#26970) 2026-06-02 16:48:18 -07:00
Alison ShaoandCheng Wan 76c9899da7 Fix hybrid linear attention misrouting plain-RadixAttention linear layers to the full backend (Ring-2.5-1T) (#26623)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
2026-06-02 16:24:49 -07:00
Hubert Lu 72929c7000 [AMD] Enable AITER custom all-gather on ROCm (#25093) 2026-06-02 15:57:37 -07:00
Khoa Pham a711c57a32 [Spec] Fix Gemma 4 MTP with trtllm_mha crash issue (#26966) 2026-06-02 14:37:51 -07:00
Alison Shao 365cc2ade5 jit_kernel tests: bump multiprocess_test timeout 90s -> 240s (cold JIT cache) (#26994) 2026-06-02 17:33:52 -04:00
Baizhou Zhang 68caf49154 Update sgl-deep-gemm to 0.1.2 (#26993) 2026-06-02 13:51:54 -07:00
Hanming Lu b603f08c0c [DP] Fix FlashInfer dispatcher workspace sizing and set_dp_buffer_len (#26643) 2026-06-02 13:25:20 -07:00
Bi Xue 9e717cae46 [sglang] Fix Mamba COW over-releasing SWA locks (cascade-evict assert crash) (#27038) 2026-06-03 01:42:04 +08:00
Muqi Li 6ba31e33e6 feat(api): add require_reasoning field for engine's generate api (#27019) 2026-06-02 17:36:29 +00:00
Cheng WanandClaude Opus 4.7 99da43b900 [refactor] init_forward_metadata 3-method ABC + side-channel removal + ForwardMetadata type rename (#26735)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-02 10:33:33 -07:00
Ilia Yastrebov 6c69756fa8 NIXL: use prep+make API to improve performance (#26406) 2026-06-02 18:36:35 +02:00
Mick 22043b917b [diffusion] CI: ad lingbot case (#27055) 2026-06-02 22:26:51 +08:00
Ke Bao 38d4c9ba88 Improve type annotations in unified radix cache (#26948) 2026-06-02 21:17:50 +08:00
Mick 64a1dec8b6 [diffusion] feat: add realtime webui super resolution controls (#27026) 2026-06-02 20:29:21 +08:00
Mick ce7da7397a [diffusion] optimize: optimize cosmos3 (#27041) 2026-06-02 19:47:42 +08:00
Librua c2eea4d7b3 [Bugfix] Fix orphaned aborted prefill bootstrap requests in PP disaggregation (#27028) 2026-06-02 19:03:46 +08:00
Mick 3394931044 [diffusion] optimize: optimize lingbot performance (#27023) 2026-06-02 18:33:06 +08:00
Mick a777672939 [diffusion] feat: enable parallel decode for cosmos3(#27037) 2026-06-02 18:18:18 +08:00
84e1108312 Optimize ngram decode id computation (#24757)
Co-authored-by: Codex <codex@example.com>
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
2026-06-02 17:37:34 +08:00
pengduriceandgithub-actions[bot] f651b48764 Apply apply_group_norm_silu to LTX-2 latent upsampler (#26045)
Signed-off-by: pengdurice <pengduhit@gmail.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-06-02 17:28:57 +08:00
Xiaoyu ZhangandBBuf 3ea1ba5b15 [GDN] Optimize prefill QKV split dispatch (#26206)
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
2026-06-02 16:48:31 +08:00
Xiaoyu ZhangandBBuf 559581b383 [codex] Centralize Triton utility kernels (#26000)
Co-authored-by: BBuf <xiaoyu.zhang@radixark.net>
2026-06-02 16:47:45 +08:00
Bruce Changlong XuandKe Bao 172bd8e6b9 [scheduler] Zero gen_throughput and flush KV events on pause (#24003)
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
2026-06-02 16:43:04 +08:00
cctryandgemini-code-assist[bot] b55570d38e [PD] Optimistic prefill (#26780)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-06-02 01:16:14 -07:00
Charles Chen 5ae8d286d2 perf(gemma4): single-launch fused router (topk + softmax + scale) (#26502) 2026-06-02 16:00:17 +08:00
3e993f6140 [PD]: Support HiCache prefetching and pd-incremental transfer on decode side (#26227)
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
2026-06-02 15:40:10 +08:00
Jinyan ChenandJinyan Chen 301bcf0872 Add FP4 Indexer for DeepSeek V4 (#26209)
Co-authored-by: Jinyan Chen <jinyanc@nvidia.com>
2026-06-02 00:14:38 -07:00
Mick 1033d835ff [diffusion] optimize: reduce cosmos3 denoise overhead (#26973) 2026-06-02 14:23:02 +08:00
Mick 3b26644bc4 [diffusion] misc: add realtime-webui (#26959) 2026-06-02 14:13:02 +08:00
Mick 2fc548f250 [diffusion] model: support lingot-world (#26954) 2026-06-02 13:52:49 +08:00