 Liangsheng YinandAlison Shao
|
ac99794e64
|
Reland spec v2 tree drafting (eagle topk>1) with page_size==1 (#26866) (#26997)
Co-authored-by: Alison Shao <54658187+alisonshao@users.noreply.github.com>
|
2026-06-03 15:40:05 -04:00 |
|
Mohammad Miadh Angkad
|
7f706f4cfb
|
[Deps] Bump FI to 0.6.12 and cutedsl to 4.5.2 (#26854)
|
2026-06-03 12:09:18 -07:00 |
|
Liangsheng Yin
|
578f232e5e
|
Fix trace_modules gate disabling default trace contexts (#27173)
|
2026-06-03 14:03:42 -04:00 |
|
 Cheng WanandClaude Opus 4.8
|
45604a0f4a
|
[refactor] Unify CUDA graph runner input buffers behind CudaGraphBufferRegistry (#26742)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-03 10:54:10 -07:00 |
|
Mick
|
b0f78bef97
|
[diffusion] improve: improve realtime webui playback pacing (#27148)
|
2026-06-04 00:33:56 +08:00 |
|
 Lijuan TangandXiaodong Ye
|
9d0e6a2df4
|
fix(mlx): set canary_manager and materialize overlap-loop inputs on Apple Silicon (#26882)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
Signed-off-by: LijuanTang94 <tang.lij@northeastern.edu>
Co-authored-by: Xiaodong Ye <yeahdongcn@gmail.com>
|
2026-06-04 00:03:46 +08:00 |
|
Xinyuan Tong
|
fa5c8a3101
|
[model] support encoder-free unified Text/Vision/Audio model (#27167)
|
2026-06-03 23:58:06 +08:00 |
|
shuwenn
|
f65aae8493
|
[HiCache] fix: truncate prefetch key on degraded allocation (#25991)
|
2026-06-03 22:06:18 +08:00 |
|
Mick
|
33f943fbf5
|
[diffusion] optimize: batch usp replicated kv prefix all-to-all (#27143)
|
2026-06-03 21:22:39 +08:00 |
|
Ye (Charlotte) Qi
|
03c77dc33d
|
[PD] Deduplicate PD logprob normalization (#27085)
|
2026-06-03 19:08:15 +08:00 |
|
Ke Bao
|
44d4a25a07
|
Type hicache transfer hook kwargs in unified cache (#27071)
|
2026-06-03 18:57:25 +08:00 |
|
 
|
e67810bea7
|
[SGLang Tracing] Add pd disaggregation mooncake backend tracing (#23755)
Co-authored-by: Mu Huai <tianbowen.tbw@antgroup.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-06-03 16:43:29 +08:00 |
|
Cheng Wan
|
73b53e7a87
|
Revert "Support NextN = 2/4 in DSV32" (#27138)
|
2026-06-03 01:29:52 -07:00 |
|
 Vladislav NosivskoyandZhangheng
|
63dc20ae6c
|
[UnifiedTree] Add CP sync (#25395)
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-06-03 16:10:27 +08:00 |
|
    
|
93173b27e8
|
integrate flash_mla_sparse_fwd (#25418)
Co-authored-by: Yuan Luo <yuan.luo@hotmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: laixinn <q865809639@gmail.com>
Co-authored-by: MeowGrange <276466210+MeowGrange@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-03 01:09:25 -07:00 |
|
ybyang
|
f790674ad8
|
fix(moe): avoid unpacking None from masked deep_gemm without overlap when sbo enabled (#26839)
|
2026-06-03 00:44:28 -07:00 |
|
Mick
|
9450696aa5
|
[diffusion] CI: add cosmos3 nano t2v gpu test (#26963)
|
2026-06-03 15:42:51 +08:00 |
|
Bingxu Chen
|
8e77af1afc
|
[AMD] fix(triton-mla): cap max_kv_splits at 256 on gfx942 (Kimi-K2.6 hang) (#24762)
|
2026-06-03 00:13:18 -07:00 |
|
 inkcherryandAlex Nails
|
e5b8e3a66a
|
Optimize streaming detokenizer updates (#24659)
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-06-02 23:44:21 -07:00 |
|
Kevin Flansburg
|
52f2fe456a
|
fix(disagg): correct DSA/SWA state-page transfer mismatch in PD disaggregation (#27004)
|
2026-06-03 14:33:41 +08:00 |
|
Mick
|
dae86f51f5
|
[diffusion] chore: polish realtime webui waiting state (#27068)
|
2026-06-03 14:29:26 +08:00 |
|
Khoa Pham
|
b5560ffc36
|
Fix flashinfer autotune oom glm51 (#24195)
|
2026-06-02 23:28:57 -07:00 |
|
Cheng Wan
|
202e618898
|
Revert "Fix hybrid linear attention misrouting plain-RadixAttention linear layers to the full backend (Ring-2.5-1T)" (#27116)
|
2026-06-02 23:27:57 -07:00 |
|
Jimmy Shong
|
0ef39784ef
|
[Bugfix] Gate DP-attention even-token padding to CP-enabled configs (#26911)
|
2026-06-03 02:06:52 -04:00 |
|
Duyi-Wang
|
ab7c4ab6bb
|
[AMD] Fix correctness for AITER MLA backend with --page-size > 1 (#25556)
|
2026-06-02 23:01:19 -07:00 |
|
Shaun Kotek
|
b8d7351a74
|
Feat/add w4a16 moe support to nemotron (#25655)
|
2026-06-02 22:42:26 -07:00 |
|
gaopengff
|
aa510bda45
|
Support specific pass of bias_grouped_topk for xpu (#26349)
|
2026-06-03 13:13:48 +08:00 |
|
 
|
f4e7a98fe5
|
[HiCache] feat: support draft offload for mooncake (#24984)
Co-authored-by: huangtingwei9988 <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: stmatengss <11641725+stmatengss@users.noreply.github.com>
|
2026-06-02 21:42:04 -07:00 |
|
CrazyCoder
|
c3aaafc5f2
|
[Bugfix] Clean up failed NIXL sender state (#27011)
|
2026-06-03 12:15:15 +08:00 |
|
 zhaoshangandShangming Cai
|
3e681d7fff
|
Add per-rank staggered weight loading for improved TP I/O concurrency (#26937)
Signed-off-by: zhaoshang <zhaoshangsjtu@linux.alibaba.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-06-03 11:25:21 +08:00 |
|
 zijiexiaandYihao Wang
|
1ebc7438ac
|
model: support Command A plus (#26106)
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
|
2026-06-03 11:23:04 +08:00 |
|
Mick
|
71a747cf15
|
[diffusion] fix: fix lingbot realtime consistency gt pin (#27080)
|
2026-06-03 10:29:37 +08:00 |
|
Chi McIsaac
|
3d54056391
|
[diffusion] optimize: optimize Cosmos3 i2v latent prep (#27084)
|
2026-06-03 10:23:37 +08:00 |
|
Mick
|
3715a07e87
|
[diffusion] improve: clamp wanvae decode output in place (#27086)
|
2026-06-03 10:16:10 +08:00 |
|
Mick
|
f0e18be0ba
|
[diffusion] improve: use conv2d width padding in wanvae (#27081)
|
2026-06-03 10:14:50 +08:00 |
|
 Brayden Zhongandb8zhong
|
13852d3f31
|
Support NextN = 2/4 in DSV32 (#24870)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
|
2026-06-02 19:06:06 -07:00 |
|
MingxuZh
|
b678448b8a
|
ci(xeon): merge 2 partitions into 1 job to reduce runner contention (#26904)
|
2026-06-03 09:46:27 +08:00 |
|
gaopengff
|
eda21f6839
|
Add fused_rope and for xpu (#25773)
|
2026-06-03 09:41:42 +08:00 |
|
Mick
|
83bc776612
|
[diffusion] optimize: preserve dtype in wanvae nearest upsample (#27077)
|
2026-06-03 08:32:28 +08:00 |
|
Qiaolin Yu
|
c55548ba11
|
[perf] Replicate embed_tokens to drop the post-embed all-reduce (#26970)
|
2026-06-02 16:48:18 -07:00 |
|
 Alison ShaoandCheng Wan
|
76c9899da7
|
Fix hybrid linear attention misrouting plain-RadixAttention linear layers to the full backend (Ring-2.5-1T) (#26623)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-06-02 16:24:49 -07:00 |
|
Hubert Lu
|
72929c7000
|
[AMD] Enable AITER custom all-gather on ROCm (#25093)
|
2026-06-02 15:57:37 -07:00 |
|
Khoa Pham
|
a711c57a32
|
[Spec] Fix Gemma 4 MTP with trtllm_mha crash issue (#26966)
|
2026-06-02 14:37:51 -07:00 |
|
Alison Shao
|
365cc2ade5
|
jit_kernel tests: bump multiprocess_test timeout 90s -> 240s (cold JIT cache) (#26994)
|
2026-06-02 17:33:52 -04:00 |
|
Baizhou Zhang
|
68caf49154
|
Update sgl-deep-gemm to 0.1.2 (#26993)
|
2026-06-02 13:51:54 -07:00 |
|
Hanming Lu
|
b603f08c0c
|
[DP] Fix FlashInfer dispatcher workspace sizing and set_dp_buffer_len (#26643)
|
2026-06-02 13:25:20 -07:00 |
|
Bi Xue
|
9e717cae46
|
[sglang] Fix Mamba COW over-releasing SWA locks (cascade-evict assert crash) (#27038)
|
2026-06-03 01:42:04 +08:00 |
|
Muqi Li
|
6ba31e33e6
|
feat(api): add require_reasoning field for engine's generate api (#27019)
|
2026-06-02 17:36:29 +00:00 |
|
 Cheng WanandClaude Opus 4.7
|
99da43b900
|
[refactor] init_forward_metadata 3-method ABC + side-channel removal + ForwardMetadata type rename (#26735)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
2026-06-02 10:33:33 -07:00 |
|
Ilia Yastrebov
|
6c69756fa8
|
NIXL: use prep+make API to improve performance (#26406)
|
2026-06-02 18:36:35 +02:00 |
|