 Cheng WanandClaude Opus 4.8
|
10ab7c919f
|
[refactor] Retire DecodeInputBuffers / PrefillInputBuffers in favor of CudaGraphBufferRegistry (#27192)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-03 20:52:56 -07:00 |
|
 Zaili WangandMa Mingfei
|
3b7a258f63
|
[CPU] upgrade dependent torch ver to PT2.12 (#21456)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-06-04 11:04:11 +08:00 |
|
 Zhanghengand晟海
|
736263f3dc
|
[UnifiedTree]: Support l3 storage for swa and deepseek v4 (#26881)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
|
2026-06-04 10:17:34 +08:00 |
|
shuwenn
|
1a57145975
|
[codex] Fix adaptive metrics test flake (#27135)
|
2026-06-03 19:07:17 -07:00 |
|
 Yuzhen ZhouandJiajun Li
|
e03dfa8182
|
[3/N][Sync sglang-miles] TITO Support (#23751)
Co-authored-by: Jiajun Li <48857426+guapisolo@users.noreply.github.com>
|
2026-06-03 21:45:33 -04:00 |
|
 
|
f6cd1a9822
|
Add num_waiting_uncached_tokens load metric (#27174)
Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
|
2026-06-03 18:29:49 -07:00 |
|
 ybyangandLianmin Zheng
|
687baf9471
|
fix(load-snapshot): avoid duplicate zmq bind in multi-tokenizer mode (#27145)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
|
2026-06-03 18:24:56 -07:00 |
|
 
|
14ed9b448e
|
Add ZMQ IPv6 support, bench_serving sampling params, and reduce routed_dp_rank log noise (#27180)
Co-authored-by: Hanming Lu <69857889+hanming-lu@users.noreply.github.com>
Co-authored-by: Grigory Sizov <grisha.sizov@gmail.com>
|
2026-06-03 17:49:34 -07:00 |
|
 Clintandclintg6
|
cfb7fb4fad
|
[AMD] Fix TP2 DeepSeek-R1 nhead=64 MLA decode crash and add nightly coverage (#27188)
Co-authored-by: clintg6 <7388379+clintg6@users.noreply.github.com>
|
2026-06-03 16:56:05 -07:00 |
|
 Cheng WanandClaude Opus 4.8
|
c9ca56da8c
|
Unify full→SWA index translation in init_forward_metadata; drop pool caches (#27091)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-06-03 16:12:27 -07:00 |
|
Cheng Wan
|
61aa3293d3
|
Revert "Fix TokenizerManager crash on top_logprobs with tensor values" (#27187)
|
2026-06-03 14:53:28 -07:00 |
|
ishandhanani
|
978fb6ed1a
|
hicache kv events: publish split write-through fragments (#27072)
|
2026-06-03 14:52:25 -07:00 |
|
Kevin Flansburg
|
7716fa00e0
|
Fix TokenizerManager crash on top_logprobs with tensor values (#26825)
|
2026-06-03 13:55:02 -07:00 |
|
YC Yen-Ching Tseng
|
d1bc06b63b
|
[AMD] Disable AITER custom all-gather in DeepSeek-R1-MXFP4 8-GPU test (#27163)
|
2026-06-03 13:38:34 -07:00 |
|
 
|
293816ab14
|
[AMD][MXFP4] Online MXFP4 quantization 1/N - dense and MOE models w. original BF16 weight (#18005)
Co-authored-by: Bowen Bao <bowenbao@amd.com>
Co-authored-by: Colin Zeng <Colin.Zeng@amd.com>
|
2026-06-03 12:55:24 -07:00 |
|
 Hanming LuandYAMY
|
e0b692600f
|
[Mamba] extra buffer lazy support (#27118)
Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com>
|
2026-06-03 12:42:11 -07:00 |
|
 Liangsheng YinandAlison Shao
|
ac99794e64
|
Reland spec v2 tree drafting (eagle topk>1) with page_size==1 (#26866) (#26997)
Co-authored-by: Alison Shao <54658187+alisonshao@users.noreply.github.com>
|
2026-06-03 15:40:05 -04:00 |
|
Liangsheng Yin
|
578f232e5e
|
Fix trace_modules gate disabling default trace contexts (#27173)
|
2026-06-03 14:03:42 -04:00 |
|
 Cheng WanandClaude Opus 4.8
|
45604a0f4a
|
[refactor] Unify CUDA graph runner input buffers behind CudaGraphBufferRegistry (#26742)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-03 10:54:10 -07:00 |
|
 Lijuan TangandXiaodong Ye
|
9d0e6a2df4
|
fix(mlx): set canary_manager and materialize overlap-loop inputs on Apple Silicon (#26882)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
Signed-off-by: LijuanTang94 <tang.lij@northeastern.edu>
Co-authored-by: Xiaodong Ye <yeahdongcn@gmail.com>
|
2026-06-04 00:03:46 +08:00 |
|
Ye (Charlotte) Qi
|
03c77dc33d
|
[PD] Deduplicate PD logprob normalization (#27085)
|
2026-06-03 19:08:15 +08:00 |
|
Bingxu Chen
|
d7013b6537
|
[AMD] [CI] Remove hardcoded model/cache paths from MI35x nightly tests (#27001)
|
2026-06-03 02:14:32 -07:00 |
|
 
|
e67810bea7
|
[SGLang Tracing] Add pd disaggregation mooncake backend tracing (#23755)
Co-authored-by: Mu Huai <tianbowen.tbw@antgroup.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-06-03 16:43:29 +08:00 |
|
 Vladislav NosivskoyandZhangheng
|
63dc20ae6c
|
[UnifiedTree] Add CP sync (#25395)
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-06-03 16:10:27 +08:00 |
|
Bingxu Chen
|
8e77af1afc
|
[AMD] fix(triton-mla): cap max_kv_splits at 256 on gfx942 (Kimi-K2.6 hang) (#24762)
|
2026-06-03 00:13:18 -07:00 |
|
Kevin Flansburg
|
52f2fe456a
|
fix(disagg): correct DSA/SWA state-page transfer mismatch in PD disaggregation (#27004)
|
2026-06-03 14:33:41 +08:00 |
|
 Khoa PhamandClaude Opus 4.8
|
6d53615699
|
[Gemma4] Use hard GSM8K accuracy floor for 31B MTP test (#27101)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-03 01:57:12 -04:00 |
|
Shaun Kotek
|
b8d7351a74
|
Feat/add w4a16 moe support to nemotron (#25655)
|
2026-06-02 22:42:26 -07:00 |
|
gaopengff
|
aa510bda45
|
Support specific pass of bias_grouped_topk for xpu (#26349)
|
2026-06-03 13:13:48 +08:00 |
|
 
|
f4e7a98fe5
|
[HiCache] feat: support draft offload for mooncake (#24984)
Co-authored-by: huangtingwei9988 <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: stmatengss <11641725+stmatengss@users.noreply.github.com>
|
2026-06-02 21:42:04 -07:00 |
|
CrazyCoder
|
c3aaafc5f2
|
[Bugfix] Clean up failed NIXL sender state (#27011)
|
2026-06-03 12:15:15 +08:00 |
|
 Alison ShaoandCheng Wan
|
76c9899da7
|
Fix hybrid linear attention misrouting plain-RadixAttention linear layers to the full backend (Ring-2.5-1T) (#26623)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-06-02 16:24:49 -07:00 |
|
Hubert Lu
|
72929c7000
|
[AMD] Enable AITER custom all-gather on ROCm (#25093)
|
2026-06-02 15:57:37 -07:00 |
|
Khoa Pham
|
22bb9a6421
|
test: disable test_gemma4_mtp_26b_a4b_extra from CI (#27082)
|
2026-06-02 14:00:48 -07:00 |
|
 Cheng WanandClaude Opus 4.7
|
99da43b900
|
[refactor] init_forward_metadata 3-method ABC + side-channel removal + ForwardMetadata type rename (#26735)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
2026-06-02 10:33:33 -07:00 |
|
Ke Bao
|
28f9c1ff24
|
Relax mamba unified cache kl threshold (#27070)
|
2026-06-02 23:28:22 +08:00 |
|
Ke Bao
|
b5e154dc73
|
Fix stale import after kl_nightly rename (#27064)
|
2026-06-02 21:46:21 +08:00 |
|
Zhangheng
|
ee4bf0a9d3
|
[UnifiedTree]: Add HiCache Nightly CI For GLM5 (#26927)
|
2026-06-02 19:02:50 +08:00 |
|
 Bruce Changlong XuandKe Bao
|
172bd8e6b9
|
[scheduler] Zero gen_throughput and flush KV events on pause (#24003)
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
|
2026-06-02 16:43:04 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) cctryandgemini-code-assist[bot]
|
b55570d38e
|
[PD] Optimistic prefill (#26780)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-06-02 01:16:14 -07:00 |
|
Charles Chen
|
5ae8d286d2
|
perf(gemma4): single-launch fused router (topk + softmax + scale) (#26502)
|
2026-06-02 16:00:17 +08:00 |
|
fzyzcjy
|
8cea0473ea
|
Fix dp-attention token alignment in the dumper comparator e2e test (#26996)
|
2026-06-02 00:50:45 -07:00 |
|
  
|
3e993f6140
|
[PD]: Support HiCache prefetching and pd-incremental transfer on decode side (#26227)
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
|
2026-06-02 15:40:10 +08:00 |
|
 Hsiu-Chun, HungandHung
|
2582134a59
|
[AMD] Add amd ci mamba state scatter test (#26677)
Co-authored-by: Hung <Emmanuel0612@users.noreply.github.com>
|
2026-06-02 00:24:59 -07:00 |
|
 
|
4226a6f13a
|
[AMD] Fix GPT-OSS MXFP4 accuracy on ROCm AITER path (#26884)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-06-01 22:30:43 -07:00 |
|
 Ethan ZHUandZhangheng
|
594ec6335d
|
[Bug Fix][HiCache] Drop @lru_cache on UnifiedTreeNode.get_prefix_hash_values (#26939)
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-06-02 12:38:28 +08:00 |
|
 
|
0574d2b8a5
|
[NVIDIA] [GDN] Enable FlashInfer MTP verify on SM100+ (Blackwell) (#23273)
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-06-01 18:56:42 -07:00 |
|
Liangsheng Yin
|
f6d0beaca8
|
Revert "Support spec v2 tree drafting (eagle topk>1) with page_size==1" (#26981)
|
2026-06-01 17:16:44 -07:00 |
|
Qiaolin Yu
|
4151a04d1a
|
[Perf][Spec Decoding] Skip cat/topk/sort/gather in draft_forward for topk=1 (#26424)
|
2026-06-01 15:37:47 -07:00 |
|
Liangsheng Yin
|
1d4ee060c2
|
Support spec v2 tree drafting (eagle topk>1) with page_size==1 (#26866)
|
2026-06-01 15:37:20 -07:00 |
|