Commit Graph
2765 Commits
Author SHA1 Message Date
Ye (Charlotte) Qi 03c77dc33d [PD] Deduplicate PD logprob normalization (#27085) 2026-06-03 19:08:15 +08:00
Bingxu Chen d7013b6537 [AMD] [CI] Remove hardcoded model/cache paths from MI35x nightly tests (#27001) 2026-06-03 02:14:32 -07:00
e67810bea7 [SGLang Tracing] Add pd disaggregation mooncake backend tracing (#23755)
Co-authored-by: Mu Huai <tianbowen.tbw@antgroup.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-06-03 16:43:29 +08:00
Vladislav NosivskoyandZhangheng 63dc20ae6c [UnifiedTree] Add CP sync (#25395)
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-06-03 16:10:27 +08:00
Bingxu Chen 8e77af1afc [AMD] fix(triton-mla): cap max_kv_splits at 256 on gfx942 (Kimi-K2.6 hang) (#24762) 2026-06-03 00:13:18 -07:00
Kevin Flansburg 52f2fe456a fix(disagg): correct DSA/SWA state-page transfer mismatch in PD disaggregation (#27004) 2026-06-03 14:33:41 +08:00
Khoa PhamandClaude Opus 4.8 6d53615699 [Gemma4] Use hard GSM8K accuracy floor for 31B MTP test (#27101)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-03 01:57:12 -04:00
Shaun Kotek b8d7351a74 Feat/add w4a16 moe support to nemotron (#25655) 2026-06-02 22:42:26 -07:00
gaopengff aa510bda45 Support specific pass of bias_grouped_topk for xpu (#26349) 2026-06-03 13:13:48 +08:00
f4e7a98fe5 [HiCache] feat: support draft offload for mooncake (#24984)
Co-authored-by: huangtingwei9988 <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: stmatengss <11641725+stmatengss@users.noreply.github.com>
2026-06-02 21:42:04 -07:00
CrazyCoder c3aaafc5f2 [Bugfix] Clean up failed NIXL sender state (#27011) 2026-06-03 12:15:15 +08:00
Alison ShaoandCheng Wan 76c9899da7 Fix hybrid linear attention misrouting plain-RadixAttention linear layers to the full backend (Ring-2.5-1T) (#26623)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
2026-06-02 16:24:49 -07:00
Hubert Lu 72929c7000 [AMD] Enable AITER custom all-gather on ROCm (#25093) 2026-06-02 15:57:37 -07:00
Khoa Pham 22bb9a6421 test: disable test_gemma4_mtp_26b_a4b_extra from CI (#27082) 2026-06-02 14:00:48 -07:00
Cheng WanandClaude Opus 4.7 99da43b900 [refactor] init_forward_metadata 3-method ABC + side-channel removal + ForwardMetadata type rename (#26735)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-02 10:33:33 -07:00
Ke Bao 28f9c1ff24 Relax mamba unified cache kl threshold (#27070) 2026-06-02 23:28:22 +08:00
Ke Bao b5e154dc73 Fix stale import after kl_nightly rename (#27064) 2026-06-02 21:46:21 +08:00
Zhangheng ee4bf0a9d3 [UnifiedTree]: Add HiCache Nightly CI For GLM5 (#26927) 2026-06-02 19:02:50 +08:00
Bruce Changlong XuandKe Bao 172bd8e6b9 [scheduler] Zero gen_throughput and flush KV events on pause (#24003)
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
2026-06-02 16:43:04 +08:00
cctryandgemini-code-assist[bot] b55570d38e [PD] Optimistic prefill (#26780)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-06-02 01:16:14 -07:00
Charles Chen 5ae8d286d2 perf(gemma4): single-launch fused router (topk + softmax + scale) (#26502) 2026-06-02 16:00:17 +08:00
fzyzcjy 8cea0473ea Fix dp-attention token alignment in the dumper comparator e2e test (#26996) 2026-06-02 00:50:45 -07:00
3e993f6140 [PD]: Support HiCache prefetching and pd-incremental transfer on decode side (#26227)
Co-authored-by: huangtingwei <141888744+huangtingwei9988@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
2026-06-02 15:40:10 +08:00
Hsiu-Chun, HungandHung 2582134a59 [AMD] Add amd ci mamba state scatter test (#26677)
Co-authored-by: Hung <Emmanuel0612@users.noreply.github.com>
2026-06-02 00:24:59 -07:00
4226a6f13a [AMD] Fix GPT-OSS MXFP4 accuracy on ROCm AITER path (#26884)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
2026-06-01 22:30:43 -07:00
Ethan ZHUandZhangheng 594ec6335d [Bug Fix][HiCache] Drop @lru_cache on UnifiedTreeNode.get_prefix_hash_values (#26939)
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-06-02 12:38:28 +08:00
0574d2b8a5 [NVIDIA] [GDN] Enable FlashInfer MTP verify on SM100+ (Blackwell) (#23273)
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-01 18:56:42 -07:00
Liangsheng Yin f6d0beaca8 Revert "Support spec v2 tree drafting (eagle topk>1) with page_size==1" (#26981) 2026-06-01 17:16:44 -07:00
Qiaolin Yu 4151a04d1a [Perf][Spec Decoding] Skip cat/topk/sort/gather in draft_forward for topk=1 (#26424) 2026-06-01 15:37:47 -07:00
Liangsheng Yin 1d4ee060c2 Support spec v2 tree drafting (eagle topk>1) with page_size==1 (#26866) 2026-06-01 15:37:20 -07:00
Yongfei Xu 5700790c05 DeepSeek V4: Support context parallelism with fused MoE (non-DeepEP) (#24947) 2026-06-01 14:25:43 -07:00
eeechoandClaude Opus 4.6 524ba10eda feat: SM120 (Blackwell Desktop) support for DeepSeek-V4 inference (#24692)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-06-01 14:05:20 -07:00
a0670b5ba3 [SPEC] feat: add adaptive speculative decoding metrics (#25940)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Jarrod Barnes <jbarnes850@gmail.com>
2026-06-01 13:53:30 -07:00
86afa21ca7 feat: optional caller-supplied mm_hashes on GenerateReqInput (#25300)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
2026-06-01 20:04:37 +02:00
f6a5a1b59c [RL+VLM] Avoid retokenization drift for pre-tokenized (token-id) VLM requests (#26555)
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
Co-authored-by: root <root@slurm-h200-209-231.slurm-compute.tenant-slurm.svc.cluster.local>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-06-01 09:58:14 -07:00
Ke Bao f59bbef841 Split SWA leaf to one window on insert (#26919) 2026-06-01 23:46:55 +08:00
Lukas HumbelandClaude Opus 4.7 d8a5a25c36 Refactor NIXL hicache. Add O_DIRECT support (#25173)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-01 17:28:53 +02:00
Ke Bao 693adabff7 Fix Mamba2Metadata dropping has_mamba_track_mask (#26877) 2026-06-01 22:03:59 +08:00
Bi Xue 6965fe0eec [sgl] Window-aware LRU refresh for SWA prefix cache in unified cache (#26615) 2026-06-01 19:35:18 +08:00
Liangsheng Yin 1bff7a290f Refactor EAGLE infer tests: shared fixture + kits + overlap matrix (#26871) 2026-06-01 03:55:01 -07:00
Yuan Luoandluoyuan.luo bc36231d65 [KDA] Support KDA packed decode (#26586)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-06-01 16:52:01 +08:00
Ke Bao cdd06011a1 Make unified tree SWA hicache tests faithful to write-through backup (#26870) 2026-06-01 16:08:57 +08:00
Lianmin Zheng 53b8378307 Fix weights_checker checksum for 0-dim tensors and multi-GPU (#26863) 2026-05-31 21:27:03 -07:00
liuxianglong17 3aaf8f115e fix test cases failed in nightly pipeline (#26714) 2026-06-01 11:33:47 +08:00
Ke Bao 972fbf7711 Skip flaky mamba extra_buffer disagg test (#26838) 2026-05-31 15:59:03 +08:00
fzyzcjy f220c72929 Add periodic KV-canary stats logging and kernel-run-counter health check (#26821) 2026-05-31 10:00:19 +08:00
fzyzcjy 7dd19ae3d8 Add a sliding-window-attention divergence reporter for the KV-canary (#26820) 2026-05-31 09:59:28 +08:00
fzyzcjy ae9db7ff4b Add the KV-canary perturb modes and PD-disaggregation e2e tests (#26819) 2026-05-31 09:59:09 +08:00
fzyzcjy 6be4b32d8d Add token-id verification to the KV-canary (#26818) 2026-05-31 09:58:51 +08:00
fzyzcjy 0ca610a6df Add real-data KV verification to the KV-canary (#26817) 2026-05-31 09:58:32 +08:00