18675 Commits
Author SHA1 Message Date
Yuxuan Zhang 7ef06bfc06 GLM-4.7-Flash: standalone MLA impl and MLA NextN/MTP (#26088) 2026-05-26 13:17:39 +08:00
xutizhou 59cad671e2 Support DeepSeek V4 DeepEP Waterfill (#25391) 2026-05-25 21:04:26 -07:00
mispa-msandMick 3142278c5f [diffusion] feat: layerwise NVTX markers for Nsight Systems profiling (#25683)
Co-authored-by: Mick <mickjagger19@icloud.com>
2026-05-26 11:24:27 +08:00
Liangsheng Yin 1953565ba1 Signal CUDA coredumps to tracker issue (#26338) 2026-05-25 20:21:19 -07:00
Chetan Kumar Verma 156d1af23a [Intel GPU] Fix incorrect KV-cache page table for local attention when page_size > 1 (#23757) 2026-05-26 11:02:21 +08:00
Baizhou Zhang 29e245e6a7 [misc] Update permission (#26336) 2026-05-25 19:50:06 -07:00
Mick 8f2a4e70f8 [SRT] minor: reuse req input id array for unpadded ids (#26232) 2026-05-26 08:58:32 +08:00
Liangsheng Yin 8805f4cf16 Fail-fast on PD subprocess exit and scheduler exception (#26298) 2026-05-25 16:42:50 -07:00
Ke Bao e7b12fe6fa Fix stale forward_metadata leak in DP attn unpadded idle batch (#26313) 2026-05-25 16:04:00 -07:00
Ziang Li 2b9dd9c8b3 [FlashInfer v0.6.10] [RL] [DSv32] [GLM-5] Add --dsa-topk-backend and integrate FlashInfer and pytorch topk (#22851) 2026-05-25 13:08:03 -07:00
Ke Bao b13d3d18c6 Refactor HiCache stack dispatch into strategies (#26295) 2026-05-26 00:06:17 +08:00
Shangming Cai 2aa6995308 [CI] Enable EPD CI for EPD architecture enhancements (#26281) 2026-05-25 23:52:58 +08:00
Xiaoyu ZhangandBBuf 121cc09405 [diffusion] Add CFG gating for denoising (#25848)
Co-authored-by: BBuf <bbuf@example.com>
2026-05-25 22:57:09 +08:00
Xiaoyu ZhangandBBuf 85f9522e36 [diffusion] Cache fp32 layernorm params (#25847)
Co-authored-by: BBuf <bbuf@example.com>
2026-05-25 22:56:39 +08:00
Makcum888eandronnie_zheng 0801cc05ed [Diffusion][NPU] Disaggregation diffusion stages support for NPU (#25895)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-05-25 13:51:25 +03:00
Liangsheng Yin 3e67398a96 Zero req_pool_indices padding in cuda-graph populate (#26292) 2026-05-25 03:29:05 -07:00
Xiaoyu Zhang 533ef41112 [Diffusion] Default NVFP4 backend to FlashInfer TRTLLM (#25523) 2026-05-25 18:14:06 +08:00
Mick c05756da7a [SRT] fix flashInfer allreduce fusion not used on blackwell (#26197) 2026-05-25 18:07:15 +08:00
Zhanghengand晟海 a4db563c87 [hisparse]: update user guide (#26249)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
2026-05-25 17:54:55 +08:00
Shangming Cai bc8d64bf36 [CI] Align score threshold in dsv4 disaggregation test (#26268) 2026-05-25 17:48:45 +08:00
Qiaolin Yu a77449f86d [perf][spec decoding] Skip full-vocab softmax in EAGLE draft when topk == 1 (#26235) 2026-05-25 02:06:48 -07:00
Kangyan-ZhouandClaude Opus 4.7 7c04b9e942 fix(docker): generate Cargo.lock in chef stage for sgl-router build (#26279)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-25 16:54:47 +08:00
Shangming Cai 7f2829af39 chore: bump mooncake version to 0.3.11.post1 (#25989) 2026-05-25 16:51:03 +08:00
xdtbynd 0942011665 [NPU] Add torchaudio dependency for NPU platform (#26267) 2026-05-25 16:30:12 +08:00
Kangyan-ZhouandClaude Opus 4.7 81704ad602 ci: add nightly Docker workflow for experimental sgl-router (#26273)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-25 16:19:07 +08:00
Jincong Chen e27d4fb70f [Perf][Qwen3.5] Add case 512 to topkGatingSoftmaxKernelLauncher, (#25775) 2026-05-25 16:08:21 +08:00
fzyzcjy b0cf01eb85 Lazy-load speculative-naming via skill instead of always-on rule (#26270) 2026-05-25 15:58:29 +08:00
Kangyan-ZhouandClaude Opus 4.7 6e8fe176be sgl-router: experimental Rust HTTP router for SGLang worker pools (#25851)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 15:34:05 +08:00
Junlin Wu aae04b1241 📝 docs(diffusion): add MXFP4 quantization docs (#25904) 2026-05-25 10:24:30 +03:00
Yuhui Liang ca029e816b Fix missing idle-batch handling in prepare_mlp_sync_batch_raw (#25404) 2026-05-24 23:53:54 -07:00
Erik Wijmans 87e69d57c4 [lora] Fix overlap loading for cancelled requests (#25413) 2026-05-25 15:18:36 +09:00
Xia WeiwenandMa Mingfei 2bd3ac0b5d [XPU] fix correctness issue of GDN triton kernel for XPU (#26065)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-05-25 13:18:11 +08:00
Qiaolin Yu ec6fcb93cb [perf][spec decoding] Skip common_template in TRTLLMMLAMultiStepDraftBackend init (#26241) 2026-05-24 21:36:16 -07:00
Ethan ZHUandZhangheng e86fdf3a3c [Bug Fix][HiCache] TreeNode.get_prefix_hash_values @lru_cache can return mutated list (#26177)
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-05-25 11:15:10 +08:00
Ma Mingfei 821d5f4a5b [CPU] add faster KV-cache writes (#25874) 2026-05-25 10:28:52 +08:00
Mick e1463bb2c2 [VLM] try to reuse precomputed padded input ids in scheduler instead of padding (#26097) 2026-05-25 10:27:04 +08:00
hanwlax de3f6fb02e Fix attr err (#25856) 2026-05-25 10:26:19 +08:00
Liangsheng Yin 850887dc63 [Spec] fix EAGLE v2 verify metadata init order on non-cuda-graph path (#26244) 2026-05-24 18:49:33 -07:00
Mick 64e2b54a8f [VLM] feat: accept grid_thws from preprocessed metadata for kimi (#26149) 2026-05-25 09:10:25 +08:00
Mick 72c1582d4e [VLM] fix: fix only the grids from last split mm item is collected for qwen-vl (#26094) 2026-05-25 09:09:46 +08:00
Liangsheng Yin ed179bf9b2 [dsv4] fix multi-step draft on non-cuda-graph path (#26239) 2026-05-24 17:04:18 -07:00
Liangsheng Yin d7e3e54148 [Test] split test/registered/distributed/ into topic folders (#26240) 2026-05-24 17:02:07 -07:00
Ming Yang 85471d253d Add --disable-attn-tp-gather opt-out for model-managed SP (#26047) 2026-05-24 15:53:51 -07:00
Lianmin Zheng 93fa577bb9 Clean up server startup log noise (#26205) 2026-05-24 14:35:15 -07:00
Liangsheng Yin 030bd5d3ed [Test] test_session_latency: assert streaming tail/head stability (#26230) 2026-05-24 14:06:41 -07:00
Polisetty V R K Jyothendra Varma fd94bd30b8 [Intel GPU] DeepSeek V4 2/N: Fix tvm ffi import (#26118) 2026-05-24 12:59:49 -07:00
Cheng WanandCheng Wan 44922de48a fix(swa): downgrade translate_loc_from_full_to_swa key-change log from warning to debug (#26225)
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
2026-05-24 11:37:46 -07:00
Siyuan Chenandxutizhou 7f45bcdd2a [dsv4] support eplb (#25948)
Co-authored-by: xutizhou <xutingz@nvidia.com>
2026-05-24 10:09:41 -07:00
Mick 5c3775823e [diffusion] chore: use model-aware vae channels_last_3d policy (#26214) 2026-05-25 00:25:34 +08:00
shuwenn 36eb72bf12 [UnifiedTree] fix: backup SWA-split parent before child under write-through (#25065) 2026-05-25 00:21:57 +08:00