Commit Graph
13098 Commits
Author SHA1 Message Date
Kangyan-ZhouandClaude Opus 4.7 6e8fe176be sgl-router: experimental Rust HTTP router for SGLang worker pools (#25851)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 15:34:05 +08:00
Junlin Wu aae04b1241 📝 docs(diffusion): add MXFP4 quantization docs (#25904) 2026-05-25 10:24:30 +03:00
Yuhui Liang ca029e816b Fix missing idle-batch handling in prepare_mlp_sync_batch_raw (#25404) 2026-05-24 23:53:54 -07:00
Erik Wijmans 87e69d57c4 [lora] Fix overlap loading for cancelled requests (#25413) 2026-05-25 15:18:36 +09:00
Xia WeiwenandMa Mingfei 2bd3ac0b5d [XPU] fix correctness issue of GDN triton kernel for XPU (#26065)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-05-25 13:18:11 +08:00
Qiaolin Yu ec6fcb93cb [perf][spec decoding] Skip common_template in TRTLLMMLAMultiStepDraftBackend init (#26241) 2026-05-24 21:36:16 -07:00
Ethan ZHUandZhangheng e86fdf3a3c [Bug Fix][HiCache] TreeNode.get_prefix_hash_values @lru_cache can return mutated list (#26177)
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-05-25 11:15:10 +08:00
Ma Mingfei 821d5f4a5b [CPU] add faster KV-cache writes (#25874) 2026-05-25 10:28:52 +08:00
Mick e1463bb2c2 [VLM] try to reuse precomputed padded input ids in scheduler instead of padding (#26097) 2026-05-25 10:27:04 +08:00
hanwlax de3f6fb02e Fix attr err (#25856) 2026-05-25 10:26:19 +08:00
Liangsheng Yin 850887dc63 [Spec] fix EAGLE v2 verify metadata init order on non-cuda-graph path (#26244) 2026-05-24 18:49:33 -07:00
Mick 64e2b54a8f [VLM] feat: accept grid_thws from preprocessed metadata for kimi (#26149) 2026-05-25 09:10:25 +08:00
Mick 72c1582d4e [VLM] fix: fix only the grids from last split mm item is collected for qwen-vl (#26094) 2026-05-25 09:09:46 +08:00
Liangsheng Yin ed179bf9b2 [dsv4] fix multi-step draft on non-cuda-graph path (#26239) 2026-05-24 17:04:18 -07:00
Liangsheng Yin d7e3e54148 [Test] split test/registered/distributed/ into topic folders (#26240) 2026-05-24 17:02:07 -07:00
Ming Yang 85471d253d Add --disable-attn-tp-gather opt-out for model-managed SP (#26047) 2026-05-24 15:53:51 -07:00
Lianmin Zheng 93fa577bb9 Clean up server startup log noise (#26205) 2026-05-24 14:35:15 -07:00
Liangsheng Yin 030bd5d3ed [Test] test_session_latency: assert streaming tail/head stability (#26230) 2026-05-24 14:06:41 -07:00
Polisetty V R K Jyothendra Varma fd94bd30b8 [Intel GPU] DeepSeek V4 2/N: Fix tvm ffi import (#26118) 2026-05-24 12:59:49 -07:00
Cheng WanandCheng Wan 44922de48a fix(swa): downgrade translate_loc_from_full_to_swa key-change log from warning to debug (#26225)
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
2026-05-24 11:37:46 -07:00
Siyuan Chenandxutizhou 7f45bcdd2a [dsv4] support eplb (#25948)
Co-authored-by: xutizhou <xutingz@nvidia.com>
2026-05-24 10:09:41 -07:00
Mick 5c3775823e [diffusion] chore: use model-aware vae channels_last_3d policy (#26214) 2026-05-25 00:25:34 +08:00
shuwenn 36eb72bf12 [UnifiedTree] fix: backup SWA-split parent before child under write-through (#25065) 2026-05-25 00:21:57 +08:00
Dongjun Na 9d50cd9742 [observability] add ServerArgs.stat_loggers for pluggable metrics backend (#24610)
Signed-off-by: Dongjun Na <kmu5544616@gmail.com>
2026-05-24 22:41:28 +08:00
Mick b6f71d5850 [VLM] avoid extra cuda-ipc staging for preprocessed input (#26096) 2026-05-24 19:48:06 +08:00
Mick 6447596501 [VLM] feat: replace small H2D calls with a single one for qwen-vl (#26167) 2026-05-24 18:27:17 +08:00
Xiaoyu Zhang 0b65588c18 [diffusion] Clean up VSA attention hot path (#25514) 2026-05-24 16:46:03 +08:00
Mick 4c2b32bfbf [VLM] accept precomputed multimodal metadata (#26101) 2026-05-24 15:43:21 +08:00
Mick 826a4de062 [srt] store req input ids as arrays (#26165) 2026-05-24 15:09:36 +08:00
Mick d6d9f12444 [VLM] adopt simplified get_rope_index for image-only requests (#26100) 2026-05-24 11:51:24 +08:00
+2 af8f66940e [AMD] Dsv4/pr1 fix run time issue (#25898)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: amd-danli103 <danli103@amd.com>
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
Co-authored-by: yichiche@amd.com <jacky.cheng>
Co-authored-by: yctseng0211 <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
2026-05-23 16:04:14 -07:00
Qiaolin Yu 982f67d9a6 Suppress cutlass-dsl noisy warning (#26169) 2026-05-23 13:19:14 -07:00
2de74035a5 [FIX][2/2] fix step3-vl/deepseek-ocr image processor error (#25403)
Co-authored-by: kousakawang <wanghanpei@bytedance.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-24 01:36:39 +08:00
Hanming Lu a5a64a311a [Spec] trtllm mha supports overlap plan stream (#25925) 2026-05-23 03:25:57 -07:00
Qiaolin Yu cb7b57955d fix tokenspeed_mla attn kernel jit (#26170) 2026-05-23 03:24:33 -07:00
Charles Chen 89ff2bc111 [bug fix] Fix 3 issues when using Gemma4 MTP (#26026) 2026-05-23 03:16:47 -07:00
Khoa PhamandClaude Opus 4.7 b0ce16d0c5 [CP] 1/N: Support MLA Prefill Context Parallel (#23292)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-23 03:07:09 -07:00
zijiexiaandClaude Opus 4.7 81cd338fcc [docs] DeepSeek-V4 cookbook: balanced MegaMoE cap, H200 Pro FP4 mem-frac, nsa-* compat, PD-disagg fixes (#26164)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-23 02:42:41 -07:00
Liangsheng Yin 1e59ed7443 compile _resolve_spec_extras gather kernels (#26129) 2026-05-23 02:34:41 -07:00
Cheng WanandCheng Wan 83a18e687d Revert "[refactor] unify cuda-graph capture/replay across attention backends (#26134)" (#26166)
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
2026-05-23 02:32:08 -07:00
liuxianglong17 8c78424701 Reduce excessively long logs caused by transformer version updates. (#26033) 2026-05-23 17:12:08 +08:00
Mick 774b29dade [VLM] feat: early-return in mm processor if the input is preprocessed (#26117) 2026-05-23 17:02:44 +08:00
Mick 19b60a4f9e [VLM] reuse pretokenized ids from preprocessed input for qwen-vl (#26116) 2026-05-23 16:01:04 +08:00
Zheng Wengang 8b9fb13c4a [BugFix][EPD] adapt for qwen3.5-mtp & del duplicated logs (#24144) 2026-05-23 15:32:57 +08:00
Cheng WanandCheng Wan 5964d30233 fix(swa): eliminate spurious translate_loc_from_full_to_swa warning in BCG and CG paths (#26152)
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
2026-05-23 00:01:12 -07:00
Xiaoyu Zhang 75427c9ca4 Route concat MLA to JIT and remove unused downcast (#25843) 2026-05-23 14:30:43 +08:00
xiaobochen-amdandfanxingran fd3e11973b [AMD][aiter] Fix cuda_graph_kv_indices OOB under page_size>1 (#24587)
Co-authored-by: fanxingran <fanxingran@amd.com>
2026-05-22 23:19:46 -07:00
longxin9715 c69844f043 [NPU]Ascend NPU Performance Profiling Guide and Ascend NPU Operator Development Guide (#26069) 2026-05-23 10:50:56 +08:00
Baizhou Zhangandyhyang201 7b7f1067bd Add non-MTP DSV4 test coverage (#26141)
Co-authored-by: yhyang201 <yhyang201@gmail.com>
2026-05-22 19:43:17 -07:00
Shangming Cai a241659d18 [PD] Consolidate shared logic into common backend (#25979)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
2026-05-23 10:41:20 +08:00