Shangming Cai
|
7f2829af39
|
chore: bump mooncake version to 0.3.11.post1 (#25989)
|
2026-05-25 16:51:03 +08:00 |
|
xdtbynd
|
0942011665
|
[NPU] Add torchaudio dependency for NPU platform (#26267)
|
2026-05-25 16:30:12 +08:00 |
|
Yuhui Liang
|
ca029e816b
|
Fix missing idle-batch handling in prepare_mlp_sync_batch_raw (#25404)
|
2026-05-24 23:53:54 -07:00 |
|
Erik Wijmans
|
87e69d57c4
|
[lora] Fix overlap loading for cancelled requests (#25413)
|
2026-05-25 15:18:36 +09:00 |
|
 Xia WeiwenandMa Mingfei
|
2bd3ac0b5d
|
[XPU] fix correctness issue of GDN triton kernel for XPU (#26065)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-05-25 13:18:11 +08:00 |
|
Qiaolin Yu
|
ec6fcb93cb
|
[perf][spec decoding] Skip common_template in TRTLLMMLAMultiStepDraftBackend init (#26241)
|
2026-05-24 21:36:16 -07:00 |
|
 Ethan ZHUandZhangheng
|
e86fdf3a3c
|
[Bug Fix][HiCache] TreeNode.get_prefix_hash_values @lru_cache can return mutated list (#26177)
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-05-25 11:15:10 +08:00 |
|
Ma Mingfei
|
821d5f4a5b
|
[CPU] add faster KV-cache writes (#25874)
|
2026-05-25 10:28:52 +08:00 |
|
Mick
|
e1463bb2c2
|
[VLM] try to reuse precomputed padded input ids in scheduler instead of padding (#26097)
|
2026-05-25 10:27:04 +08:00 |
|
Liangsheng Yin
|
850887dc63
|
[Spec] fix EAGLE v2 verify metadata init order on non-cuda-graph path (#26244)
|
2026-05-24 18:49:33 -07:00 |
|
Mick
|
64e2b54a8f
|
[VLM] feat: accept grid_thws from preprocessed metadata for kimi (#26149)
|
2026-05-25 09:10:25 +08:00 |
|
Mick
|
72c1582d4e
|
[VLM] fix: fix only the grids from last split mm item is collected for qwen-vl (#26094)
|
2026-05-25 09:09:46 +08:00 |
|
Liangsheng Yin
|
ed179bf9b2
|
[dsv4] fix multi-step draft on non-cuda-graph path (#26239)
|
2026-05-24 17:04:18 -07:00 |
|
Ming Yang
|
85471d253d
|
Add --disable-attn-tp-gather opt-out for model-managed SP (#26047)
|
2026-05-24 15:53:51 -07:00 |
|
Lianmin Zheng
|
93fa577bb9
|
Clean up server startup log noise (#26205)
|
2026-05-24 14:35:15 -07:00 |
|
Polisetty V R K Jyothendra Varma
|
fd94bd30b8
|
[Intel GPU] DeepSeek V4 2/N: Fix tvm ffi import (#26118)
|
2026-05-24 12:59:49 -07:00 |
|
 Cheng WanandCheng Wan
|
44922de48a
|
fix(swa): downgrade translate_loc_from_full_to_swa key-change log from warning to debug (#26225)
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
|
2026-05-24 11:37:46 -07:00 |
|
 Siyuan Chenandxutizhou
|
7f45bcdd2a
|
[dsv4] support eplb (#25948)
Co-authored-by: xutizhou <xutingz@nvidia.com>
|
2026-05-24 10:09:41 -07:00 |
|
Mick
|
5c3775823e
|
[diffusion] chore: use model-aware vae channels_last_3d policy (#26214)
|
2026-05-25 00:25:34 +08:00 |
|
shuwenn
|
36eb72bf12
|
[UnifiedTree] fix: backup SWA-split parent before child under write-through (#25065)
|
2026-05-25 00:21:57 +08:00 |
|
Dongjun Na
|
9d50cd9742
|
[observability] add ServerArgs.stat_loggers for pluggable metrics backend (#24610)
Signed-off-by: Dongjun Na <kmu5544616@gmail.com>
|
2026-05-24 22:41:28 +08:00 |
|
Mick
|
b6f71d5850
|
[VLM] avoid extra cuda-ipc staging for preprocessed input (#26096)
|
2026-05-24 19:48:06 +08:00 |
|
Mick
|
6447596501
|
[VLM] feat: replace small H2D calls with a single one for qwen-vl (#26167)
|
2026-05-24 18:27:17 +08:00 |
|
Xiaoyu Zhang
|
0b65588c18
|
[diffusion] Clean up VSA attention hot path (#25514)
|
2026-05-24 16:46:03 +08:00 |
|
Mick
|
4c2b32bfbf
|
[VLM] accept precomputed multimodal metadata (#26101)
|
2026-05-24 15:43:21 +08:00 |
|
Mick
|
826a4de062
|
[srt] store req input ids as arrays (#26165)
|
2026-05-24 15:09:36 +08:00 |
|
Mick
|
d6d9f12444
|
[VLM] adopt simplified get_rope_index for image-only requests (#26100)
|
2026-05-24 11:51:24 +08:00 |
|
+2        
|
af8f66940e
|
[AMD] Dsv4/pr1 fix run time issue (#25898)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: amd-danli103 <danli103@amd.com>
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
Co-authored-by: yichiche@amd.com <jacky.cheng>
Co-authored-by: yctseng0211 <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-05-23 16:04:14 -07:00 |
|
Qiaolin Yu
|
982f67d9a6
|
Suppress cutlass-dsl noisy warning (#26169)
|
2026-05-23 13:19:14 -07:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) 
|
2de74035a5
|
[FIX][2/2] fix step3-vl/deepseek-ocr image processor error (#25403)
Co-authored-by: kousakawang <wanghanpei@bytedance.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-05-24 01:36:39 +08:00 |
|
Hanming Lu
|
a5a64a311a
|
[Spec] trtllm mha supports overlap plan stream (#25925)
|
2026-05-23 03:25:57 -07:00 |
|
Qiaolin Yu
|
cb7b57955d
|
fix tokenspeed_mla attn kernel jit (#26170)
|
2026-05-23 03:24:33 -07:00 |
|
Charles Chen
|
89ff2bc111
|
[bug fix] Fix 3 issues when using Gemma4 MTP (#26026)
|
2026-05-23 03:16:47 -07:00 |
|
 Khoa PhamandClaude Opus 4.7
|
b0ce16d0c5
|
[CP] 1/N: Support MLA Prefill Context Parallel (#23292)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-23 03:07:09 -07:00 |
|
Liangsheng Yin
|
1e59ed7443
|
compile _resolve_spec_extras gather kernels (#26129)
|
2026-05-23 02:34:41 -07:00 |
|
 Cheng WanandCheng Wan
|
83a18e687d
|
Revert "[refactor] unify cuda-graph capture/replay across attention backends (#26134)" (#26166)
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
|
2026-05-23 02:32:08 -07:00 |
|
Mick
|
774b29dade
|
[VLM] feat: early-return in mm processor if the input is preprocessed (#26117)
|
2026-05-23 17:02:44 +08:00 |
|
Mick
|
19b60a4f9e
|
[VLM] reuse pretokenized ids from preprocessed input for qwen-vl (#26116)
|
2026-05-23 16:01:04 +08:00 |
|
Zheng Wengang
|
8b9fb13c4a
|
[BugFix][EPD] adapt for qwen3.5-mtp & del duplicated logs (#24144)
|
2026-05-23 15:32:57 +08:00 |
|
 Cheng WanandCheng Wan
|
5964d30233
|
fix(swa): eliminate spurious translate_loc_from_full_to_swa warning in BCG and CG paths (#26152)
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
|
2026-05-23 00:01:12 -07:00 |
|
Xiaoyu Zhang
|
75427c9ca4
|
Route concat MLA to JIT and remove unused downcast (#25843)
|
2026-05-23 14:30:43 +08:00 |
|
 xiaobochen-amdandfanxingran
|
fd3e11973b
|
[AMD][aiter] Fix cuda_graph_kv_indices OOB under page_size>1 (#24587)
Co-authored-by: fanxingran <fanxingran@amd.com>
|
2026-05-22 23:19:46 -07:00 |
|
Shangming Cai
|
a241659d18
|
[PD] Consolidate shared logic into common backend (#25979)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-05-23 10:41:20 +08:00 |
|
Mick
|
c8cea6d4aa
|
[diffusion] feat: auto-select vae channels_last_3d (#26121)
|
2026-05-23 10:20:30 +08:00 |
|
weireweire
|
629b6c6a85
|
correct allreduce fusion and dummy_run alignment in SCATTERED MLP mode (moe_dense_tp_size=1) (#19918)
|
2026-05-22 19:18:25 -07:00 |
|
 
|
d226f75669
|
[refactor] unify cuda-graph capture/replay across attention backends (#26134)
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-22 18:51:13 -07:00 |
|
Yueming Yuan
|
208397affc
|
[RL] [Spec v2] Use stop-aware seqlen for returned topk metadata (#26126)
|
2026-05-22 18:13:16 -07:00 |
|
Qiaolin Yu
|
c112f7623a
|
Skip init_mha_chunk_metadata in trtllm_mla when not needed (#26017)
|
2026-05-22 16:34:16 -07:00 |
|
nvjullin
|
cadfa2d025
|
Support piecewise CUDA graph with NSA (#23351)
|
2026-05-22 14:39:50 -07:00 |
|
maocheng23
|
2df9e8b4b3
|
[perf] DeepSeekV3: drop redundant FP32 upcasts in trtllm MoE paths (#25189)
|
2026-05-22 14:23:57 -07:00 |
|