2394b231c2
Fix Mistral3 retaining every vision-tower layer to read one ( #39185 )
...
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com >
2026-09-18 16:11:54 -07:00
fa521e2758
[MM] Copy placeholder ids to CUDA asynchronously ( #40010 )
...
Co-authored-by: metamergebot <metamergebot@users.noreply.github.com >
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com >
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com >
2026-09-18 16:10:37 -07:00
Zhiqiang Xie and cctry
ceb1d2e580
[PD] Enable optimistic prefill with buffer-only L3 write-through HiCache ( #40043 )
...
Co-authored-by: cctry <csycfl@gmail.com >
2026-09-18 16:04:15 -07:00
5e4b94b134
[MM] Skip VMM error gathers for text-only requests ( #40005 )
...
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com >
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com >
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com >
2026-09-18 16:00:43 -07:00
hunhokim and Hun-ho Kim
803f0c93d2
Fix DSA partial DP-TP mode log to use derived attn_tp_size ( #39871 )
...
Co-authored-by: Hun-ho Kim <hunho.kim@samsung.com >
2026-09-18 15:54:45 -07:00
Jialin Ouyang
f3851486cb
[Perf] Fuse SWA page lookup and mapping clear ( #38948 )
2026-09-18 15:51:07 -07:00
81a199f56a
[HiCache] Read the in-flight buffer backup's node id from its snapshot in sanity_check ( #40013 )
...
Co-authored-by: Pranjal Shankhdhar <pranjalssh@meta.com >
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com >
2026-09-18 15:47:56 -07:00
0e5347db82
Support MXFP8 and deferred route weighting in DeepEP v2 ( #40030 )
...
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com >
Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com >
Co-authored-by: pranjalssh <14260275+pranjalssh@users.noreply.github.com >
2026-09-18 15:40:38 -07:00
Liangsheng Yin
6cc9090d1f
[mem_cache] Release up to owned_kv_len on radix cache insert ( #40075 )
2026-09-18 15:37:52 -07:00
Zhiqiang Xie
a0534f8cca
[HiCache] Stop arming a prefetch retry for a too-short storage span ( #40042 )
2026-09-18 15:33:58 -07:00
Sam (Kesen Li)
d346b214fb
feat(kv-cache): support SM100 NVFP4 GenMHA and speculative decoding ( #36340 )
2026-09-18 14:50:46 -07:00
Lianmin Zheng and Jialin Ouyang
f5a1434700
[HiCache] Document transfer arguments ( #40239 )
...
Clarify the legacy host_indices argument and label Mamba test arguments,
including the current staging_tokens parameter. Executable code is unchanged.
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com >
2026-09-18 14:08:19 -07:00
amd-danli103
cd4dd81c22
[AMD][DSV4] fix: drop shadowing local get_exec import that breaks model startup on ROCm ( #40186 )
2026-09-18 13:28:38 -07:00
metamergebot and cctry
6a9c7001d3
[Logprob] Borrow graph-pool memory for input logprob logits construction ( #40007 )
...
Co-authored-by: cctry <csycfl@gmail.com >
2026-09-18 13:18:01 -07:00
Jialin Ouyang
da2f434951
[Spec] Add explicit prefill shared-read capability for plugins ( #39502 )
2026-09-18 13:11:01 -07:00
Lianmin Zheng and raghotham
248c202b46
Use runtime token widths for Triton speculative verification ( #39859 )
...
Co-authored-by: raghotham <853234+raghotham@users.noreply.github.com >
2026-09-18 11:07:55 -07:00
Lianmin Zheng
6bd1a0af1d
Add registration for external model configurations ( #39452 )
2026-09-18 10:06:24 -07:00
Shuwen Wang
191172fa74
[Unified Tree] fix: exempt host-locked aux nodes from the sanity_check host-LRU check ( #39980 )
2026-09-18 15:29:29 +00:00
Xiaoyu Zhang
7714b182f2
[Bugfix] Fix top-1 MoE routing with non-unit scaling ( #40187 )
2026-09-18 22:53:13 +08:00
81363bf8cb
[kernel] Share the warp vectorized copy and enforce its alignment ( #36176 )
...
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com >
Co-authored-by: BBuf <1182563586@qq.com >
2026-09-18 22:40:54 +08:00
a6cf05817f
dsv4.1: remaining model and runtime integration ( #38798 )
...
Co-authored-by: BBuf <1182563586@qq.com >
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com >
Co-authored-by: Xiaoyu Zhang <xiaoyu.zhang@radixark.ai >
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com >
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai >
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com >
Co-authored-by: Zhichen Zeng <zczeng@uw.edu >
Co-authored-by: Ke Bao <ispobaoke@gmail.com >
2026-09-18 02:55:30 -07:00
Liangsheng Yin
1b200ffaaa
[Quant] Serve 32-wide-K ue8m0 block-FP8 linears through the FlashInfer MXFP8 GEMMs ( #40039 )
2026-09-18 02:51:12 -07:00
iridiumine
6c7c5e78de
[NPU] Run arch35 block-FP8 dense linears on the native MXFP8 GEMM ( #39823 )
2026-09-18 17:19:31 +08:00
iridiumine
7e6d5cbfac
[NPU] Gate DFlash replay metadata refresh behind spec_algorithm check ( #39879 )
2026-09-18 17:07:07 +08:00
iridiumine
d6090f92bf
[NPU] Fuse MXFP4 W4A8 MoE gmm1 + swiglu + requant into one kernel ( #39881 )
2026-09-18 16:50:49 +08:00
2dee23a876
[ROCm][diffusion] Enable fused qk norm and rope on ROCm ( #35573 )
...
Co-authored-by: jacky.cheng <yichiche@amd.com >
Co-authored-by: Claude <noreply@anthropic.com >
2026-09-18 16:40:32 +08:00
Shu Wang
1e8699fda3
[NVIDIA][comm] Merge EP+MoE-TP post-experts all-reduces into one _TP reduction ( #32963 )
2026-09-18 01:35:10 -07:00
zhaozx-cn
8ac39c66d8
[NPU] support kimi k3 on A5 and improve performance ( #39589 )
2026-09-18 16:33:54 +08:00
Rumit Desai and Xiaoyu Zhang
1fdd6c8921
[Runtime] Let out-of-tree platforms provide full graph backends ( #37969 )
...
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com >
2026-09-18 15:42:46 +08:00
Ziang Li
c46bf5e990
[MoE] Disable FlashInfer fused finalize by default for numerical accuracy ( #40105 )
2026-09-18 00:15:57 -07:00
Shuwen Wang
0be8a0af0e
[DSV4] fix: keep the TileLang JIT cache under SGLANG_CACHE_DIR ( #39364 )
2026-09-18 00:06:35 -07:00
iridiumine
f86f60081d
[NPU] Adapt hicache for K3 hybrid models ( #39415 )
2026-09-18 14:48:20 +08:00
Jensen
3ce3b4969f
[NPU] Avoid repeated BF16 wo_a weight transposes in DeepSeek-V4 decode ( #39919 )
2026-09-18 09:10:43 +03:00
Bojiang Li and Xiaoyu Zhang
3d1b9e7549
[Bugfix] Include SM121 in DeepGEMM packed-scale selection ( #39482 )
...
Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com >
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com >
2026-09-18 13:57:14 +08:00
Hanming Lu and hanminglu
6e1d9c1464
[profiler] Ignore PREBUILT batches in profile-by-stage ( #40098 )
...
Co-authored-by: hanminglu <hanminglu@fb.com >
2026-09-17 22:33:55 -07:00
Chi McIsaac and Mick Qian
6215aecd51
[diffusion] feat: add metrics support ( #19084 )
...
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com >
2026-09-18 13:22:03 +08:00
Yuwei An
65ef55e2a8
[Scheduler] Add shortest-prefill-first scheduling ( #40024 )
2026-09-17 21:18:58 -07:00
826d5170ae
[Bug] Guard FlashInfer CUTLASS MoE against 0-token inputs ( #38780 )
...
Co-authored-by: Cursor <cursoragent@cursor.com >
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com >
2026-09-18 11:03:27 +08:00
Zixin Huang and Claude Fable 5.1
f447bb7080
[Kernel] Coalesce the KDA CuTe DSL decode state transpose: ~3x faster, bit-identical ( #39680 )
...
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com >
2026-09-18 10:53:13 +08:00
db39b7f961
[Fix] Guard conditional top-logprob keys in the completions echo path ( #34776 )
...
Co-authored-by: James Liu <jamesl@modal.com >
Co-authored-by: hnyls2002 <lsyincs@gmail.com >
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com >
2026-09-17 19:47:26 -07:00
Nan Jiang
740f57a02c
[Spec] Fix CDF boundary handling in TreeSpeculativeSamplingTargetOnly ( #35798 )
2026-09-17 19:41:52 -07:00
William Arnold and ishandhanani
0214954f26
[gRPC] Stream engine state changes ( #39915 )
...
Signed-off-by: William Arnold <7565007+Aphoh@users.noreply.github.com >
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com >
2026-09-17 19:17:25 -07:00
YanbingJiang and Ma Mingfei
a407915c17
[CPU] Avoid prefill CP predicates during decode graph capture ( #39690 )
...
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com >
2026-09-18 09:21:05 +08:00
Liangsheng Yin
f65c70bb7d
[Kernel] Move CUDA and ROCm speculative kernels to JIT ( #40033 )
2026-09-17 17:27:26 -07:00
Nan Jiang
20518d8518
[Scheduler] Align RadixCache no-insert cleanup with kv_len_to_handle ( #35204 )
2026-09-17 17:00:59 -07:00
Xiaoyu Zhang
7bc9152447
[Test] Consolidate kernel tests under plural kernels tree ( #39966 )
2026-09-18 07:37:48 +08:00
Liangsheng Yin
1f0c73e9bd
[DSV4] Generalize attention metadata, sparse prefill, and KV pool over compress ratios ( #39921 )
2026-09-17 15:55:19 -07:00
Ren Yuzhou and Ren Yuzhou
4f52a27563
[AMD] Fix DSV4 FP4 dequant path for AITER on ROCm ( #35123 )
...
Co-authored-by: Ren Yuzhou <yuzhouo7@users.noreply.github.com >
2026-09-17 13:51:22 -07:00
Thomas Wang
6c73368c32
[AMD][DSV4] Enable hicache on deepseek-v4 fp8 unified attn ( #37778 )
2026-09-17 11:29:23 -07:00
1f60ddef5d
[PD] Introduce runtime role switching between prefill and decode ( #28403 )
...
Signed-off-by: huanglong <huanglong@linux.alibaba.com >
Signed-off-by: inkcherry <mingzhi.liu@amd.com >
Co-authored-by: huanglong <huanglong@linux.alibaba.com >
Co-authored-by: Shangming Cai <csmthu@gmail.com >
Co-authored-by: Huang Long <121648372+LLLL114@users.noreply.github.com >
2026-09-18 01:45:12 +08:00