Xinyuan Tong
|
39955d5314
|
[MoE] Make DeepEP auto serve flashinfer_cutedsl FP4 (coerce to low_latency) + guard (#29523)
|
2026-07-24 06:59:18 +00:00 |
|
Qiaolin Yu
|
15d73f1e03
|
Fix dynamo recompile limit in allreduce and bf16 gemm (#32239)
|
2026-07-23 22:24:01 -07:00 |
|
 Polisetty V R K Jyothendra VarmaandMa Mingfei
|
319055c191
|
[Intel GPU] Add XPU Platform support (#31949)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-07-24 12:49:29 +08:00 |
|
ANSHUMAN TRIPATHY
|
f0f78a6c93
|
Add deterministic inference for eagle parity test (#30026)
|
2026-07-24 12:48:21 +08:00 |
|
Liangsheng Yin
|
99b29bf188
|
[Fix] Support ENCODER_ONLY target-verify in the trtllm_mha backend (#32178)
|
2026-07-23 20:33:39 -07:00 |
|
Cheng Wan
|
eac7c7d7cd
|
fix(attention): read per-runner kv cache dtype off model_runner (#32251)
|
2026-07-23 20:08:57 -07:00 |
|
  
|
1e10ec93b3
|
[XPU] Add XPU device support for LMCache radix cache integration (#23534)
Co-authored-by: Christopher Manteuffel <christopher.manteuffel@intel.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-07-24 08:41:01 +08:00 |
|
 Polisetty V R K Jyothendra VarmaandMa Mingfei
|
2f823a2eee
|
[Intel GPU] calculate free memory based on allocated memory for XPU (#32044)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-07-24 08:38:52 +08:00 |
|
Yihao Wang
|
433429b16a
|
diffusion: skip _save_gt_output for 3d/mesh (#32117)
|
2026-07-24 08:14:05 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
d4a0dfbc31
|
[Fix] Two root causes of the H100 deepep TBO CI break: scale-tensor use-after-free + missing non-finite quant sanitization (#32188)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-24 07:29:28 +08:00 |
|
Liangsheng Yin
|
59ef3b15cc
|
[Fix] Reserve the mamba pool's +1 padding slot in the memory budget solve (#32184)
|
2026-07-23 14:55:35 -07:00 |
|
Baizhou Zhang
|
3d0c6bf57f
|
[Fix] Fix trtllm_mla backend + fp8 kv cache without rope (#32181)
|
2026-07-23 14:55:18 -07:00 |
|
Qiaolin Yu
|
a0728ea502
|
[spec decoding] fix inkling multi layer mtp draft extend cuda graph (#32254)
|
2026-07-23 14:52:30 -07:00 |
|
Siyuan Chen
|
71fe41b6b3
|
[DSV4] Support megamoe for CP (#29569)
|
2026-07-23 14:46:32 -07:00 |
|
Qiaolin Yu
|
378aea1385
|
Fix nvfp4 online scale with pcg (#32246)
|
2026-07-23 14:42:37 -07:00 |
|
Jimmy Shong
|
410ab4fde5
|
Add return_token_ids support to completions and chat completions APIs (#30917)
|
2026-07-23 14:41:52 -07:00 |
|
Yongfei Xu
|
ebe3ab29e4
|
[DeepSeek V4] CP decode opt: slice repeat attention weights to local TP partition (#27657)
|
2026-07-23 14:06:25 -07:00 |
|
Mohammad Miadh Angkad
|
a2ddf92e61
|
[CI] Fix Mamba ServerArgs namespace (#32211)
|
2026-07-23 13:15:50 -07:00 |
|
Sam Shleifer
|
1f9d778d1b
|
Skip dist_init/nccl port prechecks when the dist init method is overridden (#31410)
|
2026-07-23 12:40:41 -07:00 |
|
Rain Jiang
|
7fe82dd02e
|
create rust workspace (#32014)
|
2026-07-23 12:02:41 -07:00 |
|
Hongkuan Zhou
|
d0b9689805
|
[Fix] Include disagg prefill waiting queue in FPM (#32122)
Signed-off-by: hongkuanz <hongkuanz@nvidia.com>
|
2026-07-23 09:30:27 -07:00 |
|
 Jialin Ouyangandhzh0425
|
70ac0c4b0e
|
[UnifiedRadixCache][mamba] Fix mamba state corruption and slot leak when load_back aborts (#30986)
Co-authored-by: hzh0425 <hzh0425@apache.org>
|
2026-07-23 18:48:31 +08:00 |
|
YAMY
|
c18919f8f3
|
[Mamba] Add a per-path cap for cached states (#31230)
|
2026-07-23 17:58:36 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
62aa85d9aa
|
[Kernel] Sweep missed dedicated kernels into kernels.ops (moe/quant siblings + dspark) (RFC #29630) (#32160)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-23 17:07:16 +08:00 |
|
Liangsheng Yin
|
f35411ee81
|
[Spec] Enable grammar overlap scheduling for STANDALONE speculative decoding (#32110)
|
2026-07-23 01:47:27 -07:00 |
|
Liangsheng Yin
|
9b853e6832
|
[Scheduler] Enable decode retraction ordering under speculative decoding (#32023)
|
2026-07-23 00:42:56 -07:00 |
|
heziiop
|
235a488c87
|
[NPU] ascend fuseep use moe ep group (#32040)
|
2026-07-23 14:07:14 +08:00 |
|
silencejade
|
09071be105
|
[NPU] [FIX] Fix performance degradation of Qwen3.5-397B-A17B (#32130)
|
2026-07-23 14:05:10 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
11b0e5c5ad
|
[Kernel] Classification cleanup: unify _jit_ naming, drop empty/model groups, add elementwise (RFC #29630) (#32148)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-23 13:47:02 +08:00 |
|
Liangsheng Yin
|
1b63155efe
|
[Perf] Skip blocks past per-request live length in full-width Triton kernels (#32109)
|
2026-07-22 22:00:25 -07:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
2d1a7be8c4
|
[Kernel] Reclassify kernel tests by ops group + move helpers out of the package (RFC #29630) (#32128)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-23 12:18:27 +08:00 |
|
 jiayisunxandMa Mingfei
|
a2935ce329
|
[XPU][GDN] add XPU path for causal_conv1d_fn and causal_conv1d_update (#31250)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-07-23 10:50:20 +08:00 |
|
Артем Савкин
|
108182cb81
|
[Bugfix] [NPU] Fix w4a8 MoE performance degradation (#32113)
|
2026-07-23 10:31:32 +08:00 |
|
zhaozx-cn
|
926530b0cf
|
[NPU]remove duplicate code (#31863)
|
2026-07-23 10:24:31 +08:00 |
|
huangtingwei
|
86e1bb584d
|
[HiCache] Add model-aware key isolation to Mooncake Store (#31920)
|
2026-07-23 10:20:13 +08:00 |
|
YanbingJiang
|
9a7ac3ecef
|
Fix unnecessary gather/scatter on CPU for non-contiguous Mamba statepool (#31754)
|
2026-07-23 09:08:45 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
99f636a86f
|
[Kernel] RFC #29630 finale: retire sglang.jit_kernel into sglang.kernels (#32072)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-23 08:35:09 +08:00 |
|
 danielafrimiandDaniel Afrimi
|
24da0e51b6
|
Warn on small Mamba chunked prefill size (#30938)
Co-authored-by: Daniel Afrimi <dafrimi@aws-dfw-cs-001-login-01.cm.cluster>
|
2026-07-22 15:09:53 -07:00 |
|
 Mohammad Miadh AngkadandBrayden Zhong
|
0c29c8fece
|
Bump FlashInfer to 0.6.15.post1 (#31927)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-07-22 14:21:59 -07:00 |
|
  
|
c20c48b8fd
|
Add 'anyOf' schema support for qwen3_coder tool call parser (#30832)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-07-22 14:14:47 -07:00 |
|
Brayden Zhong
|
e7511141ea
|
Support CuteDSL GEMM BF16 on SM100 on by default when allowed by heuristic (#30567)
|
2026-07-22 14:13:12 -07:00 |
|
Liangsheng Yin
|
98cc8d91cf
|
[Fix] Evict only the KV shortfall in evict_from_tree_cache (#32016)
|
2026-07-22 12:41:06 -07:00 |
|
Liangsheng Yin
|
5c6a29b3c3
|
[Fix] Unify pinned host pool release on graceful shutdown (#32029)
|
2026-07-22 12:29:34 -07:00 |
|
Cheng Wan
|
f5dcbe8f14
|
Revert RuntimeContext config-namespace reads/roles (#31813–#31817) (#32100)
|
2026-07-22 11:52:41 -07:00 |
|
Ke Bao
|
40ac197fb2
|
Fix get_server_args import lint error (#32096)
|
2026-07-23 00:09:25 +08:00 |
|
Raghavendra Vedula
|
4eaa5ca651
|
Treat partial_json_parser AssertionError as incomplete JSON (#31975)
|
2026-07-22 23:25:25 +08:00 |
|
Raghavendra Vedula
|
a9497e8d73
|
Guard min_new_tokens penalizer against None eos_token_id (#31973)
|
2026-07-22 23:24:25 +08:00 |
|
Chengze Fan
|
40b2119b23
|
[AMD] Cache AITER expert mask across decode (#31889)
|
2026-07-22 07:50:23 -07:00 |
|
  
|
e8e765b9d6
|
[AMD] Add fused all-reduce RMSNorm per-group quant for Qwen3.5 FP8 (#24651)
Co-authored-by: jacky.cheng <yichiche@amd.com>
Co-authored-by: yctseng0211 <yctseng@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-07-22 07:33:03 -07:00 |
|
Mohammad Miadh Angkad
|
b855efd9e6
|
Fix Inkling kernel imports after migration (#32076)
|
2026-07-22 22:09:01 +08:00 |
|