Commit Graph
18253 Commits
Author SHA1 Message Date
ashwini rathiandClaude Opus 4.7 b510881157 [Fix][Qwen-VL] Normalize <image> sentinel on artifact fast path (#39278)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-09-15 13:19:59 +08:00
6c514ab025 Force reasoning mode for GLM-5.3 chat templates (#39227)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-09-15 13:16:26 +08:00
Yanbin Jiang ebd37705e4 [PD][LoRA] Gate decode admission on adapter slots (#39332) 2026-09-15 13:11:04 +08:00
ishandhanani 3f871a246c feat(agent sessions): attribute stored KV cache blocks to sessions (#37482)
Signed-off-by: Ishan Dhanani <ishandhanani@gmail.com>
2026-09-14 21:34:49 -07:00
Byron HsuandByron Hsu c9fbe5f655 [MM] Add flag to force Kimi image preprocessing onto CPU (#39148)
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
2026-09-14 21:19:39 -07:00
45b511b2ef [Router] Honor KV-event storage tiers in the cache-aware tree (1/4) (#39108)
Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 21:18:15 -07:00
Shangming Cai 4b186cfea5 [CI][PD] Skip the flaky decode HiCache file-backend disaggregation test (#39368) 2026-09-15 11:19:03 +08:00
amote-i bdf8886ad3 [NPU] [DOC] Rename NPU hardware to Ascend A2/A3 Series product (#39389) 2026-09-15 10:33:14 +08:00
Aurick QiaoandAurick Qiao 5dde6e8f02 Fix MegaMoE buffer allocation and caching for effective SM budgets (#39223)
Co-authored-by: Aurick Qiao <6137920+aurickq@users.noreply.github.com>
2026-09-15 10:24:02 +08:00
Yuan Luoandluoyuan.luo 99060191e7 [KDA] Support ReplaySSM ring-write in the fused chain-verify kernel (#36821)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-09-15 10:22:41 +08:00
JoeandXiaoyu Zhang a23fd557ed [Kernel] Add OOT dispatch for clamp position (#38687)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-09-15 10:22:21 +08:00
Kan WuandClaude Fable 5.1 e89d8facab [sgl-router] Prepare dynamo-render dependencies (#39457)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 19:08:16 -07:00
a25f213bc4 [diffusion] docs: give the RTX 5090 its own H3 recipe, measured on a physical desktop (#39373)
Co-authored-by: Mick Qian <mickqian@radixark.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 09:15:13 +08:00
Charles xu fc4193a63c [HiCache] Remove duplicate benchmark result fields (#38097) 2026-09-14 17:56:26 -07:00
Kan Wu 07e1918924 [Rust] Use Dynamo native renderers when chat templates are missing (#38939) 2026-09-14 17:29:10 -07:00
Xingyu Liu 8874c51a96 [Benchmark] Add an opt-out for the token-capacity check (#39284) 2026-09-14 17:17:12 -07:00
raghothamandLianmin Zheng c0b8725f5a Fix /model_info serialization when a config value is a class (#39237)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-09-14 17:10:44 -07:00
Yang Liuandyangliu991 276663a79d [model-loader] Split weight loading from postprocessing (#34981)
Co-authored-by: yangliu991 <yangliu991@fb.com>
2026-09-14 17:10:19 -07:00
Lianmin Zheng a2b4e8888f Allow CUDA VMM feature transport with the Rust frontend (#39347) 2026-09-14 17:09:49 -07:00
faceless void 6e755e4114 [Logging] Downgrade missing TokenizerManager request state log to warning (#36625)
Signed-off-by: syd520zy <529477025@qq.com>
2026-09-14 16:18:56 -07:00
AMD-yanfeiwang 0163f8ff74 [AMD] Fix registered HiCache host pointer aliases (#35233) 2026-09-14 15:49:05 -07:00
Baizhou Zhang 2fca6d69aa bumping sgl-deep-gemm to 0.2.0 (#39371) 2026-09-14 15:30:31 -07:00
Richard WangandRichard Wang 710a044caf [dLLM] Add rwang5203 as code owner and grant CI permissions (#39483)
Co-authored-by: Richard Wang <11150595+rwang5203@users.noreply.github.com>
2026-09-14 15:27:15 -07:00
cctry dad8c074e7 Scope prefetch cache state to the request attempt (#39318) 2026-09-14 15:03:03 -07:00
Qiaolin Yuandehuaa d72e59508b Reland fix(qsa): clamp the compress gather to the rows (#38346) (#39446)
Co-authored-by: ehuaa <ehuamail1@gmail.com>
2026-09-14 14:05:41 -07:00
paulzhang-tm 9128d57966 [Spec] Allow speculative workers to stage prefill shared reads (#38554) 2026-09-14 14:03:27 -07:00
Byron HsuandByron Hsu 5c2de3f355 [PD] Preserve the prefill rank during rebootstrap (#39357)
Co-authored-by: Byron Hsu <byronhsu@users.noreply.github.com>
2026-09-14 09:20:37 -07:00
pllimax 7465e42b7a [NPU][CI] Add CANN 9.1.0 and Ascend a5 nightly suites (#38833) 2026-09-14 22:39:42 +08:00
pllimax 433c999dd0 [NPU][CI] Fix sglang.test.ascend import failure in multi-node e2e pods (#39403) 2026-09-14 22:23:47 +08:00
ChangLiu0709 242d8a70c0 [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260913, enable TOPK_V2 (#39406) 2026-09-14 21:50:52 +08:00
pllimax d5f1c593c1 [NPU] Remove temperature/top_p from Qwen3.5-397B-A17B perf test (#39047) 2026-09-14 21:48:50 +08:00
Kurt Shuster f81fbc749a [Fix] Merge adjacent KV-row frees so a mid-page split under DCP cannot double-free (#38941) 2026-09-14 21:44:32 +08:00
Aurick QiaoandAurick Qiao 2123aca87e [Fix] Wait for PDL before reading DeepSeek V4 K cache locations (#38409)
Co-authored-by: Aurick Qiao <6137920+aurickq@users.noreply.github.com>
2026-09-14 21:04:28 +08:00
iridiumine 87db743021 [NPU] Fix device mismatch in SWA mask for DSpark verify graph capture (#39353) 2026-09-14 19:36:33 +08:00
Liangsheng Yin 66c7bc838e [misc] Revert #38346, #33426, #39061 and #39219 (#39405) 2026-09-14 03:04:07 -07:00
amd-danli103 5aa9b8fb3e [AMD][DSV4] feat: enable fp8 two-pool unified_kv on gfx950 (#37413) 2026-09-14 02:49:11 -07:00
Theresa Shan 95140a7b0c [docs] DeepSeek-V4: MI355X PD disaggregation recipes for all three strategies (#39396) 2026-09-14 02:11:51 -07:00
Jacob0226andThomas Wang 5200508b0f [AMD][gfx95] Fill the chunked-prefill compute budget exactly (#32888)
Co-authored-by: Thomas Wang <thomawan@amd.com>
2026-09-14 00:35:15 -07:00
3eeb7d37f9 [AMD] gfx950 assembly attention: length-aware split-KV for dynamic workload (#39172)
Co-authored-by: Zijie Chen <300606707+zijiecode@users.noreply.github.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>
2026-09-13 23:48:03 -07:00
ZeyuanChen2000 a4781c9fe5 [NPU] Fix error due to missing parameter quant_linear passing (#37384) 2026-09-14 14:37:44 +08:00
YC Yen-Ching Tseng ce66ba2844 [AMD] Fix AITER FP8-Q unified-attention Test (#39360) 2026-09-13 23:35:24 -07:00
Bingxu Chen edf9584be9 [AMD][CI] Disable Wave attention test file on ROCm 10 (#39356) 2026-09-13 23:34:18 -07:00
Shuwen Wang 0d95a9c1ff [HiCache] Keep the file backend temp file name within NAME_MAX (#38925) 2026-09-14 05:12:28 +00:00
jacky.cheng 2f5cc8e33e [AMD] Align Qwen3.5 MI355X cookbook with AttnFP8-V2 and HiCache direct / page_first_direct (#39358) 2026-09-14 12:46:17 +08:00
Mohammad Miadh Angkadandmmangkad 2fd835b9c1 [Fix] Don't write conv state from the fused KDA verify kernel (#39219)
Co-authored-by: mmangkad <mohammad.angkad@radixark.ai>
2026-09-13 21:26:57 -07:00
Jan Bernlöhr 60f6f03409 Fix MUSA detection in compiled prefill path (#39061) 2026-09-13 21:23:51 -07:00
Kurt ShusterandBaizhou Zhang f2111715cd [fa] Make the FlashAttention backend extensible by subclasses (#33426)
Signed-off-by: Kurt Shuster <kurt@thinkingmachines.ai>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-09-13 21:21:25 -07:00
39e147443b [Session] Fix image append positions and parent metadata (#39145)
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
Co-authored-by: Manik Singhal <manikvsinghal.pub@gmail.com>
2026-09-13 21:06:47 -07:00
Jackey Hua 5a132c061b [Perf] trtllm_mla: reuse the fused fp8 KV/Q prepare on target verify (#39232) 2026-09-13 20:59:00 -07:00
42b5af8c62 [diffusion] docs: refresh the VDN-H3 on b200 numbers in cookbook (#39244)
Co-authored-by: haochengxi <xihc@berkeley.edu>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Kevin Mi <mikevin920@yahoo.com>
2026-09-14 11:50:40 +08:00