Commit Graph
18260 Commits
Author SHA1 Message Date
amote-i 575163ff32 [NPU] [DOC] delete unsupported models in npu docs (#39555) 2026-09-15 14:55:17 +08:00
06992c92ad [Fix] Aggregate all IPC weight update responses (#39534)
Signed-off-by: bojiang-li <327132355+bojiang-li@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-09-14 23:25:46 -07:00
Liangsheng Yin 860fa83a9f [CI] Wait for a killed test server's GPU memory before the next launch (#39545) 2026-09-14 23:24:37 -07:00
Cui Lily e687b8d6af [CPU] Implement fused QK Norm and RoPE kernels (#37748)
Signed-off-by: Cui, Lily <lily.cui@intel.com>
2026-09-15 14:17:57 +08:00
Khoa PhamandClaude Fable 5.1 2c37b90ad6 [Fix][DCP] Localize widened KV ids in MLA retraction CPU backup/restore (#39487)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 13:33:41 +08:00
ddd4600197 [Fix] HiCache startup ImportError on the pinned kernel wheel (#39516)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-09-14 22:31:19 -07:00
Shuwen Wang 5298d85218 fix: stop shadowing the DSpark shared-experts fusion guard (#39366) 2026-09-14 22:24:49 -07:00
ashwini rathiandClaude Opus 4.7 b510881157 [Fix][Qwen-VL] Normalize <image> sentinel on artifact fast path (#39278)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-09-15 13:19:59 +08:00
6c514ab025 Force reasoning mode for GLM-5.3 chat templates (#39227)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-09-15 13:16:26 +08:00
Yanbin Jiang ebd37705e4 [PD][LoRA] Gate decode admission on adapter slots (#39332) 2026-09-15 13:11:04 +08:00
ishandhanani 3f871a246c feat(agent sessions): attribute stored KV cache blocks to sessions (#37482)
Signed-off-by: Ishan Dhanani <ishandhanani@gmail.com>
2026-09-14 21:34:49 -07:00
Byron HsuandByron Hsu c9fbe5f655 [MM] Add flag to force Kimi image preprocessing onto CPU (#39148)
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
2026-09-14 21:19:39 -07:00
45b511b2ef [Router] Honor KV-event storage tiers in the cache-aware tree (1/4) (#39108)
Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 21:18:15 -07:00
Shangming Cai 4b186cfea5 [CI][PD] Skip the flaky decode HiCache file-backend disaggregation test (#39368) 2026-09-15 11:19:03 +08:00
amote-i bdf8886ad3 [NPU] [DOC] Rename NPU hardware to Ascend A2/A3 Series product (#39389) 2026-09-15 10:33:14 +08:00
Aurick QiaoandAurick Qiao 5dde6e8f02 Fix MegaMoE buffer allocation and caching for effective SM budgets (#39223)
Co-authored-by: Aurick Qiao <6137920+aurickq@users.noreply.github.com>
2026-09-15 10:24:02 +08:00
Yuan Luoandluoyuan.luo 99060191e7 [KDA] Support ReplaySSM ring-write in the fused chain-verify kernel (#36821)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-09-15 10:22:41 +08:00
JoeandXiaoyu Zhang a23fd557ed [Kernel] Add OOT dispatch for clamp position (#38687)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-09-15 10:22:21 +08:00
Kan WuandClaude Fable 5.1 e89d8facab [sgl-router] Prepare dynamo-render dependencies (#39457)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-14 19:08:16 -07:00
a25f213bc4 [diffusion] docs: give the RTX 5090 its own H3 recipe, measured on a physical desktop (#39373)
Co-authored-by: Mick Qian <mickqian@radixark.ai>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 09:15:13 +08:00
Charles xu fc4193a63c [HiCache] Remove duplicate benchmark result fields (#38097) 2026-09-14 17:56:26 -07:00
Kan Wu 07e1918924 [Rust] Use Dynamo native renderers when chat templates are missing (#38939) 2026-09-14 17:29:10 -07:00
Xingyu Liu 8874c51a96 [Benchmark] Add an opt-out for the token-capacity check (#39284) 2026-09-14 17:17:12 -07:00
raghothamandLianmin Zheng c0b8725f5a Fix /model_info serialization when a config value is a class (#39237)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-09-14 17:10:44 -07:00
Yang Liuandyangliu991 276663a79d [model-loader] Split weight loading from postprocessing (#34981)
Co-authored-by: yangliu991 <yangliu991@fb.com>
2026-09-14 17:10:19 -07:00
Lianmin Zheng a2b4e8888f Allow CUDA VMM feature transport with the Rust frontend (#39347) 2026-09-14 17:09:49 -07:00
faceless void 6e755e4114 [Logging] Downgrade missing TokenizerManager request state log to warning (#36625)
Signed-off-by: syd520zy <529477025@qq.com>
2026-09-14 16:18:56 -07:00
AMD-yanfeiwang 0163f8ff74 [AMD] Fix registered HiCache host pointer aliases (#35233) 2026-09-14 15:49:05 -07:00
Baizhou Zhang 2fca6d69aa bumping sgl-deep-gemm to 0.2.0 (#39371) 2026-09-14 15:30:31 -07:00
Richard WangandRichard Wang 710a044caf [dLLM] Add rwang5203 as code owner and grant CI permissions (#39483)
Co-authored-by: Richard Wang <11150595+rwang5203@users.noreply.github.com>
2026-09-14 15:27:15 -07:00
cctry dad8c074e7 Scope prefetch cache state to the request attempt (#39318) 2026-09-14 15:03:03 -07:00
Qiaolin Yuandehuaa d72e59508b Reland fix(qsa): clamp the compress gather to the rows (#38346) (#39446)
Co-authored-by: ehuaa <ehuamail1@gmail.com>
2026-09-14 14:05:41 -07:00
paulzhang-tm 9128d57966 [Spec] Allow speculative workers to stage prefill shared reads (#38554) 2026-09-14 14:03:27 -07:00
Byron HsuandByron Hsu 5c2de3f355 [PD] Preserve the prefill rank during rebootstrap (#39357)
Co-authored-by: Byron Hsu <byronhsu@users.noreply.github.com>
2026-09-14 09:20:37 -07:00
pllimax 7465e42b7a [NPU][CI] Add CANN 9.1.0 and Ascend a5 nightly suites (#38833) 2026-09-14 22:39:42 +08:00
pllimax 433c999dd0 [NPU][CI] Fix sglang.test.ascend import failure in multi-node e2e pods (#39403) 2026-09-14 22:23:47 +08:00
ChangLiu0709 242d8a70c0 [AMD] GLM-5.2 MI355X MXFP4: bump image to 20260913, enable TOPK_V2 (#39406) 2026-09-14 21:50:52 +08:00
pllimax d5f1c593c1 [NPU] Remove temperature/top_p from Qwen3.5-397B-A17B perf test (#39047) 2026-09-14 21:48:50 +08:00
Kurt Shuster f81fbc749a [Fix] Merge adjacent KV-row frees so a mid-page split under DCP cannot double-free (#38941) 2026-09-14 21:44:32 +08:00
Aurick QiaoandAurick Qiao 2123aca87e [Fix] Wait for PDL before reading DeepSeek V4 K cache locations (#38409)
Co-authored-by: Aurick Qiao <6137920+aurickq@users.noreply.github.com>
2026-09-14 21:04:28 +08:00
iridiumine 87db743021 [NPU] Fix device mismatch in SWA mask for DSpark verify graph capture (#39353) 2026-09-14 19:36:33 +08:00
Liangsheng Yin 66c7bc838e [misc] Revert #38346, #33426, #39061 and #39219 (#39405) 2026-09-14 03:04:07 -07:00
amd-danli103 5aa9b8fb3e [AMD][DSV4] feat: enable fp8 two-pool unified_kv on gfx950 (#37413) 2026-09-14 02:49:11 -07:00
Theresa Shan 95140a7b0c [docs] DeepSeek-V4: MI355X PD disaggregation recipes for all three strategies (#39396) 2026-09-14 02:11:51 -07:00
Jacob0226andThomas Wang 5200508b0f [AMD][gfx95] Fill the chunked-prefill compute budget exactly (#32888)
Co-authored-by: Thomas Wang <thomawan@amd.com>
2026-09-14 00:35:15 -07:00
3eeb7d37f9 [AMD] gfx950 assembly attention: length-aware split-KV for dynamic workload (#39172)
Co-authored-by: Zijie Chen <300606707+zijiecode@users.noreply.github.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>
2026-09-13 23:48:03 -07:00
ZeyuanChen2000 a4781c9fe5 [NPU] Fix error due to missing parameter quant_linear passing (#37384) 2026-09-14 14:37:44 +08:00
YC Yen-Ching Tseng ce66ba2844 [AMD] Fix AITER FP8-Q unified-attention Test (#39360) 2026-09-13 23:35:24 -07:00
Bingxu Chen edf9584be9 [AMD][CI] Disable Wave attention test file on ROCm 10 (#39356) 2026-09-13 23:34:18 -07:00
Shuwen Wang 0d95a9c1ff [HiCache] Keep the file backend temp file name within NAME_MAX (#38925) 2026-09-14 05:12:28 +00:00