Commit Graph
1622 Commits
Author SHA1 Message Date
Guangda LiuandGuangda Liu 04c0913434 [HiSparse] Add MHA hisparse support for MiniMax M3 (#31446)
Co-authored-by: Guangda Liu <bingps@users.noreply.github.com>
2026-09-22 13:28:03 +08:00
YAMY 9b59fc5db5 [ModelOpt][PP] Keep BF16 shared experts out of the NVFP4 fusion so TP1 pipeline stages can load (#40628) 2026-09-21 21:45:58 -07:00
e1daf68304 [AMD] [GLM-5.3-Flash Day 0] Honor fused and per-expert names in quark exclude (#39317)
Co-authored-by: Yikai Zhang <ykzhang12@gmail.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: Kevin Mi <mikevin920@yahoo.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 21:26:42 -07:00
Khoa PhamandQiaolin Yu 018b73c7a0 [PD] Pack draft KV head slices for DCP transfers (#40500)
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
2026-09-21 21:11:27 -07:00
b44e248682 [AMD] [GLM-5.3-Flash Day 0] Enable FP8 and Quark MXFP4 MoE on gfx950 (#38546)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Kevin Mi <mikevin920@yahoo.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 21:07:37 -07:00
jthomson04 15ba54bd5d perf(engine): avoid timed waits for Engine responses (#39486)
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
2026-09-21 20:42:03 -07:00
Cheng Wan 1d025491f3 [Test] Set DP size in the mocked Metal profiler test (#40667) 2026-09-21 19:57:52 -07:00
jacky.cheng bc30fa1759 [AMD][Fix] AgentX HIP TPOT regression when SGLANG_SIMULATE_ACC_LEN is set (#40598) 2026-09-21 19:16:15 -07:00
Cheng Wan 9eda772a21 [Test] Handle tied top-k indices in graph-pool logprob regression (#40661) 2026-09-21 19:02:43 -07:00
Khoa Pham c4d3770a68 [Kimi K3] Fix CUDA graph stream explosion (#40640) 2026-09-21 17:49:50 -07:00
042b6a488f [AMD] [GLM-5.3-Flash Day 0] Enable zero-RoPE MHA prefill on ROCm (#39338)
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
2026-09-21 17:38:07 -07:00
avalliappan-nvidia 61d0cf2074 [Spec] Windowed draft-decode attention for built-in EAGLE / MTP drafts (#32673) 2026-09-22 08:17:57 +08:00
Vedant V JhaveriandCopilot 9fdb71732a Avoid materializing GDN QKV tensors during target verification (#33778)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-21 17:04:18 -07:00
Cheng Wan 506698761d [unified-memory] Hierarchical cache for every unified pool shape (#37507) 2026-09-21 16:50:37 -07:00
YAMY 0229025127 [Spec][PP] Launch extend microbatches before the spec output exchange (#40499) 2026-09-21 15:47:03 -07:00
Cheng Wan acac4dd9d9 [Refactor] Clean up parallel runtime comments (#40632) 2026-09-21 14:32:22 -07:00
Cheng Wan bccf691b22 Bringing the parallel runtime up becomes a phase, not a side effect (#40345) 2026-09-21 12:29:50 -07:00
Cheng Wan 1d3243d05f Take the parallel getters off the package's public surface (#40344) 2026-09-21 12:27:50 -07:00
Cheng Wan 970e946e4f Retire the per-runner parallel record (#40343) 2026-09-21 12:26:40 -07:00
Cheng Wan 73f071db52 Deprecate the parallel getters the context answers, and ratchet them shut (#40342) 2026-09-21 12:25:32 -07:00
Cheng Wan 65be3fa71a A runner and the objects it builds freeze the placement they describe (#40341) 2026-09-21 12:24:17 -07:00
Cheng Wan 2d0e94e3a3 Check the topology identities where the layout is written, and build at the published widths (#40340) 2026-09-21 12:22:59 -07:00
Cheng Wan 0db1a93adb State the draft's whole topology in its scope, and read the rest from the context (#40339) 2026-09-21 12:19:38 -07:00
ae7a516ba7 feat: use XGrammar V4.1 DSML parameter constraints (#39026)
Co-authored-by: yuchuan <yuchuan.7streams@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-21 12:12:28 -07:00
cctry 7a6191c4b9 Preallocate HiCache MHA staging before post-capture KV sizing (#40256) 2026-09-21 10:44:29 -07:00
Sage 14e9c40a72 [Observability] Expose python/rust frontend identity in /server_info (#39993)
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com>
2026-09-21 23:32:52 +08:00
b63f8416b3 [Feature] Gigachat 3.5 support (#29189)
Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>
Co-authored-by: Viacheslav Barinov <vvadbarinov@sberbank.ru>
Co-authored-by: Viacheslav <viacheslav.teh@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-09-21 14:37:55 +08:00
kpjeeja b54d5b7c7b disaggregation: Fix FakeKVSender queue accumulation (#28652)
Signed-off-by: KP, Jeeja <jeeja.kp@intel.com>
2026-09-21 14:27:05 +08:00
skyler-apdx f5f3c38aad [Fix] Preserve YaRN scaling when extending rotary caches (#38786) 2026-09-21 14:14:05 +08:00
AMRUTHA MandMa Mingfei d20cd9d77f [XPU]Enable HiSparse hierarchical sparse KV cache on Intel XPU (#32792)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-09-21 14:07:49 +08:00
jianzhao-xu 62ba964848 Fix: post-load staging regression breaks offload meta/sharded_gpu modes (#38779) 2026-09-21 11:19:07 +08:00
Xueshen Liu ab03a8e7eb [Perf] Fork-safe import: no CUDA context at import time, lighter argument parsing (#40201) 2026-09-21 10:48:35 +08:00
Liangsheng Yin 76a9065bef [Fix] Raise on undelivered embeddings in send_with_url, fix broken tests (#40502) 2026-09-20 17:35:27 -07:00
42875bcd2a fix(modelopt): dispatch NVFP4 MoE on the cached backend, not the live global (#38932)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
2026-09-20 19:28:54 -04:00
f31a7bd45c Use pinned memory for asynchronous sampling metadata transfers (#39777)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-09-20 15:31:31 -07:00
Harmya Bhatt 95521da18d [DeepSeek-V4.1] Bound dense prefill indexer memory (#40217) 2026-09-20 14:43:19 -07:00
d229952e25 [Fix] Preserve model runner contracts in prefill CUDA graphs (#35452)
Co-authored-by: Oasis-Git <ayw.sirius19@gmail.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
2026-09-20 13:53:14 -07:00
Khoa Pham b3e4d198af [PD] Bound cached-prefix DCP transfers by pack capacity (#40376) 2026-09-21 01:17:15 +08:00
Liangsheng Yin 0024efa0de [CI] Derive registered-test kind from the registry call instead of the path (#40294) 2026-09-20 02:01:22 -07:00
Liangsheng Yin dc002c85fc [Test] Fix OOT DFlash hook test resolving the draft config over the network (#40427) 2026-09-20 01:25:34 -07:00
Shuwen WangandSeokhoon Kang 9f3d275940 [HiCache] Fix sparse hybrid transfer layer IDs (#37870)
Co-authored-by: Seokhoon Kang <sh.kang@postech.ac.kr>
2026-09-20 16:24:24 +08:00
amd-danli103andHAI e54009240a [AMD][DSV4] feat: enable DSpark with fp8 unified_kv on gfx950 (#38901)
Co-authored-by: HAI <hixiao@gmail.com>
2026-09-20 01:16:39 -07:00
Ke Bao a8a4d86be9 Remove swa and mamba radix cache (#40313) 2026-09-20 16:16:27 +08:00
Qiaolin Yu f4c256354c [kimi k3][pd disagg] support pp prefill + dcp decode with dspark (#40045) 2026-09-20 00:15:28 -07:00
Mohammad Miadh AngkadandMohammad Angkad 22f02cc339 [Test] Fix scheduler fixtures after prefill burst counting (#40411)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
2026-09-20 00:08:04 -07:00
99a44c88d4 Add out-of-tree DFlash extension points (#38740)
Co-authored-by: Yuhan Chen <yuhanc@fb.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-09-20 14:54:46 +08:00
Yuxuan ZhangandXinyuan Tong 9f21fbc34b [GLM-5.3-Flash] Reduce KPool planning synchronization and overlap indexer preparation (#39695)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-19 23:51:17 -07:00
Yuxuan ZhangandXinyuan Tong c8eb54c41d Fuse GLM-5.3-Flash KDA projections and prefill metadata (#39688)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-19 23:46:33 -07:00
chaijiacheng888 2d216a11f8 [Model] Serve DeepSeek-OCR-2 with its official 768px local-crop geometry (#38996) 2026-09-20 14:17:21 +08:00
Sasha Sidorov 59dd2fc734 [2/N] [Kernel] Fuse padding-preserving HiSparse slot translation (#39837) 2026-09-20 12:04:03 +08:00