Commit Graph
18551 Commits
Author SHA1 Message Date
ZY Y 5f017ffabb Update test cases and performance testing framework (#40392) 2026-09-20 22:42:00 +08:00
chenyang08056032 404dee10c0 [NPU] add coverage-based precision test selection pipeline (#38339) 2026-09-20 22:40:55 +08:00
Mohammad Miadh AngkadandMohammad Angkad 8923f4d779 [Test] Fix optimistic prefill disaggregation test after mamba radix cache removal (#40469)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
2026-09-20 22:39:34 +08:00
faceless voidandronnie_zheng 791c7850d0 [Diffusion] Enable shared RMSNorm dispatch for SenseNova-U1 (#39705)
Signed-off-by: syd520zy <529477025@qq.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-09-20 17:08:02 +03:00
7b1c2ed0a4 [rust-renderer] Standalone preprocessing (#36718)
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: Rain Jiang <96632942+rainj-me@users.noreply.github.com>
2026-09-20 22:03:12 +08:00
MickandMick Qian 6880a47955 [diffusion] docs: simplify Qwen-Image 2.1 cookbook (#40455)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-20 20:39:26 +08:00
Kan Wu c610c40399 [sgl-router] refactor - config and organize CLI options (#39867) 2026-09-20 19:40:04 +08:00
Ruiyan Ma efa7be2091 [Simulator][Compatibility] Adapt to latest KV cache pool interfaces (#40418) 2026-09-20 18:00:06 +08:00
Liangsheng Yin 0024efa0de [CI] Derive registered-test kind from the registry call instead of the path (#40294) 2026-09-20 02:01:22 -07:00
Kan WuandCursor 671630abf1 [sgl-router] refactor - main startup logic (#39861)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-20 16:55:34 +08:00
HuangJi 2a0cb2f04e [Diffusion][MiniMax-H3] Add SM120 Sage compute for SubBlock sparse attention (#40116) 2026-09-20 16:40:28 +08:00
MickandMick Qian 414adef060 [CI] skip srt rust extension builds for diffusion-only PRs (#40293)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-20 16:33:23 +08:00
Liangsheng Yin dc002c85fc [Test] Fix OOT DFlash hook test resolving the draft config over the network (#40427) 2026-09-20 01:25:34 -07:00
Shuwen WangandSeokhoon Kang 9f3d275940 [HiCache] Fix sparse hybrid transfer layer IDs (#37870)
Co-authored-by: Seokhoon Kang <sh.kang@postech.ac.kr>
2026-09-20 16:24:24 +08:00
amd-danli103andHAI e54009240a [AMD][DSV4] feat: enable DSpark with fp8 unified_kv on gfx950 (#38901)
Co-authored-by: HAI <hixiao@gmail.com>
2026-09-20 01:16:39 -07:00
Ke Bao a8a4d86be9 Remove swa and mamba radix cache (#40313) 2026-09-20 16:16:27 +08:00
amote-i 5c69e32abe [NPU] [DOC] fix typos, heading levels and terminology in NPU docs (#40402) 2026-09-20 15:15:58 +08:00
Qiaolin Yu f4c256354c [kimi k3][pd disagg] support pp prefill + dcp decode with dspark (#40045) 2026-09-20 00:15:28 -07:00
Mohammad Miadh AngkadandMohammad Angkad 22f02cc339 [Test] Fix scheduler fixtures after prefill burst counting (#40411)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
2026-09-20 00:08:04 -07:00
99a44c88d4 Add out-of-tree DFlash extension points (#38740)
Co-authored-by: Yuhan Chen <yuhanc@fb.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-09-20 14:54:46 +08:00
Yuxuan ZhangandXinyuan Tong 9f21fbc34b [GLM-5.3-Flash] Reduce KPool planning synchronization and overlap indexer preparation (#39695)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-19 23:51:17 -07:00
Yuxuan ZhangandXinyuan Tong c8eb54c41d Fuse GLM-5.3-Flash KDA projections and prefill metadata (#39688)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-19 23:46:33 -07:00
sglang-botandsglang-bot c1a1eb5f66 docs: sync LMSYS SGLang blog cards (#40276)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-09-20 06:35:20 +00:00
chaijiacheng888 2d216a11f8 [Model] Serve DeepSeek-OCR-2 with its official 768px local-crop geometry (#38996) 2026-09-20 14:17:21 +08:00
MickandMick Qian 031bff5dd3 [diffusion] chore: batch qwen-image 2.1 targets and document measured deployment recipes (#40408)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-20 13:59:36 +08:00
Kan Wu 99d53fe0c2 [sgl-router] refactor - chat_completions() into modules (#39848) 2026-09-20 13:51:25 +08:00
Yuan Luoandluoyuan.luo dd83b54611 Update linear attention code owner directory (#40389)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-09-19 21:13:47 -07:00
Sasha Sidorov 59dd2fc734 [2/N] [Kernel] Fuse padding-preserving HiSparse slot translation (#39837) 2026-09-20 12:04:03 +08:00
Mohammad Miadh AngkadandMohammad Angkad df0dc44931 [Fix] Forward SWA prealloc reclaim through the DSV4 HiSparse allocator (#40354)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
2026-09-19 20:27:51 -07:00
huangtingweiandChao Shi 020703923d [PP + HiCache] Add PP Prefetch Tickets for eager cross-stage storage prefetch (#36700)
Co-authored-by: Chao Shi <stepinto@live.com>
2026-09-20 11:19:31 +08:00
113f6f080e [PD] Enter the custom mem pool once when allocating DCP pack buffers (#40284)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
2026-09-19 19:34:58 -07:00
Tri Dao d82d653f96 Enable optimistic prefill for Mamba radix-cache models (#40184) 2026-09-19 19:28:11 -07:00
huangtingweiandZhangheng e9300f643e [Unified Cache][9/N] add opt-in MLA load deduplication for Mooncake Linker (#39565)
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-09-20 09:56:56 +08:00
AndyLi429andAndyLi429 d903351a66 [NPU][bugfix] update low latency quantization input and update MXFP8 tests (#38831)
Co-authored-by: AndyLi429 <AndyLi429@noreply.gitcode.com>
2026-09-20 09:55:30 +08:00
f9c2791460 [diffusion] model: support qwen-image-2.1 (#39983)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: BBuf <1182563586@qq.com>
2026-09-20 09:46:09 +08:00
黄孝君 ee5fcdf0d9 Update SGLANG_KERNEL_NPU_TAG to version 2026.9.0.post5 (#40157) 2026-09-20 09:30:25 +08:00
BourneSun0527andEven Zhou d2f291c934 [NPU][DSV4]dsv4 enable cpp (#39820)
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
2026-09-20 09:17:27 +08:00
YAMY 9cc7da2ab0 [MegaMoE] Wire Qwen MoE blocks to DeepGEMM MegaMoE (MXFP4 and NVFP4 experts) (#38080) 2026-09-19 16:00:38 -07:00
3a64faa1f2 Fix disagg PP MTP for GLM-5.2 (#39378)
Co-authored-by: Julien Lin <jullin@nvidia.com>
Co-authored-by: YAMY1234 <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
2026-09-19 13:55:07 -07:00
8139a1740e [Scheduler] Count complete prefill bursts and their tokens (#40006)
Co-authored-by: pranjalssh <pranjalssh@fb.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: Jialin Ouyang <jialino@meta.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-09-19 12:56:15 -07:00
Zhiqiang Xie 7a6c652c77 [HiCache] Auto-size the host pool to fit available host memory (#40135) 2026-09-19 12:50:43 -07:00
amd-danli103 2305242f51 [AMD][DSV4] fix: skip compressed-KV metadata on the draft worker in the HIP radix backend (#40205) 2026-09-19 12:13:49 -07:00
metamergebotandcctry 9e5a62a767 [Logprob] Serve input-logprob temporaries from CUDA-graph-pool dead space (#40038)
Co-authored-by: cctry <csycfl@gmail.com>
2026-09-19 12:03:52 -07:00
kkandwunhuang c5326d28a3 [AMD] dsv4: pick kv_splits per index stream, not by occupancy alone (#39968)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-09-19 11:53:54 -07:00
7b67a96640 [DSV4] Chunk the indexer MQA logits by query rows under a free-memory budget (#39095)
Signed-off-by: Shiki Wu <shikiw@nvidia.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
2026-09-19 11:53:20 -07:00
993d1fccba [ROCm] Widen the HiCache JIT copy rounds and enable the K-only host pool (#37152)
Co-authored-by: Xiaobo Chen <xiaobche@smci355-ccs-aus-n05-33.prov.aus.ccs.cpe.ice.amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
2026-09-19 08:58:10 -07:00
Xiaoyu Zhang 76f9213a41 [Fix] Keep mHC context out of non-V4 compiled MoE forwards (#40353) 2026-09-19 21:41:51 +08:00
Xiaoyu Zhang 9e2298e913 [CI] Propagate full-run fast-fail policy to reusable workflows (#40349) 2026-09-19 21:07:06 +08:00
Benjamin TruongandXiaoyu Zhang 83e29d6c5a [perf] Optimize w4a8 MoE for glm5.2 on H200 (#38220)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-09-19 20:34:25 +08:00
Xiaoyu Zhang 7fac84b639 [DSV4.1] Reduce mHC, metadata and small-batch router overhead (#39704) 2026-09-19 19:51:21 +08:00