Commit Graph
18690 Commits
Author SHA1 Message Date
minke.yu 92632a60ba ci: drop gha cache backend (runner cache server 404s on this instance)
build-sglang-image / build (push) Successful in 50m51s
2026-09-22 18:49:20 +08:00
minke.yu 4b3b367b63 ci: hardcode Gitea registry (vars context falls back to docker.io on this instance)
build-sglang-image / build (push) Failing after 24m13s
2026-09-22 18:23:35 +08:00
minke.yu a78da9b524 ci: Gitea Actions workflow to build sglang image on push to dsv41-pd
build-sglang-image / build (push) Failing after 27m21s
Overseas node: original docker.io base (lmsysorg/sglang:dev-dsv41) and
default pypi. Tags follow the ymkymx convention <branch>-<sha9>-<datetime>
plus a moving <branch>-latest. Registry defaults to the Gitea instance's
container registry, override with the REGISTRY variable.
2026-09-22 16:59:13 +08:00
minke.yu b081dd3d23 Merge branch 'main' into dsv41-pd 2026-09-22 14:56:35 +08:00
Kevin Mi 8ac19cc19f [AMD][Kimi-K3] Fix deferred KDA gate projection and update DCP cookbook (#39066)
PR Test (XPU) / finish (push) Blocked by required conditions
PR Test (Arm64) / check-changes (push) Successful in 10s
PR Test (NPU) / set-image-config (push) Successful in 1s
PR Test (NPU) / Recommend tests from coverage (push) Skipped
PR Test (NPU) / check-changes (push) Successful in 12s
PR Test (sgl-router) / gate (push) Successful in 8s
PR Test (Xeon) / check-changes (push) Successful in 9s
PR Test (XPU) / check-changes (push) Successful in 16s
pr-test-arm64.yml / pr-gate (push) Successful in 3s
PR Test (Arm64) / pr-gate (push) Successful in 3s
PR Test (Arm64) / build-test (push) Waiting to run
pr-test-npu.yml / pr-gate (push) Successful in 2s
PR Test (NPU) / pr-gate (push) Successful in 2s
PR Test (sgl-router) / tier-1 — lint (push) Failing after 33s
PR Test (sgl-router) / tier-2 — build + test (push) Skipped
PR Test (sgl-router) / tier-3 — docker (placeholder) (push) Skipped
PR Test (sgl-router) / tier-3 — k8s integration (push) Skipped
PR Test (sgl-router) / tier-3 — e2e (push) Skipped
pr-test-xpu.yml / pr-gate (push) Successful in 2s
PR Test (XPU) / pr-gate (push) Successful in 2s
pr-test-xeon.yml / pr-gate (push) Successful in 3s
PR Test (XPU) / stage-a-test-1-gpu-xpu (push) Waiting to run
PR Test (XPU) / multimodal-gen-test-1-gpu-xpu (push) Waiting to run
PR Test (Xeon) / pr-gate (push) Successful in 3s
PR Test (sgl-router) / finish (push) Successful in 1s
PR Test (Xeon) / build-test (gnr, gnr, xeon-gnr, stage-a-tp-test-cpu-intel) (push) Waiting to run
PR Test (Xeon) / build-test (spr1, 0, 3, spr, xeon-spr, stage-a-test-cpu-intel,stage-b-test-cpu-intel) (push) Waiting to run
PR Test (Xeon) / build-test (spr2, 1, 3, spr, xeon-spr, stage-a-test-cpu-intel,stage-b-test-cpu-intel) (push) Waiting to run
PR Test (Xeon) / build-test (spr3, 2, 3, spr, xeon-spr, stage-a-test-cpu-intel,stage-b-test-cpu-intel) (push) Waiting to run
Lint / lint (push) Failing after 2m51s
PR Test (NPU) / base-a-test-1-npu-a2 (push) Canceled after 0s
PR Test (NPU) / base-b-test-1-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-b-test-2-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-b-test-4-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-b-test-8-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-b-test-16-npu-a3 (push) Canceled after 0s
PR Test (NPU) / multimodal-gen-test-1-npu-a3 (0) (push) Canceled after 0s
PR Test (NPU) / multimodal-gen-test-1-npu-a3 (1) (push) Canceled after 0s
PR Test (NPU) / multimodal-gen-test-4-npu-a3 (0) (push) Canceled after 0s
PR Test (NPU) / multimodal-gen-test-4-npu-a3 (1) (push) Canceled after 0s
PR Test (NPU) / base-c-test-acc-2-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-c-test-acc-16-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-c-test-perf-2-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-c-test-perf-16-npu-a3 (push) Canceled after 0s
PR Test (NPU) / Analyze failure report (push) Canceled after 0s
PR Test (NPU) / setup-covstub (push) Canceled after 0s
PR Test (NPU) / pr-test-npu-finish (push) Canceled after 0s
pr-test-npu.yml / run (${{ fromJson(inputs.partitions).arr }}) (push) Canceled after 0s
2026-09-22 06:35:12 +00:00
Mohammad Miadh AngkadandMohammad Angkad 4c81cd1b09 [KDA] Fix missing beta sigmoid in PTX prefill (#40685)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
2026-09-21 23:24:38 -07:00
Khoa Pham bc22e1de9e [DSpark] Fix draft CUDA graph stream explosion (#40658) 2026-09-21 23:05:45 -07:00
Brayden ZhongandXinyuan Tong a0781f2714 [Docs] GLM-5.3/5.3-Flash cookbooks: enable reasoning/tool-call parsers by default via auto (#40497)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-22 13:54:08 +08:00
Guangda LiuandGuangda Liu 04c0913434 [HiSparse] Add MHA hisparse support for MiniMax M3 (#31446)
Co-authored-by: Guangda Liu <bingps@users.noreply.github.com>
2026-09-22 13:28:03 +08:00
095e45100b [AMD] [GLM-5.3-Flash Day 0] Route mHC through AITER on gfx950 (#38545)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: Kevin Mi <mikevin920@yahoo.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 22:24:29 -07:00
kangwangamd 264da63319 [AMD] Update ROCm AITER pin to acf8fdf9 (#39965) 2026-09-21 22:01:17 -07:00
Piotr Mazurek b01961e295 [LFM2-VL] Add DSpark speculative decoding (#40651) 2026-09-21 21:49:51 -07:00
YAMY 9b59fc5db5 [ModelOpt][PP] Keep BF16 shared experts out of the NVFP4 fusion so TP1 pipeline stages can load (#40628) 2026-09-21 21:45:58 -07:00
2032f3a071 [Router] Abort the engine when a client disconnects mid-request (#39461)
Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Kan Wu <wukanustc@gmail.com>
2026-09-22 12:35:56 +08:00
a9f02b0fa4 [NPU] Fix xgrammar apply_vocab_mask device dispatch to use torch.ops.npu (#36120)
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
2026-09-21 21:34:15 -07:00
877a293d6d [Benchmark] Optionally clear HiCache storage between cases (#40659)
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Pengchao Wang <wpc@fb.com>
2026-09-21 21:31:21 -07:00
e1daf68304 [AMD] [GLM-5.3-Flash Day 0] Honor fused and per-expert names in quark exclude (#39317)
Co-authored-by: Yikai Zhang <ykzhang12@gmail.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: Kevin Mi <mikevin920@yahoo.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-21 21:26:42 -07:00
Khoa PhamandQiaolin Yu 018b73c7a0 [PD] Pack draft KV head slices for DCP transfers (#40500)
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
2026-09-21 21:11:27 -07:00
b44e248682 [AMD] [GLM-5.3-Flash Day 0] Enable FP8 and Quark MXFP4 MoE on gfx950 (#38546)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: andyluo7 <andy.luo@amd.com>
Co-authored-by: Kevin Mi <mikevin920@yahoo.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-21 21:07:37 -07:00
jthomson04 15ba54bd5d perf(engine): avoid timed waits for Engine responses (#39486)
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
2026-09-21 20:42:03 -07:00
Jan BernlöhrandPo-Han Huang 56fee88e23 fix(moe): support Llama4 NVFP4 router input weights on SM120 (#35504)
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
2026-09-21 20:18:11 -07:00
Kan Wu 59a723ef1e [sgl-router] refactor - SLO ordering for bucket selection (#40292) 2026-09-22 11:16:10 +08:00
Kan Wu 27f796ca6c [sgl-router] Fix readiness, IPv6 discovery, logging, and model validation (#40604) 2026-09-22 11:13:06 +08:00
Cheng Wan 1d025491f3 [Test] Set DP size in the mocked Metal profiler test (#40667) 2026-09-21 19:57:52 -07:00
90cf471723 [AMD] [GLM-5.3-Flash Day 0] Support non-2048 top-k widths in the DSA page-table transform (#39340)
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: Kevin Mi <mikevin920@yahoo.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-21 19:50:17 -07:00
5f9c6b9eb0 [diffusion] fix: separate a use-scoped layerwise release from release_all (#40590)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 10:38:44 +08:00
MickandMick Qian a1b2b976fe [diffusion] CI: restore public Qwen-Image 2.1 TP2 E2E coverage (#40507)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-22 10:38:00 +08:00
Yuhan Zhou 15eba3b464 Feat: Add TensorCast storage as a new HiCache backend (#27265) 2026-09-22 10:19:04 +08:00
jacky.cheng bc30fa1759 [AMD][Fix] AgentX HIP TPOT regression when SGLANG_SIMULATE_ACC_LEN is set (#40598) 2026-09-21 19:16:15 -07:00
Cheng Wan 9eda772a21 [Test] Handle tied top-k indices in graph-pool logprob regression (#40661) 2026-09-21 19:02:43 -07:00
Mohammad Miadh Angkadandmmangkad e332e1b84e [Fix] Don't write conv state from the fused KDA verify kernel (#39524)
Co-authored-by: mmangkad <mohammad.angkad@radixark.ai>
2026-09-21 18:33:03 -07:00
Dayananda VandClaude Opus 5 35eb7cf8d6 [Intel][XPU][KVCanary] Enable KV Canary on Intel XPU (#33520)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 09:19:01 +08:00
Ma Mingfei 046cd6f4ea [XPU][ci]: disable XPU NIXL disaggregation test (#40540) 2026-09-22 09:02:22 +08:00
Khoa Pham c4d3770a68 [Kimi K3] Fix CUDA graph stream explosion (#40640) 2026-09-21 17:49:50 -07:00
jacky.cheng 31b577bb08 [AMD] Pad QSA MQA decode Q-heads to 16 for ROCm MFMA (#38875) 2026-09-21 17:43:46 -07:00
Cheng Wan 98c8dee23b Fix lint failure from draft-decode window test location (#40654) 2026-09-21 17:41:00 -07:00
042b6a488f [AMD] [GLM-5.3-Flash Day 0] Enable zero-RoPE MHA prefill on ROCm (#39338)
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
2026-09-21 17:38:07 -07:00
Cheng Wan 582389cec5 [Fix] Keep diffusion encoder TP context bindings consistent (#40646) 2026-09-21 17:23:15 -07:00
Tao LiandXiaoyu Zhang c53cc8e1eb [NPU][BugFix] Avoid M-RoPE recompilation for variable sequence lengths (#40371)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-09-22 08:19:25 +08:00
avalliappan-nvidia 61d0cf2074 [Spec] Windowed draft-decode attention for built-in EAGLE / MTP drafts (#32673) 2026-09-22 08:17:57 +08:00
Vedant V JhaveriandCopilot 9fdb71732a Avoid materializing GDN QKV tensors during target verification (#33778)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
2026-09-21 17:04:18 -07:00
Cheng Wan 506698761d [unified-memory] Hierarchical cache for every unified pool shape (#37507) 2026-09-21 16:50:37 -07:00
Cheng Wan 22587fb15c [Fix] Run KV canary hooks for context-parallel prefill (#40642) 2026-09-21 16:46:22 -07:00
Yuxuan ZhangandXinyuan Tong 00986c81be Support GLM-5.3-Flash hybrid attention CPU offload and PD index mapping (#40310)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-09-21 16:03:03 -07:00
YAMY 0229025127 [Spec][PP] Launch extend microbatches before the spec output exchange (#40499) 2026-09-21 15:47:03 -07:00
RuibinCheung 0c53fec476 [ROCm] fix: remove extra bf16 -> fp32 cast in jit grouped topk kernel path (#39775) 2026-09-21 15:43:21 -07:00
Kan Wu d47b8c454c [sgl-router] Release cancelled circuit-breaker probes (#40603) 2026-09-21 15:39:50 -07:00
Liangsheng Yin a5c2cc517c [CI] Split the CI control labels into four axes and resolve them live (#40527) 2026-09-21 15:37:28 -07:00
Zhang, Jiejing 66f19f5c46 [AMD] Enable HiCache for GLM-5.2 MI355X throughput recipe (#40570) 2026-09-21 15:28:27 -07:00
8bde82c0ad [AMD] [GLM-5.3-Flash Day 0] Build the fused DSA k-pool top-k JIT kernel on HIP (#39339)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Kevin Mi <mikevin920@yahoo.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-21 15:24:44 -07:00