minke.yu
b91137ab98
scheduler: CP-symmetric idle check for health-check admission
...
build-sglang-image / build (push) Successful in 27m52s
Health-check admit/skip used is_fully_idle(), which includes rank-local
hicache drain queues; ranks diverge right after activity, so one rank
dispatched the health-check generate while others piggyback-skipped,
deadlocking CP (hicache drain all_reduce vs CP request broadcast).
Seen on cp2/cp4 + hicache L3 after router health checks.
Recovered from b300 /data/ymk/sglang working copy (uncommitted WIP).
2026-09-24 12:01:57 +08:00
minke.yu
db7d2cb7db
fix: vision check uses get_parallel().pp_group (get_pp_group undefined in this tree)
build-sglang-image / build (push) Successful in 27m58s
2026-09-23 18:05:57 +08:00
Xinyuan Tong
6833498646
model: prune comments and redundant tests in dsv41 vision CP
2026-09-23 14:36:01 +08:00
Xinyuan Tong
b48e2cb1eb
model: TP-wide single-owner image encoding for DeepSeek V4.1
...
ViT and Aligner are replicated per TP rank, encoding each image eight
times with TP8 on both CP1 and CP8. Elect one owner per image and use
ordered full-span broadcasts with a six-phase agreement protocol.
Both CP1 and CP8 benefit while local cache hits preserve collective order.
2026-09-23 14:36:01 +08:00
Xinyuan Tong
bfeb7cd9b2
model: support DeepSeek V4.1 vision with interleave prefill CP
...
The CP runner bypassed the vision merge and used bare text embeddings.
Merge image features before sharding so request-global offsets stay valid.
Canonicalize model IDs separately to preserve scheduler hash IDs.
Keep unsupported combinations guarded and isolate embedding overrides
from multimodal prefills without starving queued FCFS requests.
2026-09-23 14:36:00 +08:00
Yuwei An
ddf5207630
[Fix] Handle chunked paged MQA metadata in DSV4.1 eager forwards ( #40637 )
build-sglang-image / build (push) Successful in 32m15s
2026-09-23 13:35:55 +08:00
minke.yu
b081dd3d23
Merge branch 'main' into dsv41-pd
2026-09-22 14:56:35 +08:00
Kevin Mi
8ac19cc19f
[AMD][Kimi-K3] Fix deferred KDA gate projection and update DCP cookbook ( #39066 )
PR Test (Arm64) / check-changes (push) Successful in 10s
PR Test (NPU) / set-image-config (push) Successful in 1s
PR Test (NPU) / Recommend tests from coverage (push) Skipped
PR Test (NPU) / check-changes (push) Successful in 12s
PR Test (sgl-router) / gate (push) Successful in 8s
PR Test (Xeon) / check-changes (push) Successful in 9s
PR Test (XPU) / check-changes (push) Successful in 16s
pr-test-arm64.yml / pr-gate (push) Successful in 3s
PR Test (Arm64) / pr-gate (push) Successful in 3s
pr-test-npu.yml / pr-gate (push) Successful in 2s
PR Test (NPU) / pr-gate (push) Successful in 2s
PR Test (sgl-router) / tier-1 — lint (push) Failing after 33s
PR Test (sgl-router) / tier-2 — build + test (push) Skipped
PR Test (sgl-router) / tier-3 — docker (placeholder) (push) Skipped
PR Test (sgl-router) / tier-3 — k8s integration (push) Skipped
PR Test (sgl-router) / tier-3 — e2e (push) Skipped
pr-test-xpu.yml / pr-gate (push) Successful in 2s
PR Test (XPU) / pr-gate (push) Successful in 2s
pr-test-xeon.yml / pr-gate (push) Successful in 3s
PR Test (Xeon) / pr-gate (push) Successful in 3s
PR Test (sgl-router) / finish (push) Successful in 1s
Lint / lint (push) Failing after 2m51s
PR Test (NPU) / base-a-test-1-npu-a2 (push) Canceled after 0s
PR Test (NPU) / base-b-test-1-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-b-test-2-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-b-test-4-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-b-test-8-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-b-test-16-npu-a3 (push) Canceled after 0s
PR Test (NPU) / multimodal-gen-test-1-npu-a3 (0) (push) Canceled after 0s
PR Test (NPU) / multimodal-gen-test-1-npu-a3 (1) (push) Canceled after 0s
PR Test (NPU) / multimodal-gen-test-4-npu-a3 (0) (push) Canceled after 0s
PR Test (NPU) / multimodal-gen-test-4-npu-a3 (1) (push) Canceled after 0s
PR Test (NPU) / base-c-test-acc-2-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-c-test-acc-16-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-c-test-perf-2-npu-a3 (push) Canceled after 0s
PR Test (NPU) / base-c-test-perf-16-npu-a3 (push) Canceled after 0s
PR Test (NPU) / Analyze failure report (push) Canceled after 0s
PR Test (NPU) / setup-covstub (push) Canceled after 0s
PR Test (NPU) / pr-test-npu-finish (push) Canceled after 0s
pr-test-npu.yml / run (${{ fromJson(inputs.partitions).arr }}) (push) Canceled after 0s
PR Test (XPU) / finish (push) Canceled after 0s
PR Test (Arm64) / build-test (push) Canceled after 0s
PR Test (XPU) / stage-a-test-1-gpu-xpu (push) Canceled after 0s
PR Test (XPU) / multimodal-gen-test-1-gpu-xpu (push) Canceled after 0s
PR Test (Xeon) / build-test (gnr, gnr, xeon-gnr, stage-a-tp-test-cpu-intel) (push) Canceled after 0s
PR Test (Xeon) / build-test (spr1, 0, 3, spr, xeon-spr, stage-a-test-cpu-intel,stage-b-test-cpu-intel) (push) Canceled after 0s
PR Test (Xeon) / build-test (spr2, 1, 3, spr, xeon-spr, stage-a-test-cpu-intel,stage-b-test-cpu-intel) (push) Canceled after 0s
PR Test (Xeon) / build-test (spr3, 2, 3, spr, xeon-spr, stage-a-test-cpu-intel,stage-b-test-cpu-intel) (push) Canceled after 0s
2026-09-22 06:35:12 +00:00
Mohammad Miadh Angkad and Mohammad Angkad
4c81cd1b09
[KDA] Fix missing beta sigmoid in PTX prefill ( #40685 )
...
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai >
2026-09-21 23:24:38 -07:00
Khoa Pham
bc22e1de9e
[DSpark] Fix draft CUDA graph stream explosion ( #40658 )
2026-09-21 23:05:45 -07:00
Guangda Liu and Guangda Liu
04c0913434
[HiSparse] Add MHA hisparse support for MiniMax M3 ( #31446 )
...
Co-authored-by: Guangda Liu <bingps@users.noreply.github.com >
2026-09-22 13:28:03 +08:00
095e45100b
[AMD] [GLM-5.3-Flash Day 0] Route mHC through AITER on gfx950 ( #38545 )
...
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com >
Co-authored-by: Thomas Wang <thomawan@amd.com >
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com >
Co-authored-by: Kevin Mi <mikevin920@yahoo.com >
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com >
2026-09-21 22:24:29 -07:00
Piotr Mazurek
b01961e295
[LFM2-VL] Add DSpark speculative decoding ( #40651 )
2026-09-21 21:49:51 -07:00
YAMY
9b59fc5db5
[ModelOpt][PP] Keep BF16 shared experts out of the NVFP4 fusion so TP1 pipeline stages can load ( #40628 )
2026-09-21 21:45:58 -07:00
a9f02b0fa4
[NPU] Fix xgrammar apply_vocab_mask device dispatch to use torch.ops.npu ( #36120 )
...
Co-authored-by: Even Zhou <even.y.zhou@outlook.com >
Co-authored-by: sglang-npu-bot <sglangnpu@163.com >
2026-09-21 21:34:15 -07:00
877a293d6d
[Benchmark] Optionally clear HiCache storage between cases ( #40659 )
...
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com >
Co-authored-by: Pengchao Wang <wpc@fb.com >
2026-09-21 21:31:21 -07:00
e1daf68304
[AMD] [GLM-5.3-Flash Day 0] Honor fused and per-expert names in quark exclude ( #39317 )
...
Co-authored-by: Yikai Zhang <ykzhang12@gmail.com >
Co-authored-by: Thomas Wang <thomawan@amd.com >
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com >
Co-authored-by: Kevin Mi <mikevin920@yahoo.com >
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com >
2026-09-21 21:26:42 -07:00
Khoa Pham and Qiaolin Yu
018b73c7a0
[PD] Pack draft KV head slices for DCP transfers ( #40500 )
...
Co-authored-by: Qiaolin Yu <liin1211@outlook.com >
2026-09-21 21:11:27 -07:00
b44e248682
[AMD] [GLM-5.3-Flash Day 0] Enable FP8 and Quark MXFP4 MoE on gfx950 ( #38546 )
...
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com >
Co-authored-by: Thomas Wang <thomawan@amd.com >
Co-authored-by: andyluo7 <andy.luo@amd.com >
Co-authored-by: Kevin Mi <mikevin920@yahoo.com >
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com >
2026-09-21 21:07:37 -07:00
jthomson04
15ba54bd5d
perf(engine): avoid timed waits for Engine responses ( #39486 )
...
Signed-off-by: jthomson04 <jwillthomson19@gmail.com >
2026-09-21 20:42:03 -07:00
Jan Bernlöhr and Po-Han Huang
56fee88e23
fix(moe): support Llama4 NVFP4 router input weights on SM120 ( #35504 )
...
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com >
2026-09-21 20:18:11 -07:00
90cf471723
[AMD] [GLM-5.3-Flash Day 0] Support non-2048 top-k widths in the DSA page-table transform ( #39340 )
...
Co-authored-by: Thomas Wang <thomawan@amd.com >
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com >
Co-authored-by: Kevin Mi <mikevin920@yahoo.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-09-21 19:50:17 -07:00
5f9c6b9eb0
[diffusion] fix: separate a use-scoped layerwise release from release_all ( #40590 )
...
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com >
Co-authored-by: Claude Opus 5 <noreply@anthropic.com >
2026-09-22 10:38:44 +08:00
Mick and Mick Qian
a1b2b976fe
[diffusion] CI: restore public Qwen-Image 2.1 TP2 E2E coverage ( #40507 )
...
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com >
2026-09-22 10:38:00 +08:00
Yuhan Zhou
15eba3b464
Feat: Add TensorCast storage as a new HiCache backend ( #27265 )
2026-09-22 10:19:04 +08:00
jacky.cheng
bc30fa1759
[AMD][Fix] AgentX HIP TPOT regression when SGLANG_SIMULATE_ACC_LEN is set ( #40598 )
2026-09-21 19:16:15 -07:00
Mohammad Miadh Angkad and mmangkad
e332e1b84e
[Fix] Don't write conv state from the fused KDA verify kernel ( #39524 )
...
Co-authored-by: mmangkad <mohammad.angkad@radixark.ai >
2026-09-21 18:33:03 -07:00
Dayananda V and Claude Opus 5
35eb7cf8d6
[Intel][XPU][KVCanary] Enable KV Canary on Intel XPU ( #33520 )
...
Co-authored-by: Claude Opus 5 <noreply@anthropic.com >
2026-09-22 09:19:01 +08:00
Khoa Pham
c4d3770a68
[Kimi K3] Fix CUDA graph stream explosion ( #40640 )
2026-09-21 17:49:50 -07:00
jacky.cheng
31b577bb08
[AMD] Pad QSA MQA decode Q-heads to 16 for ROCm MFMA ( #38875 )
2026-09-21 17:43:46 -07:00
042b6a488f
[AMD] [GLM-5.3-Flash Day 0] Enable zero-RoPE MHA prefill on ROCm ( #39338 )
...
Co-authored-by: Thomas Wang <thomawan@amd.com >
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com >
2026-09-21 17:38:07 -07:00
Cheng Wan
582389cec5
[Fix] Keep diffusion encoder TP context bindings consistent ( #40646 )
2026-09-21 17:23:15 -07:00
Tao Li and Xiaoyu Zhang
c53cc8e1eb
[NPU][BugFix] Avoid M-RoPE recompilation for variable sequence lengths ( #40371 )
...
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com >
2026-09-22 08:19:25 +08:00
avalliappan-nvidia
61d0cf2074
[Spec] Windowed draft-decode attention for built-in EAGLE / MTP drafts ( #32673 )
2026-09-22 08:17:57 +08:00
Vedant V Jhaveri and Copilot
9fdb71732a
Avoid materializing GDN QKV tensors during target verification ( #33778 )
...
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com >
2026-09-21 17:04:18 -07:00
Cheng Wan
506698761d
[unified-memory] Hierarchical cache for every unified pool shape ( #37507 )
2026-09-21 16:50:37 -07:00
Cheng Wan
22587fb15c
[Fix] Run KV canary hooks for context-parallel prefill ( #40642 )
2026-09-21 16:46:22 -07:00
Yuxuan Zhang and Xinyuan Tong
00986c81be
Support GLM-5.3-Flash hybrid attention CPU offload and PD index mapping ( #40310 )
...
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com >
2026-09-21 16:03:03 -07:00
YAMY
0229025127
[Spec][PP] Launch extend microbatches before the spec output exchange ( #40499 )
2026-09-21 15:47:03 -07:00
RuibinCheung
0c53fec476
[ROCm] fix: remove extra bf16 -> fp32 cast in jit grouped topk kernel path ( #39775 )
2026-09-21 15:43:21 -07:00
8bde82c0ad
[AMD] [GLM-5.3-Flash Day 0] Build the fused DSA k-pool top-k JIT kernel on HIP ( #39339 )
...
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com >
Co-authored-by: Thomas Wang <thomawan@amd.com >
Co-authored-by: Kevin Mi <mikevin920@yahoo.com >
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-09-21 15:24:44 -07:00
Cheng Wan
acac4dd9d9
[Refactor] Clean up parallel runtime comments ( #40632 )
2026-09-21 14:32:22 -07:00
William Hu
f532ad1f9a
Fix GLM-5.3 forget-gate shape for nvCUTEDSL verify ( #40607 )
2026-09-21 13:52:26 -07:00
jacky.cheng
e0c2e8dc4d
[AMD] Tune Qwen3.5 TP4 GDN recurrent launch on gfx950 ( #39987 )
2026-09-21 13:20:54 -07:00
Liangsheng Yin
b18ca9ca44
[CI] Bump sgl-eval to 0.1.2 ( #40620 )
2026-09-21 13:00:54 -07:00
11e661fd45
[Fix] Don't free the multi-CTAs KV counter the decode graphs captured ( #39175 )
...
Co-authored-by: mmangkad <mohammad.angkad@radixark.ai >
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai >
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com >
2026-09-21 12:58:48 -07:00
Cheng Wan
bccf691b22
Bringing the parallel runtime up becomes a phase, not a side effect ( #40345 )
2026-09-21 12:29:50 -07:00
Cheng Wan
1d3243d05f
Take the parallel getters off the package's public surface ( #40344 )
2026-09-21 12:27:50 -07:00
Cheng Wan
970e946e4f
Retire the per-runner parallel record ( #40343 )
2026-09-21 12:26:40 -07:00
Cheng Wan
73f071db52
Deprecate the parallel getters the context answers, and ratchet them shut ( #40342 )
2026-09-21 12:25:32 -07:00