This website requires JavaScript.
104218d9ed
ci: tag builds with -ci-<UTC datetime> suffix (naming rule: <branch>-<sha9>-<ci|local>-<UTC>)
dsv41-pd
minke.yu
2026-09-22 20:37:57 +08:00
92632a60ba
ci: drop gha cache backend (runner cache server 404s on this instance)
minke.yu
2026-09-22 18:49:20 +08:00
4b3b367b63
ci: hardcode Gitea registry (vars context falls back to docker.io on this instance)
minke.yu
2026-09-22 18:23:35 +08:00
a78da9b524
ci: Gitea Actions workflow to build sglang image on push to dsv41-pd
minke.yu
2026-09-22 16:59:13 +08:00
b081dd3d23
Merge branch 'main' into dsv41-pd
minke.yu
2026-09-22 14:56:35 +08:00
8ac19cc19f
[AMD][Kimi-K3] Fix deferred KDA gate projection and update DCP cookbook (#39066 )
main
Kevin Mi
2026-09-21 23:35:12 -07:00
4c81cd1b09
[KDA] Fix missing beta sigmoid in PTX prefill (#40685 )
Mohammad Miadh Angkad and Mohammad Angkad
2026-09-22 06:24:38 +00:00
bc22e1de9e
[DSpark] Fix draft CUDA graph stream explosion (#40658 )
Khoa Pham
2026-09-21 23:05:45 -07:00
a0781f2714
[Docs] GLM-5.3/5.3-Flash cookbooks: enable reasoning/tool-call parsers by default via auto (#40497 )
Brayden Zhong and Xinyuan Tong
2026-09-22 01:54:08 -04:00
04c0913434
[HiSparse] Add MHA hisparse support for MiniMax M3 (#31446 )
Guangda Liu and Guangda Liu
2026-09-22 13:28:03 +08:00
095e45100b
[AMD] [GLM-5.3-Flash Day 0] Route mHC through AITER on gfx950 (#38545 )
2026-09-21 22:24:29 -07:00
264da63319
[AMD] Update ROCm AITER pin to acf8fdf9 (#39965 )
kangwangamd
2026-09-22 13:01:17 +08:00
b01961e295
[LFM2-VL] Add DSpark speculative decoding (#40651 )
Piotr Mazurek
2026-09-21 21:49:51 -07:00
9b59fc5db5
[ModelOpt][PP] Keep BF16 shared experts out of the NVFP4 fusion so TP1 pipeline stages can load (#40628 )
YAMY
2026-09-21 23:45:58 -05:00
2032f3a071
[Router] Abort the engine when a client disconnects mid-request (#39461 )
2026-09-21 21:35:56 -07:00
a9f02b0fa4
[NPU] Fix xgrammar apply_vocab_mask device dispatch to use torch.ops.npu (#36120 )
2026-09-22 12:34:15 +08:00
877a293d6d
[Benchmark] Optionally clear HiCache storage between cases (#40659 )
2026-09-21 21:31:21 -07:00
e1daf68304
[AMD] [GLM-5.3-Flash Day 0] Honor fused and per-expert names in quark exclude (#39317 )
2026-09-21 21:26:42 -07:00
018b73c7a0
[PD] Pack draft KV head slices for DCP transfers (#40500 )
Khoa Pham and Qiaolin Yu
2026-09-21 21:11:27 -07:00
b44e248682
[AMD] [GLM-5.3-Flash Day 0] Enable FP8 and Quark MXFP4 MoE on gfx950 (#38546 )
2026-09-21 21:07:37 -07:00
15ba54bd5d
perf(engine): avoid timed waits for Engine responses (#39486 )
jthomson04
2026-09-21 20:42:03 -07:00
56fee88e23
fix(moe): support Llama4 NVFP4 router input weights on SM120 (#35504 )
Jan Bernlöhr and Po-Han Huang
2026-09-22 05:18:11 +02:00
59a723ef1e
[sgl-router] refactor - SLO ordering for bucket selection (#40292 )
Kan Wu
2026-09-21 20:16:10 -07:00
27f796ca6c
[sgl-router] Fix readiness, IPv6 discovery, logging, and model validation (#40604 )
Kan Wu
2026-09-21 20:13:06 -07:00
1d025491f3
[Test] Set DP size in the mocked Metal profiler test (#40667 )
Cheng Wan
2026-09-21 19:57:52 -07:00
90cf471723
[AMD] [GLM-5.3-Flash Day 0] Support non-2048 top-k widths in the DSA page-table transform (#39340 )
2026-09-22 10:50:17 +08:00
5f9c6b9eb0
[diffusion] fix: separate a use-scoped layerwise release from release_all (#40590 )
2026-09-22 10:38:44 +08:00
a1b2b976fe
[diffusion] CI: restore public Qwen-Image 2.1 TP2 E2E coverage (#40507 )
Mick and Mick Qian
2026-09-22 10:38:00 +08:00
15eba3b464
Feat: Add TensorCast storage as a new HiCache backend (#27265 )
Yuhan Zhou
2026-09-22 10:19:04 +08:00
bc30fa1759
[AMD][Fix] AgentX HIP TPOT regression when SGLANG_SIMULATE_ACC_LEN is set (#40598 )
jacky.cheng
2026-09-22 10:16:15 +08:00
9eda772a21
[Test] Handle tied top-k indices in graph-pool logprob regression (#40661 )
Cheng Wan
2026-09-21 19:02:43 -07:00
e332e1b84e
[Fix] Don't write conv state from the fused KDA verify kernel (#39524 )
Mohammad Miadh Angkad and mmangkad
2026-09-22 01:33:03 +00:00
35eb7cf8d6
[Intel][XPU][KVCanary] Enable KV Canary on Intel XPU (#33520 )
Dayananda V and Claude Opus 5
2026-09-22 06:49:01 +05:30
046cd6f4ea
[XPU][ci]: disable XPU NIXL disaggregation test (#40540 )
Ma Mingfei
2026-09-22 09:02:22 +08:00
c4d3770a68
[Kimi K3] Fix CUDA graph stream explosion (#40640 )
Khoa Pham
2026-09-21 17:49:50 -07:00
31b577bb08
[AMD] Pad QSA MQA decode Q-heads to 16 for ROCm MFMA (#38875 )
jacky.cheng
2026-09-22 08:43:46 +08:00
98c8dee23b
Fix lint failure from draft-decode window test location (#40654 )
Cheng Wan
2026-09-21 17:41:00 -07:00
042b6a488f
[AMD] [GLM-5.3-Flash Day 0] Enable zero-RoPE MHA prefill on ROCm (#39338 )
2026-09-22 08:38:07 +08:00
582389cec5
[Fix] Keep diffusion encoder TP context bindings consistent (#40646 )
Cheng Wan
2026-09-21 17:23:15 -07:00
c53cc8e1eb
[NPU][BugFix] Avoid M-RoPE recompilation for variable sequence lengths (#40371 )
Tao Li and Xiaoyu Zhang
2026-09-22 08:19:25 +08:00
61d0cf2074
[Spec] Windowed draft-decode attention for built-in EAGLE / MTP drafts (#32673 )
avalliappan-nvidia
2026-09-21 17:17:57 -07:00
9fdb71732a
Avoid materializing GDN QKV tensors during target verification (#33778 )
Vedant V Jhaveri and Copilot
2026-09-21 17:04:18 -07:00
506698761d
[unified-memory] Hierarchical cache for every unified pool shape (#37507 )
Cheng Wan
2026-09-21 16:50:37 -07:00
22587fb15c
[Fix] Run KV canary hooks for context-parallel prefill (#40642 )
Cheng Wan
2026-09-21 16:46:22 -07:00
00986c81be
Support GLM-5.3-Flash hybrid attention CPU offload and PD index mapping (#40310 )
Yuxuan Zhang and Xinyuan Tong
2026-09-22 07:03:03 +08:00
0229025127
[Spec][PP] Launch extend microbatches before the spec output exchange (#40499 )
YAMY
2026-09-21 17:47:03 -05:00
0c53fec476
[ROCm] fix: remove extra bf16 -> fp32 cast in jit grouped topk kernel path (#39775 )
RuibinCheung
2026-09-22 06:43:21 +08:00
d47b8c454c
[sgl-router] Release cancelled circuit-breaker probes (#40603 )
Kan Wu
2026-09-21 15:39:50 -07:00
a5c2cc517c
[CI] Split the CI control labels into four axes and resolve them live (#40527 )
Liangsheng Yin
2026-09-21 15:37:28 -07:00
66f19f5c46
[AMD] Enable HiCache for GLM-5.2 MI355X throughput recipe (#40570 )
Zhang, Jiejing
2026-09-21 15:28:27 -07:00
8bde82c0ad
[AMD] [GLM-5.3-Flash Day 0] Build the fused DSA k-pool top-k JIT kernel on HIP (#39339 )
2026-09-22 06:24:44 +08:00
2261c2e618
Add MiMo-V2.6 cookbook (#40622 )
2026-09-22 05:58:53 +08:00
acac4dd9d9
[Refactor] Clean up parallel runtime comments (#40632 )
Cheng Wan
2026-09-21 14:32:22 -07:00
f532ad1f9a
Fix GLM-5.3 forget-gate shape for nvCUTEDSL verify (#40607 )
William Hu
2026-09-21 16:52:26 -04:00
e0c2e8dc4d
[AMD] Tune Qwen3.5 TP4 GDN recurrent launch on gfx950 (#39987 )
jacky.cheng
2026-09-22 04:20:54 +08:00
1ed6822039
[Test] Anchor basic_perf thresholds to each metric's measured spread (#40617 )
Liangsheng Yin
2026-09-21 13:02:10 -07:00
b18ca9ca44
[CI] Bump sgl-eval to 0.1.2 (#40620 )
Liangsheng Yin
2026-09-21 13:00:54 -07:00
11e661fd45
[Fix] Don't free the multi-CTAs KV counter the decode graphs captured (#39175 )
2026-09-21 19:58:48 +00:00
44bdf225d8
Fix lint failure from MXFP8 reserved-slot test location (#40618 )
Cheng Wan
2026-09-21 12:31:11 -07:00
bccf691b22
Bringing the parallel runtime up becomes a phase, not a side effect (#40345 )
Cheng Wan
2026-09-21 12:29:50 -07:00
1d3243d05f
Take the parallel getters off the package's public surface (#40344 )
Cheng Wan
2026-09-21 12:27:50 -07:00
970e946e4f
Retire the per-runner parallel record (#40343 )
Cheng Wan
2026-09-21 12:26:40 -07:00
73f071db52
Deprecate the parallel getters the context answers, and ratchet them shut (#40342 )
Cheng Wan
2026-09-21 12:25:32 -07:00
65be3fa71a
A runner and the objects it builds freeze the placement they describe (#40341 )
Cheng Wan
2026-09-21 12:24:17 -07:00
2d0e94e3a3
Check the topology identities where the layout is written, and build at the published widths (#40340 )
Cheng Wan
2026-09-21 12:22:59 -07:00
d5fdab7022
chore: add NIXL owners and CI access (#40602 )
ishandhanani and Kangyan-Zhou
2026-09-21 12:21:30 -07:00
0db1a93adb
State the draft's whole topology in its scope, and read the rest from the context (#40339 )
Cheng Wan
2026-09-21 12:19:38 -07:00
ae7a516ba7
feat: use XGrammar V4.1 DSML parameter constraints (#39026 )
2026-09-21 15:12:28 -04:00
f0940fe3a6
Update DeepSeek-V4 Pro for B200 FP4 agentic PD disaggregation (#40610 )
Faradawn Yang
2026-09-21 12:10:03 -07:00
90b3f8544c
[AMD] Use Triton softmax routing for Qwen3.5 on gfx950 (#39986 )
jacky.cheng
2026-09-22 03:02:29 +08:00
f702a0be29
Revert "[Diffusion] migrate the whole _register_configs from registry.py to the model own config file" (#40611 )
ronnie_zheng
2026-09-21 21:05:19 +03:00
632919e498
[AMD] Fix DeepSeek-R1-MXFP4 accuracy with AITER FP8 (#37762 )
Bingxu Chen
2026-09-22 01:51:37 +08:00
e6931ca889
[Diffusion] migrate the whole _register_configs from registry.py to the model own config file (#40475 )
ronnie_zheng
2026-09-21 20:45:53 +03:00
7a6191c4b9
Preallocate HiCache MHA staging before post-capture KV sizing (#40256 )
cctry
2026-09-21 10:44:29 -07:00
7ad55e4386
[HiCache] TMA-staged host<->device KV transfer kernel (sm_90+) (#40278 )
cctry
2026-09-21 10:38:23 -07:00
0cb37c018c
[KDA] Enable ReplaySSM for GLM-5.3 Flash (#40517 )
William Hu
2026-09-21 13:34:08 -04:00
3c71bb018a
[AMD] Enable GLM DSA prefill top-k to the v2 kernel (#37889 )
Eric.Chin.AMD and Thomas Wang
2026-09-22 01:31:07 +08:00
5a6a1bb883
[mxfp8-kv] Skip writes to the reserved CUDA-graph padding slot (#35351 )
2026-09-21 13:21:25 -04:00
008470abd8
[sgl-router] Bound streaming lifetimes and release guards on idle disconnect (#40391 )
Kan Wu and Shangming Cai
2026-09-21 10:07:07 -07:00
800613a74b
[Test] Split the serving perf tests by topic into basic_perf/ and route their thresholds through a kit (#40505 )
Liangsheng Yin
2026-09-21 10:05:57 -07:00
14e9c40a72
[Observability] Expose python/rust frontend identity in /server_info (#39993 )
Sage
2026-09-21 18:32:52 +03:00
50ec9702d0
[diffusion] docs: update ComfyUI sections, trimmed examples, and the RTX 5090 DiT-resident recipe (1.42x) for Qwen-Image-2.1 cookbook (#40573 )
2026-09-21 21:23:11 +08:00
69d1e5cfe0
[Docs][NPU] Add MiMo-V2.5-Pro FP4 DFlash best practice on Ascend NPU (#40577 )
iridiumine
2026-09-21 20:32:52 +08:00
b410010087
[NPU] [DOC] Add kimi k3 cookbook for 950PR/DT Series (#40575 )
amote-i
2026-09-21 20:18:03 +08:00
0f6761b54f
[sgl-router] Add SGLang-compatible DeepSeek V4.1 Flash rendering (#40532 )
Kan Wu
2026-09-21 03:35:09 -07:00
0abb251a20
[sgl-router] Match DeepSeek V4 rendering to SGLang (#40530 )
2026-09-21 02:56:41 -07:00
2016f5e7a1
[sgl-router] Add Kimi-K3 rendering with SGLang parity (#40390 )
Kan Wu and Claude Fable 5.1
2026-09-21 02:38:49 -07:00
b86a30afba
[AMD][DI][CI] Add a SPUR cluster profile to AMD DI CI (#40113 )
Zhaoyi Li
2026-09-21 03:56:46 -05:00
8faa2d6731
[NPU][Diffusion] Disable loading latency checks in Ascend fixtures (#40544 )
hanwlax
2026-09-21 16:56:35 +08:00
630b1ef322
[sgl-router] Fix reorg admission proxy test build after BucketResolver::new (#40537 )
2026-09-21 16:42:29 +08:00
8d08dfdab7
[Fix] Add gigachat35 to the tool-call and reasoning parser name lists (#40554 )
Mohammad Miadh Angkad and Mohammad Angkad
2026-09-21 08:34:43 +00:00
a9871012ac
[sgl-router] refactor - generalized admission policy definitions (#40271 )
Kan Wu and Claude Fable 5.1
2026-09-21 01:26:30 -07:00
70b5b03e78
[NPU][CI] Fix paths-filter negation that makes every PR run the NPU tier (#40549 )
Shangming Cai
2026-09-21 16:07:27 +08:00
c2f860af1c
[ci][xpu] Record device time in the multimodal_gen perf lane (#39956 )
ashwini rathi
2026-09-21 13:09:29 +05:30
b63f8416b3
[Feature] Gigachat 3.5 support (#29189 )
2026-09-21 09:37:55 +03:00
b54d5b7c7b
disaggregation: Fix FakeKVSender queue accumulation (#28652 )
kpjeeja
2026-09-21 11:57:05 +05:30
f5f3c38aad
[Fix] Preserve YaRN scaling when extending rotary caches (#38786 )
skyler-apdx
2026-09-21 14:14:05 +08:00
d20cd9d77f
[XPU]Enable HiSparse hierarchical sparse KV cache on Intel XPU (#32792 )
AMRUTHA M and Ma Mingfei
2026-09-21 11:37:49 +05:30
fcb080bd40
[sgl-router] refactor - cache-aware policy (#40366 )
Kan Wu and Claude Fable 5.1
2026-09-20 22:53:57 -07:00
11ecdbf39f
Clean up startup logging and streamline log audits (#40526 )
Lianmin Zheng
2026-09-20 22:28:35 -07:00