This website requires JavaScript.
292e3ccc0c
[ci][xpu] Re-seed the wan2_1_t2v_1.3b perf baseline on Arc Pro B60 (#39955 )
ashwini rathi
2026-09-21 10:34:35 +05:30
c74a4037fb
mooncake: probe pointers via cuda.bindings instead of ctypes libcudart
minke.yu
2026-09-21 12:46:25 +08:00
b963295489
mooncake: support host-memory payloads with intra-node NVLink transport
minke.yu
2026-09-21 11:53:26 +08:00
1da8ac10e1
Update to the cookbook for XPU-supported models (#33649 )
Ziru and Niu Ziru
2026-09-21 11:26:09 +08:00
62ba964848
Fix: post-load staging regression breaks offload meta/sharded_gpu modes (#38779 )
jianzhao-xu
2026-09-21 11:19:07 +08:00
ab03a8e7eb
[Perf] Fork-safe import: no CUDA context at import time, lighter argument parsing (#40201 )
Xueshen Liu
2026-09-20 22:48:35 -04:00
176dbcb85d
[npu]add chunk gdn kernel and unify ssm state layout for ascend gdn backend (#36187 )
zhaozx-cn
2026-09-21 09:51:19 +08:00
6ad78f2281
[diffusion] docs: add verified DGX Spark recipe for Qwen-Image 2.1 (#40487 )
Mick and Mick Qian
2026-09-21 08:49:35 +08:00
b912db67ea
[diffusion] fix: keep Qwen-Image 2.1 prefix KV per layer under Cache-DiT (#40472 )
WenhaoZhang
2026-09-21 08:48:31 +08:00
501b7851e4
[diffusion] CI: guard E2E/loading latency with runner-aware baselines (#39206 )
Mick and Mick Qian
2026-09-21 08:43:22 +08:00
3a0324fb9b
[diffusion] optimization: reduce Qwen-Image 2.1 vae and graph warmup memory (#40481 )
Mick and Mick Qian
2026-09-21 08:42:02 +08:00
76a9065bef
[Fix] Raise on undelivered embeddings in send_with_url, fix broken tests (#40502 )
Liangsheng Yin
2026-09-20 17:35:27 -07:00
4027740569
[sgl-router] refactor - session-aware policy (#40379 )
Kan Wu and Claude Fable 5.1
2026-09-20 17:22:36 -07:00
aedda8377e
[sgl-router] refactor - layout BucketResolver, Bucket, EngineGroup and implement PowerOfTwo (#40241 )
Kan Wu and Claude Fable 5.1
2026-09-20 16:57:07 -07:00
4a9dc5c4af
[sgl-router] refactor - move policy-required states under src/state (#40272 )
Kan Wu and Claude Fable 5.1
2026-09-20 16:49:23 -07:00
acd20a516e
[CI] Give the kernel lane a 5090 suite and move kernel-only tests off the general lane (#40496 )
Liangsheng Yin
2026-09-20 16:46:53 -07:00
42875bcd2a
fix(modelopt): dispatch NVFP4 MoE on the cached backend, not the live global (#38932 )
2026-09-20 16:28:54 -07:00
2fa6b94e34
[Perf] Fuse the glm5_next mHC attn->MLP boundary (#39200 )
Mohammad Miadh Angkad and mmangkad
2026-09-20 23:20:37 +00:00
983e643854
[Feature] support bf16 MoE router and mxfp4 MoE for MiMo V2 (#40448 )
ollybbmonster
2026-09-21 07:13:59 +08:00
2e2d8a2fda
[CI] Drive per-commit stage jobs from a runner table instead of copied job blocks (#40495 )
Liangsheng Yin
2026-09-20 15:43:46 -07:00
d97aed2c90
Fix TopK v2 fallback when 16-block cluster capacity is zero (#40163 )
luoroger37 and Hank Han
2026-09-21 06:41:41 +08:00
f31a7bd45c
Use pinned memory for asynchronous sampling metadata transfers (#39777 )
2026-09-21 06:31:31 +08:00
95521da18d
[DeepSeek-V4.1] Bound dense prefill indexer memory (#40217 )
Harmya Bhatt
2026-09-20 16:43:19 -05:00
c2c3629f2d
[Kimi-K3] O(1) expert weight lookup in load_weights (#38805 )
2026-09-21 05:39:45 +08:00
f6483e479f
[Test] Drop cause-less disabled tests, fix XPU lane, demote quality gates off base-c (#40288 )
Liangsheng Yin
2026-09-20 14:36:02 -07:00
d229952e25
[Fix] Preserve model runner contracts in prefill CUDA graphs (#35452 )
2026-09-20 13:53:14 -07:00
745de73ba3
Add CODEOWNERS entry for sglang-renderer (#40483 )
Shangming Cai
2026-09-21 01:37:50 +08:00
b3e4d198af
[PD] Bound cached-prefix DCP transfers by pack capacity (#40376 )
Khoa Pham
2026-09-20 10:17:15 -07:00
80da4432d0
[Simulator] Fix meta host memory budgets on constrained runners (#40440 )
Shuwen Wang
2026-09-21 00:42:55 +08:00
3dbdd700e2
[CI] update CI permissions (#40474 )
WenhaoZhang
2026-09-21 00:07:37 +08:00
e97614d10c
[Qwen4-Exp] Build the offloaded PLE table on the meta device so --ple-offload-embedding never materialises it on the accelerator (#39928 )
Jimmy Shong and Yangmin Li
2026-09-20 08:28:25 -07:00
5f017ffabb
Update test cases and performance testing framework (#40392 )
ZY Y
2026-09-20 22:42:00 +08:00
404dee10c0
[NPU] add coverage-based precision test selection pipeline (#38339 )
chenyang08056032
2026-09-20 22:40:55 +08:00
8923f4d779
[Test] Fix optimistic prefill disaggregation test after mamba radix cache removal (#40469 )
Mohammad Miadh Angkad and Mohammad Angkad
2026-09-20 14:39:34 +00:00
fa826e08b1
Reuse shared compressed KV dequantization in DeepSeek V4.1 CP prefill
abing
2026-09-19 10:24:01 -07:00
3810f531a8
Support V4.1 decode vision MegaMoE
SYChen123
2026-09-18 19:13:10 +08:00
2580c24d1b
Support V4.1 DP attention in DP-only DSpark PD
SYChen123
2026-09-18 19:11:24 +08:00
791c7850d0
[Diffusion] Enable shared RMSNorm dispatch for SenseNova-U1 (#39705 )
faceless void and ronnie_zheng
2026-09-20 22:08:02 +08:00
4f22146e51
Avoid host synchronization in DeepSeek V4.1 CP prefill
abing
2026-09-18 02:28:12 -07:00
12e3b82e52
run pass cp+megamoe+bcg
abing
2026-09-17 04:04:38 -07:00
8305f66fc8
fix bug
abing
2026-09-19 02:32:27 -07:00
21a4a16b4b
update code
abing
2026-09-17 19:43:41 -07:00
fc954b7e08
add test
abing
2026-09-16 00:26:02 -07:00
c2059c4fb2
run pass llm cp
abing
2026-09-15 23:41:20 -07:00
7b1c2ed0a4
[rust-renderer] Standalone preprocessing (#36718 )
2026-09-20 17:03:12 +03:00
6880a47955
[diffusion] docs: simplify Qwen-Image 2.1 cookbook (#40455 )
Mick and Mick Qian
2026-09-20 20:39:26 +08:00
c610c40399
[sgl-router] refactor - config and organize CLI options (#39867 )
Kan Wu
2026-09-20 04:40:04 -07:00
efa7be2091
[Simulator][Compatibility] Adapt to latest KV cache pool interfaces (#40418 )
Ruiyan Ma
2026-09-20 18:00:06 +08:00
0024efa0de
[CI] Derive registered-test kind from the registry call instead of the path (#40294 )
Liangsheng Yin
2026-09-20 02:01:22 -07:00
671630abf1
[sgl-router] refactor - main startup logic (#39861 )
Kan Wu and Cursor
2026-09-20 01:55:34 -07:00
2a0cb2f04e
[Diffusion][MiniMax-H3] Add SM120 Sage compute for SubBlock sparse attention (#40116 )
HuangJi
2026-09-20 16:40:28 +08:00
414adef060
[CI] skip srt rust extension builds for diffusion-only PRs (#40293 )
Mick and Mick Qian
2026-09-20 16:33:23 +08:00
dc002c85fc
[Test] Fix OOT DFlash hook test resolving the draft config over the network (#40427 )
Liangsheng Yin
2026-09-20 01:25:34 -07:00
9f3d275940
[HiCache] Fix sparse hybrid transfer layer IDs (#37870 )
Shuwen Wang and Seokhoon Kang
2026-09-20 16:24:24 +08:00
e54009240a
[AMD][DSV4] feat: enable DSpark with fp8 unified_kv on gfx950 (#38901 )
amd-danli103 and HAI
2026-09-20 16:16:39 +08:00
a8a4d86be9
Remove swa and mamba radix cache (#40313 )
Ke Bao
2026-09-20 16:16:27 +08:00
5c69e32abe
[NPU] [DOC] fix typos, heading levels and terminology in NPU docs (#40402 )
amote-i
2026-09-20 15:15:58 +08:00
f4c256354c
[kimi k3][pd disagg] support pp prefill + dcp decode with dspark (#40045 )
Qiaolin Yu
2026-09-20 00:15:28 -07:00
22f02cc339
[Test] Fix scheduler fixtures after prefill burst counting (#40411 )
Mohammad Miadh Angkad and Mohammad Angkad
2026-09-20 07:08:04 +00:00
99a44c88d4
Add out-of-tree DFlash extension points (#38740 )
2026-09-19 23:54:46 -07:00
9f21fbc34b
[GLM-5.3-Flash] Reduce KPool planning synchronization and overlap indexer preparation (#39695 )
Yuxuan Zhang and Xinyuan Tong
2026-09-20 14:51:17 +08:00
c8eb54c41d
Fuse GLM-5.3-Flash KDA projections and prefill metadata (#39688 )
Yuxuan Zhang and Xinyuan Tong
2026-09-20 14:46:33 +08:00
c1a1eb5f66
docs: sync LMSYS SGLang blog cards (#40276 )
sglang-bot and sglang-bot
2026-09-19 23:35:20 -07:00
2d216a11f8
[Model] Serve DeepSeek-OCR-2 with its official 768px local-crop geometry (#38996 )
chaijiacheng888
2026-09-20 14:17:21 +08:00
031bff5dd3
[diffusion] chore: batch qwen-image 2.1 targets and document measured deployment recipes (#40408 )
Mick and Mick Qian
2026-09-20 13:59:36 +08:00
99d53fe0c2
[sgl-router] refactor - chat_completions() into modules (#39848 )
Kan Wu
2026-09-19 22:51:25 -07:00
dd83b54611
Update linear attention code owner directory (#40389 )
Yuan Luo and luoyuan.luo
2026-09-19 21:13:47 -07:00
59dd2fc734
[2/N] [Kernel] Fuse padding-preserving HiSparse slot translation (#39837 )
Sasha Sidorov
2026-09-19 21:04:03 -07:00
df0dc44931
[Fix] Forward SWA prealloc reclaim through the DSV4 HiSparse allocator (#40354 )
Mohammad Miadh Angkad and Mohammad Angkad
2026-09-20 03:27:51 +00:00
020703923d
[PP + HiCache] Add PP Prefetch Tickets for eager cross-stage storage prefetch (#36700 )
huangtingwei and Chao Shi
2026-09-20 11:19:31 +08:00
113f6f080e
[PD] Enter the custom mem pool once when allocating DCP pack buffers (#40284 )
2026-09-19 19:34:58 -07:00
d82d653f96
Enable optimistic prefill for Mamba radix-cache models (#40184 )
Tri Dao
2026-09-19 22:28:11 -04:00
e9300f643e
[Unified Cache][9/N] add opt-in MLA load deduplication for Mooncake Linker (#39565 )
huangtingwei and Zhangheng
2026-09-20 09:56:56 +08:00
d903351a66
[NPU][bugfix] update low latency quantization input and update MXFP8 tests (#38831 )
AndyLi429 and AndyLi429
2026-09-20 09:55:30 +08:00
f9c2791460
[diffusion] model: support qwen-image-2.1 (#39983 )
2026-09-20 09:46:09 +08:00
ee5fcdf0d9
Update SGLANG_KERNEL_NPU_TAG to version 2026.9.0.post5 (#40157 )
黄孝君
2026-09-20 09:30:25 +08:00
d2f291c934
[NPU][DSV4]dsv4 enable cpp (#39820 )
BourneSun0527 and Even Zhou
2026-09-20 09:17:27 +08:00
9cc7da2ab0
[MegaMoE] Wire Qwen MoE blocks to DeepGEMM MegaMoE (MXFP4 and NVFP4 experts) (#38080 )
YAMY
2026-09-19 16:00:38 -07:00
3a64faa1f2
Fix disagg PP MTP for GLM-5.2 (#39378 )
2026-09-20 04:55:07 +08:00
8139a1740e
[Scheduler] Count complete prefill bursts and their tokens (#40006 )
2026-09-19 12:56:15 -07:00
7a6c652c77
[HiCache] Auto-size the host pool to fit available host memory (#40135 )
Zhiqiang Xie
2026-09-19 12:50:43 -07:00
2305242f51
[AMD][DSV4] fix: skip compressed-KV metadata on the draft worker in the HIP radix backend (#40205 )
amd-danli103
2026-09-20 03:13:49 +08:00
9e5a62a767
[Logprob] Serve input-logprob temporaries from CUDA-graph-pool dead space (#40038 )
metamergebot and cctry
2026-09-19 12:03:52 -07:00
c5326d28a3
[AMD] dsv4: pick kv_splits per index stream, not by occupancy alone (#39968 )
kk and wunhuang
2026-09-20 02:53:54 +08:00
7b67a96640
[DSV4] Chunk the indexer MQA logits by query rows under a free-memory budget (#39095 )
2026-09-20 02:53:20 +08:00
993d1fccba
[ROCm] Widen the HiCache JIT copy rounds and enable the K-only host pool (#37152 )
2026-09-19 08:58:10 -07:00
76f9213a41
[Fix] Keep mHC context out of non-V4 compiled MoE forwards (#40353 )
Xiaoyu Zhang
2026-09-19 21:41:51 +08:00
9e2298e913
[CI] Propagate full-run fast-fail policy to reusable workflows (#40349 )
Xiaoyu Zhang
2026-09-19 21:07:06 +08:00
83e29d6c5a
[perf] Optimize w4a8 MoE for glm5.2 on H200 (#38220 )
Benjamin Truong and Xiaoyu Zhang
2026-09-19 19:34:25 +07:00
7fac84b639
[DSV4.1] Reduce mHC, metadata and small-batch router overhead (#39704 )
Xiaoyu Zhang
2026-09-19 19:51:21 +08:00
d1acbe0746
[DSV4.1] Big fused wo_a quant (#39957 )
DarkSharpness and BBuf
2026-09-19 19:46:53 +08:00
cb22f2451e
[Cleanup] Deduplicate kernel tests, diffusion fixtures and benchmark helpers (#40265 )
Xiaoyu Zhang
2026-09-19 19:45:38 +08:00
0b0d2c257a
[Fix] Repair CI fixtures and ROCm speculative tree device checks (#40325 )
Xiaoyu Zhang
2026-09-19 18:16:05 +08:00
567d5925fe
Fix mxfp4 padding test stubbing an accessor the module no longer imports (#40308 )
Cheng Wan
2026-09-19 00:43:56 -07:00
3a5f52e144
Record a process's placement at publish, not at group build (#40071 )
Cheng Wan
2026-09-18 23:52:37 -07:00
36aa8479ef
[Test] Fix fusion-group mocks after runtime context migration (#40290 )
Xiaoyu Zhang
2026-09-19 14:45:16 +08:00
5d703de9e4
[HiCache] Size MHA host pools from device row width (#40304 )
Lianmin Zheng
2026-09-18 23:37:58 -07:00
6533223502
[Lint] Fix logits processor formatting on main (#40303 )
Xiaoyu Zhang
2026-09-19 14:31:56 +08:00
111aeedd37
[Runtime] Add decode CUDA graph hooks for eager logits processing (#40222 )
Xiaozhu Meng and mxz
2026-09-18 23:27:03 -07:00
6e1338dd1e
Fix prefetch attempt cleanup on abort (#40262 )
2026-09-18 22:46:17 -07:00