This website requires JavaScript.
55b1c09e73
[core] Consolidate compiled-kernel caches under SGLANG_CACHE_DIR (#32434 )
Shu Wang
2026-08-05 13:54:27 -07:00
717a559f02
[Scheduler] Align WAR fences with CUDA graph metadata reads (#33587 )
Jialin Ouyang
2026-08-05 13:52:20 -07:00
ea65f8ddc9
Feat/spectrum (#31491 )
2026-08-05 23:23:17 +03:00
5c4f72f92a
[Build] Add srt_empty extra group for device-agnostic install (#31300 )
2026-08-06 04:17:46 +08:00
de34dd11e9
[CI] Fold duplicate-server suites and prune the retract matrix on 1-gpu-5090 (#33745 )
Liangsheng Yin
2026-08-05 12:41:51 -07:00
36853b8ffc
[Spec] Support logprobs with DFlash (#33459 )
Jason Mancuso
2026-08-05 15:37:12 -04:00
1a045669e4
[CI] Merge tokenizer worker tests and drop redundant triton attention e2e (#33641 )
Liangsheng Yin
2026-08-05 11:55:11 -07:00
5f79cf3511
[DCP] Match the replicated draft KV pool's page granularity to its allocator (#33348 )
Khoa Pham
2026-08-05 11:40:28 -07:00
b1bd871df5
[Unified Radix Cache] Complete the tree-core interface boundary (#33580 )
Jialin Ouyang
2026-08-05 11:39:28 -07:00
96c89863a3
Measure prefill busy time between launches (#33595 )
cctry
2026-08-05 11:26:27 -07:00
acaab22d09
[diffusion] feat: add SageAttention packed varlen path for minimax-h3 (#33703 )
WenhaoZhang
2026-08-06 01:19:58 +08:00
b3cdd016ba
Add Ling-3.0-flash cookbook (#33556 )
Xinyuan Tong and Zijie Xia
2026-08-05 22:53:34 +08:00
4e7209caa8
[NPU] Add causal conv1d (#28267 )
zhaozx-cn
2026-08-05 22:22:49 +08:00
3425c93666
[diffusion] Wan VAE RMSNorm+SiLU fusion behind quality=high (H200 FastWan2.2 e2e 9.611 -> 9.125 s) (#33546 )
Xiaoyu Zhang and Claude Fable 5
2026-08-05 21:33:35 +08:00
593777c046
[FIX] [benchmark] Fix flush_cache failure after warmup by waiting for server idle (#33527 )
silencejade
2026-08-05 21:27:43 +08:00
3b4fac5b99
[XPU] DeepSeek V4: use sgl-kernel-xpu implemetation of flash_mla_sparse_fwd for prefill (#31865 )
Xuan Liao and Ma Mingfei
2026-08-05 21:05:53 +08:00
99709f734d
[VLM] split multimodal scheduling from mm_utils (#32415 )
Mick
2026-08-05 20:24:12 +08:00
a5888c956f
[diffusion] Pack Ulysses Q/K/V input all-to-all into one collective + reusable a2a staging buffers (#33667 )
Xiaoyu Zhang and Claude Fable 5
2026-08-05 19:15:46 +08:00
2f22ed58ea
[NPU] Adding a fast layernorm for diffusion models and fix BSA (#29027 )
2026-08-05 04:06:00 -07:00
22d558b103
[Feature] Add GLM Image usage report (#33378 )
2026-08-05 18:59:49 +08:00
8279702e0b
[AMD] Stop publishing the K3 MI35X nightly image (#33689 )
YC Yen-Ching Tseng and Chen
2026-08-05 17:53:23 +08:00
6fa3f9df11
[Bugfix] Treat unsharded model.safetensors as HF weights in Mistral-native format detection (#33671 )
Alex Nails and Claude Fable 5
2026-08-05 01:54:46 -07:00
4a3d6ca88c
[CI] Skip sglang-kernel and sgl-deep-gemm reinstall on version match (#33637 )
Liangsheng Yin
2026-08-05 01:46:16 -07:00
a6e5fa7081
[Scheduler] Honor explicit min-free-slots thresholds (#33403 )
2026-08-05 04:44:18 -04:00
c0d5ebd6c4
[CI] Move CPU-only unit tests to the CPU suite and trim dead 5090 registrations (#33654 )
Liangsheng Yin
2026-08-05 01:43:29 -07:00
98ed5554bb
Stop testing cu129 DeepGEMM wheels (#33675 )
Mohammad Miadh Angkad
2026-08-05 16:38:28 +08:00
81c7a54ecd
[NVIDIA] Use sm_100f instead of sm_100a for sgl-kernel and FlashMLA (#33433 )
Trevor Morris
2026-08-05 01:36:46 -07:00
1478cdec9f
[AMD] Fuse Kimi-K3 attn-residual aggregation (#33599 )
Xinyi Song
2026-08-04 23:20:04 -07:00
d96df7bed5
[Diffusion] Batch GLM-Image AR requests (#30683 )
Артем Савкин and Xiaoyu Zhang
2026-08-05 08:47:06 +03:00
eac1f78568
[CI] Free hosted-runner disk space only when it is low (#33644 )
Liangsheng Yin
2026-08-04 21:55:06 -07:00
059269594c
[DSV4] Add official DSV4 reasoning effort support (#33140 )
2026-08-05 12:50:41 +08:00
198a3bc29b
[Test] Route GEMM backend UTs through real layer modules and weight loaders (#33615 )
Liangsheng Yin
2026-08-04 20:53:26 -07:00
1033cae8d5
[CI] Speed up dependency install: dual-ABI Rust ext cache and prevalidation pruning (#33619 )
Liangsheng Yin
2026-08-04 20:33:48 -07:00
6c05aaae7e
[trtllm_mha] perf: Stop allocating per-layer scratch inside the decode CUDA graph (#33063 )
Kaixi
2026-08-05 04:27:26 +02:00
4949b5fccf
[XPU] Add qknorm_rope support for Flux (#30883 )
2026-08-05 10:12:24 +08:00
5dc4102960
[npu] [bugfix] Fix PD‑disaggregation error (#33523 )
gjsheu
2026-08-05 09:54:00 +08:00
a0b04dbe4c
feat(grpc): add generation request semantics (#32588 )
2026-08-04 18:53:45 -07:00
29831d58ef
fix mm-chunk-embedding test suite (#32895 )
Zaili Wang and Ma Mingfei
2026-08-05 09:49:39 +08:00
d2c405f19d
[Intel GPU] DeepSeek V4 8/N: use sgl-kernel implementation of fused_k_norm_rope_flashmla on XPU (#28040 )
Polisetty V R K Jyothendra Varma and Ma Mingfei
2026-08-05 06:58:43 +05:30
9303e26f03
[ci] add qwen 3.5 mtp + replayssm + flashinfer gdn test (#33607 )
Qiaolin Yu
2026-08-04 18:08:42 -07:00
211ee64249
[rotary] Rebuild the shared RoPE cache entry when its buffers are dead (#33575 )
2026-08-04 17:24:42 -07:00
87ed82ff7e
Remove custom all-reduce disable from Kimi-K3 B300 recipe (#33612 )
Baizhou Zhang
2026-08-04 16:07:31 -07:00
76dc89f5aa
[Test] Replace NVFP4 MoE runner backend e2e matrix with a layer-level unit test (#33611 )
Liangsheng Yin
2026-08-04 16:03:29 -07:00
0d99d91e49
[CI] Make B200 base-b suites single-GPU as prep for 1-gpu B200 runners (#33605 )
Liangsheng Yin
2026-08-04 15:53:01 -07:00
a0b3f1dde6
[Test] Replace GEMM backend e2e matrices with layer-level unit tests (#33596 )
Liangsheng Yin
2026-08-04 15:50:41 -07:00
b0fd31ba07
Multiple flexibility fixes for DP attention (#33537 )
Lianmin Zheng
2026-08-04 15:40:40 -07:00
6808c6d571
[Tiny] Little enhancement of Kimi-K3 test (#33609 )
Baizhou Zhang
2026-08-04 15:36:41 -07:00
58da9859c4
[CI] Extract download-rust-ext and give every install step a cache fallback (#33597 )
Liangsheng Yin
2026-08-04 15:33:34 -07:00
b327d76682
[AMD] Bump mori to latest in sglang (#33462 )
Zhaoyi Li
2026-08-04 17:12:40 -05:00
c8822fd990
Clarify post-capture KV reservation logs (#33598 )
cctry
2026-08-04 15:06:11 -07:00
19d3f86895
[LoRA] Laguna: per-layer LoRA hidden-dim resolution for packed attention (#30298 )
Filip and Claude Opus 4.8
2026-08-04 22:48:05 +01:00
a9c3b55435
[Refactor] Keep chat template validation out of ServerArgs dispatcher (#33392 )
Xinyuan Tong and Alex Nails
2026-08-05 05:45:47 +08:00
e510dc58ba
Add @Jiminator as codeowner for Laguna model and config (#33472 )
Jimmy Shong and Claude Opus 5
2026-08-04 14:10:19 -07:00
34af3ff386
Allow optimistic prefill with L2 hierarchical cache and write-back policy (#33545 )
Lianmin Zheng and cctry
2026-08-04 13:23:08 -07:00
abddb1c7e9
[Kimi] Support kimi-k3 (#32541 )
+26
2026-08-04 13:22:49 -07:00
0753663b8e
[CI] Trim redundant B200 test registrations (#33586 )
Liangsheng Yin
2026-08-04 13:22:00 -07:00
aa06433709
Avoid TRTLLM prefill output copy (#33306 )
Xingyu Liu
2026-08-04 12:54:04 -07:00
38dc2d6cf8
[metrics] Split tokenizer request metrics by stream (#32734 )
2026-08-04 12:53:52 -07:00
4794b401d5
[Observability] Add startup, memory, and hybrid SWA diagnostics (#33375 )
Lianmin Zheng
2026-08-04 12:50:09 -07:00
5081c063c0
fix(metrics): clear forward occupancy on idle (#33562 )
Lianmin Zheng and Jialin Ouyang
2026-08-04 12:49:17 -07:00
dea2be5ae3
[CUDA Graph] Allow custom decode graph runners (#33553 )
Lianmin Zheng and Itai Gat
2026-08-04 12:48:56 -07:00
e76d0acdc9
migrate NPU PR/nightly test cases to a3-560T (#33346 )
2026-08-05 01:17:58 +08:00
95d0e57e83
[diffusion] Fuse DiT FFN tanh-GELU into up-proj GEMM (cublasLt epilogue) behind quality=high (Qwen-Image 1024^2 denoise 12.36 -> 12.05 s on H200) (#33536 )
Xiaoyu Zhang and Claude Fable 5
2026-08-04 23:49:43 +08:00
0d0c7d853f
[diffusion] FLUX.2 VAE decoder fast path behind quality=high (H200: 1024^2 97.6->29.2 ms, 2048^2 437.2->168.5 ms) (#33451 )
Xiaoyu Zhang
2026-08-04 23:48:19 +08:00
7adf2f4a9a
Inline _set_gc into _set_envs_and_config (#33538 )
Lianmin Zheng
2026-08-04 07:31:56 -07:00
d257b58e67
[Router] Report accelerator count in /v1/loads (#33548 )
Lianmin Zheng and Yinghai Lu
2026-08-04 05:55:35 -07:00
8f2a3ad6d7
[mem_cache] Label HiCache host pools and clarify post-capture KV sizing logs (#33445 )
Lianmin Zheng
2026-08-04 04:21:36 -07:00
723c277640
[AMD] [Fix] Enable aiter hd256 FP8 prefill FMHA on gfx950 (#33399 )
jacky.cheng
2026-08-04 17:40:33 +08:00
5e6c37f2b4
[cuda_graph] Gate breakable-CG capture_inputs retention to DP-gather paths (#32678 )
Xingyu Liu
2026-08-04 02:23:34 -07:00
b57721ccf7
Enable post-capture KV sizing with DP attention (#33427 )
2026-08-04 02:20:24 -07:00
16d3b118a2
Reduce startup log noise and fix Dynamo / CUDA-graph edge cases (#33428 )
Lianmin Zheng
2026-08-04 02:19:56 -07:00
b6d548afd7
[Fix] Resolve VLM test image placeholders from the model's own chat template (#33509 )
Liangsheng Yin
2026-08-04 02:01:39 -07:00
26a542722f
[CI] Replace the rust-ext-build venv instead of failing when it exists (#33512 )
2026-08-04 01:56:16 -07:00
eaf5c29cc5
[AMD] Enable block-fp8 + quick INT4 all-reduce in MiniMax-M3 MI35x nightly Test (#33402 )
YC Yen-Ching Tseng
2026-08-04 16:55:06 +08:00
53804d609c
[CI][XPU] Stabilize XPU CI: pin UMD/IGC, retry infra flakes, right-size EAGLE3 (#32438 )
ashwini rathi and Ma Mingfei
2026-08-04 13:58:52 +05:30
17d19081d9
[mm] sglang-mm: server vision pipeline core (fetch/driver/pipeline) + Qwen VL (#32364 )
Kan Wu and Claude Fable 5
2026-08-04 00:47:00 -07:00
154f0ac662
Fix DSpark and DP/EP (#33098 )
Vladislav Nosivskoy and Xinyuan Tong
2026-08-04 10:35:57 +03:00
bfa4e4a57b
[Nemotron] Hoist mamba track-mask host syncs out of the per-layer prefill path (#32589 )
Alex Nails and Claude Opus 5
2026-08-03 23:58:28 -07:00
23ea7b6481
Prewarm DSV4 MHC post kernel at model load (#30741 )
weireweire and weireweire
2026-08-04 14:41:56 +08:00
157401f050
[CI] Build the Rust extensions on the 5090 pool and seed the cache from main (#33460 )
Liangsheng Yin
2026-08-03 23:33:20 -07:00
101bb2327c
[diffusion] fix: fix local-path detection for MiniMax-H3 and other non-diffusers models (#33365 )
TobyMint and Mick
2026-08-04 14:03:35 +08:00
4494fb96b2
refactor the tcp listener binding logic (#33420 )
Rain Jiang
2026-08-03 22:10:52 -07:00
48dcadc770
[AMD][DI][CI] 5/N Add DSV4 wide-EP16 4-node 2P1D nightly recipes (#31500 )
Zhaoyi Li and Chen
2026-08-04 00:08:00 -05:00
afc868517b
[Perf] Speed up the Kimi-K2.5 vision path and match PIL bicubic in the GPU resize (#33349 )
Liangsheng Yin and Mick
2026-08-03 21:57:49 -07:00
7ba393dd15
[CI] Dispatch base-a-test-cpu through its own reusable stage workflow (#33461 )
Liangsheng Yin
2026-08-03 20:49:23 -07:00
c6f2a9c1d4
[diffusion] Restrict request-level quality to two validated tiers: lossless (default) and high (#33453 )
Xiaoyu Zhang and Claude Fable 5
2026-08-04 11:43:16 +08:00
614825fd38
[vla] fix: pi05 models does not apply scale factor for language embeddings (#33367 )
Jinchen Han and Mick
2026-08-04 11:25:54 +08:00
b058dc9106
[diffusion] fix: reject ring parallelism where it would silently miscompute (#33353 )
Mick
2026-08-04 11:15:23 +08:00
7dd8a3d5ea
[CI] Fix scheduler max new tokens test fixture (#33467 )
Mohammad Miadh Angkad
2026-08-04 11:12:21 +08:00
1e64fc1563
[Fix] Honor FlashMLA natural-log LSE in DCP reduction (#33065 )
EchO
2026-08-04 10:55:52 +08:00
91fae8a72c
[DCP] Bound a request by the aggregate KV pool, not one rank's share (#33448 )
Khoa Pham
2026-08-03 19:36:25 -07:00
c113ead98a
Bump helion version to 1.4 (#32562 )
Oguz Ulgen and Cheng Wan
2026-08-03 19:10:31 -07:00
92087ef4d2
fix(mem_cache): state the MLA KV bound in the DCP index space (#33432 )
Yuwei An and Claude Opus 5
2026-08-03 19:07:49 -07:00
eb31a53338
Revert "Add flashinfer rmsnorm + quant fusion support SM90, SM100, SM120" (#33455 )
Baizhou Zhang
2026-08-03 19:03:56 -07:00
cdff33d738
[CI] Build the Rust extension modules once per run instead of in every CUDA job (#33384 )
Liangsheng Yin
2026-08-03 18:56:48 -07:00
572924634b
[mem_cache] Build empty-prefix last_loc sentinel on-device to avoid per-call H2D sync (#32575 )
Xingyu Liu
2026-08-03 18:55:03 -07:00
a84e70eb1e
[AMD][DI][CI] 6/N Add Kimi-K2.6 MXFP4 wide-EP16 2P1D nightly recipes (#33333 )
Zhaoyi Li
2026-08-03 20:51:55 -05:00
03f44c978a
[AMD][DI][CI] Auto-mount latest host ionic/ibverbs userspace in the MI355X nightly launcher (fix ABI mismatch) (#33374 )
Zhaoyi Li and Chen
2026-08-03 20:14:14 -05:00
7f6a2e2b50
[Refactor] Clean up and split DSA indexer (#33443 )
Baizhou Zhang
2026-08-03 18:05:01 -07:00
3960983753
Add flashinfer rmsnorm + quant fusion support SM90, SM100, SM120 (#32994 )
2026-08-03 17:37:42 -07:00