-
4a78031a71
[ROCM] Optimized deepseek-r1 model with rmsnorm + fp8 quant fusion (#12689)
yctseng0211
2025-11-11 18:59:10 +08:00
-
ea10a9d165
[bug][rocm]fix qr when variable inp (#11609)
haoyangli-amd
2025-11-11 17:43:48 +08:00
-
71aea45c41
[Fix] Add TPOT back to bench_serving (#12976)
elvischenv
2025-11-11 17:32:24 +08:00
-
e3b38d7194
[Bug] TypeError: maybe_executor_submit() (#13050)
Johnsonms
2025-11-11 00:55:13 -08:00
-
5c0cadd0c6
Remove duplicate import (#12980)
LHXuuu
2025-11-11 16:53:39 +08:00
-
39b1d048a0
[AMD] Add PD test for AMD CI (#11938)
michael-amd
2025-11-11 00:49:44 -08:00
-
8db7fc4186
disable overlap schedule if mamba radix cache open (#13057)
Yi Zhang
2025-11-11 16:35:38 +08:00
-
fe92d4d88e
[CI] Auto format code (#13053)
Xiaoyu Zhang
2025-11-11 15:30:29 +08:00
-
fc8cda14cf
Sglang Tracing: optimize trace_event_batch() (#13036)
Feng Su
2025-11-11 15:02:20 +08:00
-
6a7322ffbc
[diffusion] doc: add support_new_models.md (#13043)
Mick
2025-11-11 12:49:48 +08:00
-
c751cb38b0
[AMD CI] Update CI Version Logic. (#13029)
Sai Enduri
2025-11-10 20:42:00 -08:00
-
3594815a8b
Re-enable Flashinfer TRTLLM GEN MHA and Add Unit Test (#12885)
Sam
2025-11-11 12:17:43 +08:00
-
9caca6a45c
[PieceWise CUDA Graph] Support awq/gptq model in piecewise cudagraph (#12518)
Xiaoyu Zhang
2025-11-11 11:56:15 +08:00
-
08c805a85f
fix(ci): workflow id in permission rate limit (#13035)
yinghui
2025-11-10 19:06:13 -08:00
-
f18ec927f3
fix tuning_fused_moe_triton_sep tool per_channel_quant bug (#13027)
Xiaoyu Zhang
2025-11-11 10:33:54 +08:00
-
aea88fa7af
[AMD CI] Update docker release workflows docker file name. (#13028)
Sai Enduri
2025-11-10 18:28:39 -08:00
-
2fe4e69fca
[router] add postgres databases data connector (#12218)
rongfu.leng
2025-11-11 08:51:50 +08:00
-
012bfc4fdc
[9/n] decouple quantization impl from vllm dependency - adjust ci (#12753)
Peng Zhang
2025-11-11 06:55:19 +08:00
-
0493775b06
[router][ci] Quick Improvement to make CI more stable (#12869)
Keyang Ru
2025-11-10 13:56:35 -08:00
-
40b26b456b
Simplify the BatchMultimodalOutput in io_struct.py (#12993)
Lianmin Zheng
2025-11-10 13:54:56 -08:00
-
9840bf4f84
[router][ci] Fix maturin build (#13012)
Keyang Ru
2025-11-10 12:33:32 -08:00
-
303cc957e6
chore: bump SGLang version to 0.5.5.post1 (#13000)
sglang-bot
2025-11-11 03:53:43 +08:00
-
661c1c97ad
Add pre-suffle weight for new aiter MoE support. (#12908)


2025-11-11 03:42:55 +08:00
-
c022107f8b
Resolve HF download issue and download models before CI run starts for 8-gpu-h200 runners (#12952)
Kangyan-Zhou
2025-11-10 11:32:05 -08:00
-
56c83e0fb3
[CI] Limit the CI trigger frequency of low-privilege actors (#13010)
Liangsheng Yin
2025-11-11 03:27:10 +08:00
-
b51d46d092
[AMD CI] Remove SRT docker build. (#11850)

Sai EnduriandHubert Lu
2025-11-10 11:25:37 -08:00
-
1086473111
Enhance retract test (page cases, long output cases) (#12781)
Liangsheng Yin
2025-11-11 03:03:26 +08:00
-
665416f6dd
Unify memory management across
(overlap, non-overlap) x (page>=1) x (spec, non-spec, spec v2) x (retract, finished) (#12224)
Liangsheng Yin
2025-11-11 02:56:22 +08:00
-
838bcb0d93
[misc][ci] Add run-ci after auto-labeler (#13013)
Chang Su
2025-11-10 10:41:54 -08:00
-
f1f4c451ab
Add
process_prefill_chunk back to fix PP event loop (#13009)
Liangsheng Yin
2025-11-11 00:51:36 +08:00
-
b0ee99dd03
Super tiny fix typo (#13001)
fzyzcjy
2025-11-11 00:47:45 +08:00
-
ddfcb7c8ab
minor: fix notebook bug with new model_info fields added for warmup (#13005)
Mick
2025-11-11 00:46:12 +08:00
-
58b12ccb46
Support piecewise cuda graph for deepseek v3 (#12996)
Ke Bao
2025-11-10 23:18:03 +08:00
-
547de8c774
[1 / 2] register weak_ref_tensor in sgl-kernel (#12999)
Xiaoyu Zhang
2025-11-10 22:12:59 +08:00
-
37c40a87a8
chore: bump sgl-kernel version to 0.3.17 (#12966)
sglang-bot
2025-11-10 21:50:58 +08:00
-
1240ac13b8
vlm: fix tiny multimodal cache bug (#12984)
Yuhao Yang
2025-11-10 21:30:19 +08:00
-
5639145fac
diffusion: reduce effort of supporting new model (#12982)
Mick
2025-11-10 21:20:33 +08:00
-
afee2843d5
feat(metrics): add scheduler and hiradix cache metrics (#10218) (#10225)

ShawnKungandZhiqiang Xie
2025-11-10 18:14:47 +08:00
-
6f08488042
fix missing output_token_logprobs when using ngram speculative decoding (#10702)

Zhihao Zhanganda4zhangfei
2025-11-10 18:04:42 +08:00
-
611a4fd08b
[router] bucket policy (#11719)
syy-hw
2025-11-10 18:02:53 +08:00
-
9ea2c686c7
[Auto Sync] Update batch_invariant_ops.py (20251109) (#12916)

![github-actions[bot]](/assets/img/avatar_default.png)
2025-11-10 01:51:39 -08:00
-
e2a784ecda
[RadixTree] Reduce Syscalls, Optimize Collection Filtering and Align with cpp (#12239)
PiteXChen
2025-11-10 17:47:59 +08:00
-
05559a4a90
Support hidden_dim % 4 == 0 in per_token_quant_fp8 (#12883)
Xiaoyu Zhang
2025-11-10 17:13:14 +08:00
-
7bffc5dc25
[Fix] Add validation for served model name to reserve
: for LoRA adapter syntax (#12912)

Neelabh Sinhaandneelabhsinha
2025-11-09 23:42:02 -08:00
-
a5e5088dfb
Fix errors of page head kernels in sgl-kernel for ROCm (#12604)
huangtingwei
2025-11-10 15:15:50 +08:00
-
ac19ce7efb
[PP] put pp assert in model runner (#12934)
Xuchun Shang
2025-11-10 14:58:19 +08:00
-
95876d75cb
chore: include a minimum image for vlms when warming-up (#9528)
Mick
2025-11-10 14:56:59 +08:00
-
a30f190762
chore: bump sgl-kernel version to 0.3.17 (#12931)
sglang-bot
2025-11-10 14:54:02 +08:00
-
dc8a5a1ce7
[Refactor / Style] Unify all event loops (except for PP) (#12959)
Liangsheng Yin
2025-11-10 14:49:12 +08:00
-
e123648b36
diffusion: fix wan-2.2-TI2V and support sp (#12926)
Mick
2025-11-10 14:37:57 +08:00
-
90401cf7d2
Fix the run-time error when calling fused_rms_mxfp4_quant that change return output number (#12803)

kkandwunhuang
2025-11-10 13:48:36 +08:00
-
9cfe78dd30
clean redundant code in previous PR (#12957)
Atream
2025-11-10 13:40:32 +08:00
-
307e7a6128
diffusion: fix detected file changes rule in CI (#12943)
Mick
2025-11-10 13:37:16 +08:00
-
ddd1440d0f
Refactor KTransformers heterogeneous compute with unified GPU-quantization backend (#12834)




2025-11-10 13:06:32 +08:00
-
d1be60c3c5
[Refactor] rename set_index_k_and_scale_buffer to set_index_k_scale_b… (#12956)
Edwin Gao
2025-11-09 23:35:19 -05:00
-
583bb1804e
[Docs] Add docs for Qwen3-VL image and video support (#12554)

Adarsh ShirawalmathandUbuntu
2025-11-10 09:46:04 +05:30
-
1f2a6c691b
Bugfix: LMCache Connector with Sglang (#12946)
![gemini-code-assist[bot]](/assets/img/avatar_default.png)
MMuzzammil1andgemini-code-assist[bot]
2025-11-10 13:13:37 +09:00
-
61c7fe7aed
Minor code cleanup / improvement for
PREBUILT_EXTEND mode (#12948)
Liangsheng Yin
2025-11-10 11:59:54 +08:00
-
83f89cc615
diffusion: skip full CI suite for multimodal_gen changes (#12940)
Mick
2025-11-10 10:46:24 +08:00
-
db24d34603
Support piecewise cuda graph for MLA (#11812)
Ke Bao
2025-11-10 09:13:48 +08:00
-
885cfca273
ci: try to fix gpg error during kernel build (#12928)
ishandhanani
2025-11-09 11:45:29 -08:00
-
4e916f9840
[lint] tiny fix unimported packages. (#12927)
Liangsheng Yin
2025-11-10 00:59:10 +08:00
-
4f65a64666
Refactor / Unify event loop across PD-Disagg, Overlap, DP-Attn cases (#12839)

Liangsheng Yinandcctry
2025-11-10 00:42:50 +08:00
-
f5b3ccd9a5
feat: basic support for server-level multimodal cache (#10775)
Mick
2025-11-10 00:27:50 +08:00
-
bb00e24f87
Adjust server launch time in ci (#12917)
Ke Bao
2025-11-09 20:41:29 +08:00
-
210a9cab6d
[CI] Fix
matrix.part in pr-test. (#12920)
Liangsheng Yin
2025-11-09 18:53:18 +08:00
-
b8ac4fcb51
[PD] feat: refactor custom mem pool and add barex pd support (#12332)
Teng Ma
2025-11-09 18:51:34 +08:00
-
877cb52840
[CI] increase ut buckets & adjust estimation time. (#12919)
Liangsheng Yin
2025-11-09 18:17:50 +08:00
-
3633f8b0cf
Add Jet-Nemotron (#12448)
Zijian Zhang
2025-11-09 17:32:47 +08:00
-
93cf60fc64
Fix Deepseek nightly tests (#12906)
Kangyan-Zhou
2025-11-09 00:26:06 -08:00
-
b5e0417392
Add kimi k2 thinking to ci (#12907)
Ke Bao
2025-11-09 16:10:32 +08:00
-
8a821af793
fallback to triton mm_persistent kernel when deepGemm fail (#12911)
Minglei Zhu
2025-11-08 23:42:26 -08:00
-
4b1d163bfd
Add HF cleanup logic in ci_install_dependency.sh (#12895)
Kangyan-Zhou
2025-11-08 22:56:25 -08:00
-
c21a3ec299
Fix duplicate nightly test name (#12905)
Kangyan-Zhou
2025-11-08 21:38:48 -08:00
-
b142831a26
Fix empty server args in marlin moe test (#12904)
Ke Bao
2025-11-09 13:30:47 +08:00
-
d134096319
Add Deepseek models into nightly tests (#12865)
Kangyan-Zhou
2025-11-08 21:17:49 -08:00
-
f290e8016e
Revert "Fix spec decoding acc length for dpsk-r1-fp4 tp8" (#12900)
Qiaolin Yu
2025-11-08 20:42:42 -08:00
-
9299a62fcb
Fix spec decoding acc length for dpsk-r1-fp4 tp8 (#12896)
Qiaolin Yu
2025-11-08 19:55:45 -08:00
-
49543be9da
Tiny simplify
can_run_dp_cuda_graph gather logic (#12891)
Liangsheng Yin
2025-11-09 11:15:36 +08:00
-
5236290399
Update CODEOWNERS (#12897)
Ke Bao
2025-11-09 09:37:11 +08:00
-
b2b26d4324
chore: bump sgl-kernel version to 0.3.16.post6 (#12889)
sglang-bot
2025-11-09 09:16:28 +08:00
-
d3a03aeef8
Refs/heads/add nightly test multi gpu configs (#12870)
alisonshao
2025-11-08 15:14:50 -08:00
-
5f02b918ec
[Fix] Fix trtllm-mla backend when chunked prefix cache is disabled (#12361)
Baizhou Zhang
2025-11-08 15:10:25 -08:00
-
49653c8896
use fast stream instead of torch.cuda.current_stream in llama 4 shared experts overlap (#12811)
b8zhong
2025-11-08 15:04:37 -08:00
-
44f594d832
Apply moe_reduce_sum kernel for fused_marlin_moe (#12888)
Ke Bao
2025-11-09 01:31:05 +08:00
-
2b6c4257a0
Fix sending all requests to the first rank in DP attention (#12832)
fzyzcjy
2025-11-09 00:18:53 +08:00
-
243ea585fc
[DP-Attn] Clarify MLP sync / idle batch preparation logic (#12843)
Liangsheng Yin
2025-11-08 23:23:14 +08:00
-
6fee2c535c
[CI] Tiny adjust CI esitmation time (#12886)
Liangsheng Yin
2025-11-08 23:02:19 +08:00
-
f1a9c72de3
Support capturing aux_hidden_states for minimax m2. (#12798)
Charles Chen
2025-11-08 01:54:22 -08:00
-
190002c613
[Docs][DeepseekV3.2] Update deepseekv3.2 docs for mha short seq prefill (#12868)
YAMY
2025-11-08 00:11:02 -08:00
-
0296f1cdad
[Auto Sync] Update activation.py, logits_processor.py, rota... (20251107) (#12853)

![github-actions[bot]](/assets/img/avatar_default.png)
2025-11-07 22:07:51 -08:00
-
e039ff382c
[CI] Fix huggingface access for test_flash_attention_4.py (#12846)
Baizhou Zhang
2025-11-07 20:07:06 -08:00
-
b8ddc296f4
[sgl-kernel][Deepseek V3.2] Add row_starts to topk kernel (#12582)
hlu1
2025-11-07 18:33:27 -08:00
-
0b88d520a0
Add nightly performance test for GPT-OSS 4GPU models (#12805)
alisonshao
2025-11-07 16:54:07 -08:00
-
d3d7f960b5
[router] Switch MCP tests from DeepWiki to self-hosted Brave search server (#12849)
Keyang Ru
2025-11-07 16:16:07 -08:00
-
e434187289
[router][grpc] Move all error logs to their call sites (#12859)
Chang Su
2025-11-07 15:55:34 -08:00
-
fe19a580fb
[router][grpc] Refactor: Add builders for chat and responses (#12852)
Chang Su
2025-11-07 15:43:32 -08:00
-
55e8e3999c
add back flashinfer jit cache to dev docker (#12851)

b8zhongandBrayden Zhong
2025-11-07 14:51:24 -08:00
-
32f7982800
sglang diffusion announcement (#12856)
Mingyi
2025-11-07 14:35:25 -08:00
-
0f76976c3c
remove the fa4 page_size hardcode to 128 restriction on mla model arch (#12801)
Rain Jiang
2025-11-07 13:30:30 -08:00