Commit Graph
9732 Commits
Author SHA1 Message Date
karverma-amdandsogalin_codegen 20b2817bdf [AMD] Enable BCG on ROCm + route aiter prefill via MHA during PCG/BCG capture for Kimi-2.5 (#27833)
Co-authored-by: sogalin_codegen <39478626+sogalin@users.noreply.github.com>
2026-06-23 18:22:41 -07:00
Polisetty V R K Jyothendra VarmaandMa Mingfei 5338e44483 [Intel GPU] fix triton-mla attention on XPU by limiting max_kv_splits to 8 which is default (#28646)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-06-24 09:06:13 +08:00
Raiden MakotoandRaiden-Makoto 7454735be9 [AMD] [GLM5] Add opt-in Triton fp8 sparse-MLA prefill kernel for gfx950 (#28975)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
2026-06-23 18:04:07 -07:00
Raiden MakotoandRaiden-Makoto 5e6d7c1615 [AMD] Fix DeepSeek-V4 fp8 KV path on gfx942 (e4m3fnuz) (#28455)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
2026-06-23 17:45:13 -07:00
b2c8f7a22e [AMD] Support triton backend decode context parallel for Qwen3.5 (#25090)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
Co-authored-by: zhengyao <zayao@amd.com>
2026-06-23 17:35:56 -07:00
Liangsheng Yin 9ef1830701 [Scheduler] Extract DFlash prefill refill into a standalone MinFreeSlotsDelayer (#29089) 2026-06-23 17:04:55 -07:00
Jackey Hua f444b5897b [Spec][1/N] Decoupled speculative decoding: IPC protocol + cross-process request id + server flags (#27634) 2026-06-23 17:04:45 -07:00
Lianmin Zheng 34dd9c28ca [Refactor] Introduce sock_send/sock_recv wrappers for zmq IPC (#29012) 2026-06-23 15:54:36 -07:00
Lianmin Zheng ecab3f322e Revert "Improve MFU metrics for prefill and verify timing" (#29079) 2026-06-23 15:46:22 -07:00
Liangsheng Yin 11e7c9e0e6 [misc] Add sglang.bench_one_batch deprecation shim (#29082) 2026-06-23 14:57:35 -07:00
Trevor Morris f74a1722e6 [NVIDIA] Support TF32 matmul to improve MiniMax gate gemm performance (#22744) 2026-06-23 14:54:54 -07:00
Liangsheng Yin c864c8d9c2 [misc] Move bench_one_batch into sglang/benchmark/ with a back-compat shim (#28687) 2026-06-23 14:48:35 -07:00
Liangsheng Yin 6c5f466023 [server_args] Reland FA4 page_size auto-force for combined --attention-backend fa4 (#28976) 2026-06-23 14:21:59 -07:00
YAMY 93015a9e6b fix(runner): autotune flashinfer MoE on a decode-shaped buffer (#29069) 2026-06-23 13:31:47 -07:00
Jan Bernlöhr fc27ce0666 Add GB10 FP8 fused MoE Triton config (#25665) 2026-06-23 13:30:51 -07:00
Mohammad Miadh Angkad cedb43d522 [DeepEP] Gate DeepEP MNNVL on fabric support (#28942) 2026-06-23 13:19:47 -07:00
Lianmin ZhengandPranjal Shankhdhar b60185c41c Improve MFU metrics for prefill and verify timing (#29000)
Co-authored-by: Pranjal Shankhdhar <pranjal.ssh@gmail.com>
2026-06-23 12:26:56 -07:00
Xinyuan Tong 0c6e8e9477 Expand parser auto detection coverage (#28449) 2026-06-23 12:26:37 -07:00
karverma-amdandCursor e0dc8b7137 [AMD] Fuse topk padded-token masking into a single Triton kernel (#28084)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-23 11:15:12 -07:00
ed26a109ee [Fix] Return streaming logprobs when reasoning/tool parser is active (#28601)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-06-23 09:56:09 -07:00
Xiaoyu Zhang 31d71c47af [Diffusion] Fix SANA VAE dtype and TurboWan backend selection (#28769) 2026-06-23 22:57:08 +08:00
Thomasandronnie_zheng 12b08e620b [Diffusion] [NPU] enable Helios on npu (#29011)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-06-23 17:25:10 +03:00
Thomas c67d338637 [Diffusion] enable cache-dit for ERNIE-Image model (#28266) 2026-06-23 16:08:03 +03:00
Liangsheng Yin ed0a62e4dd [Mem] Add KV-page double-free checks to the invariant checker (#27731) 2026-06-23 02:29:40 -07:00
kkandwunhuang af9027f6c9 [AMD] Improve performance of dsv4 in high concurrency (#28938)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-06-23 02:20:48 -07:00
ishandhanani e67b228d4c feat: session radix cache (#27058)
Signed-off-by: Ishan Dhanani <ishandhanani@gmail.com>
2026-06-23 01:51:41 -07:00
ybyang 349a6af6b8 [HiCache] Fix hicache host memory leak by bounding PP-sync work_list (#28916) 2026-06-23 16:39:03 +08:00
Mohammad Miadh Angkad 7b1a20344c Re-enable SM90 FlashInfer allreduce fusion with safe backend defaults (#28789) 2026-06-23 01:29:19 -07:00
Liangsheng Yin 854c688121 [Spec] Unify decode KV-commit bookkeeping across spec-v2 workers (#28754) 2026-06-23 00:53:05 -07:00
cctryandcctry 743ce88bc5 Fix flaky optimistic prefill retry test (#28995)
Co-authored-by: cctry <cctry@fb.com>
2026-06-23 00:18:01 -07:00
Mick 219742c394 [diffusion] optimize: optimize realtime causal attention fastpath (#28760) 2026-06-23 15:15:30 +08:00
Cheng Wan c4376aaa88 [Refactor] Remove dead out_cache_loc_swa buffers (#28968) 2026-06-23 00:03:30 -07:00
vikram singh shekhawatandClaude Sonnet 4.6 e63b57da0b [Fix] model init / XPU / transformers-v5 / bench-image fixes (#28292)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
2026-06-23 12:58:13 +08:00
Liangsheng Yin 4740f23e1f Revert "[server_args] compute mem_fraction_static after dp chunked-prefill division" (#28991) 2026-06-22 20:48:16 -07:00
Yuan Luoandluoyuan.luo abb0717174 [CI] Fix lint brought by #27527 (#28988)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-06-22 20:40:06 -07:00
Chetan Kumar Verma 6cd8d2869b Vectorize _create_custom_4d_mask in CustomQwen2Decoder (#27527) 2026-06-23 10:56:35 +08:00
62f7ffc492 feat: add Mooncake group semantics (#26574)
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: Teng Ma <sima.mt@alibaba-inc.com>
2026-06-22 19:38:58 -07:00
Terry-UVandhnyls2002 a17753e449 Fix EAGLE draft graph seq_lens_sum padding (#26880)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-06-22 18:00:25 -07:00
Xinyuan Tong de3ec2c437 [server_args] compute mem_fraction_static after dp chunked-prefill division (#28884) 2026-06-22 17:59:05 -07:00
Brayden Zhong ba9d5aed98 Fix nightly CI test for Kimi K2.5 INT4 + H200 (#28746) 2026-06-23 00:27:33 +00:00
Liangsheng Yin d70726a31a Revert "[server_args] fix FA4 page_size auto-force for combined --attention-backend fa4" (#28972) 2026-06-22 17:01:37 -07:00
Ting SUN e00703bb2a fix(frontend): return 400 for missing completions json_schema (#28090)
Signed-off-by: Ting Sun <suntcrick@gmail.com>
2026-06-22 15:47:05 -07:00
6c212a5d6b [server_args] fix FA4 page_size auto-force for combined --attention-backend fa4 (#28825)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-06-22 15:35:10 -07:00
Liangsheng Yin 770d6b2825 [Spec] Add sync-free fast_prefill_plan for EAGLE draft-extend CUDA graph (#28854) 2026-06-22 15:15:15 -07:00
Kevin FlansburgandYuhao Yang 4f60378ff5 Fix Kimi-VL GPU image preprocessing crash on non-RGB images (#28647)
Signed-off-by: Kevin Flansburg <kflansburg@cloudflare.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
2026-06-22 14:23:33 -07:00
Maxwill LinandClaude Opus 4.8 bbc853df46 fix(schedule_batch): trim stop string when EOS matches in the same step (#28802)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 13:14:00 -07:00
Xinyuan Tong 7c23d2255a [minimax-m3] Split 1/4: sparse attention ops + JIT kernels + config foundation (#28712) 2026-06-22 13:10:43 -07:00
Alex NailsandClaude Opus 4.7 b5e4e289b1 [gRPC] Native server: Python bridge entrypoint (2/4) (#23507)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-22 13:08:33 -07:00
Wenyao Gao 8dc27f6326 [MoE] dedup triton_kernels backend quant-arg asserts and fill weight dtype guard (#28689) 2026-06-22 13:01:21 -07:00
zijiexiaandClaude Opus 4.8 669be5448b [cuda graph] Enable prefill piecewise CUDA graph for Cohere2Vision (text path) (#28686)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 19:20:14 +00:00