5085 Commits
Author SHA1 Message Date
Stella-17andxinyue.fan 6842335fcf [MUSA][24/N] CI:Fix LLM server smoke test (#28934)
Co-authored-by: xinyue.fan <xinyue.fan@mthreads.com>
2026-06-23 22:58:59 -07:00
toufupiand“toufupi” c394f812d1 [MLX] Fix Apple Silicon server startup; align MLX tests with upstream (#28770)
Co-authored-by: “toufupi” <“byte2016@outlook.com”>
2026-06-23 22:58:49 -07:00
Alex Tumanov 33373cbb12 [misc] Move bench_serving into sglang.benchmark (#28996) 2026-06-23 19:34:11 -07:00
Liangsheng Yin b448b08401 [misc] Move bench_offline_throughput into sglang/benchmark/ with a back-compat shim (#28747) 2026-06-23 18:37:52 -07:00
karverma-amdandsogalin_codegen 20b2817bdf [AMD] Enable BCG on ROCm + route aiter prefill via MHA during PCG/BCG capture for Kimi-2.5 (#27833)
Co-authored-by: sogalin_codegen <39478626+sogalin@users.noreply.github.com>
2026-06-23 18:22:41 -07:00
Polisetty V R K Jyothendra VarmaandMa Mingfei 5338e44483 [Intel GPU] fix triton-mla attention on XPU by limiting max_kv_splits to 8 which is default (#28646)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-06-24 09:06:13 +08:00
b2c8f7a22e [AMD] Support triton backend decode context parallel for Qwen3.5 (#25090)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
Co-authored-by: zhengyao <zayao@amd.com>
2026-06-23 17:35:56 -07:00
Liangsheng Yin 9ef1830701 [Scheduler] Extract DFlash prefill refill into a standalone MinFreeSlotsDelayer (#29089) 2026-06-23 17:04:55 -07:00
Jackey Hua f444b5897b [Spec][1/N] Decoupled speculative decoding: IPC protocol + cross-process request id + server flags (#27634) 2026-06-23 17:04:45 -07:00
Lianmin Zheng 34dd9c28ca [Refactor] Introduce sock_send/sock_recv wrappers for zmq IPC (#29012) 2026-06-23 15:54:36 -07:00
Liangsheng Yin c864c8d9c2 [misc] Move bench_one_batch into sglang/benchmark/ with a back-compat shim (#28687) 2026-06-23 14:48:35 -07:00
Liangsheng Yin 6c5f466023 [server_args] Reland FA4 page_size auto-force for combined --attention-backend fa4 (#28976) 2026-06-23 14:21:59 -07:00
Xinyuan Tong 0c6e8e9477 Expand parser auto detection coverage (#28449) 2026-06-23 12:26:37 -07:00
karverma-amdandCursor e0dc8b7137 [AMD] Fuse topk padded-token masking into a single Triton kernel (#28084)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-23 11:15:12 -07:00
ed26a109ee [Fix] Return streaming logprobs when reasoning/tool parser is active (#28601)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-06-23 09:56:09 -07:00
Shangming Cai 0460f277b7 Fix manual chunked-prefill test to use req.fill_len after fill_ids refactor (#29032) 2026-06-23 17:52:26 +08:00
Liangsheng Yin ed0a62e4dd [Mem] Add KV-page double-free checks to the invariant checker (#27731) 2026-06-23 02:29:40 -07:00
Shangming Cai a65c68c1fb Revert "[CI] Fix flaky optimistic test by adding contention handling" (#29023) 2026-06-23 16:53:32 +08:00
ishandhanani e67b228d4c feat: session radix cache (#27058)
Signed-off-by: Ishan Dhanani <ishandhanani@gmail.com>
2026-06-23 01:51:41 -07:00
ybyang 349a6af6b8 [HiCache] Fix hicache host memory leak by bounding PP-sync work_list (#28916) 2026-06-23 16:39:03 +08:00
Mohammad Miadh Angkad 7b1a20344c Re-enable SM90 FlashInfer allreduce fusion with safe backend defaults (#28789) 2026-06-23 01:29:19 -07:00
Liangsheng Yin 854c688121 [Spec] Unify decode KV-commit bookkeeping across spec-v2 workers (#28754) 2026-06-23 00:53:05 -07:00
Yuan Luoandluoyuan.luo abb0717174 [CI] Fix lint brought by #27527 (#28988)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-06-22 20:40:06 -07:00
Chetan Kumar Verma 6cd8d2869b Vectorize _create_custom_4d_mask in CustomQwen2Decoder (#27527) 2026-06-23 10:56:35 +08:00
62f7ffc492 feat: add Mooncake group semantics (#26574)
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
Co-authored-by: Teng Ma <sima.mt@alibaba-inc.com>
2026-06-22 19:38:58 -07:00
Michael 28d5627fd8 [AMD] register kv_canary + mock_model e2e tests to extra-a (1-gpu-small + 2-gpu-large) (#28850) 2026-06-22 19:06:31 -07:00
Terry-UVandhnyls2002 a17753e449 Fix EAGLE draft graph seq_lens_sum padding (#26880)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-06-22 18:00:25 -07:00
Liangsheng Yin d70726a31a Revert "[server_args] fix FA4 page_size auto-force for combined --attention-backend fa4" (#28972) 2026-06-22 17:01:37 -07:00
Ting SUN e00703bb2a fix(frontend): return 400 for missing completions json_schema (#28090)
Signed-off-by: Ting Sun <suntcrick@gmail.com>
2026-06-22 15:47:05 -07:00
6c212a5d6b [server_args] fix FA4 page_size auto-force for combined --attention-backend fa4 (#28825)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-06-22 15:35:10 -07:00
Liangsheng Yin 770d6b2825 [Spec] Add sync-free fast_prefill_plan for EAGLE draft-extend CUDA graph (#28854) 2026-06-22 15:15:15 -07:00
Yuan Luoandluoyuan.luo c0198fc277 [CI] Refactor int checkpoint tests style (#28813)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-06-22 14:59:06 -07:00
Maxwill LinandClaude Opus 4.8 bbc853df46 fix(schedule_batch): trim stop string when EOS matches in the same step (#28802)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-22 13:14:00 -07:00
Xinyuan Tong 7c23d2255a [minimax-m3] Split 1/4: sparse attention ops + JIT kernels + config foundation (#28712) 2026-06-22 13:10:43 -07:00
zijiexiaandClaude Opus 4.8 669be5448b [cuda graph] Enable prefill piecewise CUDA graph for Cohere2Vision (text path) (#28686)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 19:20:14 +00:00
Shangming Cai ab21dc984a [CI] Fix flaky optimistic test by adding contention handling (#28947) 2026-06-23 00:35:13 +08:00
Zhangheng 70883cb1b0 [UnifiedTree]: Rollback mamba hicache test to direct io backend (#28904) 2026-06-22 23:48:48 +08:00
Lianmin Zheng b28e990161 Migrate all ServerArgs fields to Annotated style, reduce add_cli_args by ~2400 lines (#28919) 2026-06-22 08:34:37 -07:00
Thomas Wang 04d952ea10 [AMD] deepseek-v4 clean env vars (#28920) 2026-06-22 07:32:43 -07:00
Yuwei An 2ce32366a0 [Fix][BCG][Spec] Restore EAGLE prefill plumbing dropped by #23906 (#28870) 2026-06-22 01:54:54 -07:00
Xinyuan Tong db12bfcdc8 [JIT] Add kpool_topk_transform JIT kernel (#28670) 2026-06-22 01:04:21 -07:00
Liangsheng Yin 64e455d4bf Fix lint break on main (#28886) 2026-06-21 22:46:08 -07:00
Bingxu Chen e2540188ce [AMD] Clean up DeepSeek-R1-MXFP4 TP2/TP4 MLA GSM8K tests (#27243) 2026-06-21 21:41:19 -07:00
Lianmin Zheng 886b96621d Migrate more server args to annotated style (#28830) 2026-06-21 20:50:17 -07:00
cctryandcctry 0c065671c9 [Spec] Redo: split init_backends; account draft weights in --mem-fraction-static (#28855)
Co-authored-by: cctry <cctry@fb.com>
2026-06-21 20:45:16 -07:00
Bingxu Chen fd7874d11b [AMD] Register DP attention test (#28495) 2026-06-21 20:21:41 -07:00
Liangsheng Yin e6722c751b [Feature] Add graceful scheduler shutdown; free hisparse host buffer on exit (#28779) 2026-06-21 15:08:10 -07:00
Lianmin Zhengandhnyls2002 a4d0ff3def [misc] Make NaN-logit sanitization opt-in (default off) (#28829)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-06-21 14:31:35 -07:00
cctry 6d4ca9bc54 Cap SWA pool sizing with chunk cache (#28755) 2026-06-21 01:06:59 -07:00
Jairo David Campaña RoseroandXinyuan Tong b4dda8b3ce fix(anthropic): handle mid-conversation system messages (#26773)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-06-21 04:34:47 +00:00