Commit Graph
11858 Commits
Author SHA1 Message Date
Kangyan-ZhouandClaude Opus 4.7 77fd86f89e [ci] split stage-c-test-4-gpu-b200 to enable a low-disk runner pool (#23417)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 18:33:33 -07:00
Piotr MazurekandPiotr Mazurek 6cf0b004ca [MoE] Add LFM2 MoE tuning support + tuned configs for H100/B200/MI325X (#22791)
Co-authored-by: Piotr Mazurek <piotr.mazurek@liquid.ai>
2026-04-21 18:32:05 -07:00
Alison Shao 0e165ffbfc ci: enable /rerun-test for nightly test suites (#22830) 2026-04-21 18:28:10 -07:00
Byron Hsu c090f71bf2 feat: enable SGLANG_PATCH_TOKENIZER by default (#23409) 2026-04-21 17:53:43 -07:00
hlu1 415f64e763 Add MambaPool kvcache offloading during retraction (#22493) 2026-04-22 08:51:03 +08:00
zijiexia 1408d97408 [Docs] Improve SGLang Diffusion docs navigation and compatibility table (#23411) 2026-04-21 16:59:42 -07:00
kkandwunhuang 036edf2533 Fix docker build error (#23413)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-04-21 16:45:33 -07:00
Qiaolin YuandYuzhen Zhou c560326884 [perf] support return_routed_experts with overlap scheduling (#22911)
Co-authored-by: Yuzhen Zhou <82826991+zyzshishui@users.noreply.github.com>
2026-04-21 14:42:49 -07:00
Mingyi 9f37c1a9b0 Docs/add specforge redirect (#23406) 2026-04-21 14:35:38 -07:00
Yanbin Jiang 4f764dfbb8 [Lora] Support LoRA and multi-batch in bench_one_batch_server (#23047) 2026-04-21 14:20:11 -07:00
zijiexia 6b1e3b57d0 [docs] update logo images for google, qwen, wan, and zimage (#23404) 2026-04-21 14:12:31 -07:00
zijiexia d20ae9ceaa [docs] sync kimi-k2.6 from sgl-cookbook (#23394) 2026-04-21 13:59:55 -07:00
Charles Chen c396e4924b [bug] Fix cache salt and extra keys for prefix cache isolation (#23300) 2026-04-21 13:53:24 -07:00
e3782d04d2 fix: fallback to triton for attention-sink models (flashinfer unsupported) (#23139)
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-04-21 13:48:50 -07:00
Liangsheng Yin 6c2714f5ae guard adaptive speculative against unsupported configs (#23289) 2026-04-21 13:47:34 -07:00
5273f11fd8 [PD] Resolve missing bootstrap_room problem about fake-decode in load-balance method (#18399)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-04-21 13:47:01 -07:00
Mingyi 4c1d07fbdd docs: add redirects for /whl and /whl/:path* to external documentatio… (#23395) 2026-04-21 12:24:59 -07:00
Ma Mingfei 929e00eeab [CPU] expand the interface of shared_expert without scaling factor (#22933)
merge since this is CPU only change on sgl-kernel.
2026-04-21 20:03:39 +08:00
Yuan Luoandluoyuan.luo 48daa831ea [KDA] Fuse gate+cumsum and reuse chunk index for KDA (#23038)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-04-21 17:54:20 +08:00
Alan KaoandClaude Sonnet 4.6 8589b92a89 [AMD] Fused qk rmsnorm bf16 for amd/Kimi-K2.5-MXFP4 (#23186)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-21 02:35:09 -07:00
ybyang fa8993111d Fix: Add token heuristic increment in total_tokens load balancing (#22614) 2026-04-21 01:29:33 -07:00
Xiaoyu Zhang 0d69012ef8 Optimize LTX2 feed-forward tensor parallelism (#23221) 2026-04-21 16:29:23 +08:00
zijiexia e22dfe8fc2 [Docs] Update installation and TPU documentation to fix the render problem (#23344) 2026-04-21 01:12:50 -07:00
Mingyi 47c4b38257 docs: redirect /cookbook to /cookbook/intro (#23348) 2026-04-21 01:05:47 -07:00
Bingxu Chen 09b1d10d59 [AMD] prepare for MI300x PR runner pool: registry mirror, runner routing, threshold tuning (#23156) 2026-04-21 00:58:23 -07:00
YC Yen-Ching Tseng 74fdf9cd77 [AMD] CI - Fix the cancelled guard to AMD CI (#23338) 2026-04-21 15:45:26 +08:00
efa71ce5ab [HiCache]Fix hybrid model move_indices (#22940)
Co-authored-by: hzh0425 <hzh0425@apache.org>
Co-authored-by: flyerming <flyerming@163.com>
2026-04-21 00:15:39 -07:00
zijiexiaandMingyi 900aad5f72 [Docs] Sync docs_new with legacy docs and update migration redirects (#23337)
Co-authored-by: Mingyi <wisclmy0611@gmail.com>
2026-04-21 00:15:17 -07:00
f63def8510 [XPU] Fix DeepSeek-OCR tests under transformers 5.x (#23044)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-04-21 14:57:56 +08:00
c122d343ad [ROCm] Uniform docker to support AMD AINIC, BRCM Thor2 IBGDA NIC for MoRI-EP (#23263)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: Lzy17 <36555117+Lzy17@users.noreply.github.com>
2026-04-20 23:39:19 -07:00
jianan-gu 2cf3ac515b [Diffusion][CPU] Init CPU platform support for SGLang Diffusion (#20816) 2026-04-21 14:25:54 +08:00
Liangsheng Yin 2b2cad70d6 [Refactor] Move radix-cache utils onto RadixKey as methods (#23209) 2026-04-20 23:11:58 -07:00
a490632416 Opt-in strip of thinking tokens from radix cache (#23315)
Co-authored-by: ianliuy <ianl@alumni.usc.edu>
Co-authored-by: Wen-xuan-Xu <lilmeep727@gmail.com>
2026-04-20 22:59:50 -07:00
a8e3a534a4 [Score API] Add Multi-Item Scoring with pre-computed delimiter indices (#22544)
Co-authored-by: Chanh Nguyen <chanhnguyen@gmail.com>
Co-authored-by: Sundara Raman Ramachandran <sundar24295@gmail.com>
2026-04-20 22:50:40 -07:00
Divyam Agrawal cfd49e233c Fix formatting for ACM-VIT in README acknowledgements section (#23325) 2026-04-20 22:28:45 -07:00
jsheng_Linkedin 6d47dc8f6d [CI][MLA] Enable deterministic inference for MGSM MLA FP8 test (#23303) 2026-04-20 22:26:26 -07:00
Shangming Cai a58c7f381e [PD] Fix clip logic when state indices lens are mismatch (#23323) 2026-04-21 13:22:20 +08:00
Yuhao Yang 5595f6e988 Fix trtllm mla chunked-prefill zero-length bug (#22291) (#22688) 2026-04-20 22:10:13 -07:00
Mingyi 82beaf1748 Docs/url redirect (#23312) 2026-04-20 21:26:18 -07:00
Liangsheng Yin 6cc2eee50d [misc] CI hygiene: enforce __main__ entry, drop silent-skipped tests, fix rerun-test protoc (#23305) 2026-04-20 21:16:24 -07:00
Alison Shao 6b19e8a452 ci: reduce scheduled PR test from 4x to 3x daily (#23313) 2026-04-20 20:53:13 -07:00
amote-i 301604f953 [NPU] [DOC] Quick start doc for Ascend NPU (#23238) 2026-04-21 11:19:09 +08:00
Lewisand百麒 0d0405273b [Fix] Solve the error lead by _commit_transfer_to_req() when using IntraNode NVLink in PD disaggregation (#23252)
Co-authored-by: 百麒 <yaozhong.lyz@alibaba-inc.com>
2026-04-21 11:02:18 +08:00
Ke Bao 50fc2c9e23 Fix hybrid swa chunked prefill oom (#23174) 2026-04-21 10:46:45 +08:00
Zhangheng ab3ce02de9 [Hybrid-Cache]: Refactor hybrid_pool_assembler.py (#23243) 2026-04-21 10:45:23 +08:00
ishandhananiandjthomson04 3c007ee5d4 fix(hicache): emit KV events for L2 host cache insertions (#22894)
Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
Signed-off-by: Ishan Dhanani <ishandhanani@gmail.com>
Co-authored-by: jthomson04 <jwillthomson19@gmail.com>
2026-04-20 19:07:03 -07:00
ChangLiu0709 ac08ebed65 [AMD] Resolve Qwen3.5 MTP (speculative decoding) radix cache conflict. (#22908) 2026-04-20 18:17:11 -07:00
Liangsheng Yin c7a4ebf3c8 [Refactor] Replace page_align_keys helper with RadixKey.page_aligned method (#23107) 2026-04-20 18:10:42 -07:00
Mingyi 712b01d875 Update CODEOWNERS to include new documentation paths for docs and doc… (#23293) 2026-04-20 16:48:41 -07:00
Tarushii Goel 3e367f9bcd [sgl] fix incorrect behavior in cuda graph draft extend (#22832) 2026-04-20 16:29:16 -07:00