Commit Graph
11872 Commits
Author SHA1 Message Date
Shangming Cai fa85bdf4ed chore: bump mooncake version to 0.3.10.post2 (#23439) 2026-04-22 15:01:47 +08:00
Jia GuoandClaude Opus 4.6 286fba2073 ci: use rerun_failed_jobs for skipped workflows in /rerun-failed-ci (#23008)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-21 23:59:15 -07:00
Kangyan-ZhouandClaude Opus 4.7 88c9bab830 [diffusion] ci: allow using prebuilt sgl-kernel wheel for GT regeneration (#23443)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 23:44:02 -07:00
5c245d978f [Diffusion] Add mixed-resolution benchmark support (for #20762) (#20863)
Signed-off-by: Fengyuan Yu <15fengyuan@gmail.com>
Co-authored-by: Fengyuan Yu <15fengyuan@gmail.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-04-22 09:22:19 +03:00
cctryandgemini-code-assist[bot] e39f0f4ff3 Use libdevice tanh and support 2D-strided tensors in fused softcap kernel (#23157)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-21 22:54:37 -07:00
c3ea2d7b92 Rename mixed_with_decode_tokens in mixed chunk prefill adder (#6506)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-04-21 22:48:34 -07:00
Alison Shao 04b1caf75b ci: enable /rerun-test for multimodal gen PR tests (#22828) 2026-04-21 21:34:14 -07:00
Tarushii Goel 7607e4d180 py-spy without --native for ARM devices (#23410) 2026-04-21 20:45:52 -07:00
Tarushii Goel 3ebf066d13 [sgl] update specdec sampling kernel to return valid token ID (#22643) 2026-04-21 20:28:19 -07:00
Yuhao Yang f41f1a74a4 [diffusion] chore: support custom output folder name in GT generation workflow (#23422) 2026-04-22 11:18:21 +08:00
jianzhao-xuandJianzhao Xu 2f3e6a3143 [NPU] offloading docs update (#23378)
Co-authored-by: Jianzhao Xu <xujianchao@huawei.com>
2026-04-22 11:01:55 +08:00
shuwennandClaude Opus 4.6 4befc31408 fix: pass v_head_dim to MHA KV pools and validate MiMo HiCache geometry (#23173)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-21 19:48:45 -07:00
MARATRIX bf5e71dcec [MUSA][19/N] Support HiCache with pin_memory allocator (#23361)
Signed-off-by: yafeng.li <yafeng.li@mthreads.com>
2026-04-21 19:45:53 -07:00
MingxuZh 7c399f3c82 Update pr-test-xeon.yml cancel-in-progress config (#23420)
merge this one, as it fixed xeon ci blocking issue.
2026-04-22 10:12:36 +08:00
Kangyan-ZhouandClaude Opus 4.7 77fd86f89e [ci] split stage-c-test-4-gpu-b200 to enable a low-disk runner pool (#23417)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 18:33:33 -07:00
Piotr MazurekandPiotr Mazurek 6cf0b004ca [MoE] Add LFM2 MoE tuning support + tuned configs for H100/B200/MI325X (#22791)
Co-authored-by: Piotr Mazurek <piotr.mazurek@liquid.ai>
2026-04-21 18:32:05 -07:00
Alison Shao 0e165ffbfc ci: enable /rerun-test for nightly test suites (#22830) 2026-04-21 18:28:10 -07:00
Byron Hsu c090f71bf2 feat: enable SGLANG_PATCH_TOKENIZER by default (#23409) 2026-04-21 17:53:43 -07:00
hlu1 415f64e763 Add MambaPool kvcache offloading during retraction (#22493) 2026-04-22 08:51:03 +08:00
zijiexia 1408d97408 [Docs] Improve SGLang Diffusion docs navigation and compatibility table (#23411) 2026-04-21 16:59:42 -07:00
kkandwunhuang 036edf2533 Fix docker build error (#23413)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-04-21 16:45:33 -07:00
Qiaolin YuandYuzhen Zhou c560326884 [perf] support return_routed_experts with overlap scheduling (#22911)
Co-authored-by: Yuzhen Zhou <82826991+zyzshishui@users.noreply.github.com>
2026-04-21 14:42:49 -07:00
Mingyi 9f37c1a9b0 Docs/add specforge redirect (#23406) 2026-04-21 14:35:38 -07:00
Yanbin Jiang 4f764dfbb8 [Lora] Support LoRA and multi-batch in bench_one_batch_server (#23047) 2026-04-21 14:20:11 -07:00
zijiexia 6b1e3b57d0 [docs] update logo images for google, qwen, wan, and zimage (#23404) 2026-04-21 14:12:31 -07:00
zijiexia d20ae9ceaa [docs] sync kimi-k2.6 from sgl-cookbook (#23394) 2026-04-21 13:59:55 -07:00
Charles Chen c396e4924b [bug] Fix cache salt and extra keys for prefix cache isolation (#23300) 2026-04-21 13:53:24 -07:00
e3782d04d2 fix: fallback to triton for attention-sink models (flashinfer unsupported) (#23139)
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-04-21 13:48:50 -07:00
Liangsheng Yin 6c2714f5ae guard adaptive speculative against unsupported configs (#23289) 2026-04-21 13:47:34 -07:00
5273f11fd8 [PD] Resolve missing bootstrap_room problem about fake-decode in load-balance method (#18399)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-04-21 13:47:01 -07:00
Mingyi 4c1d07fbdd docs: add redirects for /whl and /whl/:path* to external documentatio… (#23395) 2026-04-21 12:24:59 -07:00
Ma Mingfei 929e00eeab [CPU] expand the interface of shared_expert without scaling factor (#22933)
merge since this is CPU only change on sgl-kernel.
2026-04-21 20:03:39 +08:00
Yuan Luoandluoyuan.luo 48daa831ea [KDA] Fuse gate+cumsum and reuse chunk index for KDA (#23038)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-04-21 17:54:20 +08:00
Alan KaoandClaude Sonnet 4.6 8589b92a89 [AMD] Fused qk rmsnorm bf16 for amd/Kimi-K2.5-MXFP4 (#23186)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-21 02:35:09 -07:00
ybyang fa8993111d Fix: Add token heuristic increment in total_tokens load balancing (#22614) 2026-04-21 01:29:33 -07:00
Xiaoyu Zhang 0d69012ef8 Optimize LTX2 feed-forward tensor parallelism (#23221) 2026-04-21 16:29:23 +08:00
zijiexia e22dfe8fc2 [Docs] Update installation and TPU documentation to fix the render problem (#23344) 2026-04-21 01:12:50 -07:00
Mingyi 47c4b38257 docs: redirect /cookbook to /cookbook/intro (#23348) 2026-04-21 01:05:47 -07:00
Bingxu Chen 09b1d10d59 [AMD] prepare for MI300x PR runner pool: registry mirror, runner routing, threshold tuning (#23156) 2026-04-21 00:58:23 -07:00
YC Yen-Ching Tseng 74fdf9cd77 [AMD] CI - Fix the cancelled guard to AMD CI (#23338) 2026-04-21 15:45:26 +08:00
efa71ce5ab [HiCache]Fix hybrid model move_indices (#22940)
Co-authored-by: hzh0425 <hzh0425@apache.org>
Co-authored-by: flyerming <flyerming@163.com>
2026-04-21 00:15:39 -07:00
zijiexiaandMingyi 900aad5f72 [Docs] Sync docs_new with legacy docs and update migration redirects (#23337)
Co-authored-by: Mingyi <wisclmy0611@gmail.com>
2026-04-21 00:15:17 -07:00
f63def8510 [XPU] Fix DeepSeek-OCR tests under transformers 5.x (#23044)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-04-21 14:57:56 +08:00
c122d343ad [ROCm] Uniform docker to support AMD AINIC, BRCM Thor2 IBGDA NIC for MoRI-EP (#23263)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: Lzy17 <36555117+Lzy17@users.noreply.github.com>
2026-04-20 23:39:19 -07:00
jianan-gu 2cf3ac515b [Diffusion][CPU] Init CPU platform support for SGLang Diffusion (#20816) 2026-04-21 14:25:54 +08:00
Liangsheng Yin 2b2cad70d6 [Refactor] Move radix-cache utils onto RadixKey as methods (#23209) 2026-04-20 23:11:58 -07:00
a490632416 Opt-in strip of thinking tokens from radix cache (#23315)
Co-authored-by: ianliuy <ianl@alumni.usc.edu>
Co-authored-by: Wen-xuan-Xu <lilmeep727@gmail.com>
2026-04-20 22:59:50 -07:00
a8e3a534a4 [Score API] Add Multi-Item Scoring with pre-computed delimiter indices (#22544)
Co-authored-by: Chanh Nguyen <chanhnguyen@gmail.com>
Co-authored-by: Sundara Raman Ramachandran <sundar24295@gmail.com>
2026-04-20 22:50:40 -07:00
Divyam Agrawal cfd49e233c Fix formatting for ACM-VIT in README acknowledgements section (#23325) 2026-04-20 22:28:45 -07:00
jsheng_Linkedin 6d47dc8f6d [CI][MLA] Enable deterministic inference for MGSM MLA FP8 test (#23303) 2026-04-20 22:26:26 -07:00