Commit Graph
11903 Commits
Author SHA1 Message Date
Alison Shao 267c2c0849 test: move test_epd_disaggregation to nightly-4-gpu (#23518) 2026-04-22 20:59:19 -07:00
Liangsheng Yin 0f21fe924a fix ngram greedy verify kwarg (#23521) 2026-04-22 20:49:54 -07:00
oriandzhiguo.qin 887d380ace [MUSA] Resolve output garbage in Context Parallel on MusaFlashAttentionBackend (#23270)
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
2026-04-22 20:22:20 -07:00
zijiexia 6490afe36e [docs] add deprecation notice banner to legacy documentation site (#23516) 2026-04-22 20:17:53 -07:00
Polisetty V R K Jyothendra Varma 214c35b031 [Intel GPU] Update xpu.Dockerfile to python 3.12 version (#23367) 2026-04-23 09:23:52 +08:00
Jia Guo b3e6cf60aa ci: build sgl-kernel wheels for both cu129 and cu130 (#23497) 2026-04-22 18:08:36 -07:00
Kangyan-ZhouandClaude Opus 4.7 c689f774a4 [CI] /rerun-stage: fix workflow-run URL lookup for sgl-kernel PRs (#23510)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 17:48:38 -07:00
Kangyan-ZhouandClaude Opus 4.7 0bee211335 [CI] Broaden stage-b-test-4-gpu-b200 runner pool to low-disk label (#23505)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 17:14:09 -07:00
Liangsheng YinandSoluMilken f611dd24f1 fix retrive -> retrieve typo (#23503)
Co-authored-by: SoluMilken <19161836+solumilken@users.noreply.github.com>
2026-04-22 16:35:04 -07:00
Yanbin Jiang 917d2aa1dc [LoRA] Fix EP + per-expert MoE LoRA illegal memory access (#23178) 2026-04-22 14:22:32 -07:00
Sam Shleifer b9e33d6a5b Dual MoE CUDA graph capture for lora/nolora batches (#22809) 2026-04-22 14:11:11 -07:00
Sahithi Chigurupati 9591033179 [CI] GB200 nightly: on-demand PR/branch image build and config filter (#23086) 2026-04-22 13:51:25 -07:00
Teng MaandShangming Cai 97fc950635 [Docker] chore: update mooncake wheel version (#20731)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-04-22 13:36:10 -07:00
jianan-guandMa Mingfei ad0fc88810 [CPU] [Quantization] Add GPTQ/AWQ 4bits quantization support for CPU (#22685)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-04-22 13:34:02 -07:00
Byron Hsu 0b77284587 [minor] Make DEFAULT_FORCE_STREAM_INTERVAL configurable via SGLANG_FORCE_STREAM_INTERVAL (#23215) 2026-04-22 13:05:40 -07:00
JasonHe-WQ f85e3140bf Fix:fix(timeout): fix timeout not propagated (#21944) 2026-04-22 12:48:48 -07:00
Chao ShiandChao Shi 94fb13db92 [devcontainer] Fix build error (#23478)
Co-authored-by: Chao Shi <chao.shi@alibaba-inc.com>
2026-04-22 12:24:15 -07:00
zijiexia 9b142df334 Docs/add sglang omni redirect (#23437) 2026-04-22 11:37:27 -07:00
Kangyan-ZhouandClaude Opus 4.7 14ac14287c [CI] /rerun-stage: auto-include wheel build when PR modifies sgl-kernel/ (#23492)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-22 11:28:06 -07:00
Xinyuan Tong de962f3274 docs(cookbook): add Qwen3.6-27B dense variant (#23486) 2026-04-23 01:22:46 +08:00
Yuxuan ZhangandXinyuan Tong 28cfd3d272 Support defer_loading field at function level for Chat Completions API (#22702)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-04-22 10:09:54 -07:00
Todobe 92f28e9ba8 [NPU]Fix GLM-4.7-Flash failed on NPU (#22509) 2026-04-23 01:06:58 +08:00
cctry 0addd185af Fix /generate endpoint crash when sampling params contain null values (#23401) 2026-04-22 09:56:10 -07:00
Aleksi Vesanto ac351c1f04 [diffusion] [AMD] model: allow AITER backends in Flux 2 pipeline (#22802) 2026-04-22 08:15:44 -07:00
Shenxiu Liu 8b78e0888c Skip mamba_pool_idx revert for session requests in _get_new_batch_prefill_raw (#23327) 2026-04-22 22:28:06 +08:00
Mick 4323fce82a fix: dot-boundary match in is_layer_skipped for FP8 modules_to_not_convert (#23467) 2026-04-22 22:16:22 +08:00
amote-i 18f3310aad [NPU] [DOC] Update Ascend NPU best practice (#23459) 2026-04-22 17:51:28 +08:00
Shangming Cai 1c06a3d072 [CI] Move disaggregation basic CI back to 2-gpu suite (#23447) 2026-04-22 17:50:33 +08:00
Lianmin Zheng 6a3c070ee3 Add 'allready' to ignore words list in .codespellrc (#23465) 2026-04-22 02:39:04 -07:00
Ming Yang 7b10f01d1c [model_runner] Label forward steps in profile traces with mode and token counts (#23419) 2026-04-22 02:31:18 -07:00
inkcherry 1e34cd0ba5 PD streaming: batch notify + SSE fast path (#22658) 2026-04-22 02:21:02 -07:00
Shangming Cai fa85bdf4ed chore: bump mooncake version to 0.3.10.post2 (#23439) 2026-04-22 15:01:47 +08:00
Jia GuoandClaude Opus 4.6 286fba2073 ci: use rerun_failed_jobs for skipped workflows in /rerun-failed-ci (#23008)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-21 23:59:15 -07:00
Kangyan-ZhouandClaude Opus 4.7 88c9bab830 [diffusion] ci: allow using prebuilt sgl-kernel wheel for GT regeneration (#23443)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 23:44:02 -07:00
5c245d978f [Diffusion] Add mixed-resolution benchmark support (for #20762) (#20863)
Signed-off-by: Fengyuan Yu <15fengyuan@gmail.com>
Co-authored-by: Fengyuan Yu <15fengyuan@gmail.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-04-22 09:22:19 +03:00
cctryandgemini-code-assist[bot] e39f0f4ff3 Use libdevice tanh and support 2D-strided tensors in fused softcap kernel (#23157)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-21 22:54:37 -07:00
c3ea2d7b92 Rename mixed_with_decode_tokens in mixed chunk prefill adder (#6506)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-04-21 22:48:34 -07:00
Alison Shao 04b1caf75b ci: enable /rerun-test for multimodal gen PR tests (#22828) 2026-04-21 21:34:14 -07:00
Tarushii Goel 7607e4d180 py-spy without --native for ARM devices (#23410) 2026-04-21 20:45:52 -07:00
Tarushii Goel 3ebf066d13 [sgl] update specdec sampling kernel to return valid token ID (#22643) 2026-04-21 20:28:19 -07:00
Yuhao Yang f41f1a74a4 [diffusion] chore: support custom output folder name in GT generation workflow (#23422) 2026-04-22 11:18:21 +08:00
jianzhao-xuandJianzhao Xu 2f3e6a3143 [NPU] offloading docs update (#23378)
Co-authored-by: Jianzhao Xu <xujianchao@huawei.com>
2026-04-22 11:01:55 +08:00
shuwennandClaude Opus 4.6 4befc31408 fix: pass v_head_dim to MHA KV pools and validate MiMo HiCache geometry (#23173)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-21 19:48:45 -07:00
MARATRIX bf5e71dcec [MUSA][19/N] Support HiCache with pin_memory allocator (#23361)
Signed-off-by: yafeng.li <yafeng.li@mthreads.com>
2026-04-21 19:45:53 -07:00
MingxuZh 7c399f3c82 Update pr-test-xeon.yml cancel-in-progress config (#23420)
merge this one, as it fixed xeon ci blocking issue.
2026-04-22 10:12:36 +08:00
Kangyan-ZhouandClaude Opus 4.7 77fd86f89e [ci] split stage-c-test-4-gpu-b200 to enable a low-disk runner pool (#23417)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 18:33:33 -07:00
Piotr MazurekandPiotr Mazurek 6cf0b004ca [MoE] Add LFM2 MoE tuning support + tuned configs for H100/B200/MI325X (#22791)
Co-authored-by: Piotr Mazurek <piotr.mazurek@liquid.ai>
2026-04-21 18:32:05 -07:00
Alison Shao 0e165ffbfc ci: enable /rerun-test for nightly test suites (#22830) 2026-04-21 18:28:10 -07:00
Byron Hsu c090f71bf2 feat: enable SGLANG_PATCH_TOKENIZER by default (#23409) 2026-04-21 17:53:43 -07:00
hlu1 415f64e763 Add MambaPool kvcache offloading during retraction (#22493) 2026-04-22 08:51:03 +08:00