 Teng MaandShangming Cai
|
97fc950635
|
[Docker] chore: update mooncake wheel version (#20731)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-04-22 13:36:10 -07:00 |
|
 jianan-guandMa Mingfei
|
ad0fc88810
|
[CPU] [Quantization] Add GPTQ/AWQ 4bits quantization support for CPU (#22685)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-22 13:34:02 -07:00 |
|
Byron Hsu
|
0b77284587
|
[minor] Make DEFAULT_FORCE_STREAM_INTERVAL configurable via SGLANG_FORCE_STREAM_INTERVAL (#23215)
|
2026-04-22 13:05:40 -07:00 |
|
JasonHe-WQ
|
f85e3140bf
|
Fix:fix(timeout): fix timeout not propagated (#21944)
|
2026-04-22 12:48:48 -07:00 |
|
 Chao ShiandChao Shi
|
94fb13db92
|
[devcontainer] Fix build error (#23478)
Co-authored-by: Chao Shi <chao.shi@alibaba-inc.com>
|
2026-04-22 12:24:15 -07:00 |
|
zijiexia
|
9b142df334
|
Docs/add sglang omni redirect (#23437)
|
2026-04-22 11:37:27 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.7
|
14ac14287c
|
[CI] /rerun-stage: auto-include wheel build when PR modifies sgl-kernel/ (#23492)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-22 11:28:06 -07:00 |
|
Xinyuan Tong
|
de962f3274
|
docs(cookbook): add Qwen3.6-27B dense variant (#23486)
|
2026-04-23 01:22:46 +08:00 |
|
 Yuxuan ZhangandXinyuan Tong
|
28cfd3d272
|
Support defer_loading field at function level for Chat Completions API (#22702)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-04-22 10:09:54 -07:00 |
|
Todobe
|
92f28e9ba8
|
[NPU]Fix GLM-4.7-Flash failed on NPU (#22509)
|
2026-04-23 01:06:58 +08:00 |
|
cctry
|
0addd185af
|
Fix /generate endpoint crash when sampling params contain null values (#23401)
|
2026-04-22 09:56:10 -07:00 |
|
Aleksi Vesanto
|
ac351c1f04
|
[diffusion] [AMD] model: allow AITER backends in Flux 2 pipeline (#22802)
|
2026-04-22 08:15:44 -07:00 |
|
Shenxiu Liu
|
8b78e0888c
|
Skip mamba_pool_idx revert for session requests in _get_new_batch_prefill_raw (#23327)
|
2026-04-22 22:28:06 +08:00 |
|
Mick
|
4323fce82a
|
fix: dot-boundary match in is_layer_skipped for FP8 modules_to_not_convert (#23467)
|
2026-04-22 22:16:22 +08:00 |
|
amote-i
|
18f3310aad
|
[NPU] [DOC] Update Ascend NPU best practice (#23459)
|
2026-04-22 17:51:28 +08:00 |
|
Shangming Cai
|
1c06a3d072
|
[CI] Move disaggregation basic CI back to 2-gpu suite (#23447)
|
2026-04-22 17:50:33 +08:00 |
|
Lianmin Zheng
|
6a3c070ee3
|
Add 'allready' to ignore words list in .codespellrc (#23465)
|
2026-04-22 02:39:04 -07:00 |
|
Ming Yang
|
7b10f01d1c
|
[model_runner] Label forward steps in profile traces with mode and token counts (#23419)
|
2026-04-22 02:31:18 -07:00 |
|
inkcherry
|
1e34cd0ba5
|
PD streaming: batch notify + SSE fast path (#22658)
|
2026-04-22 02:21:02 -07:00 |
|
Shangming Cai
|
fa85bdf4ed
|
chore: bump mooncake version to 0.3.10.post2 (#23439)
|
2026-04-22 15:01:47 +08:00 |
|
 Jia GuoandClaude Opus 4.6
|
286fba2073
|
ci: use rerun_failed_jobs for skipped workflows in /rerun-failed-ci (#23008)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-21 23:59:15 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.7
|
88c9bab830
|
[diffusion] ci: allow using prebuilt sgl-kernel wheel for GT regeneration (#23443)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-21 23:44:02 -07:00 |
|
 
|
5c245d978f
|
[Diffusion] Add mixed-resolution benchmark support (for #20762) (#20863)
Signed-off-by: Fengyuan Yu <15fengyuan@gmail.com>
Co-authored-by: Fengyuan Yu <15fengyuan@gmail.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-04-22 09:22:19 +03:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) cctryandgemini-code-assist[bot]
|
e39f0f4ff3
|
Use libdevice tanh and support 2D-strided tensors in fused softcap kernel (#23157)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-04-21 22:54:37 -07:00 |
|
 
|
c3ea2d7b92
|
Rename mixed_with_decode_tokens in mixed chunk prefill adder (#6506)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
|
2026-04-21 22:48:34 -07:00 |
|
Alison Shao
|
04b1caf75b
|
ci: enable /rerun-test for multimodal gen PR tests (#22828)
|
2026-04-21 21:34:14 -07:00 |
|
Tarushii Goel
|
7607e4d180
|
py-spy without --native for ARM devices (#23410)
|
2026-04-21 20:45:52 -07:00 |
|
Tarushii Goel
|
3ebf066d13
|
[sgl] update specdec sampling kernel to return valid token ID (#22643)
|
2026-04-21 20:28:19 -07:00 |
|
Yuhao Yang
|
f41f1a74a4
|
[diffusion] chore: support custom output folder name in GT generation workflow (#23422)
|
2026-04-22 11:18:21 +08:00 |
|
 jianzhao-xuandJianzhao Xu
|
2f3e6a3143
|
[NPU] offloading docs update (#23378)
Co-authored-by: Jianzhao Xu <xujianchao@huawei.com>
|
2026-04-22 11:01:55 +08:00 |
|
 shuwennandClaude Opus 4.6
|
4befc31408
|
fix: pass v_head_dim to MHA KV pools and validate MiMo HiCache geometry (#23173)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-21 19:48:45 -07:00 |
|
MARATRIX
|
bf5e71dcec
|
[MUSA][19/N] Support HiCache with pin_memory allocator (#23361)
Signed-off-by: yafeng.li <yafeng.li@mthreads.com>
|
2026-04-21 19:45:53 -07:00 |
|
MingxuZh
|
7c399f3c82
|
Update pr-test-xeon.yml cancel-in-progress config (#23420)
merge this one, as it fixed xeon ci blocking issue.
|
2026-04-22 10:12:36 +08:00 |
|
 Kangyan-ZhouandClaude Opus 4.7
|
77fd86f89e
|
[ci] split stage-c-test-4-gpu-b200 to enable a low-disk runner pool (#23417)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-21 18:33:33 -07:00 |
|
 Piotr MazurekandPiotr Mazurek
|
6cf0b004ca
|
[MoE] Add LFM2 MoE tuning support + tuned configs for H100/B200/MI325X (#22791)
Co-authored-by: Piotr Mazurek <piotr.mazurek@liquid.ai>
|
2026-04-21 18:32:05 -07:00 |
|
Alison Shao
|
0e165ffbfc
|
ci: enable /rerun-test for nightly test suites (#22830)
|
2026-04-21 18:28:10 -07:00 |
|
Byron Hsu
|
c090f71bf2
|
feat: enable SGLANG_PATCH_TOKENIZER by default (#23409)
|
2026-04-21 17:53:43 -07:00 |
|
hlu1
|
415f64e763
|
Add MambaPool kvcache offloading during retraction (#22493)
|
2026-04-22 08:51:03 +08:00 |
|
zijiexia
|
1408d97408
|
[Docs] Improve SGLang Diffusion docs navigation and compatibility table (#23411)
|
2026-04-21 16:59:42 -07:00 |
|
 kkandwunhuang
|
036edf2533
|
Fix docker build error (#23413)
Co-authored-by: wunhuang <wunhuang@amd.com>
|
2026-04-21 16:45:33 -07:00 |
|
 Qiaolin YuandYuzhen Zhou
|
c560326884
|
[perf] support return_routed_experts with overlap scheduling (#22911)
Co-authored-by: Yuzhen Zhou <82826991+zyzshishui@users.noreply.github.com>
|
2026-04-21 14:42:49 -07:00 |
|
Mingyi
|
9f37c1a9b0
|
Docs/add specforge redirect (#23406)
|
2026-04-21 14:35:38 -07:00 |
|
Yanbin Jiang
|
4f764dfbb8
|
[Lora] Support LoRA and multi-batch in bench_one_batch_server (#23047)
|
2026-04-21 14:20:11 -07:00 |
|
zijiexia
|
6b1e3b57d0
|
[docs] update logo images for google, qwen, wan, and zimage (#23404)
|
2026-04-21 14:12:31 -07:00 |
|
zijiexia
|
d20ae9ceaa
|
[docs] sync kimi-k2.6 from sgl-cookbook (#23394)
|
2026-04-21 13:59:55 -07:00 |
|
Charles Chen
|
c396e4924b
|
[bug] Fix cache salt and extra keys for prefix cache isolation (#23300)
|
2026-04-21 13:53:24 -07:00 |
|
 
|
e3782d04d2
|
fix: fallback to triton for attention-sink models (flashinfer unsupported) (#23139)
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
|
2026-04-21 13:48:50 -07:00 |
|
Liangsheng Yin
|
6c2714f5ae
|
guard adaptive speculative against unsupported configs (#23289)
|
2026-04-21 13:47:34 -07:00 |
|
 ![gemini-code-assist[bot]](/assets/img/avatar_default.png)
|
5273f11fd8
|
[PD] Resolve missing bootstrap_room problem about fake-decode in load-balance method (#18399)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
|
2026-04-21 13:47:01 -07:00 |
|
Mingyi
|
4c1d07fbdd
|
docs: add redirects for /whl and /whl/:path* to external documentatio… (#23395)
|
2026-04-21 12:24:59 -07:00 |
|