Alison Shao
|
267c2c0849
|
test: move test_epd_disaggregation to nightly-4-gpu (#23518)
|
2026-04-22 20:59:19 -07:00 |
|
Liangsheng Yin
|
0f21fe924a
|
fix ngram greedy verify kwarg (#23521)
|
2026-04-22 20:49:54 -07:00 |
|
 oriandzhiguo.qin
|
887d380ace
|
[MUSA] Resolve output garbage in Context Parallel on MusaFlashAttentionBackend (#23270)
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
|
2026-04-22 20:22:20 -07:00 |
|
zijiexia
|
6490afe36e
|
[docs] add deprecation notice banner to legacy documentation site (#23516)
|
2026-04-22 20:17:53 -07:00 |
|
Polisetty V R K Jyothendra Varma
|
214c35b031
|
[Intel GPU] Update xpu.Dockerfile to python 3.12 version (#23367)
|
2026-04-23 09:23:52 +08:00 |
|
Jia Guo
|
b3e6cf60aa
|
ci: build sgl-kernel wheels for both cu129 and cu130 (#23497)
|
2026-04-22 18:08:36 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.7
|
c689f774a4
|
[CI] /rerun-stage: fix workflow-run URL lookup for sgl-kernel PRs (#23510)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-22 17:48:38 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.7
|
0bee211335
|
[CI] Broaden stage-b-test-4-gpu-b200 runner pool to low-disk label (#23505)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-22 17:14:09 -07:00 |
|
 Liangsheng YinandSoluMilken
|
f611dd24f1
|
fix retrive -> retrieve typo (#23503)
Co-authored-by: SoluMilken <19161836+solumilken@users.noreply.github.com>
|
2026-04-22 16:35:04 -07:00 |
|
Yanbin Jiang
|
917d2aa1dc
|
[LoRA] Fix EP + per-expert MoE LoRA illegal memory access (#23178)
|
2026-04-22 14:22:32 -07:00 |
|
Sam Shleifer
|
b9e33d6a5b
|
Dual MoE CUDA graph capture for lora/nolora batches (#22809)
|
2026-04-22 14:11:11 -07:00 |
|
Sahithi Chigurupati
|
9591033179
|
[CI] GB200 nightly: on-demand PR/branch image build and config filter (#23086)
|
2026-04-22 13:51:25 -07:00 |
|
 Teng MaandShangming Cai
|
97fc950635
|
[Docker] chore: update mooncake wheel version (#20731)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-04-22 13:36:10 -07:00 |
|
 jianan-guandMa Mingfei
|
ad0fc88810
|
[CPU] [Quantization] Add GPTQ/AWQ 4bits quantization support for CPU (#22685)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-22 13:34:02 -07:00 |
|
Byron Hsu
|
0b77284587
|
[minor] Make DEFAULT_FORCE_STREAM_INTERVAL configurable via SGLANG_FORCE_STREAM_INTERVAL (#23215)
|
2026-04-22 13:05:40 -07:00 |
|
JasonHe-WQ
|
f85e3140bf
|
Fix:fix(timeout): fix timeout not propagated (#21944)
|
2026-04-22 12:48:48 -07:00 |
|
 Chao ShiandChao Shi
|
94fb13db92
|
[devcontainer] Fix build error (#23478)
Co-authored-by: Chao Shi <chao.shi@alibaba-inc.com>
|
2026-04-22 12:24:15 -07:00 |
|
zijiexia
|
9b142df334
|
Docs/add sglang omni redirect (#23437)
|
2026-04-22 11:37:27 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.7
|
14ac14287c
|
[CI] /rerun-stage: auto-include wheel build when PR modifies sgl-kernel/ (#23492)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-22 11:28:06 -07:00 |
|
Xinyuan Tong
|
de962f3274
|
docs(cookbook): add Qwen3.6-27B dense variant (#23486)
|
2026-04-23 01:22:46 +08:00 |
|
 Yuxuan ZhangandXinyuan Tong
|
28cfd3d272
|
Support defer_loading field at function level for Chat Completions API (#22702)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-04-22 10:09:54 -07:00 |
|
Todobe
|
92f28e9ba8
|
[NPU]Fix GLM-4.7-Flash failed on NPU (#22509)
|
2026-04-23 01:06:58 +08:00 |
|
cctry
|
0addd185af
|
Fix /generate endpoint crash when sampling params contain null values (#23401)
|
2026-04-22 09:56:10 -07:00 |
|
Aleksi Vesanto
|
ac351c1f04
|
[diffusion] [AMD] model: allow AITER backends in Flux 2 pipeline (#22802)
|
2026-04-22 08:15:44 -07:00 |
|
Shenxiu Liu
|
8b78e0888c
|
Skip mamba_pool_idx revert for session requests in _get_new_batch_prefill_raw (#23327)
|
2026-04-22 22:28:06 +08:00 |
|
Mick
|
4323fce82a
|
fix: dot-boundary match in is_layer_skipped for FP8 modules_to_not_convert (#23467)
|
2026-04-22 22:16:22 +08:00 |
|
amote-i
|
18f3310aad
|
[NPU] [DOC] Update Ascend NPU best practice (#23459)
|
2026-04-22 17:51:28 +08:00 |
|
Shangming Cai
|
1c06a3d072
|
[CI] Move disaggregation basic CI back to 2-gpu suite (#23447)
|
2026-04-22 17:50:33 +08:00 |
|
Lianmin Zheng
|
6a3c070ee3
|
Add 'allready' to ignore words list in .codespellrc (#23465)
|
2026-04-22 02:39:04 -07:00 |
|
Ming Yang
|
7b10f01d1c
|
[model_runner] Label forward steps in profile traces with mode and token counts (#23419)
|
2026-04-22 02:31:18 -07:00 |
|
inkcherry
|
1e34cd0ba5
|
PD streaming: batch notify + SSE fast path (#22658)
|
2026-04-22 02:21:02 -07:00 |
|
Shangming Cai
|
fa85bdf4ed
|
chore: bump mooncake version to 0.3.10.post2 (#23439)
|
2026-04-22 15:01:47 +08:00 |
|
 Jia GuoandClaude Opus 4.6
|
286fba2073
|
ci: use rerun_failed_jobs for skipped workflows in /rerun-failed-ci (#23008)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-21 23:59:15 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.7
|
88c9bab830
|
[diffusion] ci: allow using prebuilt sgl-kernel wheel for GT regeneration (#23443)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-21 23:44:02 -07:00 |
|
 
|
5c245d978f
|
[Diffusion] Add mixed-resolution benchmark support (for #20762) (#20863)
Signed-off-by: Fengyuan Yu <15fengyuan@gmail.com>
Co-authored-by: Fengyuan Yu <15fengyuan@gmail.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-04-22 09:22:19 +03:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) cctryandgemini-code-assist[bot]
|
e39f0f4ff3
|
Use libdevice tanh and support 2D-strided tensors in fused softcap kernel (#23157)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-04-21 22:54:37 -07:00 |
|
 
|
c3ea2d7b92
|
Rename mixed_with_decode_tokens in mixed chunk prefill adder (#6506)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
|
2026-04-21 22:48:34 -07:00 |
|
Alison Shao
|
04b1caf75b
|
ci: enable /rerun-test for multimodal gen PR tests (#22828)
|
2026-04-21 21:34:14 -07:00 |
|
Tarushii Goel
|
7607e4d180
|
py-spy without --native for ARM devices (#23410)
|
2026-04-21 20:45:52 -07:00 |
|
Tarushii Goel
|
3ebf066d13
|
[sgl] update specdec sampling kernel to return valid token ID (#22643)
|
2026-04-21 20:28:19 -07:00 |
|
Yuhao Yang
|
f41f1a74a4
|
[diffusion] chore: support custom output folder name in GT generation workflow (#23422)
|
2026-04-22 11:18:21 +08:00 |
|
 jianzhao-xuandJianzhao Xu
|
2f3e6a3143
|
[NPU] offloading docs update (#23378)
Co-authored-by: Jianzhao Xu <xujianchao@huawei.com>
|
2026-04-22 11:01:55 +08:00 |
|
 shuwennandClaude Opus 4.6
|
4befc31408
|
fix: pass v_head_dim to MHA KV pools and validate MiMo HiCache geometry (#23173)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-21 19:48:45 -07:00 |
|
MARATRIX
|
bf5e71dcec
|
[MUSA][19/N] Support HiCache with pin_memory allocator (#23361)
Signed-off-by: yafeng.li <yafeng.li@mthreads.com>
|
2026-04-21 19:45:53 -07:00 |
|
MingxuZh
|
7c399f3c82
|
Update pr-test-xeon.yml cancel-in-progress config (#23420)
merge this one, as it fixed xeon ci blocking issue.
|
2026-04-22 10:12:36 +08:00 |
|
 Kangyan-ZhouandClaude Opus 4.7
|
77fd86f89e
|
[ci] split stage-c-test-4-gpu-b200 to enable a low-disk runner pool (#23417)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-21 18:33:33 -07:00 |
|
 Piotr MazurekandPiotr Mazurek
|
6cf0b004ca
|
[MoE] Add LFM2 MoE tuning support + tuned configs for H100/B200/MI325X (#22791)
Co-authored-by: Piotr Mazurek <piotr.mazurek@liquid.ai>
|
2026-04-21 18:32:05 -07:00 |
|
Alison Shao
|
0e165ffbfc
|
ci: enable /rerun-test for nightly test suites (#22830)
|
2026-04-21 18:28:10 -07:00 |
|
Byron Hsu
|
c090f71bf2
|
feat: enable SGLANG_PATCH_TOKENIZER by default (#23409)
|
2026-04-21 17:53:43 -07:00 |
|
hlu1
|
415f64e763
|
Add MambaPool kvcache offloading during retraction (#22493)
|
2026-04-22 08:51:03 +08:00 |
|