 
|
aa74911448
|
[NPU] fix some npu error with OffloaderV2 (#19541)
Co-authored-by: Jianzhao Xu <xujianchao@huawei.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
|
2026-04-30 15:05:35 +03:00 |
|
Yaochen Han
|
577dbc4ab9
|
[4/N] Quantization Refactor: AWQ schemes and Kernel call and weight init split (#21126)
|
2026-04-30 14:51:01 +03:00 |
|
Lianmin Zheng
|
b1ef99f65f
|
[CI] Remove orphaned test/srt/ascend and test/srt/configs (#24145)
|
2026-04-30 04:43:11 -07:00 |
|
kkyyxhll
|
936c9c2355
|
fix(qwen3_5): broadcast per-tensor scale in _make_packed_weight_loader for FP8 models (#23062)
|
2026-04-30 14:16:57 +08:00 |
|
 Opher LieberandYanbin Jiang
|
c8c1c9261d
|
LoRA support for qwen3.5 and nemotron3 (#23594)
Co-authored-by: Yanbin Jiang <jybsuper@gmail.com>
|
2026-04-29 21:51:53 -07:00 |
|
Liangsheng Yin
|
c54ada994b
|
fix: rename mimo spec threshold attr to num_accepted_drafts_thres (#24118)
|
2026-04-29 21:00:55 -07:00 |
|
Qiaolin Yu
|
2bbd30a27a
|
relax the threshold in test_step3p5_flash_chain_mtp (#24105)
|
2026-04-29 16:53:35 -07:00 |
|
Jimmy Shong
|
3d31ac2672
|
[Fix] FP8 Qwen3-Next quant error by removing fallback fused shards (#23973)
|
2026-04-29 17:33:47 -04:00 |
|
Qiaolin Yu
|
79dbfe4505
|
Use spec v2 by default (#21062)
|
2026-04-29 13:40:42 -07:00 |
|
AndyLi429
|
4c1eefca4f
|
[NPU] ascend backend support qwen3 moe attention cp (#21685)
|
2026-04-29 19:25:17 +08:00 |
|
Sam (Kesen Li)
|
73e93bebd6
|
[1/4] NVFP4 KV cache: quantization strategy abstraction and kernel (#21954)
|
2026-04-29 01:45:48 -07:00 |
|
Xinyuan Tong
|
832b4f59ed
|
[Bench] fix MMMU answer-extraction regex dropping multi-line responses (#23864)
|
2026-04-29 14:48:49 +08:00 |
|
shuwenn
|
2c41ef4c93
|
[HiCache] feat: add draft KV cache backing for L2/L3 (#21125)
|
2026-04-28 23:47:31 -07:00 |
|
 MingxuZhandMa Mingfei
|
2d27c38f13
|
Update XPU Docker runtime stack & hf_home config (#23820)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-29 10:03:11 +08:00 |
|
Khoa Pham
|
ddcacaf1bd
|
Fix failing test_nvidia_nemotron_3_nano by fixing test_grouped_topk (#23874)
|
2026-04-28 15:03:58 -07:00 |
|
 Alex NailsandClaude Opus 4.7
|
345fecc547
|
fix(bench): wire request_func in bench_long_context ContextWorkloadGenerator (#23898)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-28 14:45:51 -07:00 |
|
 
|
914ef7c7f3
|
Fix multimodal /v1/embeddings Jinja chat template handling (#20835)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-04-28 13:05:45 -07:00 |
|
AlonKejzman
|
66ea0aee7f
|
tokenizer: Add fastokens support (#23753)
|
2026-04-28 11:43:10 -07:00 |
|
Yinzuo Jiang
|
71160e4ddb
|
feat(observability): add OpenTelemetry tracing for pipeline parallelism (#23169)
Signed-off-by: Yinzuo Jiang <jiangyinzuo@foxmail.com>
|
2026-04-28 17:05:23 +08:00 |
|
Xiaoyu Zhang
|
6fbad22feb
|
Remove smoke wording from tests and comments (#23355)
|
2026-04-28 12:05:27 +08:00 |
|
Alison Shao
|
b73c44b545
|
test: relax TestMLADeepseekV3.test_gsm8k threshold 0.62 -> 0.60 (#23879)
|
2026-04-27 15:27:15 -07:00 |
|
Vladislav Nosivskoy
|
28ee08c172
|
[HiCache] Add synchronization for context parallelism (#20460)
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
|
2026-04-28 02:13:07 +08:00 |
|
 Shenxiu LiuandXinyuan Tong
|
a3fc982ba7
|
[Whisper] Automatic language detection via structured generation (#22997)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-04-27 15:54:41 +08:00 |
|
 sglang-botandsglang-bot
|
da175b964d
|
chore: update CI test est_time values (#23785)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-04-26 20:17:50 -07:00 |
|
  
|
10fd0faccd
|
[CPU] Add Qwen3.5 model optimization for CPU (#19484)
Co-authored-by: Zheng, Beilei <beilei.zheng@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-04-26 10:12:36 -07:00 |
|
  
|
9003f24e2b
|
chore: bump sglang-kernel version to 0.4.1.post1 (#23733)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-25 23:23:49 -07:00 |
|
  
|
714173555c
|
chore: bump sgl-kernel version to 0.4.1.post1 (#23720)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-25 17:13:02 -07:00 |
|
 1874.andronnie_zheng
|
046c14a3ed
|
[NPU] Support GGUF quantization for Ascend NPU (dense + MoE) (#17883)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-04-25 17:16:47 +03:00 |
|
   
|
6d03861476
|
support Hy3 preview (#23533)
Co-authored-by: pengmeng <pengmeng@tencent.com>
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
Co-authored-by: chengvjiang <chengvjiang@tencent.com>
Co-authored-by: russellfeng <russellfeng@tencent.com>
|
2026-04-24 12:03:24 -07:00 |
|
 Yuwei AnandClaude Opus 4.6
|
60bbb800db
|
[Experimental] Breakable Piecewise Cuda Graph (#22218)
Signed-off-by: Oasis-Git <ayw.sirius19@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-24 04:33:05 -07:00 |
|
 Ziang LiandBrayden Zhong
|
1758856762
|
[CI] Fix mxfp8 TrtllmGenMoe test (#23125)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-04-24 09:02:11 +00:00 |
|
Sundara Raman Ramachandran
|
cf88fdcc9c
|
Expose child process PIDs from Engine for health check support (#23320)
|
2026-04-23 16:44:49 -07:00 |
|
Jie Hao
|
86ed0680d7
|
feat: add OpenTelemetry tracing to DiffGenerator (#21254)
|
2026-04-23 09:25:23 -07:00 |
|
 Bingxu ChenandYC Yen-Ching Tseng
|
fd88a1c562
|
[AMD] skip deterministic inference for MLA FP8 test (#23382)
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
|
2026-04-23 00:43:23 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.6
|
18359aadc8
|
[CI] Lower GSM8K baselines for B200 nightly after eval unification (#22136)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-22 22:30:54 -07:00 |
|
 Bingxu ChenandClaude Opus 4
|
2b4eeb8343
|
[AMD] Restore test_zimage_turbo.py and test_int4fp8_moe.py with __main__ entry (#23455)
Co-authored-by: Claude Opus 4 (1M context) <noreply@anthropic.com>
|
2026-04-22 22:28:20 -07:00 |
|
Alison Shao
|
267c2c0849
|
test: move test_epd_disaggregation to nightly-4-gpu (#23518)
|
2026-04-22 20:59:19 -07:00 |
|
Yanbin Jiang
|
917d2aa1dc
|
[LoRA] Fix EP + per-expert MoE LoRA illegal memory access (#23178)
|
2026-04-22 14:22:32 -07:00 |
|
 jianan-guandMa Mingfei
|
ad0fc88810
|
[CPU] [Quantization] Add GPTQ/AWQ 4bits quantization support for CPU (#22685)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-22 13:34:02 -07:00 |
|
 Yuxuan ZhangandXinyuan Tong
|
28cfd3d272
|
Support defer_loading field at function level for Chat Completions API (#22702)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-04-22 10:09:54 -07:00 |
|
Shangming Cai
|
1c06a3d072
|
[CI] Move disaggregation basic CI back to 2-gpu suite (#23447)
|
2026-04-22 17:50:33 +08:00 |
|
 
|
c3ea2d7b92
|
Rename mixed_with_decode_tokens in mixed chunk prefill adder (#6506)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
|
2026-04-21 22:48:34 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.7
|
77fd86f89e
|
[ci] split stage-c-test-4-gpu-b200 to enable a low-disk runner pool (#23417)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-21 18:33:33 -07:00 |
|
hlu1
|
415f64e763
|
Add MambaPool kvcache offloading during retraction (#22493)
|
2026-04-22 08:51:03 +08:00 |
|
Charles Chen
|
c396e4924b
|
[bug] Fix cache salt and extra keys for prefix cache isolation (#23300)
|
2026-04-21 13:53:24 -07:00 |
|
Ma Mingfei
|
929e00eeab
|
[CPU] expand the interface of shared_expert without scaling factor (#22933)
merge since this is CPU only change on sgl-kernel.
|
2026-04-21 20:03:39 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
48daa831ea
|
[KDA] Fuse gate+cumsum and reuse chunk index for KDA (#23038)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-04-21 17:54:20 +08:00 |
|
Bingxu Chen
|
09b1d10d59
|
[AMD] prepare for MI300x PR runner pool: registry mirror, runner routing, threshold tuning (#23156)
|
2026-04-21 00:58:23 -07:00 |
|
 
|
f63def8510
|
[XPU] Fix DeepSeek-OCR tests under transformers 5.x (#23044)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-21 14:57:56 +08:00 |
|
Liangsheng Yin
|
2b2cad70d6
|
[Refactor] Move radix-cache utils onto RadixKey as methods (#23209)
|
2026-04-20 23:11:58 -07:00 |
|