Commit Graph
12172 Commits
Author SHA1 Message Date
Mick 6fa499cd6d [diffusion] CI: remove parametrized-only from diffusion PR test (#24202) 2026-05-01 14:15:22 +08:00
Yanbin Jiang 8975479f87 [LoRA][MOE] Fix EP correctness in MoE LoRA slicing and virtual-experts kernels (#24171) 2026-04-30 22:42:10 -07:00
Baizhou ZhangandClaude Opus 4.7 7bc7775260 [CI] Fix stage-b-test-4-gpu-b200 silently skipped, hanging wait-for-stage-b (#24208)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 22:29:08 -07:00
Alison Shao 357e3789ea ci: limit nightly test parallelism to 1 job per hardware type (#23314) 2026-04-30 21:49:18 -07:00
AlecandKangyan-Zhou 9d95783603 Add Docker image provenance metadata (#24090)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2026-04-30 21:40:42 -07:00
Mick 9d84268705 [diffusion] refactor: introduce component residency manager (#23771) 2026-05-01 11:10:41 +08:00
Cheng WanandClaude Opus 4.7 108bfd8b6a [MoE] Add Aiter MoE runner backend and purge aiter.fused_moe from quant methods (#23597)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 19:50:52 -07:00
Yihao Wang 0acc569edd [Bench] extend MMMU answer extractor with explicit-commit patterns (#24084) 2026-04-30 19:08:08 -07:00
Baizhou ZhangandClaude Opus 4.7 1742bfb610 [CI] release-whl-kernel: skip musa wheels in update_kernel_whl_index (#24191)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 17:17:50 -07:00
Kangyan-ZhouandClaude Opus 4.7 cf346bb15d [CI] Skip worker-dependent SMG e2e tests pending runner-image debug (#24166)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 16:45:10 -07:00
Baizhou Zhang 5f88c8593a [Misc] Redirect default sglang nightly wheel to cuda 130 (#24183) 2026-04-30 16:32:50 -07:00
Yilong Zhao f67292539f spec: gate dp mlp sync with server args (#24177) 2026-04-30 16:29:41 -07:00
da7f890788 [Intel GPU] Integrate flash_mla_decode in Intel XPU attention backend (#23557)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-05-01 07:21:28 +08:00
e35ac95cdc [Test] Add XPU device support to unit tests (#22236)
Co-authored-by: vshekhawat-hlab <vshekhawat@habana.ai>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-05-01 07:18:51 +08:00
zijiexia 8b23d32ec1 feat: implement workflow to sync LMSYS SGLang blog (#23438) 2026-04-30 16:17:29 -07:00
Alison Shao 918f910cf0 ci: temporarily disable multimodal-gen-test-1-b200 (#24174) 2026-04-30 16:16:43 -07:00
9c5cad3914 Use device-agnostic helpers for Mamba tests and core ops (#20234)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-05-01 07:14:53 +08:00
Kalyan KumarandMa Mingfei 8a9e424faa Replace hardcoded CUDA device with get_device() for XPU support (#13599)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-05-01 07:13:46 +08:00
Kangyan-Zhou c5f1339773 Revert "ci: add rebase-required mode to check-maintenance action" (#24179) 2026-04-30 16:11:36 -07:00
Kangyan-ZhouandClaude Opus 4.7 e45b8ec1ec [CI] Publish nightly sglang wheel under both cu129 and cu130 indexes (#24176)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 16:08:59 -07:00
Alison ShaoandAlison Shao 694ef516cb Revert "[ci] split stage-c-test-4-gpu-b200 to enable a low-disk runner pool" (#24163)
Co-authored-by: Alison Shao <alisonshao@radixark.ai>
2026-04-30 15:57:19 -07:00
Alison ShaoandAlison Shao cdc4078815 ci: add rebase-required mode to check-maintenance action (#23109)
Co-authored-by: Alison Shao <alisonshao@radixark.ai>
2026-04-30 15:47:24 -07:00
Lawrence Wu f75a8b6220 fix: support HybridLinearAttnBackend in TboAttnBackend (#20114) 2026-04-30 15:40:13 -07:00
Hubert Lu d57671527a Fix LFM2 ShortConv Mamba State Indexing (#23975) 2026-04-30 15:23:39 -07:00
sglang-botandsglang-bot 2e027b1afe chore: bump sgl-kernel version to 0.4.2 (#24170)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-04-30 15:02:07 -07:00
Xinyuan Tong 989a16187d [Bench] Fix bench_serving missing reasoning_content stream chunks (#23954) 2026-04-30 15:00:27 -07:00
340efca244 [sgl-kernel] Prep for torch 2.11 upgrade and switch PyPI default to cu130 (#24162)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-04-30 14:54:39 -07:00
Erik Wijmans c04b20dc88 Fix KeyError in prepare_lora_batch when lora_ids contains None (#21974) 2026-04-30 14:50:16 -04:00
oriandzhiguo.qin 71e89e9003 [MUSA][19/N] Support qwen series models (#23654)
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
2026-04-30 11:26:47 -07:00
Alison Shao dc395bc059 ci: run setup_ld_library_path before install_sglang_kernel (#24141) 2026-04-30 10:55:45 -07:00
Alison Shao 7bb7f6049a ci: add per-host utilization view to runner-utilization report (#24102) 2026-04-30 10:05:16 -07:00
651af06a0b [Feature] Xiaomi MiMo-V2.5 day0 support (#23811)
Co-authored-by: 张袁 <zhangyuan36@xiaomi.com>
Co-authored-by: 刘安岐 <liuanqi6@xiaomi.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-05-01 00:02:26 +08:00
YC Yen-Ching Tsengandbingxche cf4f462094 [AMD] Nightly image release for deepseek v4 (#24155)
Co-authored-by: bingxche <bingxche@amd.com>
2026-04-30 23:49:13 +08:00
aa74911448 [NPU] fix some npu error with OffloaderV2 (#19541)
Co-authored-by: Jianzhao Xu <xujianchao@huawei.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
2026-04-30 15:05:35 +03:00
Yaochen Han 577dbc4ab9 [4/N] Quantization Refactor: AWQ schemes and Kernel call and weight init split (#21126) 2026-04-30 14:51:01 +03:00
Lianmin Zheng b1ef99f65f [CI] Remove orphaned test/srt/ascend and test/srt/configs (#24145) 2026-04-30 04:43:11 -07:00
Qiaolin Yu 583929c0a1 fix the compatibility between --moe-dense-tp-size 1 and piecewise cuda graph (#23972) 2026-04-30 02:12:13 -07:00
Opher LieberandEthan Su 99c0b62f1e allow requests with exactly context_len total tokens (#22546)
Co-authored-by: Ethan (Yusheng) Su <yushengsu.thu@gmail.com>
2026-04-30 01:12:06 -07:00
Ethan (Yusheng) Su 125f75db72 fix(lora): avoid CUDA graph-breaking scalar assignment in seg_indptr (#23738) 2026-04-30 01:11:45 -07:00
Brian da07b22949 [Docs] quick fix delete --enable-dp-attention in sgl-jax (#24052) 2026-04-30 15:39:03 +08:00
billishyahao 692979a8d9 [AMD] Support sdma path for moriep (#23929) 2026-04-29 23:57:00 -07:00
Shaojun Zhou 4f0b44c5c6 [fix] moss-vl: use Conv3dLayer and remove no-op flat_encoder_result (#23932) 2026-04-30 14:19:45 +08:00
kkyyxhll 936c9c2355 fix(qwen3_5): broadcast per-tensor scale in _make_packed_weight_loader for FP8 models (#23062) 2026-04-30 14:16:57 +08:00
Jay ThakurandMa Mingfei bcb34da9f9 Add deterministic mode for XPU operations (#16793)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-04-30 13:39:06 +08:00
Opher LieberandYanbin Jiang c8c1c9261d LoRA support for qwen3.5 and nemotron3 (#23594)
Co-authored-by: Yanbin Jiang <jybsuper@gmail.com>
2026-04-29 21:51:53 -07:00
Mick 0b1fbdba15 [diffusion] CI: change ground truth upload path and improve publish script (#24120) 2026-04-30 12:26:10 +08:00
Liangsheng Yin c54ada994b fix: rename mimo spec threshold attr to num_accepted_drafts_thres (#24118) 2026-04-29 21:00:55 -07:00
Yuxuan Zhang d040333c95 [Bug Fix] missing index/KV transfer for MTP layer in NSA disaggregation (#23539) 2026-04-30 11:55:45 +08:00
YC Yen-Ching Tseng 9e68c8527e [AMD] Update AMD Nightly Test checkout mechanism (#24112) 2026-04-30 11:36:18 +08:00
2d2be5d7b2 [PD][Bugfix] fix mamba cache capping (#22462)
Co-authored-by: hzh0425 <hzh0425@apache.org>
Co-authored-by: yizhang2077 <1109276519@qq.com>
2026-04-30 10:57:55 +08:00