Commit Graph
11084 Commits
Author SHA1 Message Date
Zhai FeiyueandHaiShaw daf697afda [AMD] Add SGLANG_DISAGGREGATION_NUM_PRE_ALLOCATE_REQS env var for configurable KV transfer overlap (#20410)
Co-authored-by: HaiShaw <hixiao@gmail.com>
2026-03-30 14:37:16 -07:00
d6029de6ad [Bugfix][NPU] Skip FRACTAL_NZ format for MoE weights with unaligned dimensions (#21209)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-03-30 23:22:17 +03:00
Vedant V Jhaveri 4a9ffc3ab6 fix nemotron capture for non attention layers (#21436) 2026-03-30 12:50:49 -07:00
Yuxuan Zhang ad064c2f4e [GLM-V and GLM-4.7] Cast to FP32 before gate projection for GLM model. (#21660) 2026-03-30 12:25:27 -07:00
yuefeng Wuandgemini-code-assist[bot] a20d12ae96 [diffusion][doc]: add ring sp performance benchmark page (#20998)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-30 20:26:05 +03:00
Makcum888e f4b0e9c64a [diffusion] [NPU] support ring attention on NPU with FA (#21383) 2026-03-30 20:10:55 +03:00
GXINand高鑫 752d260c77 [NPU][diffusion]: support parallel decoding of qwen-image (#20757)
Co-authored-by: 高鑫 <gaoxin@gaoxindeMacBook-Pro.local>
2026-03-30 20:03:24 +03:00
cen121212 ba6d54d0f0 [NPU] GLM-5 optimize with fused kernels (#18617) 2026-03-30 22:48:15 +08:00
xieminghe1andundefined 7119d59747 DeepSeek-R1-0528-w4a8: DeepEP Low Latency Dispatch Adopts FP8 Communication (#14162)
Co-authored-by: undefined <zhouchen.arrebol@jd.com>
2026-03-30 22:27:28 +08:00
heziiop 673ffb3116 [NPU] fix eagle3 accept rate (#21255) 2026-03-30 21:58:25 +08:00
GXINand高鑫 c5c58c3349 [NPU][Diffusion] fix sp modulate for qwen-image-edit (#20974)
Co-authored-by: 高鑫 <gaoxin@gaoxindeMacBook-Pro.local>
2026-03-30 16:18:48 +03:00
Mick 0a1fb42869 [diffusion] CI: relax pr-test threshold (#21682) 2026-03-30 20:23:46 +08:00
Mick b76730701b [diffusion] feat: enhance overlay mechanism (#21648) 2026-03-30 19:45:34 +08:00
1d6424d5ad fix: Mistral Small 4 fails to start due to config/weight format mismatch (#21620)
Co-authored-by: mengxiancheng03 <mengxiancheng03@kuaishou.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-30 01:57:35 -07:00
strgrb b246269444 fix mamba cache leak when adder fails to add a matched req. (#21404) 2026-03-30 16:45:49 +08:00
Baizhou ZhangandClaude Opus 4.6 62a63eeff7 [Fix] Fix weight_loader property assignment for qwen3-next FP8 models (#21662)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-30 01:35:59 -07:00
Hubert Lu e6071e60c0 [AMD] Support AMD MXFP4 Qwen3.5-397B-A17B model (#21234) 2026-03-30 01:14:18 -07:00
Michelle Wuandwuxue 965f03cdc2 [NPU] Update DeepSeek-V3.2 model deployment instructions in documentation (#21468)
Co-authored-by: wuxue (C) <w00964934@china.huawei.com>
2026-03-30 15:51:42 +08:00
kkandwunhuang b9a68c304e [AMD] Fused rope kv store (#21315)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-03-30 00:05:41 -07:00
Ma Mingfei af62bd9486 [CPU] Implement MXFP4 Gemm kernels for intel AMX to support GPT OSS series. (#14385) 2026-03-29 23:44:12 -07:00
blzhengandMa Mingfei ed01e1d5d6 [CPU] add kernel apply_rotary_pos_emb_cpu for Qwen3-VL and Qwen3-Omni (#13121)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-03-29 23:43:46 -07:00
Ma Mingfei 6da8f5f69e fix topk softmax performance issue (#14702) 2026-03-29 23:43:16 -07:00
Aishwarya Ramasethu c32ee48886 MFU metrics in Prometheus (#19395) 2026-03-29 23:40:06 -07:00
Ziang Li 1a4b383fac [CI] [FlashInfer v0.6.7] Use offline quantized checkpoint for MXFP8 Gemm tests (#21625) 2026-03-29 22:47:46 -07:00
Baizhou Zhangandgemini-code-assist[bot] 5b19c9a05d [Doc] Update tips for developer new-comers (#21659)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-29 22:40:36 -07:00
Polisetty V R K Jyothendra Varma f0303fd07e [Intel GPU] Enable DeepSeek R1 inference on XPU (#18461)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
2026-03-29 22:35:59 -07:00
Kangyan-ZhouandClaude Opus 4.6 d8ab41dce5 [Fix] Handle pre-release tags in nightly wheel version parsing (#21656)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-29 22:29:40 -07:00
Ying Sheng 90bdc3192b Update sponsorship details in README.md (#21658) 2026-03-29 21:42:59 -07:00
db5d9eb8ce [diffusion] CI: fix dashboard chart (nightly) display issues (#21653)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-30 12:02:01 +08:00
Feng Suandzhangxiaolei123456 9b4dd27478 [Fix] Fix Qwen3.5 MoE model loading and Mamba cache sharding in PP mode (#21448)
Co-authored-by: zhangxiaolei123456 <zhangxiaolei.666@bytedance.com>
2026-03-30 11:57:26 +08:00
Liangsheng Yinandwan4ch c06ca1526c Fix circular reference in CustomTestCase.__init_subclass__ (#21650)
Co-authored-by: wan4ch <wan4ch@gmail.com>
2026-03-29 20:38:12 -07:00
Lianmin Zheng afb32d7622 README: coding agent sponsorship for long-term contributors (#21642) 2026-03-29 16:02:51 -07:00
Lianmin ZhengandClaude Opus 4.6 9f7792415a Clean up TokenizerManager: remove dead code and improve rid validation (#21639)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-29 15:12:49 -07:00
Lianmin ZhengandClaude Opus 4.6 f3970b17ef [Cleanup] Remove unused BatchMultimodalOutput and BatchMultimodalDecodeReq (#21640)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-29 14:54:25 -07:00
Lianmin ZhengandClaude Opus 4.6 1d9c8e8c9e Simplify routed experts test and move base64 encoding to tokenizer manager (#21634)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-29 12:44:01 -07:00
Mohammad Miadh Angkad 2acdda1d85 [Fix] Remove redundant allreduce fusion block and skip TP=1 (#20621) 2026-03-29 12:30:40 -07:00
wili bda94fc779 [Fix] SGLANG_USE_CUDA_IPC_TRANSPORT=1 and SGLANG_ENABLE_MM_SPLITTING=1 do not work at the same time. (#19915) 2026-03-30 01:15:26 +08:00
saatwiknagpal d2440dcf58 [VLM] perf: optimize CUDA IPC for multimodal transfer by caching IPC pool handles (#21418) 2026-03-30 00:20:38 +08:00
wili 5bb9ca0e63 [Feature] Optimizations for JPEG input on NVIDIA GPU (#19749) 2026-03-30 00:06:14 +08:00
Bi Xue 42c46e6334 [sgl] disable piecewise cuda graph when a model doesn't have layers (#21565) 2026-03-29 23:04:20 +08:00
Hanlin Bi aa9177152e fix cuda graph capturing error in sm120 mxfp8 triton path (#19835) 2026-03-29 01:59:24 -07:00
Liangsheng Yin fec9961a1f Clean up _wait_for_scheduler_ready implementation (#21626) 2026-03-29 01:02:33 -07:00
shuwenn c34593f951 [HiCache] fix: graceful shutdown of pending async tasks in bench_mix.py (#20276) 2026-03-29 00:46:32 -07:00
d2fa8d67ba Wrap IPv6 addresses in gRPC, bench_serving, and log messages (#21236)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-03-29 00:36:31 -07:00
shuwenn 18074e25dc fix: scheduler launch hang when non-current rank dies (#20287) 2026-03-29 00:28:45 -07:00
22e4733ab9 Add subprocess liveness monitor to detect scheduler crashes (#18582)
Co-authored-by: 继优 <jiyou.ljy@alibaba-inc.com>
Co-authored-by: shuwenn <47200617+alphabetc1@users.noreply.github.com>
2026-03-29 00:09:13 -07:00
Junrong Lin 35f5a0ff35 [CI] Lossen test_return_routed_experts threshold (#21270) 2026-03-28 22:04:53 -07:00
Kangyan-ZhouandClaude Opus 4.6 9d64a82173 feat(ci): add GB300 nightly benchmark test suites (#21487)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-28 21:54:03 -07:00
Shangming Cai 166e9090ee [CI] Skip flaky elastic EP test (#21619) 2026-03-29 12:50:40 +08:00
Lianmin ZhengandClaude Opus 4.6 ba6b501f3a Clean up detokenizer and remove dead multimodal_gen code (#21588)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-28 21:44:40 -07:00