Commit Graph
11997 Commits
Author SHA1 Message Date
Kangyan-ZhouandClaude Opus 4.7 282b47fcfd [CI] release-whl-kernel: clean root-owned build artifacts before checkout (#23747)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 21:39:41 -07:00
Yuhao Yang 049f1bf6fb docs(DeepSeek-V4): add GB200 platform to cookbook recipe (#23725) 2026-04-25 20:54:55 -07:00
ba4e9d2ac2 Apply should_use_dp_reduce_scatterv guard to remaining MoE models (follow-up to #23731) (#23732)
Co-authored-by: Byron Hsu <byronhsu@noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
2026-04-25 20:36:16 -07:00
71029abd64 Fix Qwen3 MoE: also guard EP all-reduce with not use_reduce_scatter (follow-up to #23731) (#23734)
Co-authored-by: Byron Hsu <byron@periodiclabs.ai>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 20:35:52 -07:00
714173555c chore: bump sgl-kernel version to 0.4.1.post1 (#23720)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 17:13:02 -07:00
Byron HsuandByron Hsu 99b59b279c Fix Qwen3 MoE double-reduce when DP attention + EP + reduce_scatterv (#23729) (#23731)
Co-authored-by: Byron Hsu <byronhsu@noreply.github.com>
2026-04-25 15:28:28 -07:00
Kangyan-ZhouandClaude Opus 4.7 921e14dcac [CI] release-docker-deepseek-v4: select which flavors to push (#23730)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 14:09:34 -07:00
Kangyan-Zhou 0c826374a8 ci: add docker release workflow for deepseek_v4 branch (#23728) 2026-04-25 11:45:23 -07:00
Kangyan-ZhouandClaude Opus 4.7 acaa35664d [CI] sgl-kernel: prune dangling images before each wheel build (#23723)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-25 10:44:47 -07:00
AlbeeSo e0a4522370 [typo] fix typo in parallel_state (#23710) 2026-04-25 09:33:33 -07:00
Mick 03849496ad jit_kernel: tolerate FA3 kernels without out arg (#23717) 2026-04-25 23:42:33 +08:00
fzyzcjy d4c1665626 docs(DeepSeek-V4): mark h200|big|pd-disagg verified + recipe fixes (#23715) 2026-04-25 22:49:12 +08:00
1874.andronnie_zheng 046c14a3ed [NPU] Support GGUF quantization for Ascend NPU (dense + MoE) (#17883)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-04-25 17:16:47 +03:00
gjsheuandgengjinsong e708ea6d94 [diffusion] fix: restore cache-dit support for LTX2 (#23235)
Co-authored-by: gengjinsong <gengjinsong@huawei.com>
2026-04-25 18:10:43 +08:00
Aleksi Vesanto 50ce2708ca [diffusion] fix: Fix FLUX.1/2 graph breaks (#23648) 2026-04-25 17:54:52 +08:00
fzyzcjy 880599cd43 docs(DeepSeek-V4): bump GB300 Pro PD decode --mem-fraction-static 0.83 → 0.9 (#23698) 2026-04-25 16:35:44 +08:00
amote-i 11d77a60df [NPU] [DOC] Update supported models and features of npu (#23564) 2026-04-25 15:37:07 +08:00
kkandroot 393252f514 [AMD] fused qk gemma norm kernels to reduce four kernels (#23575)
Co-authored-by: root <root@smci355-ccs-aus-g12-26.cs-aus.dcgpu>
2026-04-25 00:30:01 -07:00
Артем Савкинandronnie_zheng bd523dd60d [NPU] [Bugfix] [Diffusion] Fixed gray images at the generation output (#23266)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-04-25 10:20:38 +03:00
Yujing 6175946db7 [Feature]Add MSProbe dump support in SGLang (#18349) 2026-04-25 10:12:50 +03:00
Yujun Dongandhzh0425 21835fb0af [HiCache] Prevent move_hybrid_indices from polluting radix-tree node host state (#23427)
Co-authored-by: hzh0425 <hzh0425@apache.org>
2026-04-25 14:27:42 +08:00
82254bd9c5 [JIT Kernel] Reland JIT activation (#22094)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Cheng Wan <chwan@rice.edu>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-24 23:00:28 -07:00
ishandhananiandBaizhou Zhang 0d224e5053 update: b300 container for dsv4 (#23697)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-04-24 22:59:20 -07:00
YC Yen-Ching Tseng adc59325bc [AMD] Optimize MiniMax-M2.5 - enable fused Triton kernel for FP8 KV cache write in aiter decode path (#23620) 2026-04-24 22:23:49 -07:00
YC Yen-Ching Tseng fb272d27db [AMD] Optimize MiniMax-M2.5 - use aiter biased_grouped_topk for sigmoid scoring in MoE routing (#23611) 2026-04-24 22:18:08 -07:00
Shenxiu Liu 8471c9ebe6 Skip torch.cuda.empty_cache() in weight update flush path (#22998) 2026-04-25 12:41:39 +08:00
Baizhou Zhang 69485a176c Small udpate gb300 recipe for deepseek v4 (#23690) 2026-04-24 21:35:16 -07:00
fzyzcjy 8a395994ed docs(DeepSeek-V4): mark gb300|{small,big}|{cp,pd-disagg} verified + GB300-specific fixes (#23691) 2026-04-25 12:21:57 +08:00
fzyzcjy d2c61acf25 docs(DeepSeek-V4): mark b200|small|pd-disagg + h200|small|{cp,pd-disagg} verified (#23689) 2026-04-25 11:57:25 +08:00
fzyzcjy fd401c2fb4 docs(DeepSeek-V4): note SGLANG_FIX_DSV4_BASE_MODEL_LOAD for base models (#23684) 2026-04-25 11:27:00 +08:00
Yuhao Yangandtrangdough 4a3fe2a091 model: support parakeet nemotron encoder (#23568)
Co-authored-by: trangdough <trangtdo22@gmail.com>
2026-04-25 11:00:23 +08:00
Jackey Hua 465abadd3c Add fused moe triton config for Qwen3.5-397B-A17B-FP8 (#23682) 2026-04-24 18:35:32 -07:00
Lianmin Zheng a4facdf3f6 [CI] Refactor ci_install_dependency.sh into standalone functions (#23592) 2026-04-24 17:39:39 -07:00
shuwenn f30a6f4d7e [DOC] Add DFLASH speculative decoding documentation (#23553) 2026-04-24 17:18:46 -07:00
Xinyi Song 76da28f6d6 [AMD][bugfix] add gate rocm >= 7.2 for bpreshuffle (#23671) 2026-04-24 13:26:16 -07:00
jhchouuu f7e840682c [AMD][MoRI] bump MoRI to v1.1.1 (#23642) 2026-04-24 13:12:20 -07:00
Jia GuoandClaude Opus 4.6 587fd15bd2 perf: eliminate attention DtoD copy by passing pre-allocated output to FA (#21985)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-24 12:05:16 -07:00
6d03861476 support Hy3 preview (#23533)
Co-authored-by: pengmeng <pengmeng@tencent.com>
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
Co-authored-by: chengvjiang <chengvjiang@tencent.com>
Co-authored-by: russellfeng <russellfeng@tencent.com>
2026-04-24 12:03:24 -07:00
Lianmin Zheng 6344b546c8 Deprecate --collect-tokens-histogram, auto-collect with --enable-metrics (#23595) 2026-04-24 12:00:16 -07:00
Mick 05696527ea [diffusion] feat: support LoRA for LTX2.3 (#23649) 2026-04-25 01:52:41 +08:00
baa0aa670f [HiCache & HybridModel] 3FS backend support DSA & mamba model (#23241)
Co-authored-by: 墨已 <kangyifei.kyf@alibaba-inc.com>
Co-authored-by: hzh0425 <hzh0425@apache.org>
2026-04-25 00:48:01 +08:00
Kangrui Du 92d262f710 [diffusion] RL: add per-step rollout options for SDE and trajectory capture (#23151) 2026-04-24 23:26:16 +08:00
Siju Samuel bca3dd958a [Intel GPU] Enable pipeline parallelism on XPU (#23645) 2026-04-24 19:52:44 +08:00
Yuwei AnandClaude Opus 4.6 60bbb800db [Experimental] Breakable Piecewise Cuda Graph (#22218)
Signed-off-by: Oasis-Git <ayw.sirius19@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-24 04:33:05 -07:00
Mick b3b03369a5 [diffusion] fix: unify LTX-2.3 HQ codepath gates for all LTX-2.3 variants (#23624) 2026-04-24 17:44:08 +08:00
YC Yen-Ching Tsengandbingxche b060a5ccfd [AMD] Fix nightly version tag selection (#23644)
Co-authored-by: bingxche <bingxche@amd.com>
2026-04-24 17:39:47 +08:00
Shangming Cai b8d883398d Revert "[Intel GPU] Enable pipeline parallelism on XPU" (#23641) 2026-04-24 17:36:35 +08:00
Ziang LiandBrayden Zhong 1758856762 [CI] Fix mxfp8 TrtllmGenMoe test (#23125)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2026-04-24 09:02:11 +00:00
fzyzcjy 92bb5c6bbe Update pro fp8 checkpoint in DeepSeek V4 cookbook (#23634) 2026-04-24 15:58:04 +08:00
fzyzcjy 3a620cb761 Again update DeepSeek V4 cookbook (#23622) 2026-04-24 15:12:35 +08:00