Commit Graph
841 Commits
Author SHA1 Message Date
Xinyuan Tong 30ea4c0f4b build(sgl-kernel): bump FlashMLA pin + fix cccl include for CUDA 13 (#29067) 2026-06-25 20:00:15 -07:00
Yibo Cai 4b0b0ea55f [sgl-kernel/cpu]: exclude amx gemm source from arm build (#29286) 2026-06-26 08:23:43 +08:00
Yibo Cai 4ce1c180bd [sgl-kernel/cpu]: fix arm64 w8a8 moe kernel signature (#29270) 2026-06-26 08:11:57 +08:00
R0CKSTAR 38e857f423 [MUSA][diffusion] Bump torchada version to 0.1.68 (#29236)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
2026-06-25 16:53:43 -07:00
Ma Mingfei 1ba7c79761 [CPU] add indices in chunk_gated_delta_rule (#29267) 2026-06-26 07:51:07 +08:00
Ma Mingfei 2c3f007a65 [CPU] optimize GDN prefill performance (#29117) 2026-06-25 09:04:34 +08:00
sglang-botandsglang-bot 3e97c9239f chore: bump sgl-kernel version to 0.4.4 (#28556)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-06-17 13:38:00 -07:00
Baizhou ZhangandShijin 4b817f5d7f Upgrade fa3 hash (#28394)
Co-authored-by: Shijin <dovis.zhang02@gmail.com>
2026-06-17 13:32:11 -07:00
Michaelandmichaelzhang-ai f42a093261 [AMD] Migrate 2-GPU kernel allreduce tests into the registered system (#27722)
Co-authored-by: michaelzhang-ai <michaelzhang@example.com>
2026-06-09 19:39:03 -07:00
c2eae96c56 MSCCL++ Integration (#22734)
Co-authored-by: Caio Rocha <caiorocha@microsof.com>
Co-authored-by: empyreus <rjsouza1995@gmail.com>
2026-06-08 21:13:13 -07:00
Yingchun Lai b047bb3e92 build(sgl-kernel): support configurable mirrors for restricted networks (#27387) 2026-06-08 08:52:26 -07:00
MARATRIX 61e4132bc2 [MUSA] bump torchada version to 0.1.59 and workaround PCG limitation. (#27537)
Signed-off-by: yafeng.li <yafeng.li@mthreads.com>
2026-06-08 08:45:39 -07:00
Zaili WangandMa Mingfei 3b7a258f63 [CPU] upgrade dependent torch ver to PT2.12 (#21456)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-06-04 11:04:11 +08:00
Zaili WangandMa Mingfei 512bfbb1e1 [CPU] Explicitly enable AVX512 & AMX instruction set (#26145)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-06-03 13:40:41 +08:00
Mengxuan Xiong ff642ed936 [MoE] Extend kimi_k2_moe_fused_gate to support 256 experts (MiMo V2 Flash) (#26303) 2026-06-01 16:03:58 +08:00
blzheng a722b1a437 [CPU] fix incorrect index of b_ptr in fused_sigmoid_gating_delta_rule… (#26634) 2026-05-29 16:08:12 +08:00
3ecf2c76ad [CPU] Add GPT-OSS model optimization for CPU (#16775)
Co-authored-by: mingfeima <mingfei.ma@intel.com>
Co-authored-by: jianan-gu <jianan.gu@intel.com>
2026-05-29 16:05:26 +08:00
Aditya SharmaandXiaodong Ye b2eed9e16d [Apple Silicon] Add custom Metal RoPE kernel with fused KV cache store (#22868)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
Co-authored-by: Xiaodong Ye <yeahdongcn@gmail.com>
2026-05-29 15:09:33 +08:00
be32df33b9 [MUSA] Fix startup with patched torchada (#26437)
Co-authored-by: popsiclexu <zhenxue.xu@mthreads.com>
Co-authored-by: R0CKSTAR <xiaodong.ye@mthreads.com>
2026-05-28 20:57:55 +08:00
Chunyuan WU 714fdd9723 Fix MiniMax-M2.7 on CPU (#25061) 2026-05-28 10:53:13 +08:00
sglang-botandsglang-bot 0753182b50 chore: bump sgl-kernel version to 0.4.3 (#26414)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-05-26 12:10:16 -07:00
Chunan Zeng b66f8e0b96 Sgl flashmla (#26132) 2026-05-26 12:00:23 -07:00
+2 3f5e2c7688 [AMD] Dsv4/pr2 compressor opt (#26208)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: amd-danli103 <danli103@amd.com>
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
Co-authored-by: yichiche@amd.com <jacky.cheng>
Co-authored-by: yctseng0211 <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
2026-05-25 23:54:40 -07:00
Jincong Chen e27d4fb70f [Perf][Qwen3.5] Add case 512 to topkGatingSoftmaxKernelLauncher, (#25775) 2026-05-25 16:08:21 +08:00
Ma Mingfei 821d5f4a5b [CPU] add faster KV-cache writes (#25874) 2026-05-25 10:28:52 +08:00
+2 af8f66940e [AMD] Dsv4/pr1 fix run time issue (#25898)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: amd-danli103 <danli103@amd.com>
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
Co-authored-by: yichiche@amd.com <jacky.cheng>
Co-authored-by: yctseng0211 <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
2026-05-23 16:04:14 -07:00
84ea47eb22 [CPU] Fix issues when running llama3.2-11B vision model with image tasks (#8666)
Co-authored-by: JieXin Liang <Alcanderian@users.noreply.github.com>
Co-authored-by: Yineng Zhang <me@zhyncs.com>
Co-authored-by: jianan-gu <jianan.gu@intel.com>
2026-05-21 13:09:18 +08:00
Michael 7fda3caea4 [AMD] test(sgl-kernel): seed RNG on ROCm in test_moe_topk_sigmoid to fix tie-break flake (#25356) 2026-05-19 23:01:13 -07:00
Yuan Luoandluoyuan.luo b7085d3860 [fp8] SM90 swap-AB scaled_mm dispatch (~1.16x kernel geomean, +5.8-18.5% end-to-end) (#25532)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-05-20 13:20:37 +08:00
miamiaoxyzandMa Mingfei 5147de26e4 Fix AMX GQA extend attention (#25180)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-05-18 09:30:02 +08:00
nohup 34cb8e2842 fix(sgl-kernel): sm90 compile flashmla failed (#24130) 2026-05-15 16:42:33 +08:00
897587b03a [MUSA]: Add flashinfer sampling backend (#24978)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
2026-05-14 20:23:15 -07:00
sglang-botandsglang-bot 0fde61535f chore: bump sgl-kernel version to 0.4.2.post2 (#25326)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-05-14 17:53:17 -07:00
sglang-botandClaude Opus 4.7 0a2615df24 chore: add vLLM SPDX copyright headers to ported files (#25182)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-13 15:17:30 -07:00
popsiclexu bfc2eda42d [MUSA] Use MUSA-optimized operators in piecewise CUDA graph (#23633)
Signed-off-by: popsiclexu <zhenxuexu@gmail.com>
2026-05-11 17:55:27 -07:00
R0CKSTAR 74d70af09a [Apple Silicon] Add Metal kernel support in sgl-kernel (#23449)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
2026-05-11 17:54:27 -07:00
Brayden Zhongandb8zhong 05d1ab51e8 Enable PDL for various kernels in DSV32/GLM5 (#23965)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
2026-05-09 03:42:56 -07:00
Brayden Zhongandb8zhong 8f33bee31b Reland Cute-DSL FP4 dense GEMM (#23590)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
2026-05-09 02:20:58 -07:00
Yuxuan Zhang d49fc092cb [Bug Fix] GLM-5.1: drop constexpr on page_indice_batch_offset, skip offloader post_init on draft worker, support N=32 in copy_to_gpu_no_ce (#23550) 2026-05-09 15:43:45 +08:00
Yibo Cai 55d8223c2b [sgl-kernel/cpu] support w8a8 int8 model for arm cpu (#16045)
skip gpu test as this one is not related to gpu backend.
2026-05-08 14:47:06 +08:00
cdf5771f91 [MUSA][17/N] ci: Add MUSA diffusion, sgl-kernel tests, and CI workflow support (#20672)
Co-authored-by: ximin.chen <ximin.chen@mthreads.com>
Co-authored-by: R0CKSTAR <xiaodong.ye@mthreads.com>
2026-05-07 20:45:21 -07:00
JoeyandR0CKSTAR 15e6572f21 [MUSA][18/N] Add MUSA-optimized kernel implementations for hot ops (#23255)
Signed-off-by: Joey-gvwal <joey_gvwal@yeah.net>
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
2026-05-07 20:38:33 -07:00
Mandepudi Rani ChowdaryandMa Mingfei 55224fff08 Add Arm64 CPU Phase 1A CI bootstrap (#22123)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-05-08 09:28:23 +08:00
R0CKSTAR 9cffa5ed6f [MUSA] Bump torchada to 0.1.54 (#24592)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
2026-05-07 11:45:49 -07:00
6764155914 chore: bump sgl-kernel version to 0.4.2.post1 (#24457)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-05-05 16:19:58 -07:00
Baizhou Zhang 200944b415 Update kernel installation instructions after shifting default cuda to 13 (#24181) 2026-05-02 16:07:20 -07:00
sglang-botandsglang-bot 2e027b1afe chore: bump sgl-kernel version to 0.4.2 (#24170)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-04-30 15:02:07 -07:00
340efca244 [sgl-kernel] Prep for torch 2.11 upgrade and switch PyPI default to cu130 (#24162)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-04-30 14:54:39 -07:00
oriandzhiguo.qin 71e89e9003 [MUSA][19/N] Support qwen series models (#23654)
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
2026-04-30 11:26:47 -07:00
10fd0faccd [CPU] Add Qwen3.5 model optimization for CPU (#19484)
Co-authored-by: Zheng, Beilei <beilei.zheng@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-04-26 10:12:36 -07:00