Mohammad Miadh Angkad
|
7045e0fdff
|
Seed sgl-kernel topk sigmoid tests on all backends (#30754)
|
2026-07-10 03:08:15 -07:00 |
|
Ma Mingfei
|
073b36853f
|
[CPU] update fla.cpp to support when num_head_v is not multiples of 16 (#30604)
|
2026-07-10 09:21:07 +08:00 |
|
Mick
|
5ce5e1ee3e
|
[Diffusion] Revert CPU AMX optimizations (#30716)
|
2026-07-10 09:09:38 +08:00 |
|
Ma Mingfei
|
fdf6c70e3a
|
fix arm test_norm.py error (#30597)
|
2026-07-09 14:36:45 +08:00 |
|
 Haotong ZouandValentine233
|
3b43df5b6d
|
Support speculative decoding on CPU (#27862)
Co-authored-by: Valentine233 <xuan.liao@intel.com>
|
2026-07-09 10:27:09 +08:00 |
|
 jianan-guandMa Mingfei
|
177c048c68
|
[Diffusion][CPU] Adding AMX optimizations for CPU platform (#28527)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-07-09 10:26:22 +08:00 |
|
Baizhou Zhang
|
946804e042
|
Disable FA3 sparse mask kernels by default (#30356)
|
2026-07-07 01:46:44 -07:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
1da7d3a50b
|
[MoE] Retire the AOT moe_fused_gate / kimi_k2_moe_fused_gate gate kernels (#26771) (#29997)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-07 13:53:17 +08:00 |
|
     
|
9bd02dc5b9
|
feat(sgl-kernel): add InfLLM v2 attention kernels (#29383)
Co-authored-by: Size Wang <paulgeorge13hhhhh@gmail.com>
Co-authored-by: lijiayi <lijiayi@modelbest.cn>
Co-authored-by: suhmily10 <suhmily@gmail.com>
Co-authored-by: Xiaoyue Xu <xiaoyue.xu.me@gmail.com>
Co-authored-by: hansjohn <74091612+hansjohn@users.noreply.github.com>
Co-authored-by: zhangyan <1762895426@qq.com>
|
2026-07-06 22:46:54 -07:00 |
|
Ma Mingfei
|
30fb0dd851
|
[CPU] add fused_qk_gemma_norm and refactor norm kernel implementation (#30216)
|
2026-07-07 08:52:59 +08:00 |
|
Cheng Wan
|
5f623ad24e
|
Revert "Fix wrong RMSNorm fallback to old Flashinfer CUDA kernel when in PCG" (#30083)
|
2026-07-03 19:05:20 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
1b6d1e9752
|
Fix wrong RMSNorm fallback to old Flashinfer CUDA kernel when in PCG (#29702)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-07-03 00:49:49 +08:00 |
|
Xiaoyu Zhang
|
b276a9acee
|
chore: cleanup garbage code (#29770)
|
2026-07-02 16:14:01 +08:00 |
|
Xinyuan Tong
|
9ba4b8f8ba
|
sgl-kernel: bump sgl-attn for varlen num_splits OOM fix (#29551)
|
2026-07-01 21:13:53 -07:00 |
|
Daniel Stokes
|
8ee200972e
|
[fix] Add support for flashinfer MOE A2A to Qwen3 BF16 model path (#26255)
|
2026-07-01 01:59:54 -07:00 |
|
  
|
df0dfbaa45
|
[Kernel] Strengthen kernel shape coverage (#29636)
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-07-01 15:44:20 +08:00 |
|
Ma Mingfei
|
bd3b252e0a
|
[CPU] use faster exp in silu_and_mul (#29382)
|
2026-06-29 09:38:30 +08:00 |
|
Ma Mingfei
|
3217410cf6
|
[CPU] enable fused_sigmoid_mul on CPU device (#29378)
|
2026-06-29 09:38:04 +08:00 |
|
 
|
714011a40f
|
[JIT Kernel] Migrate dsv3_router_gemm from AOT sgl-kernel to JIT kernel (#21531)
Co-authored-by: Guohao Shao <shao.gh.98@gmail.com>
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-06-26 11:52:18 -07:00 |
|
Xinyuan Tong
|
30ea4c0f4b
|
build(sgl-kernel): bump FlashMLA pin + fix cccl include for CUDA 13 (#29067)
|
2026-06-25 20:00:15 -07:00 |
|
Yibo Cai
|
4b0b0ea55f
|
[sgl-kernel/cpu]: exclude amx gemm source from arm build (#29286)
|
2026-06-26 08:23:43 +08:00 |
|
Yibo Cai
|
4ce1c180bd
|
[sgl-kernel/cpu]: fix arm64 w8a8 moe kernel signature (#29270)
|
2026-06-26 08:11:57 +08:00 |
|
R0CKSTAR
|
38e857f423
|
[MUSA][diffusion] Bump torchada version to 0.1.68 (#29236)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
|
2026-06-25 16:53:43 -07:00 |
|
Ma Mingfei
|
1ba7c79761
|
[CPU] add indices in chunk_gated_delta_rule (#29267)
|
2026-06-26 07:51:07 +08:00 |
|
Ma Mingfei
|
2c3f007a65
|
[CPU] optimize GDN prefill performance (#29117)
|
2026-06-25 09:04:34 +08:00 |
|
 sglang-botandsglang-bot
|
3e97c9239f
|
chore: bump sgl-kernel version to 0.4.4 (#28556)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-06-17 13:38:00 -07:00 |
|
 Baizhou ZhangandShijin
|
4b817f5d7f
|
Upgrade fa3 hash (#28394)
Co-authored-by: Shijin <dovis.zhang02@gmail.com>
|
2026-06-17 13:32:11 -07:00 |
|
 Michaelandmichaelzhang-ai
|
f42a093261
|
[AMD] Migrate 2-GPU kernel allreduce tests into the registered system (#27722)
Co-authored-by: michaelzhang-ai <michaelzhang@example.com>
|
2026-06-09 19:39:03 -07:00 |
|
 
|
c2eae96c56
|
MSCCL++ Integration (#22734)
Co-authored-by: Caio Rocha <caiorocha@microsof.com>
Co-authored-by: empyreus <rjsouza1995@gmail.com>
|
2026-06-08 21:13:13 -07:00 |
|
Yingchun Lai
|
b047bb3e92
|
build(sgl-kernel): support configurable mirrors for restricted networks (#27387)
|
2026-06-08 08:52:26 -07:00 |
|
MARATRIX
|
61e4132bc2
|
[MUSA] bump torchada version to 0.1.59 and workaround PCG limitation. (#27537)
Signed-off-by: yafeng.li <yafeng.li@mthreads.com>
|
2026-06-08 08:45:39 -07:00 |
|
 Zaili WangandMa Mingfei
|
3b7a258f63
|
[CPU] upgrade dependent torch ver to PT2.12 (#21456)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-06-04 11:04:11 +08:00 |
|
 Zaili WangandMa Mingfei
|
512bfbb1e1
|
[CPU] Explicitly enable AVX512 & AMX instruction set (#26145)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-06-03 13:40:41 +08:00 |
|
Mengxuan Xiong
|
ff642ed936
|
[MoE] Extend kimi_k2_moe_fused_gate to support 256 experts (MiMo V2 Flash) (#26303)
|
2026-06-01 16:03:58 +08:00 |
|
blzheng
|
a722b1a437
|
[CPU] fix incorrect index of b_ptr in fused_sigmoid_gating_delta_rule… (#26634)
|
2026-05-29 16:08:12 +08:00 |
|
 
|
3ecf2c76ad
|
[CPU] Add GPT-OSS model optimization for CPU (#16775)
Co-authored-by: mingfeima <mingfei.ma@intel.com>
Co-authored-by: jianan-gu <jianan.gu@intel.com>
|
2026-05-29 16:05:26 +08:00 |
|
 Aditya SharmaandXiaodong Ye
|
b2eed9e16d
|
[Apple Silicon] Add custom Metal RoPE kernel with fused KV cache store (#22868)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
Co-authored-by: Xiaodong Ye <yeahdongcn@gmail.com>
|
2026-05-29 15:09:33 +08:00 |
|
 
|
be32df33b9
|
[MUSA] Fix startup with patched torchada (#26437)
Co-authored-by: popsiclexu <zhenxue.xu@mthreads.com>
Co-authored-by: R0CKSTAR <xiaodong.ye@mthreads.com>
|
2026-05-28 20:57:55 +08:00 |
|
Chunyuan WU
|
714fdd9723
|
Fix MiniMax-M2.7 on CPU (#25061)
|
2026-05-28 10:53:13 +08:00 |
|
 sglang-botandsglang-bot
|
0753182b50
|
chore: bump sgl-kernel version to 0.4.3 (#26414)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-05-26 12:10:16 -07:00 |
|
Chunan Zeng
|
b66f8e0b96
|
Sgl flashmla (#26132)
|
2026-05-26 12:00:23 -07:00 |
|
+2        
|
3f5e2c7688
|
[AMD] Dsv4/pr2 compressor opt (#26208)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: amd-danli103 <danli103@amd.com>
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
Co-authored-by: yichiche@amd.com <jacky.cheng>
Co-authored-by: yctseng0211 <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-05-25 23:54:40 -07:00 |
|
Jincong Chen
|
e27d4fb70f
|
[Perf][Qwen3.5] Add case 512 to topkGatingSoftmaxKernelLauncher, (#25775)
|
2026-05-25 16:08:21 +08:00 |
|
Ma Mingfei
|
821d5f4a5b
|
[CPU] add faster KV-cache writes (#25874)
|
2026-05-25 10:28:52 +08:00 |
|
+2        
|
af8f66940e
|
[AMD] Dsv4/pr1 fix run time issue (#25898)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: amd-danli103 <danli103@amd.com>
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
Co-authored-by: yichiche@amd.com <jacky.cheng>
Co-authored-by: yctseng0211 <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-05-23 16:04:14 -07:00 |
|
  
|
84ea47eb22
|
[CPU] Fix issues when running llama3.2-11B vision model with image tasks (#8666)
Co-authored-by: JieXin Liang <Alcanderian@users.noreply.github.com>
Co-authored-by: Yineng Zhang <me@zhyncs.com>
Co-authored-by: jianan-gu <jianan.gu@intel.com>
|
2026-05-21 13:09:18 +08:00 |
|
Michael
|
7fda3caea4
|
[AMD] test(sgl-kernel): seed RNG on ROCm in test_moe_topk_sigmoid to fix tie-break flake (#25356)
|
2026-05-19 23:01:13 -07:00 |
|
 Yuan Luoandluoyuan.luo
|
b7085d3860
|
[fp8] SM90 swap-AB scaled_mm dispatch (~1.16x kernel geomean, +5.8-18.5% end-to-end) (#25532)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-05-20 13:20:37 +08:00 |
|
 miamiaoxyzandMa Mingfei
|
5147de26e4
|
Fix AMX GQA extend attention (#25180)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-05-18 09:30:02 +08:00 |
|
nohup
|
34cb8e2842
|
fix(sgl-kernel): sm90 compile flashmla failed (#24130)
|
2026-05-15 16:42:33 +08:00 |
|