Mengxuan Xiong
|
ff642ed936
|
[MoE] Extend kimi_k2_moe_fused_gate to support 256 experts (MiMo V2 Flash) (#26303)
|
2026-06-01 16:03:58 +08:00 |
|
blzheng
|
a722b1a437
|
[CPU] fix incorrect index of b_ptr in fused_sigmoid_gating_delta_rule… (#26634)
|
2026-05-29 16:08:12 +08:00 |
|
 
|
3ecf2c76ad
|
[CPU] Add GPT-OSS model optimization for CPU (#16775)
Co-authored-by: mingfeima <mingfei.ma@intel.com>
Co-authored-by: jianan-gu <jianan.gu@intel.com>
|
2026-05-29 16:05:26 +08:00 |
|
 Aditya SharmaandXiaodong Ye
|
b2eed9e16d
|
[Apple Silicon] Add custom Metal RoPE kernel with fused KV cache store (#22868)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
Co-authored-by: Xiaodong Ye <yeahdongcn@gmail.com>
|
2026-05-29 15:09:33 +08:00 |
|
 
|
be32df33b9
|
[MUSA] Fix startup with patched torchada (#26437)
Co-authored-by: popsiclexu <zhenxue.xu@mthreads.com>
Co-authored-by: R0CKSTAR <xiaodong.ye@mthreads.com>
|
2026-05-28 20:57:55 +08:00 |
|
Chunyuan WU
|
714fdd9723
|
Fix MiniMax-M2.7 on CPU (#25061)
|
2026-05-28 10:53:13 +08:00 |
|
 sglang-botandsglang-bot
|
0753182b50
|
chore: bump sgl-kernel version to 0.4.3 (#26414)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-05-26 12:10:16 -07:00 |
|
Chunan Zeng
|
b66f8e0b96
|
Sgl flashmla (#26132)
|
2026-05-26 12:00:23 -07:00 |
|
+2        
|
3f5e2c7688
|
[AMD] Dsv4/pr2 compressor opt (#26208)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: amd-danli103 <danli103@amd.com>
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
Co-authored-by: yichiche@amd.com <jacky.cheng>
Co-authored-by: yctseng0211 <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-05-25 23:54:40 -07:00 |
|
Jincong Chen
|
e27d4fb70f
|
[Perf][Qwen3.5] Add case 512 to topkGatingSoftmaxKernelLauncher, (#25775)
|
2026-05-25 16:08:21 +08:00 |
|
Ma Mingfei
|
821d5f4a5b
|
[CPU] add faster KV-cache writes (#25874)
|
2026-05-25 10:28:52 +08:00 |
|
+2        
|
af8f66940e
|
[AMD] Dsv4/pr1 fix run time issue (#25898)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: amd-danli103 <danli103@amd.com>
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
Co-authored-by: yichiche@amd.com <jacky.cheng>
Co-authored-by: yctseng0211 <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-05-23 16:04:14 -07:00 |
|
  
|
84ea47eb22
|
[CPU] Fix issues when running llama3.2-11B vision model with image tasks (#8666)
Co-authored-by: JieXin Liang <Alcanderian@users.noreply.github.com>
Co-authored-by: Yineng Zhang <me@zhyncs.com>
Co-authored-by: jianan-gu <jianan.gu@intel.com>
|
2026-05-21 13:09:18 +08:00 |
|
Michael
|
7fda3caea4
|
[AMD] test(sgl-kernel): seed RNG on ROCm in test_moe_topk_sigmoid to fix tie-break flake (#25356)
|
2026-05-19 23:01:13 -07:00 |
|
 Yuan Luoandluoyuan.luo
|
b7085d3860
|
[fp8] SM90 swap-AB scaled_mm dispatch (~1.16x kernel geomean, +5.8-18.5% end-to-end) (#25532)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-05-20 13:20:37 +08:00 |
|
 miamiaoxyzandMa Mingfei
|
5147de26e4
|
Fix AMX GQA extend attention (#25180)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-05-18 09:30:02 +08:00 |
|
nohup
|
34cb8e2842
|
fix(sgl-kernel): sm90 compile flashmla failed (#24130)
|
2026-05-15 16:42:33 +08:00 |
|
  ![github-actions[bot]](/assets/img/avatar_default.png)
|
897587b03a
|
[MUSA]: Add flashinfer sampling backend (#24978)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
|
2026-05-14 20:23:15 -07:00 |
|
 sglang-botandsglang-bot
|
0fde61535f
|
chore: bump sgl-kernel version to 0.4.2.post2 (#25326)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-05-14 17:53:17 -07:00 |
|
 sglang-botandClaude Opus 4.7
|
0a2615df24
|
chore: add vLLM SPDX copyright headers to ported files (#25182)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
2026-05-13 15:17:30 -07:00 |
|
popsiclexu
|
bfc2eda42d
|
[MUSA] Use MUSA-optimized operators in piecewise CUDA graph (#23633)
Signed-off-by: popsiclexu <zhenxuexu@gmail.com>
|
2026-05-11 17:55:27 -07:00 |
|
R0CKSTAR
|
74d70af09a
|
[Apple Silicon] Add Metal kernel support in sgl-kernel (#23449)
Signed-off-by: Xiaodong Ye <yeahdongcn@gmail.com>
|
2026-05-11 17:54:27 -07:00 |
|
 Brayden Zhongandb8zhong
|
05d1ab51e8
|
Enable PDL for various kernels in DSV32/GLM5 (#23965)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
|
2026-05-09 03:42:56 -07:00 |
|
 Brayden Zhongandb8zhong
|
8f33bee31b
|
Reland Cute-DSL FP4 dense GEMM (#23590)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
|
2026-05-09 02:20:58 -07:00 |
|
Yuxuan Zhang
|
d49fc092cb
|
[Bug Fix] GLM-5.1: drop constexpr on page_indice_batch_offset, skip offloader post_init on draft worker, support N=32 in copy_to_gpu_no_ce (#23550)
|
2026-05-09 15:43:45 +08:00 |
|
Yibo Cai
|
55d8223c2b
|
[sgl-kernel/cpu] support w8a8 int8 model for arm cpu (#16045)
skip gpu test as this one is not related to gpu backend.
|
2026-05-08 14:47:06 +08:00 |
|
 
|
cdf5771f91
|
[MUSA][17/N] ci: Add MUSA diffusion, sgl-kernel tests, and CI workflow support (#20672)
Co-authored-by: ximin.chen <ximin.chen@mthreads.com>
Co-authored-by: R0CKSTAR <xiaodong.ye@mthreads.com>
|
2026-05-07 20:45:21 -07:00 |
|
 JoeyandR0CKSTAR
|
15e6572f21
|
[MUSA][18/N] Add MUSA-optimized kernel implementations for hot ops (#23255)
Signed-off-by: Joey-gvwal <joey_gvwal@yeah.net>
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
|
2026-05-07 20:38:33 -07:00 |
|
 Mandepudi Rani ChowdaryandMa Mingfei
|
55224fff08
|
Add Arm64 CPU Phase 1A CI bootstrap (#22123)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-05-08 09:28:23 +08:00 |
|
R0CKSTAR
|
9cffa5ed6f
|
[MUSA] Bump torchada to 0.1.54 (#24592)
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
|
2026-05-07 11:45:49 -07:00 |
|
 
|
6764155914
|
chore: bump sgl-kernel version to 0.4.2.post1 (#24457)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-05-05 16:19:58 -07:00 |
|
Baizhou Zhang
|
200944b415
|
Update kernel installation instructions after shifting default cuda to 13 (#24181)
|
2026-05-02 16:07:20 -07:00 |
|
 sglang-botandsglang-bot
|
2e027b1afe
|
chore: bump sgl-kernel version to 0.4.2 (#24170)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-04-30 15:02:07 -07:00 |
|
 
|
340efca244
|
[sgl-kernel] Prep for torch 2.11 upgrade and switch PyPI default to cu130 (#24162)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-04-30 14:54:39 -07:00 |
|
 oriandzhiguo.qin
|
71e89e9003
|
[MUSA][19/N] Support qwen series models (#23654)
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
|
2026-04-30 11:26:47 -07:00 |
|
  
|
10fd0faccd
|
[CPU] Add Qwen3.5 model optimization for CPU (#19484)
Co-authored-by: Zheng, Beilei <beilei.zheng@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-04-26 10:12:36 -07:00 |
|
  
|
714173555c
|
chore: bump sgl-kernel version to 0.4.1.post1 (#23720)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-25 17:13:02 -07:00 |
|
 Jia GuoandClaude Opus 4.6
|
587fd15bd2
|
perf: eliminate attention DtoD copy by passing pre-allocated output to FA (#21985)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-04-24 12:05:16 -07:00 |
|
Ma Mingfei
|
23e4d381f0
|
[CPU] remove RECORD_FUNCTION (#23528)
|
2026-04-24 09:18:29 +08:00 |
|
 MARATRIXandAlex Nails
|
74c2e5bacd
|
[MUSA][8/N] Port CUDA kernels that are compatible with MUSA (#17946)
Signed-off-by: yafeng.li <yafeng.li@mthreads.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-04-23 18:04:58 -07:00 |
|
Jia Guo
|
6428392b6f
|
ci: fix cu129 wheel tagging + pipefail-abort in install script (follow-up to #23497) (#23587)
|
2026-04-23 14:52:58 -07:00 |
|
 oriandzhiguo.qin
|
887d380ace
|
[MUSA] Resolve output garbage in Context Parallel on MusaFlashAttentionBackend (#23270)
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
|
2026-04-22 20:22:20 -07:00 |
|
 jianan-guandMa Mingfei
|
ad0fc88810
|
[CPU] [Quantization] Add GPTQ/AWQ 4bits quantization support for CPU (#22685)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-22 13:34:02 -07:00 |
|
Tarushii Goel
|
3ebf066d13
|
[sgl] update specdec sampling kernel to return valid token ID (#22643)
|
2026-04-21 20:28:19 -07:00 |
|
Ma Mingfei
|
929e00eeab
|
[CPU] expand the interface of shared_expert without scaling factor (#22933)
merge since this is CPU only change on sgl-kernel.
|
2026-04-21 20:03:39 +08:00 |
|
 
|
fe9b9b254b
|
Fix segfault in cudaMemcpyBatchAsync on CUDA 13.0 (#23136)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
|
2026-04-20 12:20:22 -07:00 |
|
   
|
6ecd6f84db
|
[CI] Add per-job uv venv isolation and upgrade CI version to Cuda 13 (#23119)
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Alison Shao <a.shao@wustl.edu>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-04-19 05:32:36 -07:00 |
|
Lianmin Zheng
|
9c47bbad13
|
Clean up bench_one_batch warning and simplify norm dispatch (#23110)
|
2026-04-17 17:42:20 -07:00 |
|
 
|
0dcfae5553
|
[CPU] Add gemma4_rmsnorm_cpu kernel (#22842)
Co-authored-by: Copilot <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-17 13:03:16 +08:00 |
|
 Chunyuan WUandMa Mingfei
|
6c89214584
|
[CPU][sgl-kernel] extend_attention_cpu and flash_attn_varlen_func: fix nan for large seq (#22434)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-17 13:01:01 +08:00 |
|