Mandepudi Rani Chowdary
|
7e3e616159
|
Add Arm64 INT8 MoE test coverage (#25007)
|
2026-06-10 10:36:57 +08:00 |
|
 Zaili WangandMa Mingfei
|
3b7a258f63
|
[CPU] upgrade dependent torch ver to PT2.12 (#21456)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-06-04 11:04:11 +08:00 |
|
 
|
3ecf2c76ad
|
[CPU] Add GPT-OSS model optimization for CPU (#16775)
Co-authored-by: mingfeima <mingfei.ma@intel.com>
Co-authored-by: jianan-gu <jianan.gu@intel.com>
|
2026-05-29 16:05:26 +08:00 |
|
MingxuZh
|
21ba329dac
|
[Xeon] CPU CI enhancement for Intel Xeon platforms (#24649)
|
2026-05-28 10:49:04 +08:00 |
|
  
|
84ea47eb22
|
[CPU] Fix issues when running llama3.2-11B vision model with image tasks (#8666)
Co-authored-by: JieXin Liang <Alcanderian@users.noreply.github.com>
Co-authored-by: Yineng Zhang <me@zhyncs.com>
Co-authored-by: jianan-gu <jianan.gu@intel.com>
|
2026-05-21 13:09:18 +08:00 |
|
 miamiaoxyzandMa Mingfei
|
5147de26e4
|
Fix AMX GQA extend attention (#25180)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-05-18 09:30:02 +08:00 |
|
 Mandepudi Rani ChowdaryandMa Mingfei
|
55224fff08
|
Add Arm64 CPU Phase 1A CI bootstrap (#22123)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-05-08 09:28:23 +08:00 |
|
  
|
10fd0faccd
|
[CPU] Add Qwen3.5 model optimization for CPU (#19484)
Co-authored-by: Zheng, Beilei <beilei.zheng@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-04-26 10:12:36 -07:00 |
|
 jianan-guandMa Mingfei
|
ad0fc88810
|
[CPU] [Quantization] Add GPTQ/AWQ 4bits quantization support for CPU (#22685)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-22 13:34:02 -07:00 |
|
Ma Mingfei
|
929e00eeab
|
[CPU] expand the interface of shared_expert without scaling factor (#22933)
merge since this is CPU only change on sgl-kernel.
|
2026-04-21 20:03:39 +08:00 |
|
 
|
0dcfae5553
|
[CPU] Add gemma4_rmsnorm_cpu kernel (#22842)
Co-authored-by: Copilot <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-17 13:03:16 +08:00 |
|
 Chunyuan WUandMa Mingfei
|
6c89214584
|
[CPU][sgl-kernel] extend_attention_cpu and flash_attn_varlen_func: fix nan for large seq (#22434)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-17 13:01:01 +08:00 |
|
 Baizhou ZhangandClaude Opus 4.6
|
3c46ff2ac5
|
fix: restore CPU flash_attn test to use sgl_kernel directly (#22573)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-10 21:39:20 -07:00 |
|
 jianan-guandMa Mingfei
|
2ab141547d
|
[CPU] Add apply_routed_scaling_factor_on_output support for biased_grouped_topk fusion (#22413)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-10 15:16:05 +08:00 |
|
Rain Jiang
|
1a8eb890f6
|
Kernels community fa3 (#20796)
|
2026-04-07 12:48:44 -07:00 |
|
 blzhengandMa Mingfei
|
ed01e1d5d6
|
[CPU] add kernel apply_rotary_pos_emb_cpu for Qwen3-VL and Qwen3-Omni (#13121)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-03-29 23:43:46 -07:00 |
|
Cao E
|
274581fb77
|
Add support for more batch sizes in cpu_graph_runner (#13881)
|
2026-03-19 09:50:56 -07:00 |
|
 blzhengandFan Yin
|
cd22aa27a9
|
[CPU] Add FP8 Bmm support (#9744)
Co-authored-by: Fan Yin <1106310035@qq.com>
|
2026-03-18 22:19:48 -07:00 |
|
 blzhengandWu, Chunyuan
|
c2b01bd2fc
|
[CPU] fix bug in AVX512 implementation of flash_attn_softmax (#20220)
Co-authored-by: Wu, Chunyuan <chunyuan.wu@intel.com>
|
2026-03-18 22:18:47 -07:00 |
|
 Zaili WangandMa Mingfei
|
2f4babe32b
|
[CPU] support LayerNorm with 3D shape (#15075)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-03-18 22:15:24 -07:00 |
|
 blzhengandMa Mingfei
|
dc6aa26ce9
|
[CPU] Add mrope kernel for Qwen3-vl (#12531)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-03-18 22:12:48 -07:00 |
|
Baizhou Zhang
|
776709efe8
|
[3/n] deepseek_v2.py Refactor: Migrate MLA forward method in deepseek_v2.py (#19122)
|
2026-02-27 13:37:29 -08:00 |
|
SoluMilken
|
07a24f1a38
|
update pre-commit config (#18860)
|
2026-02-16 00:18:31 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) jianan-guandgemini-code-assist[bot]
|
c35aa0238c
|
[CPU][INT4] Add INT4 kernels for CPU (#8226)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-01-29 22:30:13 -08:00 |
|
Ma Mingfei
|
88f7759402
|
[CPU] optimize flash_attn_varlen_func (#15708)
|
2026-01-29 22:07:05 -08:00 |
|
blzheng
|
e27635a02d
|
[CPU] Add 4D input support for ROPE in sgl-kernel (#9337)
|
2025-12-16 17:27:39 +08:00 |
|
blzheng
|
d16ff357db
|
[CPU] Add Gemma3RMSNorm kernel in sgl-kernel and add ut (#9324)
|
2025-12-15 00:24:02 -08:00 |
|
Zaili Wang
|
d6bd2d1126
|
[CPU] layernorm & fused add-layernorm kernels (#14074)
|
2025-12-11 16:58:23 -08:00 |
|
blzheng
|
d257bf87b9
|
[CPU] add mamba fla kernels for Qwen3-next (#12324)
|
2025-12-06 14:16:23 +08:00 |
|
 jianan-guandZheng, Beilei
|
70d2587324
|
[CPU] Optimize small oc GEMM for Qwen3-next on CPU (#12446)
Co-authored-by: Zheng, Beilei <beilei.zheng@intel.com>
|
2025-12-04 00:38:47 -08:00 |
|
Ma Mingfei
|
f90b400431
|
[CPU] add support for mamba causal conv1d for qwen3-next (#12309)
|
2025-12-04 13:41:42 +08:00 |
|
blzheng
|
974c562a25
|
[CPU] add fused_qkvzba_split_reshape_cat kernel for Qwen3-next (#12330)
|
2025-12-03 23:46:08 +08:00 |
|
Xuan Liao
|
c233e9d7a9
|
[CPU] Support chunk_gated_delta_rule kernel for Qwen3-Next (#12441)
|
2025-12-03 17:03:48 +08:00 |
|
 Zaili WangandFan Yin
|
cce2d748ef
|
remove RoPE CPU fp32 tests (#13827)
Co-authored-by: Fan Yin <1106310035@qq.com>
|
2025-11-24 23:22:35 -08:00 |
|
alisonshao
|
6b262ac839
|
Test reorganization: Move tests to manual/ (#13610)
|
2025-11-20 13:41:58 -08:00 |
|
YanbingJiang
|
acde21d8d5
|
Add fused_rmsnorm_gated_cpu kernel for CPU to support Qwen3-Next (#11577)
|
2025-11-21 01:33:31 +08:00 |
|
iLeGend
|
10e0b83a4c
|
Add FP32 dtype support for RoPE - Part2 (#13328)
|
2025-11-19 21:19:53 -08:00 |
|
Kangyan-Zhou
|
49141df94a
|
Extend lint test to test/ directory (#13247)
|
2025-11-14 00:01:48 -08:00 |
|
Baizhou Zhang
|
ce86979355
|
[Fix] Set global args in cpu test (#12105)
|
2025-10-24 21:46:17 -07:00 |
|
fzyzcjy
|
20bd2271e2
|
Support true on-policy (#12058)
|
2025-10-25 10:23:42 +08:00 |
|
blzheng
|
13fb8b5489
|
[CPU] Optimize FP16 decode_attention_cpu (#10652)
|
2025-10-22 21:39:51 -07:00 |
|
Chunyuan WU
|
8fcc69e7c4
|
Turn on shm_allreduce and shm_allgather for fp16 (#10725)
|
2025-10-17 12:35:20 -07:00 |
|
 YanbingJiangandMa Mingfei
|
cbac499750
|
Split test_intel_amx_attention_backend.py to pass CI of timeout (#11370)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2025-10-15 19:22:32 -07:00 |
|
blzheng
|
d1d4074c4e
|
[CPU] Add gelu_and_mul kernel in sgl-kernel and add ut (#9300)
|
2025-09-08 23:23:13 -07:00 |
|
 
|
08f8f49016
|
[CPU][sgl-kernel] biased_grouped_topk: fix correction_bias dtype to float32 (#8212)
Co-authored-by: jianan-gu <jianan.gu@intel.com>
Co-authored-by: YanbingJiang <yanbing.jiang@intel.com>
|
2025-08-04 18:28:31 -07:00 |
|
YanbingJiang
|
1fe691a429
|
Fix FP8 block quantization when N or K is not multiples of 128 (#8648)
|
2025-08-01 15:57:19 -07:00 |
|
YanbingJiang
|
b044400dd3
|
Support non-contiguous query input for extend/decode attention (#7462)
|
2025-07-02 19:59:45 -07:00 |
|
Chunyuan WU
|
c5131f7a2f
|
[CPU] add c++ kernel to bind CPU cores and memory node (#7524)
|
2025-06-29 19:45:25 -07:00 |
|
YanbingJiang
|
0e05fe8cf4
|
Update seed in CPU UTs to avoid flaky failure with single test (#7544)
|
2025-06-25 21:25:50 -07:00 |
|
 Chunyuan WUandThien Tran
|
7eb47b0f3d
|
[CPU] [BF16] Call fused_experts_cpu, weight_packed_linear and bmm_cpu kernel in DeepSeek model (#6641)
Co-authored-by: Thien Tran <gau.nernst@yahoo.com.sg>
|
2025-06-25 01:43:33 -07:00 |
|