Cheng Wan
|
2d0e94e3a3
|
Check the topology identities where the layout is written, and build at the published widths (#40340)
|
2026-09-21 12:22:59 -07:00 |
|
maithilijoshi20
|
0dad91d50f
|
Fix int4 MoE tuner config filename (#35260)
Signed-off-by: maithilijoshi20 <maithilij2003@gmail.com>
|
2026-09-18 13:56:38 +08:00 |
|
Hank Han
|
a64be2e430
|
[Kernel] Add H20 block-FP8 MoE configs for GLM-5.3-Flash EP4/EP8 (#38913)
|
2026-09-15 23:42:18 -07:00 |
|
Rahul Vijayaraghavan
|
d4ad368ed9
|
[Intel XPU] Enable fused_moe_triton tuning on XPU and add tuned DeepSeek-OCR-2 configs (#28723)
|
2026-09-16 10:20:22 +08:00 |
|
+3        
|
52fecfdf09
|
support qwen 3.8 flash next (#37500)
Co-authored-by: ch-wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: ispobock <26454835+ispobock@users.noreply.github.com>
Co-authored-by: JustinTong0323 <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: samuellees <26428561+samuellees@users.noreply.github.com>
Co-authored-by: YAMY1234 <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: yhyang201 <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: yizhang2077 <25844240+yizhang2077@users.noreply.github.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Shinto C V <cshintov@gmail.com>
Co-authored-by: Julian Huang <huangzhilin.hzl@antgroup.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
|
2026-09-08 13:56:21 -07:00 |
|
Chan ahn
|
dcebe8c473
|
[Kernel] Add fused MoE Triton configs for Qwen3.8-Flash-Next FP8 on NVIDIA H200 NVL (TP2+EP2) (#38116)
|
2026-09-07 10:17:55 -07:00 |
|
Xiaoyu Zhang
|
397aeca376
|
fix(benchmark): support Glm4MoeLite in fused MoE tuner (#37623)
|
2026-09-03 17:21:42 +08:00 |
|
 Alex NailsandAlison Shao
|
28262c20df
|
[CI][RFC] Replace black-jupyter with ruff-format (#37210)
Co-authored-by: Alison Shao <a.shao@wustl.edu>
|
2026-09-02 19:46:08 -07:00 |
|
       
|
20621aa14b
|
[Model] Support Ling-3.0-flash (BailingMoeV3) (#33561)
Signed-off-by: JustinTong <justintong0323@gmail.com>
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: 得泽 <zhangkaihong.zkh@antgroup.com>
Co-authored-by: 翎悦 <vito.yy@antgroup.com>
Co-authored-by: 羽癫 <yudian.zy@antgroup.com>
Co-authored-by: tiwei.btw <tiwei.btw@antgroup.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: 文赋 <zibin.zb@antgroup.com>
Co-authored-by: JustinTong <justintong0323@gmail.com>
|
2026-08-26 17:27:23 -07:00 |
|
TobyMint
|
3adc70bb5e
|
[MoE] Add H20 fp8_w8a8 tuned configs for Qwen3.8 (triton 3.7.1) + fix Qwen3_5MoeForCausalLM tuning (#34795)
|
2026-08-16 19:41:03 -07:00 |
|
   
|
ef7208d41d
|
[kernel] add triton moe TMA up support (#33559)
Co-authored-by: undefined <zhouchen.arrebol@jd.com>
Co-authored-by: xq25478 <xq25478@qq.com>
Co-authored-by: xieminghe.simon <xieminghe.simon@jd.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-08-13 13:29:32 +08:00 |
|
  ![gemini-code-assist[bot]](/assets/img/avatar_default.png)
|
702bddcee8
|
[Model] Add support for JetBrains' Mellum v2 code generation model (#27375)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Jiminator <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-07-13 22:54:38 -07:00 |
|
Mick
|
a358abd651
|
chore: update vlm moe config and tune scripts (#30866)
|
2026-07-12 08:35:59 +08:00 |
|
 Aditya KamatandMick
|
1589603114
|
model: support baidu unlimited-ocr (#29186)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-06-27 23:36:19 +08:00 |
|
Xinyuan Tong
|
7c23d2255a
|
[minimax-m3] Split 1/4: sparse attention ops + JIT kernels + config foundation (#28712)
|
2026-06-22 13:10:43 -07:00 |
|
Mohammad Miadh Angkad
|
91c63aeb4d
|
Fix stale CUDA graph benchmark and docs refs (#28041)
|
2026-06-13 21:51:42 -07:00 |
|
  
|
ba2ffcf156
|
Add DeepSeekV4 fused MoE Triton autotune support (#25569)
Co-authored-by: undefined <zhouchen.arrebol@jd.com>
Co-authored-by: xq25478 <xq25478@qq.com>
Co-authored-by: xieminghe.simon <xieminghe.simon@jd.com>
|
2026-05-18 21:35:32 +08:00 |
|
 sglang-botandClaude Opus 4.7
|
0a2615df24
|
chore: add vLLM SPDX copyright headers to ported files (#25182)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
2026-05-13 15:17:30 -07:00 |
|
RunningLeon
|
335dbd60b4
|
Support Intern-S2-Preview (#24875)
|
2026-05-10 22:17:30 +08:00 |
|
hhwxw
|
d9270b8c6a
|
fix(moe): relocate orphan tuned configs after #23019 (#24004)
|
2026-04-29 02:00:13 -07:00 |
|
   
|
6d03861476
|
support Hy3 preview (#23533)
Co-authored-by: pengmeng <pengmeng@tencent.com>
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
Co-authored-by: chengvjiang <chengvjiang@tencent.com>
Co-authored-by: russellfeng <russellfeng@tencent.com>
|
2026-04-24 12:03:24 -07:00 |
|
 Piotr MazurekandPiotr Mazurek
|
6cf0b004ca
|
[MoE] Add LFM2 MoE tuning support + tuned configs for H100/B200/MI325X (#22791)
Co-authored-by: Piotr Mazurek <piotr.mazurek@liquid.ai>
|
2026-04-21 18:32:05 -07:00 |
|
 Cheng WanandClaude Opus 4.7
|
5f7aee726a
|
refactor(moe): de-duplicate triton MoE runner path into shared helpers (#23019)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-17 17:05:13 -07:00 |
|
 ![gemini-code-assist[bot]](/assets/img/avatar_default.png)    
|
2813cb6d9a
|
[New Model] Gemma 4 (#21952)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Pengyu Chen <pychen96@gmail.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Andy Luo <andy.luo@amd.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: adarshxs <adarsh.shirawalmath@gmail.com>
|
2026-04-06 20:24:44 -07:00 |
|
Polisetty V R K Jyothendra Varma
|
f0303fd07e
|
[Intel GPU] Enable DeepSeek R1 inference on XPU (#18461)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
|
2026-03-29 22:35:59 -07:00 |
|
zhangxiaolei
|
e2b8463c80
|
[fix] qwen3.5 fuse_moe_triton_tune bug (#20232)
|
2026-03-27 19:23:24 -04:00 |
|
Mook
|
abc672e717
|
[Benchmark] use flashinfer bench_gpu_time instead of triton do_bench (#20305)
|
2026-03-12 04:04:30 +00:00 |
|
 Yuan Luoandluoyuan.luo
|
751c454099
|
Add DeepSeek3.2 and GlmMoeDsa into moe tune (#18876)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-03-10 17:12:58 +08:00 |
|
RoyWang
|
a1ef8e2cc0
|
[AMD] optimize Kimi K2.5 fused_moe_triton performance by tuning (#19228)
|
2026-02-26 11:50:13 -08:00 |
|
 satyamk7054andSatyam Kumar
|
355127c2e9
|
Fix benchmark_sglang_fused_moe_triton.py (#18940)
Co-authored-by: Satyam Kumar <satyamk@linkedin.com>
|
2026-02-17 17:25:37 -05:00 |
|
 Zheng Liand瑀澈
|
27c447653d
|
model: support Qwen3.5 (#18489)
Co-authored-by: 瑀澈 <yuche.lz@alibaba-inc.com>
|
2026-02-10 00:27:59 +08:00 |
|
b8zhong
|
22498e10c0
|
[Fix] Triton TP MoE Dpsk V3/Qwen3 Coder with SwapAB (#17965)
|
2026-01-31 15:56:26 +08:00 |
|
 Julian Huangand墨楼
|
db2425a00b
|
[Fix]: correctly fetch ds32 config in tuning_fused_moe_triton (#17409)
Co-authored-by: 墨楼 <huangzhilin.hzl@antgroup.com>
|
2026-01-20 20:08:28 +08:00 |
|
Yongfei Xu
|
82a1b645ba
|
[DeepSeek V3.1/V3.2] Optimize fused moe configs for H20 & H20-3E based on swapab (#17133)
|
2026-01-17 00:10:52 +08:00 |
|
roikoren755
|
b021332339
|
[NemotronH] Add latent MoE support (#16227)
Signed-off-by: Roi Koren <roik@nvidia.com>
|
2026-01-02 22:08:58 +08:00 |
|
 
|
8428078436
|
Add Mistral Large 3 support. (#14213)
Co-authored-by: elvischenv <219235043+elvischenv@users.noreply.github.com>
Co-authored-by: Linda-Stadter <57756729+Linda-Stadter@users.noreply.github.com>
|
2025-12-04 20:00:05 +08:00 |
|
 
|
982db4ebac
|
Feat: GLM-4.6 supports shared experts fusion (#13873)
Signed-off-by: UranusSeven <109661872+UranusSeven@users.noreply.github.com>
Co-authored-by: Kevin-XiongC <kevin_xiong1997@outlook.com>
Co-authored-by: Mingyi Jin <jinmingyi1998@sina.cn>
|
2025-12-01 11:33:18 +08:00 |
|
roikoren755
|
1b48e1b974
|
Feat/nemotron nano v3 support (#12690)
|
2025-11-21 13:53:05 -08:00 |
|
Junlin Zhou
|
0779c3d148
|
docs: update fused MoE config path (#13211)
|
2025-11-13 11:14:01 -08:00 |
|
Xiaoyu Zhang
|
f18ec927f3
|
fix tuning_fused_moe_triton_sep tool per_channel_quant bug (#13027)
|
2025-11-11 10:33:54 +08:00 |
|
 
|
fc84b0730c
|
[Refactor] Refactor fused_moe_triton tuning tools: extract shared utils, add EP/MLLM support, reduce overhead (#12440)
Co-authored-by: xu-yfei <xu-yfei@users.noreply.github.com>
Co-authored-by: Yongfei Xu <xuyongfei.xyf@antgroup.com>
|
2025-11-06 20:54:42 +08:00 |
|
 Xinyuan TongandCursor Agent
|
82cfcd3bb8
|
[Refactor] tuning_fused_moe for MLLM and small refactor (#11224)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
|
2025-10-31 08:54:14 +08:00 |
|
Chen1022
|
1ed1abfd45
|
feat: add EP support in tuning (#12012)
|
2025-10-30 07:58:50 -07:00 |
|
Xiaoyu Zhang
|
04e5b6faa7
|
Revert "Triton fused_moe_kernel support ep moe tuning" (#12377)
|
2025-10-30 07:12:06 -07:00 |
|
Xiaoyu Zhang
|
52694b60da
|
Triton fused_moe_kernel support ep moe tuning (#12343)
|
2025-10-29 23:16:09 +08:00 |
|
 Liana KolevaandXiaoyu Zhang
|
1357397a34
|
feat: preview filename from tuning_fused_moe_triton.py (#12276)
Co-authored-by: Xiaoyu Zhang <35585791+BBuf@users.noreply.github.com>
|
2025-10-29 16:12:25 +08:00 |
|
 Yongfei Xuandybyang
|
d2b8c4123e
|
Opt fused triton moe: add tma for down proj kernel (#10567)
Co-authored-by: ybyang <10629930+whybeyoung@users.noreply.github.com>
|
2025-10-28 14:26:17 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
4f42c8cd3e
|
[sgl-kernel] Support float64 moe_sum_reduce cuda kernel (#11068)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2025-10-07 14:31:11 +00:00 |
|
lukec
|
77830a265e
|
Add fuse_moe per-channel tune (#10915)
|
2025-09-25 21:12:09 +08:00 |
|
Yiakwy
|
984730b732
|
add tunning files for QWEN-3-NEXT (#10794)
|
2025-09-23 12:46:30 -07:00 |
|