Cheng Wan
|
2d0e94e3a3
|
Check the topology identities where the layout is written, and build at the published widths (#40340)
|
2026-09-21 12:22:59 -07:00 |
|
Xiaoyu Zhang
|
cb22f2451e
|
[Cleanup] Deduplicate kernel tests, diffusion fixtures and benchmark helpers (#40265)
|
2026-09-19 19:45:38 +08:00 |
|
maithilijoshi20
|
0dad91d50f
|
Fix int4 MoE tuner config filename (#35260)
Signed-off-by: maithilijoshi20 <maithilij2003@gmail.com>
|
2026-09-18 13:56:38 +08:00 |
|
Hank Han
|
a64be2e430
|
[Kernel] Add H20 block-FP8 MoE configs for GLM-5.3-Flash EP4/EP8 (#38913)
|
2026-09-15 23:42:18 -07:00 |
|
Rahul Vijayaraghavan
|
d4ad368ed9
|
[Intel XPU] Enable fused_moe_triton tuning on XPU and add tuned DeepSeek-OCR-2 configs (#28723)
|
2026-09-16 10:20:22 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
99060191e7
|
[KDA] Support ReplaySSM ring-write in the fused chain-verify kernel (#36821)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-09-15 10:22:41 +08:00 |
|
+3        
|
52fecfdf09
|
support qwen 3.8 flash next (#37500)
Co-authored-by: ch-wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: ispobock <26454835+ispobock@users.noreply.github.com>
Co-authored-by: JustinTong0323 <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: samuellees <26428561+samuellees@users.noreply.github.com>
Co-authored-by: YAMY1234 <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: yhyang201 <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: yizhang2077 <25844240+yizhang2077@users.noreply.github.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Shinto C V <cshintov@gmail.com>
Co-authored-by: Julian Huang <huangzhilin.hzl@antgroup.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
|
2026-09-08 13:56:21 -07:00 |
|
 DavidLiandClaude Opus 5
|
5ae4ccb0fd
|
[Bench] Amortize the GDN ReplaySSM decode latency over the flush cycle (#38399)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
|
2026-09-08 16:38:01 +08:00 |
|
Chan ahn
|
dcebe8c473
|
[Kernel] Add fused MoE Triton configs for Qwen3.8-Flash-Next FP8 on NVIDIA H200 NVL (TP2+EP2) (#38116)
|
2026-09-07 10:17:55 -07:00 |
|
Xiaoyu Zhang
|
397aeca376
|
fix(benchmark): support Glm4MoeLite in fused MoE tuner (#37623)
|
2026-09-03 17:21:42 +08:00 |
|
 Alex NailsandAlison Shao
|
28262c20df
|
[CI][RFC] Replace black-jupyter with ruff-format (#37210)
Co-authored-by: Alison Shao <a.shao@wustl.edu>
|
2026-09-02 19:46:08 -07:00 |
|
       
|
20621aa14b
|
[Model] Support Ling-3.0-flash (BailingMoeV3) (#33561)
Signed-off-by: JustinTong <justintong0323@gmail.com>
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
Co-authored-by: 得泽 <zhangkaihong.zkh@antgroup.com>
Co-authored-by: 翎悦 <vito.yy@antgroup.com>
Co-authored-by: 羽癫 <yudian.zy@antgroup.com>
Co-authored-by: tiwei.btw <tiwei.btw@antgroup.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: 文赋 <zibin.zb@antgroup.com>
Co-authored-by: JustinTong <justintong0323@gmail.com>
|
2026-08-26 17:27:23 -07:00 |
|
TobyMint
|
3adc70bb5e
|
[MoE] Add H20 fp8_w8a8 tuned configs for Qwen3.8 (triton 3.7.1) + fix Qwen3_5MoeForCausalLM tuning (#34795)
|
2026-08-16 19:41:03 -07:00 |
|
 SII-yangdianandSII-yangdian
|
f2b2b567aa
|
perf(jit_kernel/deepseek_v4): optimize paged_mqa_metadata (#25855)
Co-authored-by: SII-yangdian <yangdian@sii.edu.cn>
|
2026-08-13 19:00:22 -07:00 |
|
   
|
ef7208d41d
|
[kernel] add triton moe TMA up support (#33559)
Co-authored-by: undefined <zhouchen.arrebol@jd.com>
Co-authored-by: xq25478 <xq25478@qq.com>
Co-authored-by: xieminghe.simon <xieminghe.simon@jd.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-08-13 13:29:32 +08:00 |
|
 
|
11d03eaeef
|
runtime: Add flashinfer rmsnorm + quant fusion support SM90, SM100, SM120- #32994 (#33471)
Signed-off-by: Devashish Lal <devcode@fb.com>
Co-authored-by: Devashish Lal <devcode@fb.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-08-09 20:15:03 +08:00 |
|
Baizhou Zhang
|
eb31a53338
|
Revert "Add flashinfer rmsnorm + quant fusion support SM90, SM100, SM120" (#33455)
|
2026-08-03 19:03:56 -07:00 |
|
 
|
3960983753
|
Add flashinfer rmsnorm + quant fusion support SM90, SM100, SM120 (#32994)
Signed-off-by: Devashish Lal <devcode@fb.com>
Co-authored-by: Devashish Lal <devcode@fb.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-08-04 08:37:42 +08:00 |
|
 Ho-Ren (Jack) ChuangandClaude Fable 5
|
e4a40a71f8
|
[DSA] Q8KV8 FP8 Sparse Prefill on GLM-5.2 & DeepSeek-V3.2: Q8-Path & Shared-Path Optimizations (#31888)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-30 15:15:11 +08:00 |
|
Xiaoyu Zhang
|
1d9c292547
|
[Kernel] Add inventory guards and clean benchmark layout (#32788)
|
2026-07-30 09:03:24 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
11b0e5c5ad
|
[Kernel] Classification cleanup: unify _jit_ naming, drop empty/model groups, add elementwise (RFC #29630) (#32148)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-23 13:47:02 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
99f636a86f
|
[Kernel] RFC #29630 finale: retire sglang.jit_kernel into sglang.kernels (#32072)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-23 08:35:09 +08:00 |
|
  
|
e8e765b9d6
|
[AMD] Add fused all-reduce RMSNorm per-group quant for Qwen3.5 FP8 (#24651)
Co-authored-by: jacky.cheng <yichiche@amd.com>
Co-authored-by: yctseng0211 <yctseng@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-07-22 07:33:03 -07:00 |
|
 
|
03342e7732
|
Delete sgl-kernel AOT router GEMM and fused A GEMM (#30280)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
Co-authored-by: root <root@sgl-b300-inference.datacrunch.io>
|
2026-07-22 08:44:59 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
4aadf94146
|
[Kernel] Relocate vendored fla and mamba kernel trees to sglang.kernels (RFC #29630, Phase 2.5, 7/7) (#30795)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-15 12:52:15 +08:00 |
|
  ![gemini-code-assist[bot]](/assets/img/avatar_default.png)
|
702bddcee8
|
[Model] Add support for JetBrains' Mellum v2 code generation model (#27375)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Jiminator <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-07-13 22:54:38 -07:00 |
|
 Brayden Zhongandroot
|
9756f768a6
|
Refactor FP4 quantization and remove deprecated JIT kernels (#30448)
Co-authored-by: root <root@sgl-b300-inference.datacrunch.io>
|
2026-07-14 09:22:07 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
e9ef06c560
|
[Kernel] Migrate top-level srt/layers stray kernels to sglang.kernels (RFC #29630, Phase 2.5, 3/7) (#30787)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-14 09:20:59 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
874fc07d9b
|
[Kernel] Migrate scattered quantization kernels to sglang.kernels (RFC #29630, Phase 2.5, 1/7) (#30784)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-13 16:17:01 +08:00 |
|
Mick
|
a358abd651
|
chore: update vlm moe config and tune scripts (#30866)
|
2026-07-12 08:35:59 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
6ed9843b57
|
[Kernel] Introduce sglang.kernels namespace and migrate scattered triton_ops kernels (RFC #29630, Phase 2) (#30044)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-10 21:41:08 +08:00 |
|
 
|
fefc1743a9
|
Cute-DSL FP8 MQA logits (#25220)
Co-authored-by: Mindy Li <11663212+limin2021@users.noreply.github.com>
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-07-06 23:34:07 -07:00 |
|
 Aditya KamatandMick
|
1589603114
|
model: support baidu unlimited-ocr (#29186)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-06-27 23:36:19 +08:00 |
|
Xinyuan Tong
|
7c23d2255a
|
[minimax-m3] Split 1/4: sparse attention ops + JIT kernels + config foundation (#28712)
|
2026-06-22 13:10:43 -07:00 |
|
giang_ng_tr
|
b36360dc5b
|
[AMD][Perf] Split-KV flash-decode attention for EAGLE target-verify (Triton backend) (#27382)
|
2026-06-18 19:11:10 -07:00 |
|
Mohammad Miadh Angkad
|
91c63aeb4d
|
Fix stale CUDA graph benchmark and docs refs (#28041)
|
2026-06-13 21:51:42 -07:00 |
|
Yuhao Yang
|
aea0e30853
|
[4/N] Qwen3.5Opt: Overlap mamba verify update with draft extend (#26924)
|
2026-06-13 20:29:20 +08:00 |
|
 
|
c2eae96c56
|
MSCCL++ Integration (#22734)
Co-authored-by: Caio Rocha <caiorocha@microsof.com>
Co-authored-by: empyreus <rjsouza1995@gmail.com>
|
2026-06-08 21:13:13 -07:00 |
|
Hubert Lu
|
72929c7000
|
[AMD] Enable AITER custom all-gather on ROCm (#25093)
|
2026-06-02 15:57:37 -07:00 |
|
  
|
ba2ffcf156
|
Add DeepSeekV4 fused MoE Triton autotune support (#25569)
Co-authored-by: undefined <zhouchen.arrebol@jd.com>
Co-authored-by: xq25478 <xq25478@qq.com>
Co-authored-by: xieminghe.simon <xieminghe.simon@jd.com>
|
2026-05-18 21:35:32 +08:00 |
|
 sglang-botandClaude Opus 4.7
|
0a2615df24
|
chore: add vLLM SPDX copyright headers to ported files (#25182)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
2026-05-13 15:17:30 -07:00 |
|
 Brayden Zhongandb8zhong
|
1d80a1a9fe
|
Use Cute-DSL NVFP4 quantization kernels (#23745)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
|
2026-05-11 00:40:02 -07:00 |
|
RunningLeon
|
335dbd60b4
|
Support Intern-S2-Preview (#24875)
|
2026-05-10 22:17:30 +08:00 |
|
hhwxw
|
d9270b8c6a
|
fix(moe): relocate orphan tuned configs after #23019 (#24004)
|
2026-04-29 02:00:13 -07:00 |
|
Muqi Li
|
69a71219cb
|
feat: tiny improve fp8_gemm tune usage (#23912)
|
2026-04-28 07:47:46 -04:00 |
|
   
|
6d03861476
|
support Hy3 preview (#23533)
Co-authored-by: pengmeng <pengmeng@tencent.com>
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
Co-authored-by: chengvjiang <chengvjiang@tencent.com>
Co-authored-by: russellfeng <russellfeng@tencent.com>
|
2026-04-24 12:03:24 -07:00 |
|
 Piotr MazurekandPiotr Mazurek
|
6cf0b004ca
|
[MoE] Add LFM2 MoE tuning support + tuned configs for H100/B200/MI325X (#22791)
Co-authored-by: Piotr Mazurek <piotr.mazurek@liquid.ai>
|
2026-04-21 18:32:05 -07:00 |
|
 Cheng WanandClaude Opus 4.7
|
5f7aee726a
|
refactor(moe): de-duplicate triton MoE runner path into shared helpers (#23019)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-04-17 17:05:13 -07:00 |
|
 Hubert LuandHAI
|
edaa5973d4
|
[AMD][No-Merge] Simplify fused allreduce + RMSNorm and remove hidden_dim allowlist (#21986)
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-04-11 23:47:08 -07:00 |
|
 satyamk7054andSatyam Kumar
|
059b287e25
|
Add offline auto-tuning for LoRA CSGMV kernel (#20391)
Co-authored-by: Satyam Kumar <satyamk@linkedin.com>
|
2026-04-10 13:10:43 -07:00 |
|