 Mohammad Miadh AngkadandMohammad Angkad
|
72d5c5bb73
|
[Kimi-K3] Accept fp32 routing weights in the fused MoE finalize (#38612)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
|
2026-09-09 01:33:39 -07:00 |
|
Yuhao Yang
|
13469c16d3
|
Store mamba prefix-cache checkpoints at the configured SSM state dtype (#34820)
|
2026-09-09 15:25:57 +08:00 |
|
sogalin_codegen
|
5998e9321c
|
[AMD][Fix] Fix regression in jit build error with FLUX.2-dev on gfx1250 (#38329)
|
2026-09-08 22:55:13 -07:00 |
|
 DayuxiaoshuiandXiaoyu Zhang
|
db1de6ff4c
|
[Diffusion] Keep the Wan VAE decoder channels_last and add a Triton NHWC nearest upsample (#38182)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-09-09 09:31:35 +08:00 |
|
Cheng Wan
|
db272201a2
|
[Config] Retire get_global_server_args, and clear the deprecated flags that have a replacement (#38375)
|
2026-09-08 16:42:12 -07:00 |
|
 Chunan ZengandKevin Mi
|
5177a3ec08
|
MiniMax-M3: Triton split-K router GEMV with in-kernel fixup (#36557)
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
|
2026-09-08 16:39:41 -07:00 |
|
 Chunan ZengandKevin Mi
|
a25bbca8ed
|
MiniMax-M3: share the sparse index top-k across layers and reuse the decode top-k buffer (#36527)
Co-authored-by: Kevin Mi <mikevin920@yahoo.com>
|
2026-09-08 16:39:01 -07:00 |
|
Baizhou Zhang
|
ed183d45ac
|
[CP V1 Deprecation 4/5] Canonicalize prefill CP API names (#36229)
|
2026-09-08 16:03:36 -07:00 |
|
+3        
|
52fecfdf09
|
support qwen 3.8 flash next (#37500)
Co-authored-by: ch-wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: ispobock <26454835+ispobock@users.noreply.github.com>
Co-authored-by: JustinTong0323 <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: samuellees <26428561+samuellees@users.noreply.github.com>
Co-authored-by: YAMY1234 <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: yhyang201 <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: yizhang2077 <25844240+yizhang2077@users.noreply.github.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Shinto C V <cshintov@gmail.com>
Co-authored-by: Julian Huang <huangzhilin.hzl@antgroup.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
|
2026-09-08 13:56:21 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
30e7a3072d
|
Keep fp32 routing weights in the fp8 block-scale and bf16 trtllm MoE (#33631)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-09-08 08:56:40 -07:00 |
|
Xiaoyu Zhang
|
554f817948
|
[Diffusion] Optimize LTX-2 QKNorm and split RoPE on Hopper (#38396)
|
2026-09-08 19:05:02 +08:00 |
|
 zijiecandZijie Chen
|
e634ba78a4
|
[AMD] gfx950 assembly attention for EAGLE verify, draft extend and decode (#37465)
Co-authored-by: Zijie Chen <300606707+zijiecode@users.noreply.github.com>
|
2026-09-08 03:10:09 -07:00 |
|
amd-danli103
|
141febf329
|
[AMD] fix: use the hardware fp8 e4m3 convert on gfx950 (#37140)
Signed-off-by: amd-danli103 <danli103@amd.com>
|
2026-09-08 03:01:53 -07:00 |
|
HZY
|
8656901504
|
[Fix][DSA] Bound prefill Triton specializations for page-table stride (#37093)
|
2026-09-08 10:23:51 +08:00 |
|
Liangsheng Yin
|
4dcecc7891
|
Revert "[kernel] add fused silu mul quant fp8" (#38381)
|
2026-09-07 17:36:10 -07:00 |
|
 
|
570087ceda
|
[AMD][DSV4] Reland unified-KV pool sizing and SWA ring accounting, fully gated (#38192)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
|
2026-09-07 13:13:04 -07:00 |
|
 Mohammad Miadh AngkadandMohammad Angkad
|
a711785475
|
[Kernel] Drop the vendored dense BF16 GEMM port in favor of FlashInfer 0.6.18 (#38124)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
|
2026-09-07 12:59:33 -07:00 |
|
Cheng Wan
|
b5766336d4
|
[Perf] Unified memory: close the DCP decode gap on Blackwell (#37926)
|
2026-09-07 01:10:44 -07:00 |
|
 Xiaoyu ZhangandMick Qian
|
4d23a4fa6d
|
[Test] Consolidate test cleanup and CI taxonomy (net -11.4K lines) (#37436)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-07 15:13:59 +08:00 |
|
Kotthagattu Meher Sai
|
97dcbf9410
|
[XPU] Re-add intel xpu on triton paths in diffusion platforms (#36654)
|
2026-09-07 13:30:22 +08:00 |
|
   
|
a8b2f36dee
|
[kernel] add fused silu mul quant fp8 (#37376)
Co-authored-by: undefined <zhouchen.arrebol@jd.com>
Co-authored-by: xq25478 <xq25478@qq.com>
Co-authored-by: xieminghe.simon <xieminghe.simon@jd.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-09-07 11:39:11 +08:00 |
|
 
|
c4e52a1051
|
XPU: Enable GLM5.1 (GlmMoeDsaForCausalLM) DSA Attention (#24959)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-09-07 09:24:45 +08:00 |
|
 
|
39a80354aa
|
[MUSA] Add installation guide and Dockerfile (#36709)
Co-authored-by: zhiguo.qin <zhiguo.qin@mthreads.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
|
2026-09-06 20:13:53 -05:00 |
|
 Yuxingwang-intelandMa Mingfei
|
707da81e84
|
[CPU] Add native CPU kernel for MurmurHash32 (#35604)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-09-07 09:10:35 +08:00 |
|
+2        
|
97c6978369
|
GLM-5.3-Flash support (#36507)
Co-authored-by: zRzRzRzRzRzRzR <Yuxuan.Zhang2@liverpool.ac.uk>
Co-authored-by: Shijin Zhang <75300765+Dovis01@users.noreply.github.com>
Co-authored-by: zanes-ops <zanes@nvidia.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Jian Chen <jianchen0311@gmail.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: andyluo7 <43718156+andyluo7@users.noreply.github.com>
Co-authored-by: Ehsan Akhgari <ehsan.akhgari@gmail.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com>
|
2026-09-06 02:27:59 -07:00 |
|
Zhang, Jiejing
|
6cee9285a3
|
[ROCm] Make DSA indexer top-k exact with cooperative selection (#37591)
|
2026-09-05 23:55:56 -07:00 |
|
Liangsheng Yin
|
f5819b09bf
|
Revert "[AMD][DSV4] Fix unified-KV pool sizing and SWA ring accounting" (#38163)
|
2026-09-05 17:28:46 -07:00 |
|
yuttian1
|
514b45fd34
|
[AMD][DSV4] Fix unified-KV pool sizing and SWA ring accounting (#30315)
|
2026-09-05 16:39:40 -07:00 |
|
 Xiaoyu ZhangandWaterpine
|
ccf9fe6590
|
[Kernel] Add KDA FP8 skinny GEMM for SM120 (#38082)
Co-authored-by: Waterpine <biansonghz@gmail.com>
|
2026-09-05 22:27:06 +08:00 |
|
Xiaoyu Zhang
|
dc2843801d
|
perf(lfm2): fuse gating and short convolution on SM90 (#37622)
|
2026-09-05 21:52:16 +08:00 |
|
 
|
bd16c22a04
|
[diffusion] fuse LingBot MoE group-limited top-k index selection (#38044)
Co-authored-by: BBuf <bbuf@users.noreply.github.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-05 18:12:30 +08:00 |
|
 Xiaoyu ZhangandBBuf
|
d49180019b
|
fix(moe): cast filtered-activation expert_ids to int32 for torch.compile (#38085)
Co-authored-by: BBuf <bbuf@users.noreply.github.com>
|
2026-09-05 17:13:07 +08:00 |
|
Xiaoyu Zhang
|
a74470e904
|
fix(mamba): unify causal_conv1d col* dtype to x (MiniCPM-V-4.6 GDN prefill bf16/fp16 mismatch) (#38039)
|
2026-09-05 17:08:20 +08:00 |
|
 Xiaoyu ZhangandBBuf
|
d6e0a8cbf4
|
[diffusion] add Helios per-token gated-residual fusion (quality-gated) (#38042)
Co-authored-by: BBuf <bbuf@users.noreply.github.com>
|
2026-09-05 10:29:09 +08:00 |
|
 
|
55bf3380e0
|
Support Hy4-preview (#36805)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: alphabetc1 <2508695655@qq.com>
|
2026-09-04 18:03:49 -07:00 |
|
   
|
3c2724c48d
|
[AMD] Optimize Kimi-K3 Triton MLA prefill on gfx950 (#35770)
Co-authored-by: clintg6 <7388379+clintg6@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
|
2026-09-04 17:44:30 -07:00 |
|
YAMY
|
db89f639ef
|
[GDN] Amortize ReplaySSM checkpoint materialization (#35544)
|
2026-09-04 15:13:20 -07:00 |
|
YAMY
|
07199fa220
|
[Performance] Optimize Qwen3.5 GDN prefill projection layouts (#36267)
|
2026-09-04 10:32:44 -07:00 |
|
 Ke Baoandadityakamat24
|
0b57847ebf
|
Fix mamba radix cache ssm state indexing (#37836)
Co-authored-by: adityakamat24 <adityakamat007@gmail.com>
|
2026-09-04 15:46:33 +08:00 |
|
 AMD-yanfeiwangandkk
|
7f89cc5286
|
[AMD] Skip unused TOPK v2 plan kernel on ROCm (#37580)
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com>
|
2026-09-04 00:00:38 -07:00 |
|
Hemanth Acharya
|
31ebd8f437
|
[AMD][DSv4] Fuse the DSv4 FP4 indexer prefill-schedule preamble into one kernel (#37764)
|
2026-09-03 23:44:49 -07:00 |
|
Xiaoyu Zhang
|
54c2c99feb
|
[Diffusion] Fuse LingBot per-token gated residual and RMSNorm modulate (#37910)
|
2026-09-04 14:35:13 +08:00 |
|
 mqhc2020andBingxu Chen
|
4e756ecc4a
|
[AMD] CI: fix Lean decode crash on the EAGLE path (#37119)
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-09-03 23:04:21 -07:00 |
|
Mohammad Miadh Angkad
|
3ad3f23ed5
|
[Comm] Drop the in-tree MNNVL CuTe DSL port in favor of FlashInfer 0.6.18 (#37206)
|
2026-09-03 21:49:48 -07:00 |
|
 
|
2bb25dc18b
|
[Speculative Decoding] Add native UNO serving support (#37667)
Co-authored-by: drproduck <drproduck@MacBook-Air-2.local>
Co-authored-by: BBuf <1182563586@qq.com>
|
2026-09-03 20:08:41 +08:00 |
|
Cheng Wan
|
a11dba1a01
|
[Feature] Unified memory: support decode context parallelism for the trtllm_mla family (#37693)
|
2026-09-03 03:37:12 -07:00 |
|
 Xinyi SongandThomas Wang
|
7ed29eba80
|
[AMD] Fix FP4 indexer OOR (#37660)
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
|
2026-09-03 01:49:46 -07:00 |
|
 YC Yen-Ching TsengandPhil Li
|
1fb85053e7
|
[AMD][Diffusion] Migrate FlyDSL fused norm kernels to the v0.3.0 stable API (#36349)
Co-authored-by: Phil Li <haicli@amd.com>
|
2026-09-02 23:03:08 -07:00 |
|
James Liu
|
4229088a48
|
feat(kernels): generalize persistent CuTe JIT cache (#33911)
|
2026-09-02 19:59:04 -07:00 |
|
 Alex NailsandAlison Shao
|
28262c20df
|
[CI][RFC] Replace black-jupyter with ruff-format (#37210)
Co-authored-by: Alison Shao <a.shao@wustl.edu>
|
2026-09-02 19:46:08 -07:00 |
|