Xiaoyu Zhang
|
cb22f2451e
|
[Cleanup] Deduplicate kernel tests, diffusion fixtures and benchmark helpers (#40265)
|
2026-09-19 19:45:38 +08:00 |
|
Xiaoyu Zhang
|
0b0d2c257a
|
[Fix] Repair CI fixtures and ROCm speculative tree device checks (#40325)
|
2026-09-19 18:16:05 +08:00 |
|
Cheng Wan
|
afe71f4b9e
|
Read process groups through the runtime context (#40068)
|
2026-09-18 17:40:32 -07:00 |
|
   
|
5931fd60ee
|
Support unified memory page-envelope transfers in PD (#39477)
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
|
2026-09-18 17:39:50 -07:00 |
|
Lifan Shen
|
8ea0ee300d
|
perf(sampling): avoid GPU syncs when applying custom logit processors (#39234)
|
2026-09-18 17:09:30 -07:00 |
|
Jialin Ouyang
|
f3851486cb
|
[Perf] Fuse SWA page lookup and mapping clear (#38948)
|
2026-09-18 15:51:07 -07:00 |
|
  
|
0e5347db82
|
Support MXFP8 and deferred route weighting in DeepEP v2 (#40030)
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>
Co-authored-by: pranjalssh <14260275+pranjalssh@users.noreply.github.com>
|
2026-09-18 15:40:38 -07:00 |
|
 
|
81363bf8cb
|
[kernel] Share the warp vectorized copy and enforce its alignment (#36176)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: BBuf <1182563586@qq.com>
|
2026-09-18 22:40:54 +08:00 |
|
       
|
a6cf05817f
|
dsv4.1: remaining model and runtime integration (#38798)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <xiaoyu.zhang@radixark.ai>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Zhichen Zeng <zczeng@uw.edu>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
|
2026-09-18 02:55:30 -07:00 |
|
Nan Jiang
|
740f57a02c
|
[Spec] Fix CDF boundary handling in TreeSpeculativeSamplingTargetOnly (#35798)
|
2026-09-17 19:41:52 -07:00 |
|
Liangsheng Yin
|
f65c70bb7d
|
[Kernel] Move CUDA and ROCm speculative kernels to JIT (#40033)
|
2026-09-17 17:27:26 -07:00 |
|
Xiaoyu Zhang
|
7bc9152447
|
[Test] Consolidate kernel tests under plural kernels tree (#39966)
|
2026-09-18 07:37:48 +08:00 |
|
Liangsheng Yin
|
1f0c73e9bd
|
[DSV4] Generalize attention metadata, sparse prefill, and KV pool over compress ratios (#39921)
|
2026-09-17 15:55:19 -07:00 |
|
  
|
13d593b6cf
|
dsv4.1: compression, KV I/O, and metadata kernels (#39652)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
|
2026-09-16 13:54:07 -07:00 |
|
+4        
|
faaff1eca8
|
dsv4.1: Top-k kernels and candidate selection (#39648)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Xiaoyu Zhang <xiaoyu.zhang@radixark.ai>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Zhichen Zeng <zczeng@uw.edu>
Co-authored-by: Ziyi Xu <ziyi.xu@radixark.ai>
|
2026-09-15 23:54:35 -07:00 |
|
  
|
ddd4600197
|
[Fix] HiCache startup ImportError on the pinned kernel wheel (#39516)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
|
2026-09-14 22:31:19 -07:00 |
|
 JoeandXiaoyu Zhang
|
a23fd557ed
|
[Kernel] Add OOT dispatch for clamp position (#38687)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-09-15 10:22:21 +08:00 |
|
AMD-yanfeiwang
|
0163f8ff74
|
[AMD] Fix registered HiCache host pointer aliases (#35233)
|
2026-09-14 15:49:05 -07:00 |
|
 Xiaoyu ZhangandMick Qian
|
3e035a3513
|
[Diffusion] Optimize Qwen-Image-Edit attention on Hopper (#38584)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-13 09:15:01 +08:00 |
|
  
|
335f6aab27
|
[DSV4] Support raw-index output in TopK v2 (#33672)
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
|
2026-09-11 07:17:36 -07:00 |
|
Jialin Ouyang
|
6182d45524
|
[Test] Fix Q8KV8 sparse-prefill pool fixture after page-size rename (#39019)
|
2026-09-10 22:22:35 -07:00 |
|
Pengyun Lin
|
d076eec427
|
[SM120] Use exact query-head widths for DeepSeek-V4 sparse MLA decode (#36655)
|
2026-09-10 15:40:17 -07:00 |
|
paulzhang-tm
|
a63efd9056
|
[Spec] Support large MTP batches in short-convolution metadata (#38558)
|
2026-09-10 15:13:17 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
c0b790cf7f
|
Delete cutlass_mla, non-Marlin GPTQ, AWQ AOT kernel, and Dual Chunk Flash Attention (#32114)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-09-10 15:12:01 +08:00 |
|
 Mohammad Miadh AngkadandMohammad Angkad
|
72d5c5bb73
|
[Kimi-K3] Accept fp32 routing weights in the fused MoE finalize (#38612)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
|
2026-09-09 01:33:39 -07:00 |
|
 DayuxiaoshuiandXiaoyu Zhang
|
db1de6ff4c
|
[Diffusion] Keep the Wan VAE decoder channels_last and add a Triton NHWC nearest upsample (#38182)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-09-09 09:31:35 +08:00 |
|
Baizhou Zhang
|
ed183d45ac
|
[CP V1 Deprecation 4/5] Canonicalize prefill CP API names (#36229)
|
2026-09-08 16:03:36 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
30e7a3072d
|
Keep fp32 routing weights in the fp8 block-scale and bf16 trtllm MoE (#33631)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-09-08 08:56:40 -07:00 |
|
Xiaoyu Zhang
|
554f817948
|
[Diffusion] Optimize LTX-2 QKNorm and split RoPE on Hopper (#38396)
|
2026-09-08 19:05:02 +08:00 |
|
amd-danli103
|
141febf329
|
[AMD] fix: use the hardware fp8 e4m3 convert on gfx950 (#37140)
Signed-off-by: amd-danli103 <danli103@amd.com>
|
2026-09-08 03:01:53 -07:00 |
|
HZY
|
8656901504
|
[Fix][DSA] Bound prefill Triton specializations for page-table stride (#37093)
|
2026-09-08 10:23:51 +08:00 |
|
Liangsheng Yin
|
4dcecc7891
|
Revert "[kernel] add fused silu mul quant fp8" (#38381)
|
2026-09-07 17:36:10 -07:00 |
|
 
|
570087ceda
|
[AMD][DSV4] Reland unified-KV pool sizing and SWA ring accounting, fully gated (#38192)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
|
2026-09-07 13:13:04 -07:00 |
|
Mohammad Miadh Angkad
|
c5367fa964
|
[CI] Fix stale ServerArgs fake in chunked-SGMV LoRA test (#38315)
|
2026-09-07 04:25:12 -07:00 |
|
Cheng Wan
|
b5766336d4
|
[Perf] Unified memory: close the DCP decode gap on Blackwell (#37926)
|
2026-09-07 01:10:44 -07:00 |
|
 Xiaoyu ZhangandMick Qian
|
4d23a4fa6d
|
[Test] Consolidate test cleanup and CI taxonomy (net -11.4K lines) (#37436)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-07 15:13:59 +08:00 |
|
   
|
a8b2f36dee
|
[kernel] add fused silu mul quant fp8 (#37376)
Co-authored-by: undefined <zhouchen.arrebol@jd.com>
Co-authored-by: xq25478 <xq25478@qq.com>
Co-authored-by: xieminghe.simon <xieminghe.simon@jd.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-09-07 11:39:11 +08:00 |
|
 sglang-botandsglang-bot
|
6252993afe
|
chore: update CI test est_time values (#38238)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-09-06 17:49:41 -07:00 |
|
+2        
|
97c6978369
|
GLM-5.3-Flash support (#36507)
Co-authored-by: zRzRzRzRzRzRzR <Yuxuan.Zhang2@liverpool.ac.uk>
Co-authored-by: Shijin Zhang <75300765+Dovis01@users.noreply.github.com>
Co-authored-by: zanes-ops <zanes@nvidia.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Jian Chen <jianchen0311@gmail.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: andyluo7 <43718156+andyluo7@users.noreply.github.com>
Co-authored-by: Ehsan Akhgari <ehsan.akhgari@gmail.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com>
|
2026-09-06 02:27:59 -07:00 |
|
Liangsheng Yin
|
f5819b09bf
|
Revert "[AMD][DSV4] Fix unified-KV pool sizing and SWA ring accounting" (#38163)
|
2026-09-05 17:28:46 -07:00 |
|
yuttian1
|
514b45fd34
|
[AMD][DSV4] Fix unified-KV pool sizing and SWA ring accounting (#30315)
|
2026-09-05 16:39:40 -07:00 |
|
 Xiaoyu ZhangandWaterpine
|
ccf9fe6590
|
[Kernel] Add KDA FP8 skinny GEMM for SM120 (#38082)
Co-authored-by: Waterpine <biansonghz@gmail.com>
|
2026-09-05 22:27:06 +08:00 |
|
Xiaoyu Zhang
|
dc2843801d
|
perf(lfm2): fuse gating and short convolution on SM90 (#37622)
|
2026-09-05 21:52:16 +08:00 |
|
 
|
bd16c22a04
|
[diffusion] fuse LingBot MoE group-limited top-k index selection (#38044)
Co-authored-by: BBuf <bbuf@users.noreply.github.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-05 18:12:30 +08:00 |
|
 DayuxiaoshuiandXiaoyu Zhang
|
50c1bf0db0
|
[Diffusion] Port the Wan VAE decoder fast paths to the Qwen-Image VAE (#38020)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-09-05 18:02:33 +08:00 |
|
 Xiaoyu ZhangandBBuf
|
d49180019b
|
fix(moe): cast filtered-activation expert_ids to int32 for torch.compile (#38085)
Co-authored-by: BBuf <bbuf@users.noreply.github.com>
|
2026-09-05 17:13:07 +08:00 |
|
 
|
55bf3380e0
|
Support Hy4-preview (#36805)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: alphabetc1 <2508695655@qq.com>
|
2026-09-04 18:03:49 -07:00 |
|
Xiaoyu Zhang
|
54c2c99feb
|
[Diffusion] Fuse LingBot per-token gated residual and RMSNorm modulate (#37910)
|
2026-09-04 14:35:13 +08:00 |
|
 
|
2bb25dc18b
|
[Speculative Decoding] Add native UNO serving support (#37667)
Co-authored-by: drproduck <drproduck@MacBook-Air-2.local>
Co-authored-by: BBuf <1182563586@qq.com>
|
2026-09-03 20:08:41 +08:00 |
|
 Xinyi SongandThomas Wang
|
7ed29eba80
|
[AMD] Fix FP4 indexer OOR (#37660)
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
|
2026-09-03 01:49:46 -07:00 |
|