Commit Graph
203 Commits
Author SHA1 Message Date
Xiaoyu Zhang cb22f2451e [Cleanup] Deduplicate kernel tests, diffusion fixtures and benchmark helpers (#40265) 2026-09-19 19:45:38 +08:00
Xiaoyu Zhang 0b0d2c257a [Fix] Repair CI fixtures and ROCm speculative tree device checks (#40325) 2026-09-19 18:16:05 +08:00
Cheng Wan afe71f4b9e Read process groups through the runtime context (#40068) 2026-09-18 17:40:32 -07:00
5931fd60ee Support unified memory page-envelope transfers in PD (#39477)
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
2026-09-18 17:39:50 -07:00
Lifan Shen 8ea0ee300d perf(sampling): avoid GPU syncs when applying custom logit processors (#39234) 2026-09-18 17:09:30 -07:00
Jialin Ouyang f3851486cb [Perf] Fuse SWA page lookup and mapping clear (#38948) 2026-09-18 15:51:07 -07:00
0e5347db82 Support MXFP8 and deferred route weighting in DeepEP v2 (#40030)
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>
Co-authored-by: pranjalssh <14260275+pranjalssh@users.noreply.github.com>
2026-09-18 15:40:38 -07:00
81363bf8cb [kernel] Share the warp vectorized copy and enforce its alignment (#36176)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: BBuf <1182563586@qq.com>
2026-09-18 22:40:54 +08:00
a6cf05817f dsv4.1: remaining model and runtime integration (#38798)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <xiaoyu.zhang@radixark.ai>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Zhichen Zeng <zczeng@uw.edu>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
2026-09-18 02:55:30 -07:00
Nan Jiang 740f57a02c [Spec] Fix CDF boundary handling in TreeSpeculativeSamplingTargetOnly (#35798) 2026-09-17 19:41:52 -07:00
Liangsheng Yin f65c70bb7d [Kernel] Move CUDA and ROCm speculative kernels to JIT (#40033) 2026-09-17 17:27:26 -07:00
Xiaoyu Zhang 7bc9152447 [Test] Consolidate kernel tests under plural kernels tree (#39966) 2026-09-18 07:37:48 +08:00
Liangsheng Yin 1f0c73e9bd [DSV4] Generalize attention metadata, sparse prefill, and KV pool over compress ratios (#39921) 2026-09-17 15:55:19 -07:00
13d593b6cf dsv4.1: compression, KV I/O, and metadata kernels (#39652)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
2026-09-16 13:54:07 -07:00
+4 faaff1eca8 dsv4.1: Top-k kernels and candidate selection (#39648)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: DarkSharpness <2040703891@qq.com>
Co-authored-by: DarkSharpness <76582120+DarkSharpness@users.noreply.github.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Xiaoyu Zhang <xiaoyu.zhang@radixark.ai>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Zhichen Zeng <zczeng@uw.edu>
Co-authored-by: Ziyi Xu <ziyi.xu@radixark.ai>
2026-09-15 23:54:35 -07:00
ddd4600197 [Fix] HiCache startup ImportError on the pinned kernel wheel (#39516)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-09-14 22:31:19 -07:00
JoeandXiaoyu Zhang a23fd557ed [Kernel] Add OOT dispatch for clamp position (#38687)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-09-15 10:22:21 +08:00
AMD-yanfeiwang 0163f8ff74 [AMD] Fix registered HiCache host pointer aliases (#35233) 2026-09-14 15:49:05 -07:00
Xiaoyu ZhangandMick Qian 3e035a3513 [Diffusion] Optimize Qwen-Image-Edit attention on Hopper (#38584)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-13 09:15:01 +08:00
335f6aab27 [DSV4] Support raw-index output in TopK v2 (#33672)
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
2026-09-11 07:17:36 -07:00
Jialin Ouyang 6182d45524 [Test] Fix Q8KV8 sparse-prefill pool fixture after page-size rename (#39019) 2026-09-10 22:22:35 -07:00
Pengyun Lin d076eec427 [SM120] Use exact query-head widths for DeepSeek-V4 sparse MLA decode (#36655) 2026-09-10 15:40:17 -07:00
paulzhang-tm a63efd9056 [Spec] Support large MTP batches in short-convolution metadata (#38558) 2026-09-10 15:13:17 -07:00
Brayden ZhongandBrayden Zhong c0b790cf7f Delete cutlass_mla, non-Marlin GPTQ, AWQ AOT kernel, and Dual Chunk Flash Attention (#32114)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-09-10 15:12:01 +08:00
Mohammad Miadh AngkadandMohammad Angkad 72d5c5bb73 [Kimi-K3] Accept fp32 routing weights in the fused MoE finalize (#38612)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
2026-09-09 01:33:39 -07:00
DayuxiaoshuiandXiaoyu Zhang db1de6ff4c [Diffusion] Keep the Wan VAE decoder channels_last and add a Triton NHWC nearest upsample (#38182)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-09-09 09:31:35 +08:00
Baizhou Zhang ed183d45ac [CP V1 Deprecation 4/5] Canonicalize prefill CP API names (#36229) 2026-09-08 16:03:36 -07:00
Brayden ZhongandBrayden Zhong 30e7a3072d Keep fp32 routing weights in the fp8 block-scale and bf16 trtllm MoE (#33631)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-09-08 08:56:40 -07:00
Xiaoyu Zhang 554f817948 [Diffusion] Optimize LTX-2 QKNorm and split RoPE on Hopper (#38396) 2026-09-08 19:05:02 +08:00
amd-danli103 141febf329 [AMD] fix: use the hardware fp8 e4m3 convert on gfx950 (#37140)
Signed-off-by: amd-danli103 <danli103@amd.com>
2026-09-08 03:01:53 -07:00
HZY 8656901504 [Fix][DSA] Bound prefill Triton specializations for page-table stride (#37093) 2026-09-08 10:23:51 +08:00
Liangsheng Yin 4dcecc7891 Revert "[kernel] add fused silu mul quant fp8" (#38381) 2026-09-07 17:36:10 -07:00
570087ceda [AMD][DSV4] Reland unified-KV pool sizing and SWA ring accounting, fully gated (#38192)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-09-07 13:13:04 -07:00
Mohammad Miadh Angkad c5367fa964 [CI] Fix stale ServerArgs fake in chunked-SGMV LoRA test (#38315) 2026-09-07 04:25:12 -07:00
Cheng Wan b5766336d4 [Perf] Unified memory: close the DCP decode gap on Blackwell (#37926) 2026-09-07 01:10:44 -07:00
Xiaoyu ZhangandMick Qian 4d23a4fa6d [Test] Consolidate test cleanup and CI taxonomy (net -11.4K lines) (#37436)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-07 15:13:59 +08:00
a8b2f36dee [kernel] add fused silu mul quant fp8 (#37376)
Co-authored-by: undefined <zhouchen.arrebol@jd.com>
Co-authored-by: xq25478 <xq25478@qq.com>
Co-authored-by: xieminghe.simon <xieminghe.simon@jd.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-09-07 11:39:11 +08:00
sglang-botandsglang-bot 6252993afe chore: update CI test est_time values (#38238)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-09-06 17:49:41 -07:00
+2 97c6978369 GLM-5.3-Flash support (#36507)
Co-authored-by: zRzRzRzRzRzRzR <Yuxuan.Zhang2@liverpool.ac.uk>
Co-authored-by: Shijin Zhang <75300765+Dovis01@users.noreply.github.com>
Co-authored-by: zanes-ops <zanes@nvidia.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Jian Chen <jianchen0311@gmail.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: andyluo7 <43718156+andyluo7@users.noreply.github.com>
Co-authored-by: Ehsan Akhgari <ehsan.akhgari@gmail.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Raiden Makoto <81530826+Raiden-Makoto@users.noreply.github.com>
2026-09-06 02:27:59 -07:00
Liangsheng Yin f5819b09bf Revert "[AMD][DSV4] Fix unified-KV pool sizing and SWA ring accounting" (#38163) 2026-09-05 17:28:46 -07:00
yuttian1 514b45fd34 [AMD][DSV4] Fix unified-KV pool sizing and SWA ring accounting (#30315) 2026-09-05 16:39:40 -07:00
Xiaoyu ZhangandWaterpine ccf9fe6590 [Kernel] Add KDA FP8 skinny GEMM for SM120 (#38082)
Co-authored-by: Waterpine <biansonghz@gmail.com>
2026-09-05 22:27:06 +08:00
Xiaoyu Zhang dc2843801d perf(lfm2): fuse gating and short convolution on SM90 (#37622) 2026-09-05 21:52:16 +08:00
bd16c22a04 [diffusion] fuse LingBot MoE group-limited top-k index selection (#38044)
Co-authored-by: BBuf <bbuf@users.noreply.github.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-05 18:12:30 +08:00
DayuxiaoshuiandXiaoyu Zhang 50c1bf0db0 [Diffusion] Port the Wan VAE decoder fast paths to the Qwen-Image VAE (#38020)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-09-05 18:02:33 +08:00
Xiaoyu ZhangandBBuf d49180019b fix(moe): cast filtered-activation expert_ids to int32 for torch.compile (#38085)
Co-authored-by: BBuf <bbuf@users.noreply.github.com>
2026-09-05 17:13:07 +08:00
55bf3380e0 Support Hy4-preview (#36805)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: alphabetc1 <2508695655@qq.com>
2026-09-04 18:03:49 -07:00
Xiaoyu Zhang 54c2c99feb [Diffusion] Fuse LingBot per-token gated residual and RMSNorm modulate (#37910) 2026-09-04 14:35:13 +08:00
2bb25dc18b [Speculative Decoding] Add native UNO serving support (#37667)
Co-authored-by: drproduck <drproduck@MacBook-Air-2.local>
Co-authored-by: BBuf <1182563586@qq.com>
2026-09-03 20:08:41 +08:00
Xinyi SongandThomas Wang 7ed29eba80 [AMD] Fix FP4 indexer OOR (#37660)
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
2026-09-03 01:49:46 -07:00