 Xiaoyu ZhangandMick Qian
|
4d23a4fa6d
|
[Test] Consolidate test cleanup and CI taxonomy (net -11.4K lines) (#37436)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-07 15:13:59 +08:00 |
|
 sglang-botandsglang-bot
|
6252993afe
|
chore: update CI test est_time values (#38238)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-09-06 17:49:41 -07:00 |
|
 
|
bd16c22a04
|
[diffusion] fuse LingBot MoE group-limited top-k index selection (#38044)
Co-authored-by: BBuf <bbuf@users.noreply.github.com>
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
|
2026-09-05 18:12:30 +08:00 |
|
 DayuxiaoshuiandXiaoyu Zhang
|
50c1bf0db0
|
[Diffusion] Port the Wan VAE decoder fast paths to the Qwen-Image VAE (#38020)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-09-05 18:02:33 +08:00 |
|
Xiaoyu Zhang
|
54c2c99feb
|
[Diffusion] Fuse LingBot per-token gated residual and RMSNorm modulate (#37910)
|
2026-09-04 14:35:13 +08:00 |
|
 YC Yen-Ching TsengandPhil Li
|
1fb85053e7
|
[AMD][Diffusion] Migrate FlyDSL fused norm kernels to the v0.3.0 stable API (#36349)
Co-authored-by: Phil Li <haicli@amd.com>
|
2026-09-02 23:03:08 -07:00 |
|
 Xiaoyu ZhangandCursor
|
f4c17fed07
|
[Diffusion] Fuse FLUX.2 NVFP4 FC1, SwiGLU, and FC2 quantization (#37096)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-09-02 08:21:14 +08:00 |
|
 
|
c593527f33
|
[Kernel] Add KDA NVFP4 GEMM for Qwen3.x on SM120 (#36865)
Co-authored-by: Song Bian <biansonghz@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-09-02 08:16:54 +08:00 |
|
 Xiaoyu ZhangandCursor
|
1c3ad92438
|
[Diffusion] Fuse FLUX.2 ModelOpt FP8 producers and QKV packing (#37162)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-09-01 16:14:29 +08:00 |
|
 
|
71cee04ebe
|
[Diffusion] Optimize Qwen-Image TP collectives and attention (#36680)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-09-01 10:28:37 +08:00 |
|
 Xiaoyu ZhangandCursor
|
6715debb2a
|
[Diffusion] Fuse Qwen-Image FP8 QKV projection and Blackwell epilogue (#37123)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-09-01 08:34:41 +08:00 |
|
Xiaoyu Zhang
|
52e1c24744
|
[Diffusion] Fuse FLUX.2 token concatenation and NVFP4 quantization (#37141)
|
2026-08-31 21:33:02 +08:00 |
|
Xiaoyu Zhang
|
771e613d96
|
[Diffusion] Fuse Qwen-Image final adaptive LayerNorm (#37144)
|
2026-08-31 18:22:20 +08:00 |
|
 Xiaoyu ZhangandCursor
|
1ed9bfac2c
|
[Diffusion] Fuse Qwen-Image FP8 norm and activation quantization (#37156)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-08-31 18:19:42 +08:00 |
|
Xiaoyu Zhang
|
bb5e619860
|
[diffusion] perf: absorb Qwen-Image output projection biases (#37116)
|
2026-08-31 08:25:37 +08:00 |
|
Xiaoyu Zhang
|
5ab97c4f44
|
[Diffusion] Cache Qwen-Image modulation across serial CFG branches (#37090)
|
2026-08-31 01:40:29 +08:00 |
|
Xiaoyu Zhang
|
96a4dcdde8
|
[CI] Slim JIT kernel unit tests (#36887)
|
2026-08-29 07:26:49 +08:00 |
|
Xiaoyu Zhang
|
eebb99c049
|
[diffusion][kernel] avoid 4D scale-shift autotuning (#36521)
|
2026-08-28 16:58:19 +08:00 |
|
Xiaoyu Zhang
|
45424d8434
|
[diffusion][kernel] support transposed residual-gate add (#36504)
|
2026-08-28 16:54:37 +08:00 |
|
Xiaoyu Zhang
|
96b31770f9
|
[diffusion] fuse Helios paired transposed RoPE (#36502)
|
2026-08-28 08:57:37 +08:00 |
|
 Alex NailsandClaude Opus 5
|
5ffdb02d0c
|
[CI] Cut repeated tokenizer loads, serial subprocesses and a double scan (#36241)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-24 18:22:43 -07:00 |
|
 Hert4andAlex Nails
|
b4bd5f91ee
|
[diffusion] Fix test_model_fast_paths import after sana_ln_modulate rename (#36175)
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-08-24 04:48:17 -07:00 |
|
Xiaoyu Zhang
|
46b92b22e2
|
[diffusion] Accelerate LingBot Video RMSNorm in quality=high (#35969)
|
2026-08-24 18:02:00 +08:00 |
|
Xiaoyu Zhang
|
6d40b8aebf
|
[diffusion] Fix Hunyuan QKV pack indexing at production video shapes (#36009)
|
2026-08-24 13:39:26 +08:00 |
|
Xiaoyu Zhang
|
8dcfb3b5e7
|
[diffusion] Fuse LongCat-Image QKNorm and interleaved RoPE (#35995)
|
2026-08-24 12:07:26 +08:00 |
|
Xiaoyu Zhang
|
e129fe21e5
|
[diffusion] Flatten Wan VAE RMSNorm row addressing (#35981)
|
2026-08-24 08:57:54 +08:00 |
|
Xiaoyu Zhang
|
83e9ece672
|
[diffusion] Fuse SANA-Video interleaved RoPE (#35695)
|
2026-08-22 12:56:45 +08:00 |
|
 R0CKSTARandAlex Nails
|
d90318b3e2
|
[MLX] Upgrade to Torch 2.13/MLX 0.32+ and redesign the Torch-MLX tensor bridge (#32984)
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-08-21 18:51:42 -07:00 |
|
Xiaoyu Zhang
|
39d4d65a51
|
[diffusion] Accelerate SANA-Video linear attention in quality=high (#35728)
|
2026-08-21 18:05:43 +08:00 |
|
Xiaoyu Zhang
|
7e80e889a2
|
[diffusion] Fuse LTX-2.5 decoder 3D RoPE (#35698)
|
2026-08-21 10:13:09 +08:00 |
|
 MichaelandCursor Agent
|
e805a8f98e
|
[AMD] Keep the PTX-inline-asm diffusion norm fusions off on ROCm (fix FLUX warmup crash) (#34481)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
|
2026-08-20 08:55:28 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 5
|
ae6945e112
|
[kernels] Reorganize ops/diffusion by operator domain behind a lazy facade (#35114)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-18 20:37:43 +08:00 |
|
Xiaoyu Zhang
|
0aa09ab40d
|
[diffusion] Reuse bit-exact modulation fast path for LTX-2.3 (#34930)
|
2026-08-17 09:04:10 +08:00 |
|
Xiaoyu Zhang
|
41abbb0d32
|
[diffusion] Accelerate Cosmos3 T2I QKNorm+RoPE (#34932)
|
2026-08-16 20:15:32 +08:00 |
|
Xiaoyu Zhang
|
095ec6c997
|
[diffusion][kernel] Accelerate Sana BCG with bit-exact conv post-processing (#34928)
|
2026-08-16 20:05:57 +08:00 |
|
Xiaoyu Zhang
|
ebca0bbde4
|
[Diffusion][ERNIE] Fuse QKNorm with full-width RoPE (#34620)
|
2026-08-13 23:23:21 +08:00 |
|
Xiaoyu Zhang
|
74c0322342
|
[Diffusion][FLUX.2] Fuse eager AdaLN and packed SwiGLU (#34616)
|
2026-08-13 19:55:53 +08:00 |
|
Xiaoyu Zhang
|
3c1791a7df
|
[Diffusion][HunyuanVideo] Fuse eager QKV packing and high-quality QKNorm (#34617)
|
2026-08-13 19:54:00 +08:00 |
|
Mick
|
ad47dde65c
|
[diffusion] optimize: fuse cosmos qk norm, rope, and kv packing (#34275)
|
2026-08-12 23:18:16 +08:00 |
|
Xiaoyu Zhang
|
1f008dc226
|
[Diffusion][LTX-2] Allocate AdaLN outputs from one contiguous slab (#34508)
|
2026-08-12 16:28:27 +08:00 |
|
Xiaoyu Zhang
|
4aff4b1822
|
[Diffusion] Improve bit-exact fusion fallback diagnostics (#34412)
|
2026-08-12 10:42:48 +08:00 |
|
 
|
fd3036523a
|
[diffusion] Clean up shared bitexact gates, helpers, and stale naming (#34180)
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-10 22:14:09 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
441910f926
|
[diffusion] LTX-2 quality=high fused RMSNorm+modulate + FFN GELU epilogue (H200 ltx23-one-stage denoise 45.85->43.24 s, ~matches torch.compile) (#34172)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-08-10 09:46:17 +08:00 |
|
Liangsheng Yin
|
7c90840bad
|
[CI] Key scheduled CUDA suites by runner_config instead of hand-written jobs (#34186)
|
2026-08-09 16:44:53 -07:00 |
|
Xiaoyu Zhang
|
38c007dfe5
|
[diffusion] FLUX.1: route the adaLN LN+modulate sites through the bit-exact fused LayerNorm+modulate kernel (H200 1024^2 lossless denoise -1.2%, e2e wall -2.9%) (#34126)
|
2026-08-09 09:52:37 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
6424fec326
|
[diffusion] Bit-exact data-movement elimination for the Wan causal VAE decoder (H200 LongLive2 704x1280x61f: decode 2.80->2.32 s lossless / 2.12->1.67 s quality=high, e2e -10.7%) (#34125)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-09 09:50:56 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
33ed5d4413
|
[diffusion] perf_logger: SYNC_STAGE_PROFILING must drain the GPU queue for stage records too (fixes 2-3x inflated DecodingStage readings) (#34124)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-09 09:49:52 +08:00 |
|
Xiaoyu Zhang
|
dc9624deb2
|
[diffusion] Clean up kernels and shared fast paths (#34085)
|
2026-08-09 00:37:00 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
891445676c
|
[diffusion] Sana: bit-exact fused aten LayerNorm+modulate under BCG (H200 denoise -4.8%) (#34015)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-08 16:05:25 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
5dffa06fe1
|
[diffusion] GLM-Image bit-exact fused aten LayerNorm+modulate / qk-LN (H200 30-step denoise -8.1%) (#34008)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-08 13:26:11 +08:00 |
|