 Xiaoyu ZhangandCursor
|
175973d834
|
[Diffusion] Fuse Qwen-Image residual norm and NVFP4 quantization (#37129)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-09-01 08:54:10 +08:00 |
|
 Xiaoyu ZhangandCursor
|
079afaffb1
|
[Diffusion] Fuse FLUX.2 gated residual normalization on Blackwell (#37112)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-09-01 08:38:33 +08:00 |
|
 Xiaoyu ZhangandCursor
|
6715debb2a
|
[Diffusion] Fuse Qwen-Image FP8 QKV projection and Blackwell epilogue (#37123)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-09-01 08:34:41 +08:00 |
|
Xiaoyu Zhang
|
52e1c24744
|
[Diffusion] Fuse FLUX.2 token concatenation and NVFP4 quantization (#37141)
|
2026-08-31 21:33:02 +08:00 |
|
 Xiaoyu ZhangandSong Bian
|
d60d658f5f
|
[Kernel] Add GB300 Triton MoE configs for GLM-4.5 FP8 (#37159)
Co-authored-by: Song Bian <biansonghz@gmail.com>
|
2026-08-31 21:31:04 +08:00 |
|
Xiaoyu Zhang
|
771e613d96
|
[Diffusion] Fuse Qwen-Image final adaptive LayerNorm (#37144)
|
2026-08-31 18:22:20 +08:00 |
|
 Xiaoyu ZhangandCursor
|
1ed9bfac2c
|
[Diffusion] Fuse Qwen-Image FP8 norm and activation quantization (#37156)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-08-31 18:19:42 +08:00 |
|
Xiaoyu Zhang
|
bb5e619860
|
[diffusion] perf: absorb Qwen-Image output projection biases (#37116)
|
2026-08-31 08:25:37 +08:00 |
|
Xiaoyu Zhang
|
5ab97c4f44
|
[Diffusion] Cache Qwen-Image modulation across serial CFG branches (#37090)
|
2026-08-31 01:40:29 +08:00 |
|
Xiaoyu Zhang
|
8c28cdd116
|
[Diffusion][Kernel] Fuse Wan2.2 NVFP4 bias + GELU on Blackwell (#37075)
|
2026-08-31 01:37:35 +08:00 |
|
Xiaoyu Zhang
|
0e1146d04f
|
[diffusion] optimization: optimize Pi0.5 inference and bounded graph serving (#34599)
|
2026-08-29 14:47:32 +08:00 |
|
Xiaoyu Zhang
|
db6f0a9d53
|
Refactor JIT kernel and expert-pack directory layout (#36704)
|
2026-08-29 07:41:25 +08:00 |
|
Xiaoyu Zhang
|
50bc1a3767
|
[diffusion] Keep Cosmos3 Nano resident on 96 GB GPUs (#36641)
|
2026-08-29 07:40:46 +08:00 |
|
Xiaoyu Zhang
|
96a4dcdde8
|
[CI] Slim JIT kernel unit tests (#36887)
|
2026-08-29 07:26:49 +08:00 |
|
Xiaoyu Zhang
|
eebb99c049
|
[diffusion][kernel] avoid 4D scale-shift autotuning (#36521)
|
2026-08-28 16:58:19 +08:00 |
|
Xiaoyu Zhang
|
45424d8434
|
[diffusion][kernel] support transposed residual-gate add (#36504)
|
2026-08-28 16:54:37 +08:00 |
|
Xiaoyu Zhang
|
96b31770f9
|
[diffusion] fuse Helios paired transposed RoPE (#36502)
|
2026-08-28 08:57:37 +08:00 |
|
Xiaoyu Zhang
|
024a7a1031
|
[diffusion] Fix native LingBot-Video text encoding (#36542)
|
2026-08-27 22:58:22 +08:00 |
|
Xiaoyu Zhang
|
a7e3f590ca
|
[Diffusion] Fuse LongCat residual gate updates (#36577)
|
2026-08-27 21:07:04 +08:00 |
|
Xiaoyu Zhang
|
e061dd1b47
|
[Diffusion][Kernel] Fuse Wan FFN GELU epilogue (#36592)
|
2026-08-27 21:06:28 +08:00 |
|
Xiaoyu Zhang
|
1af95ffded
|
[diffusion] Fuse Cosmos3 Nano T2I attention on Hopper (#36571)
|
2026-08-27 20:59:15 +08:00 |
|
Xiaoyu Zhang
|
db4125bb56
|
[diffusion] Accept mesh benchmark artifacts (#36553)
|
2026-08-27 20:53:25 +08:00 |
|
Xiaoyu Zhang
|
d42fa5e10a
|
[diffusion] align video BCG warmup frame count (#36485)
|
2026-08-27 20:52:00 +08:00 |
|
Xiaoyu Zhang
|
ad911a5ec0
|
[kernel] Tune LingBot-Video MoE TMA configs for H100 (#36543)
|
2026-08-27 02:47:32 -07:00 |
|
Xiaoyu Zhang
|
702de26310
|
[diffusion] make benchmark caches seedable and cover missing native families (#36463)
|
2026-08-26 19:04:46 +08:00 |
|
Xiaoyu Zhang
|
fa3ac61661
|
[Diffusion] Bound reusable Ulysses A2A staging buffers across shapes (#36327)
|
2026-08-26 10:57:29 +08:00 |
|
Xiaoyu Zhang
|
46b92b22e2
|
[diffusion] Accelerate LingBot Video RMSNorm in quality=high (#35969)
|
2026-08-24 18:02:00 +08:00 |
|
Xiaoyu Zhang
|
9866fe910b
|
[diffusion] Speed up LingBot high-quality VAE decode (#36024)
|
2026-08-24 14:13:03 +08:00 |
|
Xiaoyu Zhang
|
cc74aba330
|
[diffusion] Honor XDG cache for model overlays (#36019)
|
2026-08-24 14:11:26 +08:00 |
|
Xiaoyu Zhang
|
6d40b8aebf
|
[diffusion] Fix Hunyuan QKV pack indexing at production video shapes (#36009)
|
2026-08-24 13:39:26 +08:00 |
|
Xiaoyu Zhang
|
b43931e878
|
[diffusion] Refresh quality and BCG benchmark skills (#36016)
Signed-off-by: BBuf <1182563586@qq.com>
|
2026-08-24 13:37:45 +08:00 |
|
Xiaoyu Zhang
|
344613c159
|
[diffusion] Default Hunyuan VAE to tiled decode (#36012)
|
2026-08-24 13:18:06 +08:00 |
|
Xiaoyu Zhang
|
8dcfb3b5e7
|
[diffusion] Fuse LongCat-Image QKNorm and interleaved RoPE (#35995)
|
2026-08-24 12:07:26 +08:00 |
|
Xiaoyu Zhang
|
09592f5889
|
[diffusion] Keep LongLive2 components resident on large GPUs (#35993)
|
2026-08-24 12:06:52 +08:00 |
|
Xiaoyu Zhang
|
e129fe21e5
|
[diffusion] Flatten Wan VAE RMSNorm row addressing (#35981)
|
2026-08-24 08:57:54 +08:00 |
|
Xiaoyu Zhang
|
b2eb0fa51e
|
[diffusion] Keep Cosmos3 Nano resident on high-memory GPUs (#36000)
|
2026-08-24 08:51:58 +08:00 |
|
Xiaoyu Zhang
|
447048dba2
|
[diffusion] Reject unsafe quality=high BCG replay (#36008)
|
2026-08-24 08:50:26 +08:00 |
|
Xiaoyu Zhang
|
f4448e677f
|
[diffusion] Reuse SANA fast paths in SANA-Video BCG (#35961)
|
2026-08-24 08:47:04 +08:00 |
|
Xiaoyu Zhang
|
96bfd2476c
|
[diffusion] Enable SANA-Video breakable CUDA graphs (#35729)
|
2026-08-22 12:57:09 +08:00 |
|
Xiaoyu Zhang
|
83e9ece672
|
[diffusion] Fuse SANA-Video interleaved RoPE (#35695)
|
2026-08-22 12:56:45 +08:00 |
|
Xiaoyu Zhang
|
39d4d65a51
|
[diffusion] Accelerate SANA-Video linear attention in quality=high (#35728)
|
2026-08-21 18:05:43 +08:00 |
|
Xiaoyu Zhang
|
a5c52a9358
|
[diffusion] Enable LongCat breakable CUDA graphs (#35724)
|
2026-08-21 17:59:16 +08:00 |
|
Xiaoyu Zhang
|
7e80e889a2
|
[diffusion] Fuse LTX-2.5 decoder 3D RoPE (#35698)
|
2026-08-21 10:13:09 +08:00 |
|
Xiaoyu Zhang
|
04444ee352
|
[diffusion] Refresh eager optimization skills and benchmark safeguards (#35679)
|
2026-08-20 22:03:57 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 5
|
9113fc6d93
|
[docs] Add a fused-kernels page for SGLang Diffusion (#35436)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-19 16:32:42 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 5
|
ae6945e112
|
[kernels] Reorganize ops/diffusion by operator domain behind a lazy facade (#35114)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-08-18 20:37:43 +08:00 |
|
Xiaoyu Zhang
|
0aa09ab40d
|
[diffusion] Reuse bit-exact modulation fast path for LTX-2.3 (#34930)
|
2026-08-17 09:04:10 +08:00 |
|
Xiaoyu Zhang
|
41abbb0d32
|
[diffusion] Accelerate Cosmos3 T2I QKNorm+RoPE (#34932)
|
2026-08-16 20:15:32 +08:00 |
|
Xiaoyu Zhang
|
095ec6c997
|
[diffusion][kernel] Accelerate Sana BCG with bit-exact conv post-processing (#34928)
|
2026-08-16 20:05:57 +08:00 |
|
Xiaoyu Zhang
|
0761d3f3a4
|
[diffusion] Accelerate lossless Ideogram norm post-processing (#34931)
|
2026-08-16 17:20:09 +08:00 |
|
Xiaoyu Zhang
|
b752f1e533
|
[diffusion] Enable breakable CUDA graphs for LTX-2.3 (#34929)
|
2026-08-16 17:18:43 +08:00 |
|
Xiaoyu Zhang
|
0c072235f4
|
[diffusion] Bound overlong weight lock filenames (#34825)
|
2026-08-15 17:20:30 +08:00 |
|
Xiaoyu Zhang
|
9c9a3273be
|
[diffusion] Fix Helios denoising profiler stepping (#34826)
|
2026-08-14 23:21:54 +08:00 |
|
Xiaoyu Zhang
|
5f2a6d6422
|
[diffusion] Fix symbolic replicated-mode counting under torch.compile (#34824)
|
2026-08-14 23:20:29 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
704e512836
|
[GDN] Honor configured linear-attn verify backend in the kernel dispatcher (#34592)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-14 08:59:52 +08:00 |
|
Xiaoyu Zhang
|
c255fbc4fe
|
[Diffusion] Add @triple-mu as a code owner (#34748)
|
2026-08-14 00:37:49 +08:00 |
|
Xiaoyu Zhang
|
ebca0bbde4
|
[Diffusion][ERNIE] Fuse QKNorm with full-width RoPE (#34620)
|
2026-08-13 23:23:21 +08:00 |
|
Xiaoyu Zhang
|
82f7afb881
|
[Diffusion] Make auto residency decisions component-scoped (#34615)
|
2026-08-13 23:20:55 +08:00 |
|
Xiaoyu Zhang
|
74c0322342
|
[Diffusion][FLUX.2] Fuse eager AdaLN and packed SwiGLU (#34616)
|
2026-08-13 19:55:53 +08:00 |
|
Xiaoyu Zhang
|
3c1791a7df
|
[Diffusion][HunyuanVideo] Fuse eager QKV packing and high-quality QKNorm (#34617)
|
2026-08-13 19:54:00 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
a23670ddbf
|
[diffusion] Wan2.2-TI2V: fuse per-token adaLN table add into contiguous slices + hoist rope cache (denoise -13.1% H100 / -12.6% H200, bit-exact; eager beats compile) (#34584)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-13 17:26:46 +08:00 |
|
Xiaoyu Zhang
|
26579d893e
|
[Diffusion][GLM-Image] Retune QK head LayerNorm for SM103 (#34619)
|
2026-08-13 10:49:13 +08:00 |
|
Xiaoyu Zhang
|
1f008dc226
|
[Diffusion][LTX-2] Allocate AdaLN outputs from one contiguous slab (#34508)
|
2026-08-12 16:28:27 +08:00 |
|
Xiaoyu Zhang
|
daae3acb36
|
[Diffusion][Z-Image] Tune native QK RMSNorm launch for SM103 (#34507)
|
2026-08-12 16:27:44 +08:00 |
|
Xiaoyu Zhang
|
4827061247
|
[Diffusion] Make weight-only FP8 dequant cache torch.compile-safe (#34506)
|
2026-08-12 16:26:31 +08:00 |
|
Xiaoyu Zhang
|
84ce7502cf
|
[Diffusion][MiniMax H3] Extend exact QKNorm+RoPE rounding to SM103 (#34505)
|
2026-08-12 16:25:24 +08:00 |
|
Xiaoyu Zhang
|
45f7063335
|
[Diffusion] Tune QK head LayerNorm for SM103 (#34503)
|
2026-08-12 16:24:18 +08:00 |
|
Xiaoyu Zhang
|
22e4b3a81f
|
[Diffusion] Avoid slow cuBLASLt GELU epilogue on SM120 (#34350)
|
2026-08-12 12:12:00 +08:00 |
|
Xiaoyu Zhang
|
4aff4b1822
|
[Diffusion] Improve bit-exact fusion fallback diagnostics (#34412)
|
2026-08-12 10:42:48 +08:00 |
|
Xiaoyu Zhang
|
3f9d184833
|
[Diffusion] Tune QK head LayerNorm for SM120 (#34349)
|
2026-08-12 10:41:54 +08:00 |
|
Xiaoyu Zhang
|
a53d3636ce
|
[diffusion][model] Add native SANA-Video T2V support (#32921)
|
2026-08-12 10:07:24 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
37c631ef23
|
[diffusion] Ideogram-4: fuse Qwen3-style RoPE and SwiGLU silu-mul (denoise -5.1% H100 / -4.7% H200, bit-exact) (#34314)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-12 09:19:35 +08:00 |
|
Xiaoyu Zhang
|
b1b8ce715b
|
[Diffusion][MiniMax H3] Fix SM120 QKNorm+RoPE rounding (#34347)
|
2026-08-12 09:18:16 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
546965fc72
|
[diffusion] LTX-2: mount the bit-exact fused modulate at the 8 bare adaLN sites (ltx23-one-stage denoise -2.8% H100 / -2.6% H200) (#34315)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-11 18:23:08 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
071f0f1e9d
|
[diffusion] ERNIE-Image: fuse rotate-half RoPE + GELU-mul and hoist rope cos/sin (denoise -16.2% H100 / -12.7% H200, bit-exact) (#34306)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-11 18:18:37 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
ba3dc16401
|
[diffusion] weight-only FP8: dequantize linear weights once at first use (Ideogram-4 denoise -18.8% H200 / -7.8% H100, bit-exact) (#34305)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-11 18:15:26 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
f5f0c3ee7a
|
[diffusion] Z-Image single-GPU BCG: fix the replay crash and make output bit-exact vs eager (#34183) (#34210)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-10 22:18:32 +08:00 |
|
 
|
fd3036523a
|
[diffusion] Clean up shared bitexact gates, helpers, and stale naming (#34180)
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-10 22:14:09 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
441910f926
|
[diffusion] LTX-2 quality=high fused RMSNorm+modulate + FFN GELU epilogue (H200 ltx23-one-stage denoise 45.85->43.24 s, ~matches torch.compile) (#34172)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-08-10 09:46:17 +08:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
56ef810cad
|
[diffusion] BCG: auto-capture the default warmup resolution instead of hard-requiring --warmup-resolutions (H200 SANA denoise 0.73->0.457 s with a single flag) (#34174)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-08-10 08:37:32 +08:00 |
|
Xiaoyu Zhang
|
f6cbdc1dd1
|
docs(diffusion): refresh skills for latest runtime (#34143)
|
2026-08-09 10:30:46 +08:00 |
|
Xiaoyu Zhang
|
38c007dfe5
|
[diffusion] FLUX.1: route the adaLN LN+modulate sites through the bit-exact fused LayerNorm+modulate kernel (H200 1024^2 lossless denoise -1.2%, e2e wall -2.9%) (#34126)
|
2026-08-09 09:52:37 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
6424fec326
|
[diffusion] Bit-exact data-movement elimination for the Wan causal VAE decoder (H200 LongLive2 704x1280x61f: decode 2.80->2.32 s lossless / 2.12->1.67 s quality=high, e2e -10.7%) (#34125)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-09 09:50:56 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
33ed5d4413
|
[diffusion] perf_logger: SYNC_STAGE_PROFILING must drain the GPU queue for stage records too (fixes 2-3x inflated DecodingStage readings) (#34124)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-09 09:49:52 +08:00 |
|
Xiaoyu Zhang
|
dc9624deb2
|
[diffusion] Clean up kernels and shared fast paths (#34085)
|
2026-08-09 00:37:00 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
891445676c
|
[diffusion] Sana: bit-exact fused aten LayerNorm+modulate under BCG (H200 denoise -4.8%) (#34015)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-08 16:05:25 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
5dffa06fe1
|
[diffusion] GLM-Image bit-exact fused aten LayerNorm+modulate / qk-LN (H200 30-step denoise -8.1%) (#34008)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-08 13:26:11 +08:00 |
|
Xiaoyu Zhang
|
148f15b0af
|
[diffusion] FLUX.1 fused adaLN modulate (bit-exact) + RoPE cache hoist, LN-affine folding behind quality=high (H200 e2e -3.5% lossless / -6.9% high) (#34004)
|
2026-08-08 13:07:42 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
6c7498113f
|
[diffusion] Enable breakable CUDA graph for SANA (H200 1024px e2e -26%, bit-exact) (#33989)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-07 23:54:39 +08:00 |
|
Xiaoyu Zhang
|
d4be483efb
|
[diffusion] Enable breakable CUDA graph for LTX-2 (H200 two-stage e2e 10.75 s -> 6.90 s, 1.56x) (#33885)
|
2026-08-07 22:29:14 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
572434e2f6
|
[diffusion] Z-Image bit-exact fused qk-norm (H200 Turbo 1024px e2e -6.4%) (#33886)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-07 17:03:28 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
591cfb0881
|
[diffusion] FLUX.2 bit-exact residual-gate fast path (H200 klein-4B 50-step denoise -1.2%) (#33823)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 22:54:40 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
dd98c9572a
|
[diffusion] Generalize the FLUX.2 VAE decoder fast path to AutoencoderKL (Z-Image / FLUX.1) behind quality=high (#33818)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 22:53:15 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
3654740347
|
[diffusion] ERNIE-Image bit-exact fused RMSNorm+scale/shift (H200 1024^2 e2e 15.63 -> 15.00 s, denoise -3.3%) (#33854)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 19:58:44 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
295784723a
|
[diffusion] Ideogram 4: fuse RMSNorm modulate/gate chains via the Z-Image Triton suite behind quality=high (H200 e2e -2.9%/-3.4%) (#33822)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 19:57:34 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
eff6a11350
|
[diffusion] FLUX.1 bit-exact residual-gate fast path + tanh-GELU epilogue behind quality=high (H200 e2e -1.1% lossless / -4.3% high) (#33819)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 19:56:38 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
b6876fc652
|
[diffusion] ERNIE-Image bit-exact residual-gate fast path (H200 1024^2 e2e 16.17 -> 15.75 s) (#33734)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 13:43:09 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
4c0a8940fa
|
[Kernel] Unify BaseFusedOp and MultiPlatformOp dispatch (#33205)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-06 08:52:09 +08:00 |
|
 Xiaoyu ZhangandMohammad Miadh Angkad
|
ba12a16a62
|
[diffusion] Prefer cuDNN SDPA over FA4 for dense attention on sm_100 (B200) (#33655)
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-08-06 08:49:10 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
3425c93666
|
[diffusion] Wan VAE RMSNorm+SiLU fusion behind quality=high (H200 FastWan2.2 e2e 9.611 -> 9.125 s) (#33546)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-08-05 21:33:35 +08:00 |
|