Xiaoyu Zhang
|
fd79cd8d9c
|
[Skills] Refine jit_kernel and sgl-kernel skills (#20095)
|
2026-03-07 22:54:01 +08:00 |
|
 Xiaoyu ZhangandMick
|
6d22c9f369
|
[Diffusion] Move hf kernels diffusion cuda kernels skills to SGLD (#20001)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-03-06 22:16:06 +08:00 |
|
 Xiaoyu ZhangandMick
|
9795b4cd5b
|
[Diffusion] Open t5 encoder parallel folding for wan2.2 and mova video (#18493)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-03-05 10:18:00 +08:00 |
|
Xiaoyu Zhang
|
4348976f80
|
[Diffusion] Refactor diffusion benchmark/profile skill to reuse diffusion-perf skill and clarify profiling trigger (#19783)
|
2026-03-04 10:54:42 +08:00 |
|
Xiaoyu Zhang
|
115e9a1acd
|
[Diffusion] Delete useless _ulysses_input_split func (#19786)
|
2026-03-04 10:45:11 +08:00 |
|
Xiaoyu Zhang
|
145ae518ac
|
[Diffusion] Revert 18619 (#19510)
|
2026-03-03 08:15:15 +08:00 |
|
Xiaoyu Zhang
|
51ee17ce44
|
[diffusion] move skills dir (#19697)
|
2026-03-03 02:51:29 +08:00 |
|
Xiaoyu Zhang
|
53de53fb53
|
[jit_kernel] Tiny unify jit_kernel tests style (#19694)
|
2026-03-02 21:33:59 +08:00 |
|
Xiaoyu Zhang
|
e42fa009d4
|
[Diffusion] diffusion profile and opt skills (#19540)
|
2026-03-02 15:06:29 +08:00 |
|
Xiaoyu Zhang
|
fe971b620a
|
[Diffusion] Add SGL-D diffusion efficient kernel skills (#19473)
|
2026-02-27 11:42:13 +08:00 |
|
Xiaoyu Zhang
|
054bd71086
|
[sgl-kernel slimming] remove sgl-kernel moe-wna16-marlin (#19379)
|
2026-02-27 09:17:46 +08:00 |
|
Xiaoyu Zhang
|
74c8e7b215
|
refactor(jit_kernel): reduce duplication and separate test code (#19323)
|
2026-02-26 18:30:49 +08:00 |
|
Xiaoyu Zhang
|
914ed34757
|
update jit_kernel codeowners (#19385)
|
2026-02-26 10:36:34 +08:00 |
|
Xiaoyu Zhang
|
ab7071b545
|
[SKILL] Better claude skills for sgl-kernel and jit-kernel (#19302)
|
2026-02-25 15:26:55 +08:00 |
|
Xiaoyu Zhang
|
694924b878
|
Update .gitignore to remove '.claude/' (#19296)
|
2026-02-25 11:51:50 +08:00 |
|
Xiaoyu Zhang
|
9dff933164
|
[Kernel Slimming] Remove sgl-kernel AOT marlin kernels (#19241)
|
2026-02-25 10:08:22 +08:00 |
|
Xiaoyu Zhang
|
e0e0cad6bc
|
[Diffusion] Match rotary_embedding module name style (#19179)
|
2026-02-23 21:02:34 +08:00 |
|
Xiaoyu Zhang
|
2717393681
|
[Refactor] Split rotary_embedding.py into a modular package (#19144)
|
2026-02-23 20:05:29 +08:00 |
|
Xiaoyu Zhang
|
66497ab0aa
|
[Diffusion] Restruct and clean Diffusion rotary embedding (#19064)
|
2026-02-21 21:41:47 +08:00 |
|
Xiaoyu Zhang
|
19aa19b111
|
[diffusion] refactor: refactor diffusion triton kernels (#18966)
|
2026-02-19 17:03:44 +08:00 |
|
Xiaoyu Zhang
|
390c154306
|
[Tiny fix] Super tiny fix mul_add naive forward bug (#18964)
|
2026-02-18 16:18:43 +08:00 |
|
Xiaoyu Zhang
|
513c12d23f
|
Remove unused fast-hadamard-transform PyTorch extension sources (#18927)
|
2026-02-18 15:51:07 +08:00 |
|
Xiaoyu Zhang
|
d3bae71e3f
|
Add claude skills for sgl-kernel and jit-kernel (#18855)
|
2026-02-15 15:14:03 -08:00 |
|
Xiaoyu Zhang
|
4067d9487d
|
[diffusion] feat: opt vae decode with channels_last_3d (#18540)
|
2026-02-14 23:19:45 +08:00 |
|
Xiaoyu Zhang
|
c29394e3c8
|
[kernel slimming] Move fast_hadamard_transform to jit_kernel (#18475)
|
2026-02-14 23:06:21 +08:00 |
|
Xiaoyu Zhang
|
013a199bc6
|
[CI] Skip cutedsl gdn performance test in jit_kernel ci (#18783)
|
2026-02-13 15:49:30 +08:00 |
|
Xiaoyu Zhang
|
9e9e949261
|
speed up sgl-kernel build (#18586)
|
2026-02-12 23:43:22 +08:00 |
|
Xiaoyu Zhang
|
bec7fe9e65
|
[sgl-kernel] upgrade deepgemm (#18362)
|
2026-02-10 21:31:30 +08:00 |
|
 Xiaoyu Zhangandyihanc
|
baec650462
|
[Diffusion] Apply fused_norm_scale_shift to LTX2/MOVA (#18257)
Co-authored-by: yihanc <yingluosanqian@gmail.com>
|
2026-02-07 17:28:42 +08:00 |
|
Xiaoyu Zhang
|
dff3ba202a
|
[Diffusion] Support layerwise offload for mova (#18272)
|
2026-02-05 13:16:07 +08:00 |
|
Xiaoyu Zhang
|
2e9d0442e2
|
[diffusion] update code owner (#18247)
|
2026-02-04 19:12:32 +08:00 |
|
Xiaoyu Zhang
|
eedd472025
|
[Diffusion] fix serving image_edit get input image bug (#18109)
|
2026-02-03 12:17:16 +08:00 |
|
Xiaoyu Zhang
|
a1bbc892af
|
[Diffsuion & JIT_kernel] QKNorm cross heads kernel (#18073)
|
2026-02-03 10:03:17 +08:00 |
|
Xiaoyu Zhang
|
a0757c9624
|
[Diffusion] Fix Ring Parallel bug with FA4 (#18062)
|
2026-02-02 17:06:51 +08:00 |
|
Xiaoyu Zhang
|
22aad4e2c4
|
[Diffusion] Fix FLUX.1-schnell time embedding argument mismatch (#17988)
|
2026-01-31 11:47:27 +08:00 |
|
Xiaoyu Zhang
|
abf13ccc11
|
[Diffusion] Fix lora default lora_scale bug (#17982)
|
2026-01-30 22:04:54 +08:00 |
|
Xiaoyu Zhang
|
c08b54a575
|
[JIT kernel] Update jit_kernel cache and develop doc (#17842)
|
2026-01-28 15:09:47 +08:00 |
|
Xiaoyu Zhang
|
fb74e43707
|
[Diffusion] Delete sgl-kernel outdated time_embedding kernel (#17278)
|
2026-01-28 14:18:53 +08:00 |
|
Xiaoyu Zhang
|
67fb492c9a
|
[CI] Fix test_moe_fused_gate error (#17844)
|
2026-01-28 12:03:17 +08:00 |
|
Xiaoyu Zhang
|
331a22427c
|
[Diffusion] glm-image apply flashinfer rope (#17689)
|
2026-01-28 08:51:37 +08:00 |
|
Xiaoyu Zhang
|
3992a023e6
|
Move fa4 from sgl-kernel to jit kernel (#17353)
|
2026-01-24 15:25:03 +08:00 |
|
Xiaoyu Zhang
|
7a4bb0d516
|
[Diffusion] Add diffusion time embedding to jit kernel (#17658)
|
2026-01-24 14:27:08 +08:00 |
|
Xiaoyu Zhang
|
5324027007
|
[Diffusion] Make the apply_qknorm function easier to use (#17537)
|
2026-01-22 22:32:15 +08:00 |
|
Xiaoyu Zhang
|
590969ee9c
|
[Diffusion] Support select fa2 backend in hopper (#17514)
|
2026-01-22 08:23:53 +08:00 |
|
Xiaoyu Zhang
|
19089aa431
|
[Diffusion] Refactor diffusion is_cuda check (#17498)
|
2026-01-21 23:02:24 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Xiaoyu Zhangandgemini-code-assist[bot]
|
cc410a1088
|
[Diffusion] Apply qknorm to flux2 and apply lightx2v rms_norm_one_pass kernel(without residual) (#17305)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-01-19 21:25:33 +08:00 |
|
Xiaoyu Zhang
|
a3d9a21882
|
Revert "[Perf] fuse q, k norm for Flux2Attention (#17241)" (#17332)
|
2026-01-19 15:24:11 +08:00 |
|
Xiaoyu Zhang
|
330605cc88
|
[Diffusion] Apply jit qk_norm to flux1 (#17296)
|
2026-01-19 00:28:36 +08:00 |
|
Xiaoyu Zhang
|
2cdd4370bc
|
[Diffusion] Move diffusion time embedding to jit kernel (#16879)
|
2026-01-17 12:21:22 +08:00 |
|
Xiaoyu Zhang
|
6ee970a365
|
[Diffusion] Hot fix broken output_path default value (#17180)
|
2026-01-16 12:14:09 +08:00 |
|
Xiaoyu Zhang
|
0d904ef44c
|
[diffusion] fix: fix fsdp tp load make param miss parallel meta data (#17058)
|
2026-01-15 00:12:33 +08:00 |
|
Xiaoyu Zhang
|
2ab3ed3e9e
|
Fix sgl-kernel per_token_quant fp8 kernel scale shared_memory bug (#16886)
|
2026-01-13 23:22:05 +08:00 |
|
Xiaoyu Zhang
|
740d3c0b39
|
[Diffusion] Remove useless dependency in diffusion (#16967)
|
2026-01-13 17:25:53 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Xiaoyu Zhangandgemini-code-assist[bot]
|
9d4d57dbfa
|
[Diffusion] Tiny rename parallel_groups (#16743)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-01-09 15:08:25 +08:00 |
|
 Xiaoyu ZhangandMick
|
294ff71d18
|
[Diffusion] Avoid cpu2gpu sync in flashinfer rope and apply flashinfer rope to wanvideo (#16668)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-01-08 22:44:38 +08:00 |
|
Xiaoyu Zhang
|
5a5cece561
|
[Diffusion] clean useless and buggy set_seq_parallel_pg in yunchang (#16669)
|
2026-01-08 11:08:34 +08:00 |
|
 Xiaoyu ZhangandMick
|
32a6540afc
|
[Diffusion] Fix Ulysses/Ring process group construction under TP to enable correct Wan2.2 tensor parallelism (#16532)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-01-07 20:52:36 +08:00 |
|
Xiaoyu Zhang
|
62d0280f62
|
Tiny fix readme (#16654)
|
2026-01-07 20:41:07 +08:00 |
|
Xiaoyu Zhang
|
5e5b1183ed
|
[Diffusion] Ring Attention support sage backend (#16496)
|
2026-01-06 14:45:53 +08:00 |
|
Xiaoyu Zhang
|
4ea6a11c83
|
[CI] Fail wheel build when sgl-kernel artifacts are missing (#16450)
|
2026-01-04 21:46:40 -08:00 |
|
Xiaoyu Zhang
|
520c048d55
|
[diffusion] CI: add script for automatically generation ci perf baseline (#16389)
|
2026-01-05 13:18:35 +08:00 |
|
Xiaoyu Zhang
|
0fee6bc632
|
[JIT kernel] Apply jit per_tensor_quant_fp8 kernel (#15836)
|
2026-01-05 10:15:00 +08:00 |
|
 Xiaoyu ZhangandMick
|
d0fb24ee7b
|
[Diffusion] Flux2 tp support (#16219)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-01-03 17:02:27 +08:00 |
|
Xiaoyu Zhang
|
1cfd2b2ded
|
[diffusion] chore: remove redundant ulysses nccl warmup (#16301)
|
2026-01-03 00:19:35 +08:00 |
|
Xiaoyu Zhang
|
bd48ad5e6b
|
[Diffusion] Fix broken ring_attention when use upstream fa3 (#16270)
|
2026-01-02 14:33:12 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Xiaoyu Zhangandgemini-code-assist[bot]
|
733a0c1a37
|
[Diffusion] Zimage opt with qknorm and flashinfer rope (#16161)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-30 23:39:32 +08:00 |
|
![github-actions[bot]](/assets/img/avatar_default.png) Xiaoyu Zhangandgithub-actions[bot]
|
b369aaa23f
|
[Diffusion] Refine diffusion profling doc (#16163)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2025-12-30 23:38:46 +08:00 |
|
Xiaoyu Zhang
|
88f3de2514
|
[JIT kernel] Jit kernel add codeowners (#16085)
|
2025-12-29 22:32:55 +08:00 |
|
Xiaoyu Zhang
|
24616c5234
|
[Diffusion] Qwen image edit support qknorm optimization (#16062)
|
2025-12-29 21:37:24 +08:00 |
|
Xiaoyu Zhang
|
f3d73b0199
|
[Diffusion] Refactor qwen_image's rope in a single helper func (#16047)
|
2025-12-29 17:24:26 +08:00 |
|
Xiaoyu Zhang
|
8305dc1718
|
[Diffusion] Disable packed QKV for FLUX & Z-Image (#16038)
|
2025-12-29 14:33:05 +08:00 |
|
Xiaoyu Zhang
|
7d02c8e59f
|
[JIT kernel] CI support jit kernel tests (#15939)
|
2025-12-28 23:11:02 +08:00 |
|
 Xiaoyu ZhangandMick
|
51dbdb2202
|
[diffusion] improve: improve qwen-image-edit performance to align with LightX2V (#15812)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-26 22:25:10 +08:00 |
|
 Xiaoyu ZhangandMick
|
e6ce16a4c2
|
[diffusion] feat: support TP for Flux.1.dev (#15666)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-25 18:00:44 +08:00 |
|
Xiaoyu Zhang
|
de2f2880b5
|
[JIT sgl-kernel] Jit support per tensor quant (#15709)
|
2025-12-25 16:24:37 +08:00 |
|
Xiaoyu Zhang
|
d77f3fccbf
|
[Diffusion] Support peak memory record in offline generate and serving (#15610)
|
2025-12-22 21:21:21 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Xiaoyu Zhangandgemini-code-assist[bot]
|
42bff706df
|
[diffusion] profiling: simplify --perf-dump-path JSON output (remove duplicate denoise steps) (#15537)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-21 12:56:46 +08:00 |
|
Xiaoyu Zhang
|
4b351f6b95
|
Apply new moe align block size kernel (#14134)
|
2025-12-21 10:13:32 +08:00 |
|
Xiaoyu Zhang
|
7fa4906f4f
|
[sgl-kernel] Streamline kernel size report (Top 20 only) and clean up (#15552)
|
2025-12-21 10:00:47 +08:00 |
|
Xiaoyu Zhang
|
bee8ac5b88
|
[diffusion] doc: add --perf-dump-path section to profiling doc (#15533)
|
2025-12-21 00:15:27 +08:00 |
|
Xiaoyu Zhang
|
8999ce754f
|
[diffusion] perf: support zero-cost weight offload and overlap with compute for wan-series (#15511)
|
2025-12-20 22:52:40 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Xiaoyu Zhangandgemini-code-assist[bot]
|
f3705b0115
|
[diffusion] doc: add doc for attention backends (#15408)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-19 22:10:29 +08:00 |
|
Xiaoyu Zhang
|
9a7641d7bf
|
[diffusion] profiling: include per-denoising-step timings in perf-dump-path (#15397)
|
2025-12-18 18:43:33 +08:00 |
|
Xiaoyu Zhang
|
56d12b4aea
|
Fix warp illegal instruction in kimi k2 thinking PCG (#15306)
|
2025-12-18 16:58:23 +08:00 |
|
![github-actions[bot]](/assets/img/avatar_default.png) ![gemini-code-assist[bot]](/assets/img/avatar_default.png) 
|
6c4bf8a0be
|
[diffusion] profiling: enhance trace export with gzip and integrity check (#15326)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2025-12-17 20:44:55 +08:00 |
|
Xiaoyu Zhang
|
533851fbcb
|
[diffusion] ci: add flux2 tp2 test into ci to avoid breaking tensor parallel (#15237)
|
2025-12-17 20:44:21 +08:00 |
|
Xiaoyu Zhang
|
6292d97135
|
[diffusion] fix: fix pack qkv opt break tensor parallel (#15225)
|
2025-12-16 14:33:49 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) 
|
4901693110
|
[diffusion] perf: support FFN pack gate and up proj for Z-Image(#15201)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-16 01:18:47 +08:00 |
|
Xiaoyu Zhang
|
c0d94440b7
|
[diffusion] perf: support pack qkv for Z-Image (#15191)
|
2025-12-16 00:22:24 +08:00 |
|
 Xiaoyu ZhangandMick
|
7bc8b1532e
|
[diffusion] fix: fix AttributeError in _build_parallelism_config when accessing tp_group.device_group (#15196)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-15 23:41:06 +08:00 |
|
 Xiaoyu ZhangandMick
|
92c29d43ac
|
[diffusion] fix: cache dit with parallel (#15163)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-15 19:15:51 +08:00 |
|
Xiaoyu Zhang
|
4513f549ee
|
[diffusion] fix: fix default resolution 720p width from 1080 to 1280 (#15058)
|
2025-12-15 09:16:47 +08:00 |
|
 Xiaoyu ZhangandMick
|
64b5c3ab90
|
[diffusion] refactor: refactor fuse qkv with QKVParallelLinear linear (#15090)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-15 00:33:29 +08:00 |
|
 Xiaoyu ZhangandMick
|
e3f51e823e
|
[diffusion] feat: add support for additional sampling parameters in video generation API (#15062)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-14 19:44:03 +08:00 |
|
 ![github-actions[bot]](/assets/img/avatar_default.png)
|
fdfabb7afc
|
[diffusion] fix: tiny fix _templated_ring_attention bug (#15053)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-14 19:41:53 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) 
|
0c23331e2e
|
[diffusion] doc: add multimodal-gen profiling doc (#15069)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-13 22:26:20 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) 
|
3e1e71575c
|
[diffusion] docker: Tiny fix Docker Hub link in installation documentation (#14987)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-12 20:25:36 +08:00 |
|
Xiaoyu Zhang
|
12b7a4fab0
|
[diffusion] performance: refactor diffusion fuse qkv and apply to qwen-image (#14793)
|
2025-12-10 18:55:41 +08:00 |
|
Xiaoyu Zhang
|
53d170883a
|
Add fuse_marlin_moe test to ci and add new ep test (#14686)
|
2025-12-09 20:17:38 +08:00 |
|
Xiaoyu Zhang
|
03b835e7d1
|
Refactor tuning block wise kernel and opt Qwen/Qwen3-VL-32B-Instruct-FP8 (#14141)
|
2025-12-08 09:24:58 +08:00 |
|