Xiaoyu Zhang
|
a9a8b20a90
|
[codex] Optimize Z-Image packed QKV (#24117)
|
2026-05-07 07:51:22 +08:00 |
|
Xiaoyu Zhang
|
d86f2916cc
|
Fix diffusion fallback guards and validation (#23335)
|
2026-05-07 00:05:43 +08:00 |
|
 Xiaoyu ZhangandBBuf Codex
|
b67df7cd1b
|
[Codex] Diffusion handle non-contiguous CFG communication (#24332)
Co-authored-by: BBuf Codex <bbuf-codex@users.noreply.github.com>
|
2026-05-06 17:27:14 +08:00 |
|
Xiaoyu Zhang
|
d7385b575f
|
[Diffusion] Optimize Hunyuan3D shape denoising (#24287)
|
2026-05-06 10:10:09 +08:00 |
|
Xiaoyu Zhang
|
67e8bd7a80
|
[codex] Optimize Helios fused norm modulation (#24059)
|
2026-05-05 19:28:37 +08:00 |
|
Xiaoyu Zhang
|
8c703f215e
|
Add HunyuanVideo ModelOpt FP8 diffusion support (#23199)
|
2026-05-05 19:27:28 +08:00 |
|
 Xiaoyu ZhangandBBuf Codex
|
078f84d80d
|
[SKILL] Add diffusion benchmark presets for edit and Hunyuan3D models (#24288)
Co-authored-by: BBuf Codex <bbuf-codex@users.noreply.github.com>
|
2026-05-05 08:18:12 +08:00 |
|
Xiaoyu Zhang
|
4b6d44641b
|
[diffusion] chore: enable channels-last 3D VAE convs by default (#23200)
|
2026-05-04 22:59:31 +08:00 |
|
 Xiaoyu ZhangandMick
|
f2d1390909
|
[Diffusion] Add Qwen Image ModelOpt FP8 support (#23155)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-05-04 00:24:22 +08:00 |
|
Xiaoyu Zhang
|
4128f1ffe2
|
[SKILLS] Tiny upgrade diffusion skills (#24273)
|
2026-05-02 22:04:05 +08:00 |
|
Xiaoyu Zhang
|
b712dd48fe
|
[codex] diffusion: enable group norm silu fuse by default (#23148)
|
2026-05-02 20:55:51 +08:00 |
|
Xiaoyu Zhang
|
1360848ee1
|
Optimize large GroupNorm SiLU apply (#23938)
|
2026-05-02 20:54:46 +08:00 |
|
Xiaoyu Zhang
|
589f90b368
|
[diffusion] chore: use lmsys as org for modelopt checkpoints (#23924)
|
2026-05-02 17:18:58 +08:00 |
|
Xiaoyu Zhang
|
321298da75
|
[SKILL] Upgrade sglang profile and auto_benchmark skills (#24250)
|
2026-05-02 10:12:47 +08:00 |
|
Xiaoyu Zhang
|
13afe8acdf
|
[codex] Enable Qwen3-Next MoE all-reduce fusion (#23619)
|
2026-04-29 09:11:35 +08:00 |
|
Xiaoyu Zhang
|
7824903417
|
[SKILL] Sync SGLang skill docs (#23921)
|
2026-04-28 17:05:36 +08:00 |
|
Xiaoyu Zhang
|
6fbad22feb
|
Remove smoke wording from tests and comments (#23355)
|
2026-04-28 12:05:27 +08:00 |
|
Xiaoyu Zhang
|
5f47cae1a0
|
add H100 configs for GLM-4.7-Flash (#23719)
|
2026-04-27 15:07:39 +08:00 |
|
Xiaoyu Zhang
|
0d69012ef8
|
Optimize LTX2 feed-forward tensor parallelism (#23221)
|
2026-04-21 16:29:23 +08:00 |
|
Xiaoyu Zhang
|
cd6ad80c00
|
diffusion: add HunyuanVideo GroupNorm+SiLU fast path (#22814)
|
2026-04-18 23:38:49 +08:00 |
|
Xiaoyu Zhang
|
c6a45fab64
|
Qwen3next flashinfer allreduce auto enable (#22664)
|
2026-04-18 22:32:41 +08:00 |
|
Xiaoyu Zhang
|
615d6c93b2
|
[codex] Add flashinfer TRTLLM backend for diffusion NVFP4 (#22717)
|
2026-04-18 09:06:28 +08:00 |
|
Xiaoyu Zhang
|
83c5119d01
|
[diffusion] CI: fix ModelOpt B200 CI artifact coverage (#22955)
|
2026-04-17 23:33:42 +08:00 |
|
Xiaoyu Zhang
|
91679d935d
|
[codex] Update diffusion skills (#23028)
|
2026-04-17 13:29:26 +08:00 |
|
Xiaoyu Zhang
|
695ab705cb
|
[diffusion] quant: update modelopt quantization docs and CI coverage (#22772)
|
2026-04-15 21:30:28 +08:00 |
|
Xiaoyu Zhang
|
f97c608caa
|
[diffusion] quant: add FLUX.1-dev modelopt nvfp4 support (#22672)
|
2026-04-14 15:00:59 +08:00 |
|
Xiaoyu Zhang
|
fae0a2fc3c
|
[codex] Add LTX-2.3 benchmark skill recipes (#22631)
|
2026-04-13 12:23:32 +08:00 |
|
Xiaoyu Zhang
|
37fc47c645
|
diffusion: fix layerwise offload for ModelOpt quantized DiTs (#22594)
|
2026-04-13 08:01:54 +08:00 |
|
Xiaoyu Zhang
|
03a1a7b81c
|
[Diffusion] Add FLUX.1-dev ModelOpt NVFP4 support (#22574)
|
2026-04-13 07:57:41 +08:00 |
|
Xiaoyu Zhang
|
1ff51555f2
|
[Diffusion] modelopt diffusion fp8 support for flux1/flux2 and wan2.2 (#22365)
|
2026-04-10 20:56:57 +08:00 |
|
Xiaoyu Zhang
|
7603b226ce
|
Upgrade sglang-torch-profiler-analysis SKILLS (#22440)
|
2026-04-09 18:23:03 +08:00 |
|
Xiaoyu Zhang
|
30b738d3a6
|
[SKILL] add torch profiler analysis workflow (#22353)
|
2026-04-09 12:53:48 +08:00 |
|
 Xiaoyu ZhangandMick
|
b5b2dbe05f
|
[Diffusion] Add diffusion NVFP4 scaled-mm correctness test (#22127)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-04-08 22:07:24 +08:00 |
|
Xiaoyu Zhang
|
ea119adc90
|
Refactor auto benchmark unit tests and fix CI bug (#22270)
|
2026-04-08 21:54:41 +08:00 |
|
Xiaoyu Zhang
|
0f0f004f1f
|
[Benchmark] Add auto benchmark tool with YAML-driven server flag search and canonical dataset format (#21736)
|
2026-04-04 21:46:58 +08:00 |
|
Xiaoyu Zhang
|
da25b471e3
|
Align diffusion nightly presets and broaden skill discovery (#22099)
|
2026-04-04 21:43:52 +08:00 |
|
Xiaoyu Zhang
|
f3f7711dac
|
Fix Python 3.11 f-string lint error in deepgemm Blackwell benchmark (#22108)
|
2026-04-04 21:15:22 +08:00 |
|
Xiaoyu Zhang
|
82ea4906cf
|
[diffusion] Default NVFP4 to CUTLASS and add all-model shape benchmarks (#22091)
|
2026-04-04 16:14:38 +08:00 |
|
Xiaoyu Zhang
|
ee9d922f5a
|
Revert "[Kernel] Fuse temperature + softmax in sampling for decode speedup" (#22046)
|
2026-04-03 21:32:08 +08:00 |
|
Xiaoyu Zhang
|
89affff290
|
Skip broken AutoModel mapping entries when resolving Llava submodules (#21892)
|
2026-04-03 09:04:26 +08:00 |
|
Xiaoyu Zhang
|
cdd7d6a227
|
Remove obsolete sgl-kernel legacy paths (#21528)
|
2026-04-01 09:00:20 +08:00 |
|
Xiaoyu Zhang
|
505eb312ec
|
Revert "DeepSeek-R1-0528-w4a8: DeepEP Low Latency Dispatch Adopts FP8 Communication" (#21719)
|
2026-03-31 10:22:01 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Xiaoyu Zhangandgemini-code-assist[bot]
|
516cff97a3
|
[Diffusion] Align diffusion benchmark skill presets with nightly comparison cases (#21616)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-03-29 12:12:17 +08:00 |
|
Xiaoyu Zhang
|
9238bd08a2
|
[CI] Register missing jit_kernel test files (#21547)
|
2026-03-27 19:39:08 +08:00 |
|
Xiaoyu Zhang
|
d633ab7349
|
[Diffusion] Add qknorm rope fuse kernel (#21440)
|
2026-03-27 14:27:08 +08:00 |
|
Xiaoyu Zhang
|
e8d46f145c
|
Opt jit qknorm_across_heads cuda kernel (#21503)
|
2026-03-27 13:30:46 +08:00 |
|
Xiaoyu Zhang
|
7ca015fe65
|
[Diffusion] Refactor diffusion JIT kernel test layout and narrow CI triggers (#21385)
|
2026-03-26 15:02:02 +08:00 |
|
Xiaoyu Zhang
|
6f2b51ade1
|
[Diffusion] Optimize diffusion Triton rotary embedding by processing multiple heads per token (#21387)
|
2026-03-26 08:59:25 +08:00 |
|
Xiaoyu Zhang
|
68f7f00174
|
[Diffusion] Speed up Qwen select01 Triton modulation kernels (#21318)
|
2026-03-25 20:48:39 +08:00 |
|
Xiaoyu Zhang
|
689e9ef05c
|
[Diffusion] Add AKO4ALL kernel optimization skill (#21323)
|
2026-03-25 18:46:21 +08:00 |
|
Xiaoyu Zhang
|
e4ad10520b
|
[diffusion] Skip automatic Wan/MOVA DiT layerwise offload on high-end GPUs (#21248)
|
2026-03-25 18:45:30 +08:00 |
|
Xiaoyu Zhang
|
69f02e36e8
|
[diffusion] Fix torch.zeros typo in causal wan (#21250)
|
2026-03-24 14:39:16 +08:00 |
|
Xiaoyu Zhang
|
d9f97b2115
|
Refine diffusion skills and align JIT kernel docs with the new CI flow (#21283)
|
2026-03-24 14:38:36 +08:00 |
|
Xiaoyu Zhang
|
a94d67d44b
|
[SKILL] fix(bench): Support model-specific DenoisingStage variants in… (#21137)
|
2026-03-23 12:08:00 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Xiaoyu Zhangandgemini-code-assist[bot]
|
c1fe5de69c
|
[Diffusion] Clean up diffusion Triton kernels and modernize custom op registration (#21122)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-03-22 22:38:57 +08:00 |
|
Xiaoyu Zhang
|
766d225fcc
|
Add SGLang CUDA crash API logging inspired by FlashInfer (#20910)
|
2026-03-22 16:39:40 +08:00 |
|
 Xiaoyu ZhangandMick
|
1b65c0d259
|
[Diffusion] Fix torch.compile RMSNorm fallback for Z-Image (#20962)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-03-22 15:38:22 +08:00 |
|
Xiaoyu Zhang
|
c076968c52
|
[CI] Remove obsolete AOT-only jit-kernel benchmarks after sgl-kernel 4.0 (#21075)
|
2026-03-21 13:40:42 +08:00 |
|
Xiaoyu Zhang
|
cf60c5bd15
|
[CI] Fix jit_kernel benchmark ci (#20990)
|
2026-03-20 16:40:20 +08:00 |
|
Xiaoyu Zhang
|
20a23e3173
|
[SKILL] Refine kernel authoring docs and validate add-jit-kernel / add-sgl-kernel end to end with Codex (#20867)
|
2026-03-18 23:00:33 +08:00 |
|
Xiaoyu Zhang
|
6489f77733
|
[Diffusion] Fix compile graph broken by flashinfer rope (#20699)
|
2026-03-16 23:14:27 +08:00 |
|
 Xiaoyu ZhangandBaizhou Zhang
|
15097c5c3b
|
Release sglang kernel 0.4.0 (#20440)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-03-16 20:34:58 +08:00 |
|
Xiaoyu Zhang
|
3055b6906d
|
[Diffusion] Document torch.compile graph-break checks in diffusion benchmark skills (#20681)
|
2026-03-16 17:41:40 +08:00 |
|
Xiaoyu Zhang
|
e1eb25880f
|
[Diffusion] Add a benchmark for rmsnorm/fuse_add_rmsnorm (#20632)
|
2026-03-16 09:50:33 +08:00 |
|
Xiaoyu Zhang
|
45dd06f4e0
|
Remove insecure auto-format label workflow (#20629)
|
2026-03-15 21:04:29 +08:00 |
|
Xiaoyu Zhang
|
5ab2cfe9a8
|
[Diffusion] Clean upstream fa3 in hopper (#20576)
|
2026-03-14 23:41:23 +08:00 |
|
Xiaoyu Zhang
|
25e38216b6
|
[kernel slimming] Clean many useless sgl-kernel deprecated kernels (#20277)
|
2026-03-14 16:45:54 +08:00 |
|
Xiaoyu Zhang
|
f9e4221b71
|
[Diffusion] add mova and hunyuanvideo to perf skills (#20563)
|
2026-03-14 13:49:50 +08:00 |
|
Xiaoyu Zhang
|
be7a0311a0
|
[Diffusion] Fix and validate diffusion skills benchmarking/profiling workflow (#20528)
|
2026-03-13 21:11:37 +08:00 |
|
 Xiaoyu ZhangandYihan Chen
|
e00328d1e5
|
[Diffusion] Opt qwen-image-edit with fuse_residual_layernorm_scale_shift_gate_select01_kernel (#20395)
Co-authored-by: Yihan Chen <yingluosanqian@gmail.com>
|
2026-03-13 13:15:22 +08:00 |
|
Xiaoyu Zhang
|
7ecf07b8f4
|
[jit_kernel] Temporarily Skip Flaky JIT Kernel GDN Test and Add PR Label (#20436)
|
2026-03-13 09:34:22 +08:00 |
|
 Xiaoyu ZhangandBaizhou Zhang
|
680d9d98e4
|
Fix cutedsl ci error (#20309)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-03-11 16:17:35 +08:00 |
|
Xiaoyu Zhang
|
60cc06297e
|
[4/n jit_kernel restruct] speed up CI tests and add benchmark workflow (#20268)
|
2026-03-10 21:37:41 +08:00 |
|
Xiaoyu Zhang
|
51d9d34977
|
[2/n jit_kernel restruct] unify rotary embedding entrypoints under rope.py (#20247)
|
2026-03-10 17:49:57 +08:00 |
|
Xiaoyu Zhang
|
8517da5d08
|
[3/n jit_kernel restruct] Clean up benchmark naming and benchmarking helpers (#20250)
|
2026-03-10 16:39:03 +08:00 |
|
Xiaoyu Zhang
|
c812504b92
|
[1/n jit_kernel restruct] unify cache usage and clean up naming in ngram_embedding (#20244)
|
2026-03-10 15:53:43 +08:00 |
|
Xiaoyu Zhang
|
fd79cd8d9c
|
[Skills] Refine jit_kernel and sgl-kernel skills (#20095)
|
2026-03-07 22:54:01 +08:00 |
|
 Xiaoyu ZhangandMick
|
6d22c9f369
|
[Diffusion] Move hf kernels diffusion cuda kernels skills to SGLD (#20001)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-03-06 22:16:06 +08:00 |
|
 Xiaoyu ZhangandMick
|
9795b4cd5b
|
[Diffusion] Open t5 encoder parallel folding for wan2.2 and mova video (#18493)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-03-05 10:18:00 +08:00 |
|
Xiaoyu Zhang
|
4348976f80
|
[Diffusion] Refactor diffusion benchmark/profile skill to reuse diffusion-perf skill and clarify profiling trigger (#19783)
|
2026-03-04 10:54:42 +08:00 |
|
Xiaoyu Zhang
|
115e9a1acd
|
[Diffusion] Delete useless _ulysses_input_split func (#19786)
|
2026-03-04 10:45:11 +08:00 |
|
Xiaoyu Zhang
|
145ae518ac
|
[Diffusion] Revert 18619 (#19510)
|
2026-03-03 08:15:15 +08:00 |
|
Xiaoyu Zhang
|
51ee17ce44
|
[diffusion] move skills dir (#19697)
|
2026-03-03 02:51:29 +08:00 |
|
Xiaoyu Zhang
|
53de53fb53
|
[jit_kernel] Tiny unify jit_kernel tests style (#19694)
|
2026-03-02 21:33:59 +08:00 |
|
Xiaoyu Zhang
|
e42fa009d4
|
[Diffusion] diffusion profile and opt skills (#19540)
|
2026-03-02 15:06:29 +08:00 |
|
Xiaoyu Zhang
|
fe971b620a
|
[Diffusion] Add SGL-D diffusion efficient kernel skills (#19473)
|
2026-02-27 11:42:13 +08:00 |
|
Xiaoyu Zhang
|
054bd71086
|
[sgl-kernel slimming] remove sgl-kernel moe-wna16-marlin (#19379)
|
2026-02-27 09:17:46 +08:00 |
|
Xiaoyu Zhang
|
74c8e7b215
|
refactor(jit_kernel): reduce duplication and separate test code (#19323)
|
2026-02-26 18:30:49 +08:00 |
|
Xiaoyu Zhang
|
914ed34757
|
update jit_kernel codeowners (#19385)
|
2026-02-26 10:36:34 +08:00 |
|
Xiaoyu Zhang
|
ab7071b545
|
[SKILL] Better claude skills for sgl-kernel and jit-kernel (#19302)
|
2026-02-25 15:26:55 +08:00 |
|
Xiaoyu Zhang
|
694924b878
|
Update .gitignore to remove '.claude/' (#19296)
|
2026-02-25 11:51:50 +08:00 |
|
Xiaoyu Zhang
|
9dff933164
|
[Kernel Slimming] Remove sgl-kernel AOT marlin kernels (#19241)
|
2026-02-25 10:08:22 +08:00 |
|
Xiaoyu Zhang
|
e0e0cad6bc
|
[Diffusion] Match rotary_embedding module name style (#19179)
|
2026-02-23 21:02:34 +08:00 |
|
Xiaoyu Zhang
|
2717393681
|
[Refactor] Split rotary_embedding.py into a modular package (#19144)
|
2026-02-23 20:05:29 +08:00 |
|
Xiaoyu Zhang
|
66497ab0aa
|
[Diffusion] Restruct and clean Diffusion rotary embedding (#19064)
|
2026-02-21 21:41:47 +08:00 |
|
Xiaoyu Zhang
|
19aa19b111
|
[diffusion] refactor: refactor diffusion triton kernels (#18966)
|
2026-02-19 17:03:44 +08:00 |
|
Xiaoyu Zhang
|
390c154306
|
[Tiny fix] Super tiny fix mul_add naive forward bug (#18964)
|
2026-02-18 16:18:43 +08:00 |
|
Xiaoyu Zhang
|
513c12d23f
|
Remove unused fast-hadamard-transform PyTorch extension sources (#18927)
|
2026-02-18 15:51:07 +08:00 |
|
Xiaoyu Zhang
|
d3bae71e3f
|
Add claude skills for sgl-kernel and jit-kernel (#18855)
|
2026-02-15 15:14:03 -08:00 |
|
Xiaoyu Zhang
|
4067d9487d
|
[diffusion] feat: opt vae decode with channels_last_3d (#18540)
|
2026-02-14 23:19:45 +08:00 |
|