Xiaoyu Zhang
|
6292d97135
|
[diffusion] fix: fix pack qkv opt break tensor parallel (#15225)
|
2025-12-16 14:33:49 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) 
|
4901693110
|
[diffusion] perf: support FFN pack gate and up proj for Z-Image(#15201)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-16 01:18:47 +08:00 |
|
Xiaoyu Zhang
|
c0d94440b7
|
[diffusion] perf: support pack qkv for Z-Image (#15191)
|
2025-12-16 00:22:24 +08:00 |
|
 Xiaoyu ZhangandMick
|
7bc8b1532e
|
[diffusion] fix: fix AttributeError in _build_parallelism_config when accessing tp_group.device_group (#15196)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-15 23:41:06 +08:00 |
|
 Xiaoyu ZhangandMick
|
92c29d43ac
|
[diffusion] fix: cache dit with parallel (#15163)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-15 19:15:51 +08:00 |
|
Xiaoyu Zhang
|
4513f549ee
|
[diffusion] fix: fix default resolution 720p width from 1080 to 1280 (#15058)
|
2025-12-15 09:16:47 +08:00 |
|
 Xiaoyu ZhangandMick
|
64b5c3ab90
|
[diffusion] refactor: refactor fuse qkv with QKVParallelLinear linear (#15090)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-15 00:33:29 +08:00 |
|
 Xiaoyu ZhangandMick
|
e3f51e823e
|
[diffusion] feat: add support for additional sampling parameters in video generation API (#15062)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-14 19:44:03 +08:00 |
|
 ![github-actions[bot]](/assets/img/avatar_default.png)
|
fdfabb7afc
|
[diffusion] fix: tiny fix _templated_ring_attention bug (#15053)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-14 19:41:53 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) 
|
0c23331e2e
|
[diffusion] doc: add multimodal-gen profiling doc (#15069)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-13 22:26:20 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) 
|
3e1e71575c
|
[diffusion] docker: Tiny fix Docker Hub link in installation documentation (#14987)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2025-12-12 20:25:36 +08:00 |
|
Xiaoyu Zhang
|
12b7a4fab0
|
[diffusion] performance: refactor diffusion fuse qkv and apply to qwen-image (#14793)
|
2025-12-10 18:55:41 +08:00 |
|
Xiaoyu Zhang
|
53d170883a
|
Add fuse_marlin_moe test to ci and add new ep test (#14686)
|
2025-12-09 20:17:38 +08:00 |
|
Xiaoyu Zhang
|
03b835e7d1
|
Refactor tuning block wise kernel and opt Qwen/Qwen3-VL-32B-Instruct-FP8 (#14141)
|
2025-12-08 09:24:58 +08:00 |
|
Xiaoyu Zhang
|
ae6a6630e4
|
Add Expert Parallelism (EP) support for kimi-k2-thinking (#13725)
|
2025-12-07 20:28:57 +08:00 |
|
Xiaoyu Zhang
|
e5135b73f4
|
Add CUDA kernel size analysis tool for sgl-kernel optimization (#14544)
|
2025-12-07 15:29:41 +08:00 |
|
 Xiaoyu ZhangandMick
|
6d41791823
|
[diffusion] perf: add QKV fusion optimization for Flux models (#14505)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2025-12-06 20:44:16 +08:00 |
|
Xiaoyu Zhang
|
5347732219
|
[diffusion] fix: Fix profiler trace missing Python stack in diffusion pipeline (#14499)
|
2025-12-05 12:12:35 +00:00 |
|
Xiaoyu Zhang
|
c5947ecd85
|
Opt moe align block size kernel (#14133)
|
2025-12-02 19:13:55 +08:00 |
|
Xiaoyu Zhang
|
3de09aadbc
|
Add new moe wna16 marlin gemm (#14122)
|
2025-12-01 23:07:53 +08:00 |
|
Xiaoyu Zhang
|
fa9021b21f
|
fix: Increase FlashInfer workspace size for Qwen3VL models (#14173)
|
2025-12-01 17:54:23 +08:00 |
|
![github-actions[bot]](/assets/img/avatar_default.png) Xiaoyu Zhangandgithub-actions[bot]
|
9c80072845
|
Add peak output tokens per second in bench_serving (#14165)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2025-12-01 17:47:54 +08:00 |
|
![github-actions[bot]](/assets/img/avatar_default.png) Xiaoyu Zhangandgithub-actions[bot]
|
407cb3ce1e
|
[CI tiny fix] Enhance robustness of vision chunked prefill test with ROUGE-L metric (#13793)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2025-11-25 15:41:14 +08:00 |
|
Xiaoyu Zhang
|
ecefc7904f
|
[sgl-kernel Code Clean] Remove useless lightning_attention kernel (#13819)
|
2025-11-24 18:26:25 +08:00 |
|
Xiaoyu Zhang
|
9ea1953331
|
[Doc] Refine fused_moe_triton configs doc (#13820)
|
2025-11-23 19:09:41 -08:00 |
|
Xiaoyu Zhang
|
b964ce61d6
|
[DeepEP] Add SGLANG_DEEPEP_BF16_DISPATCH env var in Normal mode (#13787)
|
2025-11-23 17:32:43 +08:00 |
|
Xiaoyu Zhang
|
a34d3abb54
|
[Clean code] Compressed_tensors_moe code clean (#13719)
|
2025-11-21 18:15:46 +08:00 |
|
Xiaoyu Zhang
|
bfcf15a129
|
[opt kimi k2 4 / n] Delete useless pad kernel in sgl_moe_align_block_size (#13587)
|
2025-11-21 13:16:42 +08:00 |
|
![github-actions[bot]](/assets/img/avatar_default.png) ![gemini-code-assist[bot]](/assets/img/avatar_default.png)
|
fb04d43428
|
[kimi k2 thinking] Avoid useless torch.zeros_ (#13596)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2025-11-21 13:15:27 +08:00 |
|
Xiaoyu Zhang
|
dc69462456
|
[CI fix] Fix image download failures in VLM CI tests (#13613)
|
2025-11-20 11:18:06 +08:00 |
|
Xiaoyu Zhang
|
820e13c9c1
|
[opt kimi k2 3/n] opt kimi_k2 moe_fused_gate kernel (#13374)
|
2025-11-18 15:36:31 +08:00 |
|
Xiaoyu Zhang
|
7b44526038
|
[Ci tiny fix] Lower score threshold in evaluation test (#13443)
|
2025-11-18 00:45:20 +08:00 |
|
Xiaoyu Zhang
|
8b5e2c5368
|
[Tiny fix] Fix bench_speculative.py run bug (#13416)
|
2025-11-17 18:58:19 +08:00 |
|
 Xiaoyu ZhangandYuan Luo
|
95f43669b5
|
[2 / 2] apply sgl-kernel weak_ref_tensor (#12978)
Co-authored-by: Yuan Luo <yuan.luo@hotmail.com>
|
2025-11-16 21:24:35 +08:00 |
|
Xiaoyu Zhang
|
50691d7b49
|
[opt kimi k2 2/n] apply kimi k2 thinking moe_fused_gate (#13332)
|
2025-11-16 21:20:08 +08:00 |
|
Xiaoyu Zhang
|
1d3d42bda0
|
[opt kimi k2 1 / n] Add kimi k2 moe fused gate (#13287)
|
2025-11-15 17:14:19 +08:00 |
|
Xiaoyu Zhang
|
a1cb717d0b
|
Opt kimi_k2_thinking biased topk module (#13150)
|
2025-11-12 16:52:06 -08:00 |
|
Xiaoyu Zhang
|
fe92d4d88e
|
[CI] Auto format code (#13053)
|
2025-11-10 23:30:29 -08:00 |
|
Xiaoyu Zhang
|
9caca6a45c
|
[PieceWise CUDA Graph] Support awq/gptq model in piecewise cudagraph (#12518)
|
2025-11-11 11:56:15 +08:00 |
|
Xiaoyu Zhang
|
f18ec927f3
|
fix tuning_fused_moe_triton_sep tool per_channel_quant bug (#13027)
|
2025-11-11 10:33:54 +08:00 |
|
Xiaoyu Zhang
|
547de8c774
|
[1 / 2] register weak_ref_tensor in sgl-kernel (#12999)
|
2025-11-10 22:12:59 +08:00 |
|
Xiaoyu Zhang
|
05559a4a90
|
Support hidden_dim % 4 == 0 in per_token_quant_fp8 (#12883)
|
2025-11-10 17:13:14 +08:00 |
|
 
|
fc84b0730c
|
[Refactor] Refactor fused_moe_triton tuning tools: extract shared utils, add EP/MLLM support, reduce overhead (#12440)
Co-authored-by: xu-yfei <xu-yfei@users.noreply.github.com>
Co-authored-by: Yongfei Xu <xuyongfei.xyf@antgroup.com>
|
2025-11-06 20:54:42 +08:00 |
|
Xiaoyu Zhang
|
95191ebdca
|
Migrate weak_ref_tensor to sgl-kernel (#12505)
|
2025-11-02 10:55:39 +08:00 |
|
Xiaoyu Zhang
|
d8fcbaa38d
|
[CI Monitor] Fix ci_monitor perf analyzer bug (#12281)
|
2025-10-30 09:47:12 -07:00 |
|
Xiaoyu Zhang
|
04e5b6faa7
|
Revert "Triton fused_moe_kernel support ep moe tuning" (#12377)
|
2025-10-30 07:12:06 -07:00 |
|
Xiaoyu Zhang
|
52694b60da
|
Triton fused_moe_kernel support ep moe tuning (#12343)
|
2025-10-29 23:16:09 +08:00 |
|
Xiaoyu Zhang
|
334543ff3b
|
Add continuous_usage_stats support for streaming responses (#12241)
|
2025-10-29 10:01:23 +08:00 |
|
Xiaoyu Zhang
|
d0cff78f54
|
[CI] Add ci monitor balance workflow (#11962)
|
2025-10-25 12:14:36 -07:00 |
|
Xiaoyu Zhang
|
8470133852
|
[b200] fix piecewise cuda graph launch bug (#12067)
|
2025-10-24 22:36:39 +08:00 |
|
Xiaoyu Zhang
|
8374a96e49
|
piecewise cuda graph support qwen3-moe (#11845)
|
2025-10-21 10:55:49 +08:00 |
|
Xiaoyu Zhang
|
984fbeb16b
|
Revert "[CI Monitor] Ci monitor only deal with main branch in default" (#11846)
|
2025-10-19 22:06:40 -07:00 |
|
Xiaoyu Zhang
|
24ed3f32c0
|
fix(ci): Fix CI Monitor limit parameter and add CI Analysis to summary (#11832)
|
2025-10-19 18:08:34 -07:00 |
|
Xiaoyu Zhang
|
88a6f9dab5
|
bench_serving support PD Disaggregation (#11542)
|
2025-10-13 19:43:26 -07:00 |
|
Xiaoyu Zhang
|
8e51049f56
|
[CI Monitor] Ci monitor only deal with main branch in default (#11538)
|
2025-10-13 13:50:04 -07:00 |
|
Xiaoyu Zhang
|
6806c4e63e
|
[CI monitor] Improve CI analyzer: fix job failure tracking and add CUDA-focused filtering (#11505)
|
2025-10-13 13:31:09 +08:00 |
|
Xiaoyu Zhang
|
6f16bf9d9d
|
[Ci Monitor] Auto uploaded performance data to sglang_ci_data repo (#10976)
|
2025-09-29 16:17:27 +08:00 |
|
Xiaoyu Zhang
|
11965b0daf
|
Fix sgl-kernel benchmark dead code (#11022)
|
2025-09-29 15:06:40 +08:00 |
|
Xiaoyu Zhang
|
2387c22b56
|
Ci monitor support performance (#10965)
|
2025-09-27 09:11:21 +08:00 |
|
 
|
05a3526654
|
Restruct gpu_memory_settings in a unify function and relax max_cuda_graph_bs (#10372)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: sglang-bot <sglangbot@gmail.com>
|
2025-09-26 15:10:49 -07:00 |
|
Xiaoyu Zhang
|
c4197e99bb
|
[ci] add ci-monitor workflow (#10898)
|
2025-09-25 19:29:47 -07:00 |
|
Xiaoyu Zhang
|
c1f39013b7
|
[ci feature] add ci monitor (#10872)
|
2025-09-24 23:16:29 -07:00 |
|
Xiaoyu Zhang
|
c4e314f986
|
Restruct sgl-kernel benchmark (#10861)
|
2025-09-25 07:45:25 +08:00 |
|
Xiaoyu Zhang
|
37367da639
|
[fix CI] Fix logical condition in fused MoE layer for compressed tensor quantization (#10299)
|
2025-09-10 23:54:09 -07:00 |
|
Xiaoyu Zhang
|
b1fb7e458c
|
[benchmark] add flashinfer_allreduce_fusion benchmark (#9937)
|
2025-09-03 16:31:01 +08:00 |
|
Xiaoyu Zhang
|
a1e5d78115
|
fix parallel_state.py current_platform bug (#9919)
|
2025-09-02 03:17:15 -07:00 |
|
Xiaoyu Zhang
|
b5245064f6
|
[code style] restruct fused_moe to avoid very long single file (#9878)
|
2025-09-02 11:04:27 +08:00 |
|
Xiaoyu Zhang
|
f96413c444
|
Refactor allreduce add rmsnorm pattern (#9278)
|
2025-08-20 02:03:08 -07:00 |
|
Xiaoyu Zhang
|
63d82a776a
|
refine mxfp4 shuffling log (#9194)
|
2025-08-14 10:57:29 -07:00 |
|
Xiaoyu Zhang
|
44e86480e8
|
fuse allreduce and residual_rmsnorm (#8731)
|
2025-08-11 13:50:53 -07:00 |
|
Xiaoyu Zhang
|
a886564a18
|
fix flashinfer allreduce fusion import bug (#9007)
|
2025-08-09 13:47:05 -07:00 |
|
Xiaoyu Zhang
|
0d1e27a0c5
|
Better optimization log for gpt-oss model (#8953)
|
2025-08-08 00:11:48 -07:00 |
|
Xiaoyu Zhang
|
76915d68a8
|
Fix enable flashinfer mxfp4 moe bf16 check (#8950)
|
2025-08-07 22:52:09 -07:00 |
|
Xiaoyu Zhang
|
3ae33fcd0a
|
Fix hopper launch gpt-oss model illegal memory (#8908)
|
2025-08-07 10:02:40 -07:00 |
|
Xiaoyu Zhang
|
47824c1488
|
[Perf] Auto enable best flashinfer mxfp4 kernel in b200 (#8898)
|
2025-08-07 01:08:41 -07:00 |
|
Xiaoyu Zhang
|
4373df5525
|
add flashinfer mxfp4 (#8847)
|
2025-08-06 16:23:41 -07:00 |
|
Xiaoyu Zhang
|
f57d2dc162
|
[sgl-kernel] avoid per_token_quant_fp8.cu hardcode sm_count (#8738)
|
2025-08-04 12:55:57 +08:00 |
|
 Xiaoyu ZhangandKe Bao
|
7a4309cc8a
|
[sgl-kernel performace] fix fp8 quant kernels dispatch __nv_fp8_e4m3 bug to improve performance 10%-20% (#8499)
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
|
2025-07-29 23:31:54 +08:00 |
|
Xiaoyu Zhang
|
2262369905
|
Revert "[kernel] opt moe align block kernel by block/warp scan algorithm" (#8457)
|
2025-07-28 01:35:43 -07:00 |
|
Xiaoyu Zhang
|
9045cc1eb8
|
[torch.compile bug] avoid biased_grouped_topk_impl func repeatedly triggering torch.compile in forward pass (#8353)
|
2025-07-25 21:17:47 +08:00 |
|
Xiaoyu Zhang
|
a167fd0bcb
|
[code style] Clean dead triton kernel code in fused_moe and useless vllm_ops import (#8310)
|
2025-07-24 14:38:30 +08:00 |
|
Xiaoyu Zhang
|
aa2056091a
|
delete uselese code caused by fuse allreduce+add_rmsnorm pr (#7970)
|
2025-07-11 19:43:38 -07:00 |
|
Xiaoyu Zhang
|
49a5915f53
|
[ready b200] fuse allreduce+add_rmsnorm in prepare_attention + mlp module (#7775)
|
2025-07-10 15:12:39 -07:00 |
|
Xiaoyu Zhang
|
2e7ab862e3
|
Fix illegal memory in trtllm allreduce fusion (#7864)
|
2025-07-08 11:47:17 -07:00 |
|
Xiaoyu Zhang
|
8e64140e35
|
[b200] support trt-llm allreduce fuse rms_norm_add kernel (#7621)
|
2025-07-02 19:36:20 -07:00 |
|
Xiaoyu Zhang
|
ff2e9c9479
|
Add small requirements for benchmark/parse_result tools (#7671)
|
2025-06-30 21:52:20 -07:00 |
|
Xiaoyu Zhang
|
8ecad0b16f
|
[benchmark] fbgemm benchmark support bandwidth report and support fbgemm_cutlass_gmm (#7422)
|
2025-06-24 09:44:55 -07:00 |
|
Xiaoyu Zhang
|
0ae1e9a755
|
refine fused_moe benchmark (#7221)
|
2025-06-15 21:21:32 -07:00 |
|
Xiaoyu Zhang
|
3712abfaf9
|
Fuse routed scaling factor in deepseek (#6970)
|
2025-06-08 15:24:24 -07:00 |
|
Xiaoyu Zhang
|
fa3592cfeb
|
rebase h20 fused_moe config (#6966)
|
2025-06-08 05:01:34 -07:00 |
|
Xiaoyu Zhang
|
515ef4facb
|
Fuse routed scaling factor in topk_reduce kernel (#6220)
|
2025-06-07 11:06:50 -07:00 |
|
Xiaoyu Zhang
|
bae4fdc7ab
|
add fbgemm moe grouped gemm kernel benchmark (#6924)
|
2025-06-07 02:57:30 -07:00 |
|
Xiaoyu Zhang
|
8b5f83ed3b
|
reduce torch.zeros overhead in moe align block size kernel (#6369)
|
2025-06-07 02:47:36 -07:00 |
|
Xiaoyu Zhang
|
2a413829f4
|
Add triton version as a fused_moe_triton config search key to avoid performace decrease in different Triton version (#5955)
|
2025-06-07 02:43:50 -07:00 |
|
 Xiaoyu ZhangandJieXin Liang
|
bd75690f4e
|
fix ep_moe_reorder kernel bugs (#6858)
Co-authored-by: JieXin Liang <Alcanderian@users.noreply.github.com>
|
2025-06-04 19:13:59 +08:00 |
|
Xiaoyu Zhang
|
076103535c
|
fix log_info_on_rank0 error when run benchmark (#6260)
|
2025-05-28 00:20:01 -07:00 |
|
Xiaoyu Zhang
|
9f2c9568f0
|
[doc] add a note for --n-share-experts-fusion args (#6154)
|
2025-05-11 23:18:38 -07:00 |
|
Xiaoyu Zhang
|
d25398cbc8
|
fix custom_allreduce namespace (#6039)
|
2025-05-06 19:13:06 -07:00 |
|
Xiaoyu Zhang
|
5bb0accbcf
|
cutlass 3.9 supported to improve fp8_blockwise_gemm (#5820)
|
2025-04-28 21:52:36 -07:00 |
|
Xiaoyu Zhang
|
1cc326032d
|
simplify fused_moe config logging (#5801)
|
2025-04-28 17:04:54 -07:00 |
|