Hubert Lu
|
10d33bd77e
|
[AMD] Enable Piecewise CUDA Graph for AMD GPUs (#22299)
|
2026-06-07 16:28:20 -07:00 |
|
Hubert Lu
|
72929c7000
|
[AMD] Enable AITER custom all-gather on ROCm (#25093)
|
2026-06-02 15:57:37 -07:00 |
|
 
|
c2db19ffa4
|
[AMD] Enable EAGLE speculative decoding for Qwen3.5 FP8 and MXFP4 models with aiter's unified attention (#23146)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: sogalin <39478626+sogalin@users.noreply.github.com>
|
2026-05-05 00:09:40 -07:00 |
|
Hubert Lu
|
d57671527a
|
Fix LFM2 ShortConv Mamba State Indexing (#23975)
|
2026-04-30 15:23:39 -07:00 |
|
Hubert Lu
|
e5da200d0a
|
[AMD] Fix Aiter RMSNorm layout handling (#23974)
|
2026-04-28 19:28:46 -07:00 |
|
Hubert Lu
|
4cb0c4e1f3
|
[AMD] Fix memory access fault when --page-size > 1 with speculative decoding on AMD GPUs (#23596)
|
2026-04-23 23:56:36 -07:00 |
|
 Hubert LuandHaiShaw
|
b2af34be54
|
[AMD] Optimize _append_shared_to_topk_output by a single fused Triton kernel for Qwen3.5 (#22844)
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2026-04-14 23:50:32 -07:00 |
|
 Hubert LuandHAI
|
edaa5973d4
|
[AMD][No-Merge] Simplify fused allreduce + RMSNorm and remove hidden_dim allowlist (#21986)
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-04-11 23:47:08 -07:00 |
|
Hubert Lu
|
e6071e60c0
|
[AMD] Support AMD MXFP4 Qwen3.5-397B-A17B model (#21234)
|
2026-03-30 01:14:18 -07:00 |
|
Hubert Lu
|
7c7b2a8c97
|
[Bugfix] Lazy-import CuteDSL KDA kernel to fix AMD/ROCm startup crash (#21428)
|
2026-03-25 16:37:26 -07:00 |
|
 Hubert LuandClaude Opus 4.6
|
943f34f642
|
Add NCCL/RCCL pre-warming to reduce P99 TTFT cold-start latency (#20477)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-03-16 20:23:14 -07:00 |
|
 
|
67f02681c9
|
[AMD] Support speculative decoding v2 for aiter backend on ROCm/HIP (#17450)
Co-authored-by: kkHuang-amd <wunhuang@amd.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2026-03-11 17:01:01 -07:00 |
|
Hubert Lu
|
441045a7bf
|
[AMD] Fix EAGLE3 speculative decoding with aiter attention backend (#19362)
|
2026-03-03 16:12:13 -08:00 |
|
 Hubert Luandbingxche
|
05f68e1230
|
[AMD] Fix the hipDeviceGetName issue in ROCm based docker images (#19440)
Co-authored-by: bingxche <Bingxu.Chen@amd.com>
|
2026-03-03 11:42:48 -08:00 |
|
 Hubert Luandyctseng0211
|
17b0affbdf
|
[AMD] Support --enable-aiter-allreduce-fusion on AMD GPUs (#13747)
Co-authored-by: yctseng0211 <yctseng@amd.com>
|
2026-02-24 23:11:55 -08:00 |
|
 Hubert LuandCursor
|
8bd644765f
|
[AMD] Enable ROCm kvcache JIT path and add AMD CI coverage. (#18992)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-02-25 14:15:05 +08:00 |
|
Hubert Lu
|
93423ff780
|
[AMD] Deprecate ROCm 6.3 artifacts and standardize gfx942 on ROCm 7 (#17785)
|
2026-01-27 15:58:49 -08:00 |
|
 Hubert Luandwufann
|
df42f4d386
|
[AMD] Update dsv3.2 AMD GPU docs and unify ROCm TileLang build (#17783)
Co-authored-by: wufann <715544327@qq.com>
|
2026-01-26 21:10:32 -08:00 |
|
 Hubert Luandwufann
|
afe285f7bd
|
[AMD] enable CUDA graph for NSA backend and fix NSA FP8 fused RMSNorm group quant (#16841)
Co-authored-by: wufann <715544327@qq.com>
|
2026-01-13 17:36:01 -08:00 |
|
Hubert Lu
|
8716589826
|
[AMD][Diffusion] support timestep embedding kernel for AMD GPUs (#16766)
|
2026-01-12 22:17:07 -08:00 |
|
 Hubert LuandHAI
|
d6d5c3fdea
|
[AMD] Clean up vllm dependencies in moe_runner/triton.py (#11349)
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-01-09 00:24:04 -08:00 |
|
Hubert Lu
|
4935344fcd
|
[AMD] Fix aiter page-size handling, DeepSeek MLA tuple inputs, and HiCache/FA3 decode-backend override (#16531)
|
2026-01-07 21:14:32 -08:00 |
|
Hubert Lu
|
b86bbf841e
|
[AMD] Add 8-GPU MX35X test running DSR1-MXFP4 model for AMD CI (#13602)
|
2026-01-07 01:43:11 -08:00 |
|
Hubert Lu
|
51e2eaa458
|
[AMD] Support fast_topk kernels in sgl-kernel (#15172)
|
2025-12-19 22:19:09 -08:00 |
|
Hubert Lu
|
2ce2377740
|
[AMD] Fix AITER_MXFP4_MOE_SF setting for gfx950 (#13239)
|
2025-11-13 17:03:41 -08:00 |
|
 
|
e4b2937017
|
[AMD] Add AITER Custom All-Reduce (#13102)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
Co-authored-by: HaiShaw <hixiao@gmail.com>
|
2025-11-12 21:53:44 -08:00 |
|
Hubert Lu
|
36d147121f
|
[AMD] Apply AITER_MXFP4_MOE_SF=1 only to gfx950 in aiter build (#13092)
|
2025-11-11 12:26:09 -08:00 |
|
Hubert Lu
|
3694266051
|
Expand and update test coverage for AMD CI (#10044)
|
2025-11-04 22:15:13 -08:00 |
|
Hubert Lu
|
01a26544a3
|
[AMD] Add Tilelang and Fast Hadamard Transform builds to Dockerfile.rocm (#11114)
|
2025-09-30 20:00:37 -07:00 |
|
Hubert Lu
|
fe68c1486f
|
Fix errors of hicache kernels in sgl-kernel for ROCm (#10339)
|
2025-09-11 14:54:34 -07:00 |
|
 Hubert LuandSai Enduri
|
91b3555d2d
|
Add tests to AMD CI for MI35x (#9662)
Co-authored-by: Sai Enduri <saimanas.enduri@amd.com>
|
2025-09-10 12:50:05 -07:00 |
|
Hubert Lu
|
2c562fd2d0
|
Fix Llama 4 with MXFP4 dynamic quant on MI35x (#9993)
|
2025-09-04 00:48:58 -07:00 |
|
Hubert Lu
|
711390a971
|
[AMD] Support Hierarchical Caching on AMD GPUs (#8236)
|
2025-08-28 15:27:07 -07:00 |
|
Hubert Lu
|
f445a1d9a3
|
[AMD] Fix Llama 4 FP8 accuracy issues on MI300X (#7699)
|
2025-08-22 13:13:45 -07:00 |
|
Hubert Lu
|
704ced1b2e
|
[AMD] Remove the deprecated C10_WARP_SIZE (#9356)
|
2025-08-21 18:16:35 -07:00 |
|
Hubert Lu
|
c6c379ab31
|
[AMD] Reorganize hip-related header files in sgl-kernel (#9320)
|
2025-08-18 16:53:44 -07:00 |
|
Hubert Lu
|
9c3e95d98b
|
[AMD] Expand test coverage for AMD CI and enable apply_token_bitmask_inplace_cuda in sgl-kernel (#8268)
|
2025-08-15 12:32:51 -07:00 |
|
 
|
af4b9bae95
|
[AMD] Add silu_and_mul, gelu_and_mul, gelu_tanh_and_mul, and gelu_quick kernels for AMD GPUs (#7135)
Co-authored-by: yiakwy-xpu-ml-framework-team <961186938@qq.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2025-07-24 23:44:28 -07:00 |
|
Hubert Lu
|
e50109f2ed
|
[AMD] Remove vllm's scaled_fp8_quant and moe_sum when SGLANG_USE_AITER=1 (#7484)
|
2025-07-21 17:33:19 -07:00 |
|
Hubert Lu
|
7750b91ca8
|
[AMD] Add triton awq_dequantize kernel to support AWQ on ROCm (#7661)
|
2025-07-18 14:27:25 -07:00 |
|
Hubert Lu
|
e00715eb66
|
[AMD] Add test_fused_moe.py and test_rope_rocm.py to AMD CI (#5246)
|
2025-07-06 01:47:16 -07:00 |
|
Hubert Lu
|
b116b21a46
|
[AMD] Temporarily disable test_no_overlap_scheduler and test_vision_chunked_prefill (#7717)
|
2025-07-02 12:39:18 -07:00 |
|
Hubert Lu
|
3b3f1e3aeb
|
[AMD] Add unit-test-sgl-kernel-amd to AMD CI (#7539)
|
2025-06-29 15:50:09 -07:00 |
|
Hubert Lu
|
4740288303
|
[AMD] Add more tests to per-commit-amd (#6926)
|
2025-06-08 01:08:37 -07:00 |
|
Hubert Lu
|
198b9056d1
|
[AMD] Fix Llama 4 Scout and Maverick accuracy issues on MI300X (#6274)
|
2025-05-14 22:07:29 +00:00 |
|
Hubert Lu
|
2a936a841e
|
[AMD] switch to custom allreduce regardless of MSCCL setting on ROCm (#6097)
|
2025-05-08 13:46:58 -07:00 |
|
Hubert Lu
|
afb752bcbe
|
[AMD] Fix missing per_token_group_quant_fp8 for ROCm (#5140)
|
2025-04-07 22:38:25 -07:00 |
|
Hubert Lu
|
9cf4077294
|
Enable custom AR for AMD GPUs and maintain it in sgl-kernel (#3406)
|
2025-03-02 15:19:06 -08:00 |
|
Hubert Lu
|
f8b28e461a
|
Add CPU affinity setting to latency benchmark (#3085)
|
2025-01-25 23:52:05 -08:00 |
|