14 Commits
Author SHA1 Message Date
kangwangamd 264da63319 [AMD] Update ROCm AITER pin to acf8fdf9 (#39965) 2026-09-21 22:01:17 -07:00
kangwangamd b9b39f222d [AMD] Update ROCm AITER pin to 4ad9983 (#37784) 2026-09-06 19:22:37 -07:00
kangwangamdandbingxche ebec85f606 [AMD][DI][CI] Run MI355X disagg nightly at 7AM UTC (#35467)
Co-authored-by: bingxche <bingxche@users.noreply.github.com>
2026-08-19 16:27:20 +08:00
kangwangamd c82e928fe5 [AMD] diffusion: normalize ModelOpt-FP8 weights to e4m3fnuz on gfx942 (#35111) 2026-08-17 02:35:10 -07:00
kangwangamdandBingxu Chen 29b067245b [AMD] CI: drop the spaces from SGL_EVAL_SPEC (fixes ROCm 7.2 stage-a sgl-eval install) (#34689)
Co-authored-by: Bingxu Chen <bingxche@amd.com>
2026-08-13 12:48:58 -07:00
kangwangamd 4cdab7b4f4 [AMD] Add msgpack to ROCm diffusion deps (fix multimodal-gen unit test ModuleNotFoundError) (#31899) 2026-08-06 01:52:23 -07:00
kangwangamd 12eadf86f1 [AMD] Enable mamba JIT transfer kernel on ROCm (fix transfer_kv_mamba NameError) (#31741) 2026-08-02 16:40:17 -07:00
kangwangamd 4b52758c76 [AMD] Skip test_update_weights_from_disk on ROCm pending reload fix (#31924) (#31925) 2026-07-30 02:32:09 -07:00
kangwangamdandBingxu Chen 50c118704a [diffusion] disagg: handle numpy arrays in cross-role transfer field extraction (#31325)
Co-authored-by: Bingxu Chen <bingxche@amd.com>
2026-07-20 00:50:17 -07:00
kangwangamdandYC Yen-Ching Tseng 9f8e916131 [diffusion] post_training: run weight update under torch.inference_mode() (#31263)
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
2026-07-20 11:14:13 +08:00
kangwangamd f8eac995aa [diffusion] test: fix GLM-Image AR model-path to resolve local snapshot subfolder (#31313) 2026-07-15 23:57:34 -07:00
4d94e9471a [AMD] Relax allreduce-fusion residual accuracy tolerance to 1 bf16 ULP (#28226)
Co-authored-by: kangwangamd <kangwangamd@users.noreply.github.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
2026-06-18 19:18:53 -07:00
kangwangamdandkangwangamd 6309fb9abb [AMD] Fix Always mask padded topk_ids on HIP to prevent garbage MoE routing (DeepSeek-R1-MXFP4 accuracy regression) (#28378)
Co-authored-by: kangwangamd <kangwangamd@users.noreply.github.com>
2026-06-17 23:39:44 -07:00
7256ee9871 [AMD] Update test_aiter_allgather_amd.py data types alignment between benchmark aiter and custom all-reduce kernel (#27815)
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
2026-06-17 00:48:15 -07:00