Commit Graph
8 Commits
Author SHA1 Message Date
kangwangamd 12eadf86f1 [AMD] Enable mamba JIT transfer kernel on ROCm (fix transfer_kv_mamba NameError) (#31741) 2026-08-02 16:40:17 -07:00
kangwangamd 4b52758c76 [AMD] Skip test_update_weights_from_disk on ROCm pending reload fix (#31924) (#31925) 2026-07-30 02:32:09 -07:00
kangwangamdandBingxu Chen 50c118704a [diffusion] disagg: handle numpy arrays in cross-role transfer field extraction (#31325)
Co-authored-by: Bingxu Chen <bingxche@amd.com>
2026-07-20 00:50:17 -07:00
kangwangamdandYC Yen-Ching Tseng 9f8e916131 [diffusion] post_training: run weight update under torch.inference_mode() (#31263)
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
2026-07-20 11:14:13 +08:00
kangwangamd f8eac995aa [diffusion] test: fix GLM-Image AR model-path to resolve local snapshot subfolder (#31313) 2026-07-15 23:57:34 -07:00
4d94e9471a [AMD] Relax allreduce-fusion residual accuracy tolerance to 1 bf16 ULP (#28226)
Co-authored-by: kangwangamd <kangwangamd@users.noreply.github.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
2026-06-18 19:18:53 -07:00
kangwangamdandkangwangamd 6309fb9abb [AMD] Fix Always mask padded topk_ids on HIP to prevent garbage MoE routing (DeepSeek-R1-MXFP4 accuracy regression) (#28378)
Co-authored-by: kangwangamd <kangwangamd@users.noreply.github.com>
2026-06-17 23:39:44 -07:00
7256ee9871 [AMD] Update test_aiter_allgather_amd.py data types alignment between benchmark aiter and custom all-reduce kernel (#27815)
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
2026-06-17 00:48:15 -07:00