Zhang, Jiejing
|
66f19f5c46
|
[AMD] Enable HiCache for GLM-5.2 MI355X throughput recipe (#40570)
|
2026-09-21 15:28:27 -07:00 |
|
 
|
993d1fccba
|
[ROCm] Widen the HiCache JIT copy rounds and enable the K-only host pool (#37152)
Co-authored-by: Xiaobo Chen <xiaobche@smci355-ccs-aus-n05-33.prov.aus.ccs.cpe.ice.amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-09-19 08:58:10 -07:00 |
|
Zhang, Jiejing
|
6331e43081
|
[AMD] Prefer HIP Top-K for GLM-5.x on ROCm (#39631)
|
2026-09-15 20:55:57 -07:00 |
|
 Zhang, Jiejingandfanxingran
|
f920be4b09
|
[AMD] GLM-5.2 NextN: cast draft fused MoE to per-channel FP8 (#39155)
Co-authored-by: fanxingran <xingran.fan@amd.com>
|
2026-09-15 18:56:48 -07:00 |
|
Zhang, Jiejing
|
288627e400
|
[AMD] Document GLM-5.2 MXFP4 recipe update on MI355X (#39230)
|
2026-09-12 15:57:15 -07:00 |
|
Zhang, Jiejing
|
d7c284b894
|
[AMD] Use the triton DSA backend for GLM-5.2 MXFP4 on MI355X (#39106)
|
2026-09-11 11:48:50 -07:00 |
|
Zhang, Jiejing
|
6cee9285a3
|
[ROCm] Make DSA indexer top-k exact with cooperative selection (#37591)
|
2026-09-05 23:55:56 -07:00 |
|