Commit Graph
8 Commits
Author SHA1 Message Date
3eeb7d37f9 [AMD] gfx950 assembly attention: length-aware split-KV for dynamic workload (#39172)
Co-authored-by: Zijie Chen <300606707+zijiecode@users.noreply.github.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>
2026-09-13 23:48:03 -07:00
a26273d668 [AMD] Quantize the bf16 MTP draft experts online to MXFP4 for Qwen3.5 (#38748)
Co-authored-by: Zijie Chen <300606707+zijiecode@users.noreply.github.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>
2026-09-10 13:00:41 -07:00
92d831d3d7 [AMD] Parallelize aiter spec-decode KV index building over token blocks (#37659)
Co-authored-by: Zijie Chen <300606707+zijiecode@users.noreply.github.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>
2026-09-09 17:53:53 -07:00
zijiecandZijie Chen e634ba78a4 [AMD] gfx950 assembly attention for EAGLE verify, draft extend and decode (#37465)
Co-authored-by: Zijie Chen <300606707+zijiecode@users.noreply.github.com>
2026-09-08 03:10:09 -07:00
zijiecandjacky.cheng 7f2ee22b70 [AMD] Fix eager metadata for AITER EAGLE draft extend (#36915)
Co-authored-by: jacky.cheng <yichiche@amd.com>
2026-08-28 22:56:38 -07:00
acc918b3ec [AMD] Qwen3.5 ASM FMHA chunked-prefill context attention (#36758)
Co-authored-by: Zijie Chen <300606707+zijiecode@users.noreply.github.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>
2026-08-27 23:37:45 -07:00
ba1d980b35 [AMD] Accelerate AITER unified-attention decode with scaled FP8 Q (#31856)
Co-authored-by: jacky.cheng <yichiche@amd.com>
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
2026-08-13 23:42:25 -07:00
zijiec 93e9db5eb8 [Fix][Qwen]: fused shared-expert detection PP-safe protection (#34447) 2026-08-11 18:50:05 -07:00