10 Commits
Author SHA1 Message Date
amd-danli103andHAI e54009240a [AMD][DSV4] feat: enable DSpark with fp8 unified_kv on gfx950 (#38901)
Co-authored-by: HAI <hixiao@gmail.com>
2026-09-20 01:16:39 -07:00
amd-danli103 2305242f51 [AMD][DSV4] fix: skip compressed-KV metadata on the draft worker in the HIP radix backend (#40205) 2026-09-19 12:13:49 -07:00
amd-danli103 cd4dd81c22 [AMD][DSV4] fix: drop shadowing local get_exec import that breaks model startup on ROCm (#40186) 2026-09-18 13:28:38 -07:00
amd-danli103 5aa9b8fb3e [AMD][DSV4] feat: enable fp8 two-pool unified_kv on gfx950 (#37413) 2026-09-14 02:49:11 -07:00
amd-danli103 141febf329 [AMD] fix: use the hardware fp8 e4m3 convert on gfx950 (#37140)
Signed-off-by: amd-danli103 <danli103@amd.com>
2026-09-08 03:01:53 -07:00
amd-danli103andThomas Wang e9e9e37ddc [AMD] Restore SWA reprefill-tail on UnifiedRadixCache when HiCache is off (#32759)
Co-authored-by: Thomas Wang <thomawan@amd.com>
2026-09-07 14:27:53 -07:00
amd-danli103andThomas Wang d01812d89e [AMD] Optimize KIMI-K3 with Triton MLA decode kernel by tuning the stage-1 geometry for gfx950 (#34580)
Co-authored-by: Thomas Wang <thomawan@amd.com>
2026-08-17 21:16:14 -07:00
amd-danli103andamd-danli103 462b6171bd [AMD] Fix stale SWA ring buffer on radix prefix reuse for DeepSeek-V4 with unified_kv backend (#30339)
Co-authored-by: amd-danli103 <dan2.li@amd.com>
2026-07-09 11:28:22 -07:00
54e71506b3 [AMD][DSV4] Remove per-batch D2H syncs in MTP to avoid bubbles between 2 batches (#29420)
Co-authored-by: amd-danli103 <amd-danli103@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
2026-06-29 22:39:59 -07:00
c4ec39a785 [AMD] refactor sparse MLA decode kernel for Deepseek V4 triton backend (#28265)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: yichiche@amd.com <jacky.cheng>
2026-06-15 03:26:59 -07:00