14 Commits
Author SHA1 Message Date
90cf471723 [AMD] [GLM-5.3-Flash Day 0] Support non-2048 top-k widths in the DSA page-table transform (#39340)
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
Co-authored-by: Kevin Mi <mikevin920@yahoo.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-21 19:50:17 -07:00
042b6a488f [AMD] [GLM-5.3-Flash Day 0] Enable zero-RoPE MHA prefill on ROCm (#39338)
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Kevin Mi <45493463+kevin-mii@users.noreply.github.com>
2026-09-21 17:38:07 -07:00
8bde82c0ad [AMD] [GLM-5.3-Flash Day 0] Build the fused DSA k-pool top-k JIT kernel on HIP (#39339)
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: Kevin Mi <mikevin920@yahoo.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-21 15:24:44 -07:00
Jacob0226andThomas Wang 5200508b0f [AMD][gfx95] Fill the chunked-prefill compute budget exactly (#32888)
Co-authored-by: Thomas Wang <thomawan@amd.com>
2026-09-14 00:35:15 -07:00
7bbd0ddeb5 [AMD] Quark shared-experts gate: recognise a trailing MTP layer (#36124)
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
2026-08-24 00:53:52 -07:00
92bce3d7bb [AMD] [GLM5] fp8 MLA absorbed bmm for GLM-5.2 on gfx950 (#30519)
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: sogalin_codegen <39478626+sogalin@users.noreply.github.com>
2026-08-17 02:15:16 -07:00
8e0499bd50 [AMD] [GLM5] Fuse shared-expert append into aiter grouped-topk (skip per-layer append kernel) (#31323)
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
2026-08-16 18:22:26 -07:00
f7cb328eb7 [AMD] [GLM5] Skip DSA decode indexer when kv_len <= index_topk (dense k-only fast path) (#31324)
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
2026-08-16 16:30:54 -07:00
Jacob0226 d44584e8d8 [AMD] [CI] Add GLM-5.1 MXFP4 TP2 accuracy gate (#26396) 2026-05-27 01:49:48 -07:00
Jacob0226andCursor bf5bc23431 [AMD] [CI] Add DeepSeek-R1-0528 FP8 HiCache GSM8K test on MI35x (#26395)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-27 01:48:38 -07:00
Jacob0226 fc20f5b114 [AMD] Skip redundant CatArrayBatchedCopy in GLM-5 NSA TileLang decode (#24125) 2026-05-13 02:55:28 -07:00
Jacob0226andClaude Opus 4.6 7e4e1dcd7a [AMD] Fuse RMSNorm + FP8 per-token quant for GLM-4.7-FP8 (#21403)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-10 22:45:31 -07:00
Jacob0226andClaude Opus 4.6 dd41764487 [AMD][HIP] NSA: bf16 passthrough from RMSNorm to eliminate FP8 dequantization (#22258)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-10 01:08:32 -07:00
Jacob0226andClaude Opus 4.6 7078e385ea [AMD] Add GLM-4.7-FP8 accuracy CI test for MI35x (#21534)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-28 00:28:56 -07:00