10 Commits
Author SHA1 Message Date
Jackey Hua 5a132c061b [Perf] trtllm_mla: reuse the fused fp8 KV/Q prepare on target verify (#39232) 2026-09-13 20:59:00 -07:00
Jackey HuaandClaude Opus 5 1af761a09a [SM12x] Default the fused MHC post+pre path on (#34019)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 19:51:21 +00:00
Jackey HuaandClaude Opus 5 00a219f6c9 [Quant] Keep the flashinfer_deepgemm FP8 GEMM to 1 <= M < 32 (#32843)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 00:36:13 +08:00
Jackey HuaandClaude Opus 5 9a0bd24bed model: serve bare Qwen3Model backbone natively as an embedding model (#32457)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 15:49:58 +08:00
Jackey HuaandClaude Opus 4.8 ee77a7d330 [Fix] DSA: size cudagraph page_table to req_to_token width (#29379)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-26 12:17:34 -07:00
Jackey Hua f444b5897b [Spec][1/N] Decoupled speculative decoding: IPC protocol + cross-process request id + server flags (#27634) 2026-06-23 17:04:45 -07:00
Jackey Hua 24c5d76f74 fix: per-sequence last-token embedding in EAGLE3/MTP draft for batched multimodal spec decoding (#27846) 2026-06-11 13:33:16 -07:00
Jackey Hua 465abadd3c Add fused moe triton config for Qwen3.5-397B-A17B-FP8 (#23682) 2026-04-24 18:35:32 -07:00
jackey hua 922fbc21e2 [Perf] Tune MiniMax M2 fused moe kernel on H100 GPU (#18851) 2026-02-15 15:30:52 +08:00
jackey hua 0998de088b [Perf] Tune Llama-4-Scout-17B-16E-Instruct fused moe kernel (#17891) 2026-01-28 14:06:46 -08:00