Commit Graph
14 Commits
Author SHA1 Message Date
Sam (Kesen Li) d346b214fb feat(kv-cache): support SM100 NVFP4 GenMHA and speculative decoding (#36340) 2026-09-18 14:50:46 -07:00
Sam (Kesen Li) 8fc54d46ef Fix MoE reduce-scatterv eligibility check (#32663) 2026-07-29 14:58:55 -07:00
Sam (Kesen Li) ec6a3163b7 [Feature] Add FP4 KV Cache Design and support SM120 GPUs (#21601) 2026-07-17 14:49:43 -07:00
044649c23a feat: Support flashinfer_cutedsl MoE runner with flashinfer alltoall backend (#22669)
Co-authored-by: Trevor Morris <tmorris@nvidia.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-20 00:36:26 -07:00
Sam (Kesen Li) 73e93bebd6 [1/4] NVFP4 KV cache: quantization strategy abstraction and kernel (#21954) 2026-04-29 01:45:48 -07:00
Sam (Kesen Li)andXiaowei Wang 5b2e2750b5 Enable XQA for SM90 and SM120 (#17115)
Co-authored-by: Xiaowei Wang <100599594+xiaoweiw-nv@users.noreply.github.com>
2026-03-03 14:09:44 -08:00
Sam (Kesen Li) 5194eef88b [Fix] A followup fix for TRTLLM BF16 MoE (#15303) 2026-02-26 17:14:20 -08:00
Sam (Kesen Li) 81449b4bee Optimize GDN decode for Qwen3 Next (#17094) 2026-01-31 01:02:12 +08:00
Sam 4eda4194f2 [Fix] Disable trtllm moe backend for draft model for a qucik fix (#15002) 2025-12-12 23:58:47 -08:00
Sam d7ed8a8c24 [NVIDIA] Enable TRTLLM BF16 MoE on Blackwell GPUs (#13798) 2025-12-11 22:56:13 -08:00
Sam 922756aaa1 [FIX] trtllm-moe-fp4-renorm for Qwen series models (#14350) 2025-12-04 12:52:21 -08:00
SamandKaixi Hou 91e8dc371a [Feat][NVFP4] Enable NVFP4 MoE for Qwen series models (eg. Qwen3-Next) #13761 (#13761)
Co-authored-by: Kaixi Hou <kaixih@nvidia.com>
2025-11-26 17:53:45 -07:00
Sam e7e89349c9 Enable Flashinfer TRTLLM-GEN-MoE FP8 blockwise kernel for Qwen3-Next on Blackwell (#12543) 2025-11-13 19:44:44 +08:00
Sam 3594815a8b Re-enable Flashinfer TRTLLM GEN MHA and Add Unit Test (#12885) 2025-11-10 20:17:43 -08:00