Sam (Kesen Li)
|
d346b214fb
|
feat(kv-cache): support SM100 NVFP4 GenMHA and speculative decoding (#36340)
|
2026-09-18 14:50:46 -07:00 |
|
Sam (Kesen Li)
|
8fc54d46ef
|
Fix MoE reduce-scatterv eligibility check (#32663)
|
2026-07-29 14:58:55 -07:00 |
|
Sam (Kesen Li)
|
ec6a3163b7
|
[Feature] Add FP4 KV Cache Design and support SM120 GPUs (#21601)
|
2026-07-17 14:49:43 -07:00 |
|
 
|
044649c23a
|
feat: Support flashinfer_cutedsl MoE runner with flashinfer alltoall backend (#22669)
Co-authored-by: Trevor Morris <tmorris@nvidia.com>
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
|
2026-05-20 00:36:26 -07:00 |
|
Sam (Kesen Li)
|
73e93bebd6
|
[1/4] NVFP4 KV cache: quantization strategy abstraction and kernel (#21954)
|
2026-04-29 01:45:48 -07:00 |
|
 Sam (Kesen Li)andXiaowei Wang
|
5b2e2750b5
|
Enable XQA for SM90 and SM120 (#17115)
Co-authored-by: Xiaowei Wang <100599594+xiaoweiw-nv@users.noreply.github.com>
|
2026-03-03 14:09:44 -08:00 |
|
Sam (Kesen Li)
|
5194eef88b
|
[Fix] A followup fix for TRTLLM BF16 MoE (#15303)
|
2026-02-26 17:14:20 -08:00 |
|
Sam (Kesen Li)
|
81449b4bee
|
Optimize GDN decode for Qwen3 Next (#17094)
|
2026-01-31 01:02:12 +08:00 |
|
Sam
|
4eda4194f2
|
[Fix] Disable trtllm moe backend for draft model for a qucik fix (#15002)
|
2025-12-12 23:58:47 -08:00 |
|
Sam
|
d7ed8a8c24
|
[NVIDIA] Enable TRTLLM BF16 MoE on Blackwell GPUs (#13798)
|
2025-12-11 22:56:13 -08:00 |
|
Sam
|
922756aaa1
|
[FIX] trtllm-moe-fp4-renorm for Qwen series models (#14350)
|
2025-12-04 12:52:21 -08:00 |
|
 SamandKaixi Hou
|
91e8dc371a
|
[Feat][NVFP4] Enable NVFP4 MoE for Qwen series models (eg. Qwen3-Next) #13761 (#13761)
Co-authored-by: Kaixi Hou <kaixih@nvidia.com>
|
2025-11-26 17:53:45 -07:00 |
|
Sam
|
e7e89349c9
|
Enable Flashinfer TRTLLM-GEN-MoE FP8 blockwise kernel for Qwen3-Next on Blackwell (#12543)
|
2025-11-13 19:44:44 +08:00 |
|
Sam
|
3594815a8b
|
Re-enable Flashinfer TRTLLM GEN MHA and Add Unit Test (#12885)
|
2025-11-10 20:17:43 -08:00 |
|