Shu Wang
|
1e8699fda3
|
[NVIDIA][comm] Merge EP+MoE-TP post-experts all-reduces into one _TP reduction (#32963)
|
2026-09-18 01:35:10 -07:00 |
|
   
|
1b77f498a0
|
[NVIDIA] Support flashinfer Mega Moe (#31470)
Co-authored-by: djns99 <40156487+djns99@users.noreply.github.com>
Co-authored-by: 云挚 <ningyunxiao.nyx@antgroup.com>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
|
2026-09-10 00:22:47 -07:00 |
|
Shu Wang
|
c7c03ec53b
|
[NVIDIA] Add flashinfer MNNVL backend for allreduce only (#30700)
|
2026-08-11 16:46:46 -07:00 |
|
Shu Wang
|
32685874f3
|
Reenable MNNVL backend for FlashInfer allreduce fusion (#23402)
|
2026-06-15 20:19:15 -07:00 |
|
 
|
0574d2b8a5
|
[NVIDIA] [GDN] Enable FlashInfer MTP verify on SM100+ (Blackwell) (#23273)
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-06-01 18:56:42 -07:00 |
|
 Shu Wangandzijiexia
|
106092123f
|
Update Qwen3-Coder docs_new NVIDIA guidance (#24435)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
|
2026-06-01 13:38:34 -07:00 |
|
 Shu WangandKhoa Pham
|
c67b287056
|
Enable trtllm_mha as gemma4 default attn backend. (#25006)
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
|
2026-05-17 14:58:12 -07:00 |
|
Shu Wang
|
5638d40f3a
|
[nvidia] Gemma4 nvfp4 fix (#22079)
|
2026-04-10 08:44:34 +08:00 |
|
Shu Wang
|
efebcab43e
|
Support skip-softmax attention (#19089)
|
2026-03-28 15:55:48 -07:00 |
|
Shu Wang
|
d35fea1b2b
|
[Nvidia] Add trtllm mnnvl allreduce with unified flashinfer allreduce fusion api (#12787)
|
2026-03-17 10:02:45 -07:00 |
|
Shu Wang
|
61de303f0a
|
Fix fallback to default tactic (flashinfer autotuner) with trtllm_fp4_block_scale_moe (#19189)
|
2026-03-06 15:15:04 -08:00 |
|
 Shu WangandBaizhou Zhang
|
43bdee703e
|
Fix Fp8 MTP layer a2a backend without EP. (#18515)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-03-04 16:28:10 -08:00 |
|
Shu Wang
|
1b75d0d1a9
|
Fix BatchMLAPagedAttentionWrapper query/qo_inptr mismatch for EAGLE (#15601)
|
2026-02-27 11:35:45 -08:00 |
|
Shu Wang
|
5c02217746
|
Inclusion of nvfp4 blockscale in EPLB Rebalance (#17158)
|
2026-01-19 17:45:27 +08:00 |
|
Shu Wang
|
665cb02003
|
Fix regression caused by fa3 block_table (#15009)
|
2025-12-12 19:11:09 -08:00 |
|
Shu Wang
|
a56f770277
|
Fix global scaling factor loading hang (#13484)
|
2025-11-21 16:07:06 -08:00 |
|
Shu Wang
|
6664083522
|
Replace [silu_and_mul_]scaled_fp4_group_quant by Flashinfer equivalent (#12376)
|
2025-11-13 00:26:00 -08:00 |
|
Shu Wang
|
7aa443903d
|
Fix nan in global scaling factor for large scale nvfp4 EP (#13162)
|
2025-11-12 15:32:23 -08:00 |
|
Shu Wang
|
82f39dc11d
|
Add mm_fp4 trtllm backend (#12406)
|
2025-11-05 14:31:46 -08:00 |
|
Shu Wang
|
1ba137e98f
|
Enable trtllm mla prefix extend (#10526)
|
2025-09-17 16:44:11 -07:00 |
|
Shu Wang
|
124097fc5b
|
enable prefix cache with dp (#10459)
|
2025-09-16 18:26:58 -07:00 |
|
 Shu WangandElfie Guo
|
36acd2ff16
|
Fix chunked prefix cache for nvfp4 (#10180)
Co-authored-by: Elfie Guo <elfieg@nvidia.com>
|
2025-09-12 03:20:30 -07:00 |
|
Shu Wang
|
3df05f4d6a
|
[NVIDIA] [3/N] Nvfp4 Masked Gemm: Add flashinfer grouped_gemm_nt_masked (#9199)
|
2025-09-11 20:18:43 -07:00 |
|
Shu Wang
|
288ae41f7a
|
[NVIDIA] Fix num_experts in modelopt_quant (#8811)
|
2025-08-06 14:35:07 -07:00 |
|
Shu Wang
|
b01eeb80f8
|
[NVIDIA]Fix local_num_experts for EP (#8779)
|
2025-08-04 22:01:14 -07:00 |
|
Shu Wang
|
ad4e58bf67
|
Support fp8 gemm for blackwell (#4558)
|
2025-03-20 12:40:28 -07:00 |
|