Commit Graph
32 Commits
Author SHA1 Message Date
Ziang Li c46bf5e990 [MoE] Disable FlashInfer fused finalize by default for numerical accuracy (#40105) 2026-09-18 00:15:57 -07:00
Ziang Li dbd4302bbb Keep NVFP4 blockscale swizzle padding on the input device (#39141) 2026-09-11 22:18:02 -07:00
Ziang Li 20ca564bf7 Add zianglih as online NVFP4 and DSA Top-K code owner (#33624) 2026-09-07 16:08:11 -07:00
Ziang Li 5edcd0a445 [FlashInfer V0.6.18] feat(dsv4): support --dsa-topk-backend flashinfer with fused top-k (#33237) 2026-09-01 01:18:10 -07:00
Ziang Li 9a85473a89 [FlashInfer v0.6.18] add FlashInfer CuTe DSL NVFP4 W4A16 mode (#35120) 2026-08-31 18:47:30 -07:00
Ziang Li 4f997a432a test: re-enable FlashInfer per-token NVFP4 coverage (#36985) 2026-08-30 23:44:09 -07:00
Ziang Li 3c9febc68b [Spec][DSA] Add --speculative-dsa-topk-backend (#36313) 2026-08-25 23:35:03 -07:00
Ziang Li a9654eacc1 fix(dsa): use FlashInfer fused top-k for packed PAGED rows (#33006) 2026-08-14 15:22:18 -07:00
Ziang Li 9d34c2809f [FlashInfer v0.6.16] Support FlashInfer CuTe DSL NVFP4 MoE quantization (#28354) 2026-08-13 17:33:46 -07:00
Ziang LiandBrayden Zhong 4ad990ba7d [ModelOpt FP4] Support online MoE weight quantization (#33115)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2026-08-06 11:01:55 -07:00
Ziang Li 988c6e6aeb Pin online NVFP4 4over6 quantization settings (#33621) 2026-08-05 21:29:32 +00:00
Ziang Li 937c77cf50 [Fix] Clear stale FlashInfer BF16 MoE index cache (#33016) 2026-07-31 00:35:15 -07:00
Ziang LiandParth Chadha 0aefba7283 fix(dsa): correct packed FlashInfer top-k and backend selection semantics (#32490)
Co-authored-by: Parth Chadha <parth@humansand.ai>
2026-07-30 20:20:16 -07:00
Ziang Li 2cf2920d07 [FlashInfer v0.6.13] Use CuTe DSL backend for FlashInfer per-token NVFP4 quantization (#28220) 2026-07-13 22:37:46 +08:00
Ziang Li 3fb65ebabd [RL] Fix FlashInfer TRTLLM MXFP8 dense weight layout (#28459) 2026-06-17 10:33:35 +00:00
Ziang Li 01f10acd06 Implement online nvfp4 quantization (#26083) 2026-06-10 00:26:51 -07:00
Ziang Li 4cfebbb95f [FlashInfer v0.6.12] Support FlashInfer 4over6 NVFP4 (#25239) 2026-06-04 14:35:07 -07:00
Ziang Li 2b1e53c98d [RL] Fix FP8 skip matching for trailing-dot prefixes (#26287) 2026-05-26 20:30:08 +00:00
Ziang Li 2b9dd9c8b3 [FlashInfer v0.6.10] [RL] [DSv32] [GLM-5] Add --dsa-topk-backend and integrate FlashInfer and pytorch topk (#22851) 2026-05-25 13:08:03 -07:00
Ziang Li 78cb38ed5e [FlashInfer v0.6.11] [RL] Support FlashInfer per-token NVFP4 MoE (#22918) 2026-05-19 01:04:48 -07:00
Ziang LiandBrayden Zhong 1758856762 [CI] Fix mxfp8 TrtllmGenMoe test (#23125)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2026-04-24 09:02:11 +00:00
Ziang Li 5593539942 [RL] Refactor NVFP4 shuffling/swizzling to in-place replacement (#22204) 2026-04-12 19:08:45 -07:00
Ziang Li 31453bb76a [RL] Fix weight update for mxfp8 flashinfer_cutlass gemm backend (#22484) 2026-04-12 13:02:17 +00:00
Ziang Li 78043d4448 [Misc] [MXFP8] Drop sm100 mxfp8 warning (#21881) 2026-04-11 11:11:28 +00:00
Ziang Li 990c7590b8 [RL] Support mxfp8 DeepSeek V3 (#21280) 2026-04-03 21:57:45 -07:00
Ziang Li a19ef3a615 [FlashInver v0.6.7] Integrate flashinfer_trtllm mxfp8 gemm (#21576) 2026-04-01 15:55:06 -04:00
Ziang Li 1a4b383fac [CI] [FlashInfer v0.6.7] Use offline quantized checkpoint for MXFP8 Gemm tests (#21625) 2026-03-29 22:47:46 -07:00
Ziang Li ce0541404f [FlashInfer v0.6.6][RL] Support fp8-last-n-bf16 RL for flashinfer_trtllm_routed moe backend (#20214) 2026-03-22 11:17:01 -07:00
Ziang Li 76ee4bb98c [FlashInfer v0.6.4] [RL] Integrate FlashInfer mxfp8 gemm, MoE, and routed MoE (#19537) 2026-03-10 15:37:57 -07:00
Ziang Li 0e86977811 [RL] Support per-layer mixed FP8/BF16 serving for FP8 checkpoints (#18742) 2026-03-01 21:59:22 +08:00
Ziang Li 9469ad089b Fix nvfp4 weight update (#18085) 2026-02-27 14:55:08 -08:00
Ziang Li eddf193292 [DSv32] [GLM5] Improve Model Quality by Avoiding FP32 Precision Loss in weights_proj (#19041) 2026-02-22 16:20:51 +08:00