Commit Graph
30 Commits
Author SHA1 Message Date
Mohammad Miadh Angkad f88acf8780 [JIT Kernel] Reland NVFP4 kernels to JIT (#20012) 2026-03-07 10:31:08 +08:00
Mohammad Miadh Angkad 8cdb7e1fd4 [CI] Add GPT-OSS test for SM120 (#20056) 2026-03-06 11:43:04 -08:00
Mohammad Miadh Angkad 759700c808 Fix SM120 triton_kernels MXFP4 block_k for GPT-OSS (#20040) 2026-03-06 10:53:08 -08:00
Mohammad Miadh Angkad 41fd53fe37 Fix profile_activities parameter name in bench_one_batch_server_internal.py (#19954) 2026-03-05 10:34:06 -08:00
Mohammad Miadh Angkad 2bdd89a6cd [Kernel Slimming] Migrate NVFP4 kernels to JIT (#19437) 2026-03-05 15:22:28 +08:00
Mohammad Miadh Angkad 1b76eb9361 [Doc] Update version references and add automation (#18409) 2026-03-04 09:51:46 -08:00
Mohammad Miadh Angkad 09fa012ba7 Fix /health regression from early prebound socket listen (#19805) 2026-03-03 23:00:46 -08:00
Mohammad Miadh Angkad 6822941514 [FlashInfer] Bump FlashInfer version from 0.6.3 to 0.6.4 (#19005) 2026-03-02 16:12:09 -08:00
Mohammad Miadh Angkad 3f9fc8b848 [Qwen3.5] Fix missing quant_config in Qwen3VL (#19291) 2026-03-02 14:07:51 -08:00
Mohammad Miadh AngkadandXinyuan Tong 9c81ce4707 [Anthropic API] Preserve image content in tool_result conversion (#19233)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-02-28 12:07:22 -08:00
Mohammad Miadh Angkad 35ef38c61b Remove gpt-oss hybrid swa gate for trtllm_mha (#19079) 2026-02-27 10:30:00 -08:00
Mohammad Miadh Angkad 671b595570 Fix trtllm_mha fp8 SWA KV index translation (#19107) 2026-02-25 17:02:17 +08:00
Mohammad Miadh Angkad 7b0fb43c7a [FlashInfer] Switch FlashInfer allreduce fusion to unified API (#18341) 2026-02-22 00:07:16 +08:00
Mohammad Miadh Angkadandrainj-me f23a23cc05 Fix NSA FP8 KV cache path for both-trtllm MHA one-shot (#18931)
Co-authored-by: rainj-me <96632942+rainj-me@users.noreply.github.com>
2026-02-20 22:00:09 +08:00
Mohammad Miadh Angkad 2f592c3b18 [Doc] Add flashinfer_deepgemm to --fp8-gemm-backend (#18982) 2026-02-18 14:45:47 -05:00
Mohammad Miadh Angkad 90a0d66e1e [Tiny] Fix assert syntax warning in compressed_tensors_w4a4_mxint4_moe.py (#18899) 2026-02-17 12:54:30 +08:00
Mohammad Miadh Angkad b86c6491fa [Perf] ~9.5x faster Blackwell MXFP4 MoE weight loading (#18858) 2026-02-16 19:47:09 +08:00
Mohammad Miadh Angkad 8290171f52 [CI] Remove --mem-fraction-static 0.93 from gpt-oss test (#18869) 2026-02-16 09:24:11 +08:00
Mohammad Miadh Angkad b1b69ae0a9 Add CI permissions (#18847) 2026-02-15 08:24:36 +08:00
Mohammad Miadh Angkad 1be41e9036 [FlashInfer] Bump FlashInfer version from 0.6.2 to 0.6.3 (#18448) 2026-02-14 07:43:33 +08:00
Mohammad Miadh Angkad 071bf2ce09 [Kimi-K2.5] Fix missing quant_config in KimiK25 (#18440) 2026-02-08 12:02:45 -08:00
Mohammad Miadh Angkad 7b83659310 fix: fix NVFP4 Kimi-K2.5 weight mapping and exclude list (#18370) 2026-02-08 10:23:48 +08:00
Mohammad Miadh Angkad fddef76619 [Doc] Fix outdated --fp4-gemm-backend documentation (#18350) 2026-02-07 20:42:47 +08:00
Mohammad Miadh AngkadandBaizhou Zhang c47c2f9466 [Doc] Update CUDA 13 install guide to install torch first (#18404)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-02-07 18:04:37 +08:00
Mohammad Miadh Angkad efbf39583e Add MoE fused config for Qwen3-Coder-Next-FP8 on H100 TP=2 (#18195) 2026-02-04 13:36:35 -08:00
Mohammad Miadh Angkad 25508d11c0 [Docker] Remove hardcoded America/Los_Angeles timezone, default to UTC (#18121) 2026-02-02 23:22:15 -08:00
Mohammad Miadh Angkad 6f6b9c6e42 [Perf] Use safetensors load_file in multithread loader (#18124) 2026-02-02 23:21:13 -08:00
Mohammad Miadh Angkad 9ac4dcada4 [Tiny] Fix grammar in shared experts fusion log messages (#18043) 2026-01-31 13:14:25 -08:00
Mohammad Miadh Angkad d0d9cecd1b Fix cuBLAS >=12.9 detection for cu12/cu13 package naming (#17766) 2026-01-31 12:01:52 +08:00
Mohammad Miadh Angkad 1674b9ef44 [DeepSeek-V3.2] Fix TRT-LLM NSA in target_verify/draft_extend (#17662) 2026-01-25 13:10:14 +08:00