Commit Graph
46 Commits
Author SHA1 Message Date
Mohammad Miadh Angkad 76b1efaffd Fix SMG service discovery Clippy lint (#26034) 2026-05-22 10:51:13 +08:00
Mohammad Miadh Angkad a449ee4822 [Deps] Use cu13 extra for nvidia cutlass dsl (#25576) 2026-05-21 10:31:27 +08:00
Mohammad Miadh AngkadandBaizhou Zhang 8cb337c8ea [Bugfix] Temporarily skip TRTLLM attention on (G)B300 (SM103) to avoid high-concurrency hang (#21906)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-04-03 14:19:13 -07:00
Mohammad Miadh Angkad 883ba640b2 [CI] Remove more redundant PCG tests (#21554) 2026-03-31 16:25:30 -07:00
Mohammad Miadh Angkad dd9c9c1b8e Add explicit disable flag for FlashInfer allreduce fusion (#21446) 2026-03-31 00:15:44 -07:00
Mohammad Miadh Angkad 2acdda1d85 [Fix] Remove redundant allreduce fusion block and skip TP=1 (#20621) 2026-03-29 12:30:40 -07:00
Mohammad Miadh Angkad eaf392b9cc Remove redundant DeepSeek V3 FP4 PCG test (#21485) 2026-03-26 21:52:47 -07:00
Mohammad Miadh Angkadandelvischenv bbe25b2412 Use FlashInfer tinygemm for GPT-OSS MoE router on SM90+ (#20755)
Co-authored-by: elvischenv <219235043+elvischenv@users.noreply.github.com>
2026-03-24 15:00:18 -07:00
Mohammad Miadh Angkad 4fbb311234 [Fix][Eval] Keep --dataset-path scoped to longbench_v2 (#21156) 2026-03-24 02:25:11 -07:00
Mohammad Miadh Angkad d8a5b1dbaf [Bugfix] Work around FlashInfer unified transport issue on GB (#20039) 2026-03-22 21:10:25 -07:00
Mohammad Miadh Angkad 3d749c49ca [JIT Kernel] Fix NVFP4 multi-arch compilation failure (#20874) 2026-03-20 10:30:04 +08:00
Mohammad Miadh Angkad 29ced9c162 [UX] Suppress noisy httpx/httpcore INFO logs (#20944) 2026-03-19 10:58:41 -07:00
Mohammad Miadh Angkad 3879c466b4 [CI] Add Nemotron 3 Super 120B nightly 8-GPU tests (#20616) 2026-03-15 18:03:20 -07:00
Mohammad Miadh Angkad 3e643967e6 [CI] Add Nemotron 3 Super 120B CI tests for BF16 and NVFP4 (#20575) 2026-03-14 12:30:27 -07:00
Mohammad Miadh Angkad 75a7879fd4 [Model] Support Nemotron 3 Super NVFP4 (#20407) 2026-03-14 00:56:26 -07:00
Mohammad Miadh Angkad ca997b7ba9 Add min_p and chat-template kwargs support to run_eval (#19571) 2026-03-09 14:53:09 -07:00
Mohammad Miadh Angkad f88acf8780 [JIT Kernel] Reland NVFP4 kernels to JIT (#20012) 2026-03-07 10:31:08 +08:00
Mohammad Miadh Angkad 8cdb7e1fd4 [CI] Add GPT-OSS test for SM120 (#20056) 2026-03-06 11:43:04 -08:00
Mohammad Miadh Angkad 759700c808 Fix SM120 triton_kernels MXFP4 block_k for GPT-OSS (#20040) 2026-03-06 10:53:08 -08:00
Mohammad Miadh Angkad 41fd53fe37 Fix profile_activities parameter name in bench_one_batch_server_internal.py (#19954) 2026-03-05 10:34:06 -08:00
Mohammad Miadh Angkad 2bdd89a6cd [Kernel Slimming] Migrate NVFP4 kernels to JIT (#19437) 2026-03-05 15:22:28 +08:00
Mohammad Miadh Angkad 1b76eb9361 [Doc] Update version references and add automation (#18409) 2026-03-04 09:51:46 -08:00
Mohammad Miadh Angkad 09fa012ba7 Fix /health regression from early prebound socket listen (#19805) 2026-03-03 23:00:46 -08:00
Mohammad Miadh Angkad 6822941514 [FlashInfer] Bump FlashInfer version from 0.6.3 to 0.6.4 (#19005) 2026-03-02 16:12:09 -08:00
Mohammad Miadh Angkad 3f9fc8b848 [Qwen3.5] Fix missing quant_config in Qwen3VL (#19291) 2026-03-02 14:07:51 -08:00
Mohammad Miadh AngkadandXinyuan Tong 9c81ce4707 [Anthropic API] Preserve image content in tool_result conversion (#19233)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-02-28 12:07:22 -08:00
Mohammad Miadh Angkad 35ef38c61b Remove gpt-oss hybrid swa gate for trtllm_mha (#19079) 2026-02-27 10:30:00 -08:00
Mohammad Miadh Angkad 671b595570 Fix trtllm_mha fp8 SWA KV index translation (#19107) 2026-02-25 17:02:17 +08:00
Mohammad Miadh Angkad 7b0fb43c7a [FlashInfer] Switch FlashInfer allreduce fusion to unified API (#18341) 2026-02-22 00:07:16 +08:00
Mohammad Miadh Angkadandrainj-me f23a23cc05 Fix NSA FP8 KV cache path for both-trtllm MHA one-shot (#18931)
Co-authored-by: rainj-me <96632942+rainj-me@users.noreply.github.com>
2026-02-20 22:00:09 +08:00
Mohammad Miadh Angkad 2f592c3b18 [Doc] Add flashinfer_deepgemm to --fp8-gemm-backend (#18982) 2026-02-18 14:45:47 -05:00
Mohammad Miadh Angkad 90a0d66e1e [Tiny] Fix assert syntax warning in compressed_tensors_w4a4_mxint4_moe.py (#18899) 2026-02-17 12:54:30 +08:00
Mohammad Miadh Angkad b86c6491fa [Perf] ~9.5x faster Blackwell MXFP4 MoE weight loading (#18858) 2026-02-16 19:47:09 +08:00
Mohammad Miadh Angkad 8290171f52 [CI] Remove --mem-fraction-static 0.93 from gpt-oss test (#18869) 2026-02-16 09:24:11 +08:00
Mohammad Miadh Angkad b1b69ae0a9 Add CI permissions (#18847) 2026-02-15 08:24:36 +08:00
Mohammad Miadh Angkad 1be41e9036 [FlashInfer] Bump FlashInfer version from 0.6.2 to 0.6.3 (#18448) 2026-02-14 07:43:33 +08:00
Mohammad Miadh Angkad 071bf2ce09 [Kimi-K2.5] Fix missing quant_config in KimiK25 (#18440) 2026-02-08 12:02:45 -08:00
Mohammad Miadh Angkad 7b83659310 fix: fix NVFP4 Kimi-K2.5 weight mapping and exclude list (#18370) 2026-02-08 10:23:48 +08:00
Mohammad Miadh Angkad fddef76619 [Doc] Fix outdated --fp4-gemm-backend documentation (#18350) 2026-02-07 20:42:47 +08:00
Mohammad Miadh AngkadandBaizhou Zhang c47c2f9466 [Doc] Update CUDA 13 install guide to install torch first (#18404)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-02-07 18:04:37 +08:00
Mohammad Miadh Angkad efbf39583e Add MoE fused config for Qwen3-Coder-Next-FP8 on H100 TP=2 (#18195) 2026-02-04 13:36:35 -08:00
Mohammad Miadh Angkad 25508d11c0 [Docker] Remove hardcoded America/Los_Angeles timezone, default to UTC (#18121) 2026-02-02 23:22:15 -08:00
Mohammad Miadh Angkad 6f6b9c6e42 [Perf] Use safetensors load_file in multithread loader (#18124) 2026-02-02 23:21:13 -08:00
Mohammad Miadh Angkad 9ac4dcada4 [Tiny] Fix grammar in shared experts fusion log messages (#18043) 2026-01-31 13:14:25 -08:00
Mohammad Miadh Angkad d0d9cecd1b Fix cuBLAS >=12.9 detection for cu12/cu13 package naming (#17766) 2026-01-31 12:01:52 +08:00
Mohammad Miadh Angkad 1674b9ef44 [DeepSeek-V3.2] Fix TRT-LLM NSA in target_verify/draft_extend (#17662) 2026-01-25 13:10:14 +08:00