20 Commits
Author SHA1 Message Date
3a64faa1f2 Fix disagg PP MTP for GLM-5.2 (#39378)
Co-authored-by: Julien Lin <jullin@nvidia.com>
Co-authored-by: YAMY1234 <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
2026-09-19 13:55:07 -07:00
Po-Han Huang (NVIDIA) a2a261963f [CI] Fix SWA decode radix cache NIXL import (#39600) 2026-09-15 21:01:05 +08:00
Po-Han Huang (NVIDIA) 51c8581a26 [Rust] Gate health on startup warmup completion (#37994) 2026-09-09 15:55:13 -07:00
Po-Han Huang (NVIDIA)andYAMY 96d91ef926 [Model] Support GLM-5.3 Flash NVFP4 loading (#38621)
Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com>
2026-09-09 13:51:08 -07:00
Po-Han Huang (NVIDIA) 5097f9ac95 Disable Hopper GLM shared-expert fusion for modelopt_fp4 Marlin (#37325) 2026-09-08 06:16:36 -07:00
Po-Han Huang (NVIDIA) 991368d880 Fix Muse Glimmer ModelOpt mixed weight mapping (#37510) 2026-09-04 20:53:55 -07:00
Po-Han Huang (NVIDIA)andkpham-sgl d122ca99b2 [Bugfix] Fix Llama 4 FA3 local attention with paged KV cache (#32902)
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
2026-09-03 23:15:23 -07:00
Po-Han Huang (NVIDIA) d585cec4bd Fix nondeterministic FlashInfer GDN alignment test (#37343) 2026-09-01 23:46:58 -07:00
Po-Han Huang (NVIDIA) e57e934bcc [Bugfix] Accept int64 top-k IDs in FlashInfer routed MoE packer (#32882) 2026-09-01 14:17:46 -07:00
Po-Han Huang (NVIDIA) 6ff2fe6a6a CI: split JIT kernel unit tests into two partitions (#36775) 2026-08-28 09:32:58 -07:00
Po-Han Huang (NVIDIA) c608f9bf75 [CI] Fix Q8KV8 sparse prefill test fixture (#36639) 2026-08-27 01:00:22 -07:00
Po-Han Huang (NVIDIA) 6f69f927da [Scheduler] Add configurable decode interval after prefill (#35017) 2026-08-19 12:01:36 -07:00
Po-Han Huang (NVIDIA) bae29f716a Fix inference mode mismatch in FlashInfer warmup (#33788) 2026-08-06 16:11:41 -07:00
Po-Han Huang (NVIDIA) 7eb27372b3 fix(server): capture legal multi-request prefill CUDA graph batches (#30206) 2026-08-03 16:56:59 -07:00
Po-Han Huang (NVIDIA) 5df193b4ac [Speculative Decoding] Fix GPT-OSS EAGLE3 hidden states (#32334) 2026-07-31 11:58:49 -07:00
Po-Han Huang (NVIDIA) 5d004a20c5 Fix FlashInfer A2A top-k ID dtype (#29929) 2026-07-15 17:56:11 -07:00
Po-Han Huang (NVIDIA) cfc3d0555e Fix ModelOpt NVFP4 scalar scales for merged linears (#29151) 2026-07-13 16:14:20 -07:00
Po-Han Huang (NVIDIA) 17cce6a85f Fix shared logits buffer for reduced-vocab draft models (#29943) 2026-07-02 16:08:56 -07:00
Po-Han Huang (NVIDIA) 72812db138 Avoid dynamic Q quantization in trtllm_mha (#29423) 2026-06-26 15:51:15 -07:00
Po-Han Huang (NVIDIA)andClaude Opus 4.6 ada52e5972 [Docs] Move ptxas sm_103a workaround into For CUDA 13 section (#22852)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 22:30:21 -07:00