  
|
3a64faa1f2
|
Fix disagg PP MTP for GLM-5.2 (#39378)
Co-authored-by: Julien Lin <jullin@nvidia.com>
Co-authored-by: YAMY1234 <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
|
2026-09-19 13:55:07 -07:00 |
|
Po-Han Huang (NVIDIA)
|
a2a261963f
|
[CI] Fix SWA decode radix cache NIXL import (#39600)
|
2026-09-15 21:01:05 +08:00 |
|
Po-Han Huang (NVIDIA)
|
51c8581a26
|
[Rust] Gate health on startup warmup completion (#37994)
|
2026-09-09 15:55:13 -07:00 |
|
 Po-Han Huang (NVIDIA)andYAMY
|
96d91ef926
|
[Model] Support GLM-5.3 Flash NVFP4 loading (#38621)
Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com>
|
2026-09-09 13:51:08 -07:00 |
|
Po-Han Huang (NVIDIA)
|
5097f9ac95
|
Disable Hopper GLM shared-expert fusion for modelopt_fp4 Marlin (#37325)
|
2026-09-08 06:16:36 -07:00 |
|
Po-Han Huang (NVIDIA)
|
991368d880
|
Fix Muse Glimmer ModelOpt mixed weight mapping (#37510)
|
2026-09-04 20:53:55 -07:00 |
|
 Po-Han Huang (NVIDIA)andkpham-sgl
|
d122ca99b2
|
[Bugfix] Fix Llama 4 FA3 local attention with paged KV cache (#32902)
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
|
2026-09-03 23:15:23 -07:00 |
|
Po-Han Huang (NVIDIA)
|
d585cec4bd
|
Fix nondeterministic FlashInfer GDN alignment test (#37343)
|
2026-09-01 23:46:58 -07:00 |
|
Po-Han Huang (NVIDIA)
|
e57e934bcc
|
[Bugfix] Accept int64 top-k IDs in FlashInfer routed MoE packer (#32882)
|
2026-09-01 14:17:46 -07:00 |
|
Po-Han Huang (NVIDIA)
|
6ff2fe6a6a
|
CI: split JIT kernel unit tests into two partitions (#36775)
|
2026-08-28 09:32:58 -07:00 |
|
Po-Han Huang (NVIDIA)
|
c608f9bf75
|
[CI] Fix Q8KV8 sparse prefill test fixture (#36639)
|
2026-08-27 01:00:22 -07:00 |
|
Po-Han Huang (NVIDIA)
|
6f69f927da
|
[Scheduler] Add configurable decode interval after prefill (#35017)
|
2026-08-19 12:01:36 -07:00 |
|
Po-Han Huang (NVIDIA)
|
bae29f716a
|
Fix inference mode mismatch in FlashInfer warmup (#33788)
|
2026-08-06 16:11:41 -07:00 |
|
Po-Han Huang (NVIDIA)
|
7eb27372b3
|
fix(server): capture legal multi-request prefill CUDA graph batches (#30206)
|
2026-08-03 16:56:59 -07:00 |
|
Po-Han Huang (NVIDIA)
|
5df193b4ac
|
[Speculative Decoding] Fix GPT-OSS EAGLE3 hidden states (#32334)
|
2026-07-31 11:58:49 -07:00 |
|
Po-Han Huang (NVIDIA)
|
5d004a20c5
|
Fix FlashInfer A2A top-k ID dtype (#29929)
|
2026-07-15 17:56:11 -07:00 |
|
Po-Han Huang (NVIDIA)
|
cfc3d0555e
|
Fix ModelOpt NVFP4 scalar scales for merged linears (#29151)
|
2026-07-13 16:14:20 -07:00 |
|
Po-Han Huang (NVIDIA)
|
17cce6a85f
|
Fix shared logits buffer for reduced-vocab draft models (#29943)
|
2026-07-02 16:08:56 -07:00 |
|
Po-Han Huang (NVIDIA)
|
72812db138
|
Avoid dynamic Q quantization in trtllm_mha (#29423)
|
2026-06-26 15:51:15 -07:00 |
|
 Po-Han Huang (NVIDIA)andClaude Opus 4.6
|
ada52e5972
|
[Docs] Move ptxas sm_103a workaround into For CUDA 13 section (#22852)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-14 22:30:21 -07:00 |
|