13 Commits
Author SHA1 Message Date
eigenandYingyi Huang 516cfbd362 Fix UNO test adapter subdirectory resolution (#37872)
Co-authored-by: Yingyi Huang <averyh@nvidia.com>
2026-09-03 18:07:09 -07:00
eigenandAvery Huang 3ab9afd653 fix: piecewise_cuda_graph get correct qo_indptr (#21452)
Co-authored-by: Avery Huang <averyh@nvidia.com>
2026-03-28 15:57:29 -07:00
ac1f2928ae feat: add fast_decode_plan from flashinfer, flashinfer to 0.4.0rc3 (#10760)
Co-authored-by: Zihao Ye <yezihhhao@gmail.com>
Co-authored-by: Sleepcoo <Sleepcoo@gmail.com>
2025-10-01 02:56:13 -07:00
eigen 70c0c1f926 fix: trtllm-gen attention take zero-init workspace (#10330) 2025-09-11 14:35:23 -07:00
eigen b0fcbb74d0 [DOC]: some minor updates (#10134) 2025-09-07 14:58:15 -07:00
eigenandYineng Zhang 4dbf43601d fix: zero_init buffer (#9065)
Co-authored-by: Yineng Zhang <me@zhyncs.com>
2025-08-14 02:39:09 -07:00
eigen faa25df1ae feat: update flashinfer ar oneshot params (#8687) 2025-08-09 00:51:27 -07:00
eigenandaveryhuang 9c7e392465 bench: add attention sink op benchmark, triton and trtllm-gen [B200] (#8932)
Co-authored-by: averyhuang <averyh@nvidia.com>
2025-08-08 00:16:23 -07:00
eigen 08fab2b0c4 minor: global workspace buffer for trtllm-gen mha from flashinfer (#8952) 2025-08-08 00:12:12 -07:00
eigenandaveryhuang 6ad6c8c9e6 feat: openai oss attention sink support with trtllm-gen backend #8825 (#8834)
Co-authored-by: averyhuang <averyh@nvidia.com>
2025-08-06 19:18:27 -07:00
eigenandBaizhou Zhang 40e3b2beeb feat: add trtllm-gen mha from direct call (#8782)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2025-08-05 03:28:39 -07:00
eigen 20beb3702b feat: add return hidden_states at async generation (#7507) 2025-06-25 02:10:09 -07:00
8f783c1943 [Model Support] unsloth/Phi-4-mini bnb model (#4982)
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: Chayenne <zhaochen20@outlook.com>
Co-authored-by: Yineng Zhang <me@zhyncs.com>
2025-04-16 19:58:20 -07:00