Mohammad Miadh Angkad
|
d7207be156
|
Fix startup weight load after TorchAO removal (#34869)
|
2026-08-14 13:18:53 -07:00 |
|
 
|
6b94d39f13
|
[Model Loading] Overlap checkpoint staging with CUDA graph capture during startup (#32017)
Co-authored-by: Wenhui Zhu <wzhu59@asu.edu>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-08-13 12:26:25 -07:00 |
|
Mick
|
db75dfe10f
|
fix: always capture default prefill CUDA graph (#33352)
|
2026-08-08 19:24:49 +08:00 |
|
 Lianmin ZhengandItai Gat
|
dea2be5ae3
|
[CUDA Graph] Allow custom decode graph runners (#33553)
Co-authored-by: Itai Gat <itaigat.mail@gmail.com>
|
2026-08-04 12:48:56 -07:00 |
|
Peng Wu
|
e3d4f48e55
|
[Fix] missing max_context_len on HybridAttnBackend (#32690)
|
2026-07-31 19:43:09 +08:00 |
|
Mohammad Miadh Angkad
|
0bdd4730af
|
[CI] Fix failures on main (#32091)
|
2026-07-23 00:11:25 +08:00 |
|
 ThanhhaoandHao Phan
|
72c4ed1a3f
|
[Spec] DFlash: remove per-step host syncs so the CPU runs a full step ahead (spec-v2 overlap) (#31468)
Co-authored-by: Hao Phan <htphan@nvidia.com>
|
2026-07-17 23:22:07 -07:00 |
|
Mick
|
681c223570
|
refactor: wrap split backends once on full-attention backends (#31439)
|
2026-07-17 19:15:04 +08:00 |
|
Mick
|
24a8944e15
|
fix: enable Kimi multimodal breakable prefill cuda graph replay (#31391)
|
2026-07-17 19:13:54 +08:00 |
|
Mick
|
d9003dd452
|
fix: skip unsafe automatic prefill graph capture (#31204)
|
2026-07-16 09:27:38 +08:00 |
|
fzyzcjy
|
d15f6a9ac3
|
Introduce NgramEmbeddingManager component (#31154)
|
2026-07-14 15:58:08 +08:00 |
|