 
|
325ab245a1
|
[DCP] Allow fi_a2a on single-node systems Blackwell without MNNVL fabric ( ex B200 B300) (#37767)
Co-authored-by: Hao Phan <htphan@nvidia.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
|
2026-09-08 07:55:36 -07:00 |
|
    
|
35e25f5356
|
[Feature] DCP: A2A + FlashInfer-MNNVL comm backends and q-replicate (Helix) (#21637)
Co-authored-by: Hao Phan <htphan@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-07-24 00:21:53 -07:00 |
|
 ThanhhaoandHao Phan
|
72c4ed1a3f
|
[Spec] DFlash: remove per-step host syncs so the CPU runs a full step ahead (spec-v2 overlap) (#31468)
Co-authored-by: Hao Phan <htphan@nvidia.com>
|
2026-07-17 23:22:07 -07:00 |
|
 
|
7bc343470f
|
[Spec] DFlash: support pure-MLA targets with an fp8 KV cache (Kimi-K2.x-NVFP4) (#29218)
Co-authored-by: Hao Phan <htphan@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-07 19:52:14 -07:00 |
|
 
|
0203c60fdf
|
[CP] Consolidate decode-context-parallel (DCP) helpers into layers/dcp/ (#29365)
Co-authored-by: Hao Phan <htphan@nvidia.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-03 12:25:39 -07:00 |
|
 ThanhhaoandHao Phan
|
2e7523ddbe
|
fix(spec-dec): treat num_nextn_predict_layers=0 the same as absent for EAGLE3 drafts (#26726)
Co-authored-by: Hao Phan <htphan@nvidia.com>
|
2026-06-05 16:08:07 -07:00 |
|
  
|
1c0019da75
|
[Docs] GLM-4.7 cookbook: add NVIDIA Blackwell (B200, GB200) + NVFP4 sections (#26384)
Co-authored-by: Hao Phan <htphan@nvidia.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-06-01 20:47:51 -07:00 |
|