  
|
25ce8063f7
|
Pipeline parallelism x speculative decoding (EAGLE/MTP) compatibility (#30775)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: YAMY1234 <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
|
2026-09-16 21:41:54 -07:00 |
|
 eeechoandBaizhou Zhang
|
19c30dff56
|
[SM120] DeepSeek-V4: DeepGEMM paged-MQA indexer +FP4 MoE+ page-split (#29927)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-09-02 15:49:59 -07:00 |
|
eeecho
|
b68702be99
|
[DSV4] hc-prenorm: fuse the combine step into a Triton kernel (#35118)
|
2026-09-01 02:04:05 -07:00 |
|
eeecho
|
91e7e84ee5
|
[SM120] flash_mla: allocate the page-split buffer outside inference mode (#35116)
|
2026-08-24 17:42:23 -07:00 |
|
 eeechoandAliceChenyy
|
f0bf96390b
|
[SM120] Add FlashInfer sparse MLA decode for DSv4-Flash (#27455)
Co-authored-by: AliceChenyy <alicechenyy@users.noreply.github.com>
|
2026-06-29 16:27:16 -07:00 |
|
 eeechoandClaude Opus 4.6
|
524ba10eda
|
feat: SM120 (Blackwell Desktop) support for DeepSeek-V4 inference (#24692)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-06-01 14:05:20 -07:00 |
|