cctry
|
7a6191c4b9
|
Preallocate HiCache MHA staging before post-capture KV sizing (#40256)
|
2026-09-21 10:44:29 -07:00 |
|
cctry
|
7ad55e4386
|
[HiCache] TMA-staged host<->device KV transfer kernel (sm_90+) (#40278)
|
2026-09-21 10:38:23 -07:00 |
|
 Eric.Chin.AMDandThomas Wang
|
3c71bb018a
|
[AMD] Enable GLM DSA prefill top-k to the v2 kernel (#37889)
Co-authored-by: Thomas Wang <thomawan@amd.com>
|
2026-09-21 10:31:07 -07:00 |
|
  
|
5a6a1bb883
|
[mxfp8-kv] Skip writes to the reserved CUDA-graph padding slot (#35351)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Sam Shleifer <sam@thinkingmachines.ai>
|
2026-09-22 01:21:25 +08:00 |
|
Liangsheng Yin
|
800613a74b
|
[Test] Split the serving perf tests by topic into basic_perf/ and route their thresholds through a kit (#40505)
|
2026-09-21 10:05:57 -07:00 |
|
Sage
|
14e9c40a72
|
[Observability] Expose python/rust frontend identity in /server_info (#39993)
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com>
|
2026-09-21 23:32:52 +08:00 |
|
    
|
b63f8416b3
|
[Feature] Gigachat 3.5 support (#29189)
Co-authored-by: Stanislav Petrov <stapetrov@sberbank.ru>
Co-authored-by: Viacheslav Barinov <vvadbarinov@sberbank.ru>
Co-authored-by: Viacheslav <viacheslav.teh@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-09-21 14:37:55 +08:00 |
|
kpjeeja
|
b54d5b7c7b
|
disaggregation: Fix FakeKVSender queue accumulation (#28652)
Signed-off-by: KP, Jeeja <jeeja.kp@intel.com>
|
2026-09-21 14:27:05 +08:00 |
|
skyler-apdx
|
f5f3c38aad
|
[Fix] Preserve YaRN scaling when extending rotary caches (#38786)
|
2026-09-21 14:14:05 +08:00 |
|
 AMRUTHA MandMa Mingfei
|
d20cd9d77f
|
[XPU]Enable HiSparse hierarchical sparse KV cache on Intel XPU (#32792)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-09-21 14:07:49 +08:00 |
|
jianzhao-xu
|
62ba964848
|
Fix: post-load staging regression breaks offload meta/sharded_gpu modes (#38779)
|
2026-09-21 11:19:07 +08:00 |
|
Xueshen Liu
|
ab03a8e7eb
|
[Perf] Fork-safe import: no CUDA context at import time, lighter argument parsing (#40201)
|
2026-09-21 10:48:35 +08:00 |
|
zhaozx-cn
|
176dbcb85d
|
[npu]add chunk gdn kernel and unify ssm state layout for ascend gdn backend (#36187)
|
2026-09-21 09:51:19 +08:00 |
|
Liangsheng Yin
|
76a9065bef
|
[Fix] Raise on undelivered embeddings in send_with_url, fix broken tests (#40502)
|
2026-09-20 17:35:27 -07:00 |
|
Liangsheng Yin
|
acd20a516e
|
[CI] Give the kernel lane a 5090 suite and move kernel-only tests off the general lane (#40496)
|
2026-09-20 16:46:53 -07:00 |
|
 
|
42875bcd2a
|
fix(modelopt): dispatch NVFP4 MoE on the cached backend, not the live global (#38932)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
|
2026-09-20 19:28:54 -04:00 |
|
 Mohammad Miadh Angkadandmmangkad
|
2fa6b94e34
|
[Perf] Fuse the glm5_next mHC attn->MLP boundary (#39200)
Co-authored-by: mmangkad <mohammad.angkad@radixark.ai>
|
2026-09-20 16:20:37 -07:00 |
|
 
|
f31a7bd45c
|
Use pinned memory for asynchronous sampling metadata transfers (#39777)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
|
2026-09-20 15:31:31 -07:00 |
|
Harmya Bhatt
|
95521da18d
|
[DeepSeek-V4.1] Bound dense prefill indexer memory (#40217)
|
2026-09-20 14:43:19 -07:00 |
|
Liangsheng Yin
|
f6483e479f
|
[Test] Drop cause-less disabled tests, fix XPU lane, demote quality gates off base-c (#40288)
|
2026-09-20 14:36:02 -07:00 |
|
 
|
d229952e25
|
[Fix] Preserve model runner contracts in prefill CUDA graphs (#35452)
Co-authored-by: Oasis-Git <ayw.sirius19@gmail.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
|
2026-09-20 13:53:14 -07:00 |
|
Khoa Pham
|
b3e4d198af
|
[PD] Bound cached-prefix DCP transfers by pack capacity (#40376)
|
2026-09-21 01:17:15 +08:00 |
|
 Jimmy ShongandYangmin Li
|
e97614d10c
|
[Qwen4-Exp] Build the offloaded PLE table on the meta device so --ple-offload-embedding never materialises it on the accelerator (#39928)
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
|
2026-09-20 08:28:25 -07:00 |
|
ZY Y
|
5f017ffabb
|
Update test cases and performance testing framework (#40392)
|
2026-09-20 22:42:00 +08:00 |
|
 Mohammad Miadh AngkadandMohammad Angkad
|
8923f4d779
|
[Test] Fix optimistic prefill disaggregation test after mamba radix cache removal (#40469)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
|
2026-09-20 22:39:34 +08:00 |
|
Liangsheng Yin
|
0024efa0de
|
[CI] Derive registered-test kind from the registry call instead of the path (#40294)
|
2026-09-20 02:01:22 -07:00 |
|
HuangJi
|
2a0cb2f04e
|
[Diffusion][MiniMax-H3] Add SM120 Sage compute for SubBlock sparse attention (#40116)
|
2026-09-20 16:40:28 +08:00 |
|
Liangsheng Yin
|
dc002c85fc
|
[Test] Fix OOT DFlash hook test resolving the draft config over the network (#40427)
|
2026-09-20 01:25:34 -07:00 |
|
 Shuwen WangandSeokhoon Kang
|
9f3d275940
|
[HiCache] Fix sparse hybrid transfer layer IDs (#37870)
Co-authored-by: Seokhoon Kang <sh.kang@postech.ac.kr>
|
2026-09-20 16:24:24 +08:00 |
|
 amd-danli103andHAI
|
e54009240a
|
[AMD][DSV4] feat: enable DSpark with fp8 unified_kv on gfx950 (#38901)
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-09-20 01:16:39 -07:00 |
|
Ke Bao
|
a8a4d86be9
|
Remove swa and mamba radix cache (#40313)
|
2026-09-20 16:16:27 +08:00 |
|
Qiaolin Yu
|
f4c256354c
|
[kimi k3][pd disagg] support pp prefill + dcp decode with dspark (#40045)
|
2026-09-20 00:15:28 -07:00 |
|
 Mohammad Miadh AngkadandMohammad Angkad
|
22f02cc339
|
[Test] Fix scheduler fixtures after prefill burst counting (#40411)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
|
2026-09-20 00:08:04 -07:00 |
|
 
|
99a44c88d4
|
Add out-of-tree DFlash extension points (#38740)
Co-authored-by: Yuhan Chen <yuhanc@fb.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-09-20 14:54:46 +08:00 |
|
 Yuxuan ZhangandXinyuan Tong
|
9f21fbc34b
|
[GLM-5.3-Flash] Reduce KPool planning synchronization and overlap indexer preparation (#39695)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-09-19 23:51:17 -07:00 |
|
 Yuxuan ZhangandXinyuan Tong
|
c8eb54c41d
|
Fuse GLM-5.3-Flash KDA projections and prefill metadata (#39688)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-09-19 23:46:33 -07:00 |
|
chaijiacheng888
|
2d216a11f8
|
[Model] Serve DeepSeek-OCR-2 with its official 768px local-crop geometry (#38996)
|
2026-09-20 14:17:21 +08:00 |
|
Sasha Sidorov
|
59dd2fc734
|
[2/N] [Kernel] Fuse padding-preserving HiSparse slot translation (#39837)
|
2026-09-20 12:04:03 +08:00 |
|
 Mohammad Miadh AngkadandMohammad Angkad
|
df0dc44931
|
[Fix] Forward SWA prealloc reclaim through the DSV4 HiSparse allocator (#40354)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
|
2026-09-19 20:27:51 -07:00 |
|
 huangtingweiandChao Shi
|
020703923d
|
[PP + HiCache] Add PP Prefetch Tickets for eager cross-stage storage prefetch (#36700)
Co-authored-by: Chao Shi <stepinto@live.com>
|
2026-09-20 11:19:31 +08:00 |
|
Tri Dao
|
d82d653f96
|
Enable optimistic prefill for Mamba radix-cache models (#40184)
|
2026-09-19 19:28:11 -07:00 |
|
 huangtingweiandZhangheng
|
e9300f643e
|
[Unified Cache][9/N] add opt-in MLA load deduplication for Mooncake Linker (#39565)
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-09-20 09:56:56 +08:00 |
|
 AndyLi429andAndyLi429
|
d903351a66
|
[NPU][bugfix] update low latency quantization input and update MXFP8 tests (#38831)
Co-authored-by: AndyLi429 <AndyLi429@noreply.gitcode.com>
|
2026-09-20 09:55:30 +08:00 |
|
 
|
f9c2791460
|
[diffusion] model: support qwen-image-2.1 (#39983)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: BBuf <1182563586@qq.com>
|
2026-09-20 09:46:09 +08:00 |
|
 BourneSun0527andEven Zhou
|
d2f291c934
|
[NPU][DSV4]dsv4 enable cpp (#39820)
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
|
2026-09-20 09:17:27 +08:00 |
|
YAMY
|
9cc7da2ab0
|
[MegaMoE] Wire Qwen MoE blocks to DeepGEMM MegaMoE (MXFP4 and NVFP4 experts) (#38080)
|
2026-09-19 16:00:38 -07:00 |
|
  
|
3a64faa1f2
|
Fix disagg PP MTP for GLM-5.2 (#39378)
Co-authored-by: Julien Lin <jullin@nvidia.com>
Co-authored-by: YAMY1234 <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
|
2026-09-19 13:55:07 -07:00 |
|
Zhiqiang Xie
|
7a6c652c77
|
[HiCache] Auto-size the host pool to fit available host memory (#40135)
|
2026-09-19 12:50:43 -07:00 |
|
 metamergebotandcctry
|
9e5a62a767
|
[Logprob] Serve input-logprob temporaries from CUDA-graph-pool dead space (#40038)
Co-authored-by: cctry <csycfl@gmail.com>
|
2026-09-19 12:03:52 -07:00 |
|
 
|
7b67a96640
|
[DSV4] Chunk the indexer MQA logits by query rows under a free-memory budget (#39095)
Signed-off-by: Shiki Wu <shikiw@nvidia.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
|
2026-09-19 11:53:20 -07:00 |
|