kk and wunhuang
c5326d28a3
[AMD] dsv4: pick kv_splits per index stream, not by occupancy alone ( #39968 )
...
Co-authored-by: wunhuang <wunhuang@amd.com >
2026-09-19 11:53:54 -07:00
7b67a96640
[DSV4] Chunk the indexer MQA logits by query rows under a free-memory budget ( #39095 )
...
Signed-off-by: Shiki Wu <shikiw@nvidia.com >
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com >
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com >
2026-09-19 11:53:20 -07:00
993d1fccba
[ROCm] Widen the HiCache JIT copy rounds and enable the K-only host pool ( #37152 )
...
Co-authored-by: Xiaobo Chen <xiaobche@smci355-ccs-aus-n05-33.prov.aus.ccs.cpe.ice.amd.com >
Co-authored-by: HAI <hixiao@gmail.com >
2026-09-19 08:58:10 -07:00
Xiaoyu Zhang
76f9213a41
[Fix] Keep mHC context out of non-V4 compiled MoE forwards ( #40353 )
2026-09-19 21:41:51 +08:00
Benjamin Truong and Xiaoyu Zhang
83e29d6c5a
[perf] Optimize w4a8 MoE for glm5.2 on H200 ( #38220 )
...
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com >
2026-09-19 20:34:25 +08:00
Xiaoyu Zhang
7fac84b639
[DSV4.1] Reduce mHC, metadata and small-batch router overhead ( #39704 )
2026-09-19 19:51:21 +08:00
DarkSharpness and BBuf
d1acbe0746
[DSV4.1] Big fused wo_a quant ( #39957 )
...
Co-authored-by: BBuf <1182563586@qq.com >
2026-09-19 19:46:53 +08:00
Xiaoyu Zhang
cb22f2451e
[Cleanup] Deduplicate kernel tests, diffusion fixtures and benchmark helpers ( #40265 )
2026-09-19 19:45:38 +08:00
Xiaoyu Zhang
0b0d2c257a
[Fix] Repair CI fixtures and ROCm speculative tree device checks ( #40325 )
2026-09-19 18:16:05 +08:00
Cheng Wan
3a5f52e144
Record a process's placement at publish, not at group build ( #40071 )
2026-09-18 23:52:37 -07:00
Lianmin Zheng
5d703de9e4
[HiCache] Size MHA host pools from device row width ( #40304 )
2026-09-18 23:37:58 -07:00
Xiaoyu Zhang
6533223502
[Lint] Fix logits processor formatting on main ( #40303 )
2026-09-19 14:31:56 +08:00
Xiaozhu Meng and mxz
111aeedd37
[Runtime] Add decode CUDA graph hooks for eager logits processing ( #40222 )
...
Co-authored-by: mxz <mxz@fb.com >
2026-09-18 23:27:03 -07:00
6e1338dd1e
Fix prefetch attempt cleanup on abort ( #40262 )
...
Co-authored-by: cctry <cctry@meta.com >
Co-authored-by: cctry <cctry@fb.com >
2026-09-18 22:46:17 -07:00
677c1cbdc9
[Metrics] Propagate idle gaps across all scheduler loops ( #40004 )
...
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com >
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com >
2026-09-18 22:42:00 -07:00
8189e3896b
[PD] Skip singleton transfer-status all-reduces ( #40003 )
...
Co-authored-by: Pranjal Shankhdhar <pranjalssh@users.noreply.github.com >
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com >
2026-09-18 22:41:50 -07:00
5b42d10edf
fix: restrict SafeUnpickler to explicit globals ( #40259 )
...
Co-authored-by: yhzhuang <yhzhuang@fb.com >
Co-authored-by: Jihui Yang <16509088+jihuiyang@users.noreply.github.com >
2026-09-18 22:38:11 -07:00
5e9342d16f
[PP][DeepSeek V4] Overlap communication and optimize SM120 prefill ( #38792 )
...
Co-authored-by: Yangmin Li <yangminl@nvidia.com >
Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com >
2026-09-18 21:18:04 -07:00
929230a6f0
Fix GLM-OCR MTP multimodal embeddings and positions ( #39088 )
...
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com >
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
2026-09-18 21:10:04 -07:00
Mind Lab
111b905bd1
fix(openai): recover logprobs token bytes from token_id (UTF-8 fragments) ( #38604 )
2026-09-18 20:30:04 -07:00
DefTruth and copilot-swe-agent[bot]
f1fbbd17bb
[diffusion] chore: update Cache-DiT to 1.5.1 for DMD Calibrator, SVDQuant DQ, etc ( #40104 )
...
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com >
2026-09-19 11:14:33 +08:00
Vincent Liu
090263eff6
[VLM] avoid CUDA placement on non-CUDA platforms ( #38750 )
2026-09-19 11:13:22 +08:00
Alex and kpham-sgl
c3aa09b0db
[Kernel] Fuse hc_combine_norm for mid-size verify batches (9-96 rows) ( #40208 )
...
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai >
2026-09-18 20:03:49 -07:00
hmalgewatta and Mick Qian
c475ac5eaf
[diffusion] feat: out of tree platform support ( #37547 )
...
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com >
2026-09-19 09:35:58 +08:00
Xiaoyu Zhang
986959e3c4
[Refactor] Deduplicate kernel helpers and remove unused code ( #40197 )
2026-09-19 09:27:30 +08:00
Khoa Pham and Cursor
10b0bcfd18
[PD] Allow decode radix cache and HiCache L1/L2 with DCP ( #40263 )
...
Co-authored-by: Cursor <cursoragent@cursor.com >
2026-09-18 17:54:35 -07:00
Cheng Wan
fa7e83fd09
Name the two widths of the WORLD group ( #40070 )
2026-09-18 17:47:02 -07:00
Cheng Wan
81421b91e9
One read path for every parallel name ( #40069 )
2026-09-18 17:43:56 -07:00
Cheng Wan
afe71f4b9e
Read process groups through the runtime context ( #40068 )
2026-09-18 17:40:32 -07:00
5931fd60ee
Support unified memory page-envelope transfers in PD ( #39477 )
...
Co-authored-by: yhzhuang <yhzhuang@fb.com >
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com >
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com >
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai >
2026-09-18 17:39:50 -07:00
Cheng Wan
d0730a0e8b
Give the attention-DP width and rank one home ( #40067 )
2026-09-18 17:34:33 -07:00
Lifan Shen
8ea0ee300d
perf(sampling): avoid GPU syncs when applying custom logit processors ( #39234 )
2026-09-18 17:09:30 -07:00
Alison Shao and Xinyuan Tong
21e6c98ccb
Fix corrupted chat prompts on mistral_common tokenizers (tool_choice auto never fires) ( #39773 )
...
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
2026-09-18 16:12:06 -07:00
2394b231c2
Fix Mistral3 retaining every vision-tower layer to read one ( #39185 )
...
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com >
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com >
2026-09-18 16:11:54 -07:00
fa521e2758
[MM] Copy placeholder ids to CUDA asynchronously ( #40010 )
...
Co-authored-by: metamergebot <metamergebot@users.noreply.github.com >
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com >
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com >
2026-09-18 16:10:37 -07:00
Zhiqiang Xie and cctry
ceb1d2e580
[PD] Enable optimistic prefill with buffer-only L3 write-through HiCache ( #40043 )
...
Co-authored-by: cctry <csycfl@gmail.com >
2026-09-18 16:04:15 -07:00
5e4b94b134
[MM] Skip VMM error gathers for text-only requests ( #40005 )
...
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com >
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com >
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com >
2026-09-18 16:00:43 -07:00
hunhokim and Hun-ho Kim
803f0c93d2
Fix DSA partial DP-TP mode log to use derived attn_tp_size ( #39871 )
...
Co-authored-by: Hun-ho Kim <hunho.kim@samsung.com >
2026-09-18 15:54:45 -07:00
Jialin Ouyang
f3851486cb
[Perf] Fuse SWA page lookup and mapping clear ( #38948 )
2026-09-18 15:51:07 -07:00
81a199f56a
[HiCache] Read the in-flight buffer backup's node id from its snapshot in sanity_check ( #40013 )
...
Co-authored-by: Pranjal Shankhdhar <pranjalssh@meta.com >
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com >
2026-09-18 15:47:56 -07:00
0e5347db82
Support MXFP8 and deferred route weighting in DeepEP v2 ( #40030 )
...
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com >
Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com >
Co-authored-by: pranjalssh <14260275+pranjalssh@users.noreply.github.com >
2026-09-18 15:40:38 -07:00
Liangsheng Yin
6cc9090d1f
[mem_cache] Release up to owned_kv_len on radix cache insert ( #40075 )
2026-09-18 15:37:52 -07:00
Zhiqiang Xie
a0534f8cca
[HiCache] Stop arming a prefetch retry for a too-short storage span ( #40042 )
2026-09-18 15:33:58 -07:00
Sam (Kesen Li)
d346b214fb
feat(kv-cache): support SM100 NVFP4 GenMHA and speculative decoding ( #36340 )
2026-09-18 14:50:46 -07:00
Lianmin Zheng and Jialin Ouyang
f5a1434700
[HiCache] Document transfer arguments ( #40239 )
...
Clarify the legacy host_indices argument and label Mamba test arguments,
including the current staging_tokens parameter. Executable code is unchanged.
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com >
2026-09-18 14:08:19 -07:00
amd-danli103
cd4dd81c22
[AMD][DSV4] fix: drop shadowing local get_exec import that breaks model startup on ROCm ( #40186 )
2026-09-18 13:28:38 -07:00
metamergebot and cctry
6a9c7001d3
[Logprob] Borrow graph-pool memory for input logprob logits construction ( #40007 )
...
Co-authored-by: cctry <csycfl@gmail.com >
2026-09-18 13:18:01 -07:00
Jialin Ouyang
da2f434951
[Spec] Add explicit prefill shared-read capability for plugins ( #39502 )
2026-09-18 13:11:01 -07:00
Lianmin Zheng and raghotham
248c202b46
Use runtime token widths for Triton speculative verification ( #39859 )
...
Co-authored-by: raghotham <853234+raghotham@users.noreply.github.com >
2026-09-18 11:07:55 -07:00
Lianmin Zheng
6bd1a0af1d
Add registration for external model configurations ( #39452 )
2026-09-18 10:06:24 -07:00