Commit Graph
12886 Commits
Author SHA1 Message Date
huangtingweiandChao Shi 020703923d [PP + HiCache] Add PP Prefetch Tickets for eager cross-stage storage prefetch (#36700)
Co-authored-by: Chao Shi <stepinto@live.com>
2026-09-20 11:19:31 +08:00
113f6f080e [PD] Enter the custom mem pool once when allocating DCP pack buffers (#40284)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
2026-09-19 19:34:58 -07:00
Tri Dao d82d653f96 Enable optimistic prefill for Mamba radix-cache models (#40184) 2026-09-19 19:28:11 -07:00
huangtingweiandZhangheng e9300f643e [Unified Cache][9/N] add opt-in MLA load deduplication for Mooncake Linker (#39565)
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-09-20 09:56:56 +08:00
AndyLi429andAndyLi429 d903351a66 [NPU][bugfix] update low latency quantization input and update MXFP8 tests (#38831)
Co-authored-by: AndyLi429 <AndyLi429@noreply.gitcode.com>
2026-09-20 09:55:30 +08:00
f9c2791460 [diffusion] model: support qwen-image-2.1 (#39983)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: BBuf <1182563586@qq.com>
2026-09-20 09:46:09 +08:00
BourneSun0527andEven Zhou d2f291c934 [NPU][DSV4]dsv4 enable cpp (#39820)
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
2026-09-20 09:17:27 +08:00
YAMY 9cc7da2ab0 [MegaMoE] Wire Qwen MoE blocks to DeepGEMM MegaMoE (MXFP4 and NVFP4 experts) (#38080) 2026-09-19 16:00:38 -07:00
3a64faa1f2 Fix disagg PP MTP for GLM-5.2 (#39378)
Co-authored-by: Julien Lin <jullin@nvidia.com>
Co-authored-by: YAMY1234 <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
2026-09-19 13:55:07 -07:00
8139a1740e [Scheduler] Count complete prefill bursts and their tokens (#40006)
Co-authored-by: pranjalssh <pranjalssh@fb.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: Jialin Ouyang <jialino@meta.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-09-19 12:56:15 -07:00
Zhiqiang Xie 7a6c652c77 [HiCache] Auto-size the host pool to fit available host memory (#40135) 2026-09-19 12:50:43 -07:00
amd-danli103 2305242f51 [AMD][DSV4] fix: skip compressed-KV metadata on the draft worker in the HIP radix backend (#40205) 2026-09-19 12:13:49 -07:00
metamergebotandcctry 9e5a62a767 [Logprob] Serve input-logprob temporaries from CUDA-graph-pool dead space (#40038)
Co-authored-by: cctry <csycfl@gmail.com>
2026-09-19 12:03:52 -07:00
kkandwunhuang c5326d28a3 [AMD] dsv4: pick kv_splits per index stream, not by occupancy alone (#39968)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-09-19 11:53:54 -07:00
7b67a96640 [DSV4] Chunk the indexer MQA logits by query rows under a free-memory budget (#39095)
Signed-off-by: Shiki Wu <shikiw@nvidia.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
2026-09-19 11:53:20 -07:00
993d1fccba [ROCm] Widen the HiCache JIT copy rounds and enable the K-only host pool (#37152)
Co-authored-by: Xiaobo Chen <xiaobche@smci355-ccs-aus-n05-33.prov.aus.ccs.cpe.ice.amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
2026-09-19 08:58:10 -07:00
Xiaoyu Zhang 76f9213a41 [Fix] Keep mHC context out of non-V4 compiled MoE forwards (#40353) 2026-09-19 21:41:51 +08:00
Benjamin TruongandXiaoyu Zhang 83e29d6c5a [perf] Optimize w4a8 MoE for glm5.2 on H200 (#38220)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-09-19 20:34:25 +08:00
Xiaoyu Zhang 7fac84b639 [DSV4.1] Reduce mHC, metadata and small-batch router overhead (#39704) 2026-09-19 19:51:21 +08:00
DarkSharpnessandBBuf d1acbe0746 [DSV4.1] Big fused wo_a quant (#39957)
Co-authored-by: BBuf <1182563586@qq.com>
2026-09-19 19:46:53 +08:00
Xiaoyu Zhang cb22f2451e [Cleanup] Deduplicate kernel tests, diffusion fixtures and benchmark helpers (#40265) 2026-09-19 19:45:38 +08:00
Xiaoyu Zhang 0b0d2c257a [Fix] Repair CI fixtures and ROCm speculative tree device checks (#40325) 2026-09-19 18:16:05 +08:00
Cheng Wan 3a5f52e144 Record a process's placement at publish, not at group build (#40071) 2026-09-18 23:52:37 -07:00
Lianmin Zheng 5d703de9e4 [HiCache] Size MHA host pools from device row width (#40304) 2026-09-18 23:37:58 -07:00
Xiaoyu Zhang 6533223502 [Lint] Fix logits processor formatting on main (#40303) 2026-09-19 14:31:56 +08:00
Xiaozhu Mengandmxz 111aeedd37 [Runtime] Add decode CUDA graph hooks for eager logits processing (#40222)
Co-authored-by: mxz <mxz@fb.com>
2026-09-18 23:27:03 -07:00
6e1338dd1e Fix prefetch attempt cleanup on abort (#40262)
Co-authored-by: cctry <cctry@meta.com>
Co-authored-by: cctry <cctry@fb.com>
2026-09-18 22:46:17 -07:00
677c1cbdc9 [Metrics] Propagate idle gaps across all scheduler loops (#40004)
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
2026-09-18 22:42:00 -07:00
8189e3896b [PD] Skip singleton transfer-status all-reduces (#40003)
Co-authored-by: Pranjal Shankhdhar <pranjalssh@users.noreply.github.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
2026-09-18 22:41:50 -07:00
5b42d10edf fix: restrict SafeUnpickler to explicit globals (#40259)
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Jihui Yang <16509088+jihuiyang@users.noreply.github.com>
2026-09-18 22:38:11 -07:00
5e9342d16f [PP][DeepSeek V4] Overlap communication and optimize SM120 prefill (#38792)
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com>
2026-09-18 21:18:04 -07:00
929230a6f0 Fix GLM-OCR MTP multimodal embeddings and positions (#39088)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-18 21:10:04 -07:00
Mind Lab 111b905bd1 fix(openai): recover logprobs token bytes from token_id (UTF-8 fragments) (#38604) 2026-09-18 20:30:04 -07:00
DefTruthandcopilot-swe-agent[bot] f1fbbd17bb [diffusion] chore: update Cache-DiT to 1.5.1 for DMD Calibrator, SVDQuant DQ, etc (#40104)
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
2026-09-19 11:14:33 +08:00
Vincent Liu 090263eff6 [VLM] avoid CUDA placement on non-CUDA platforms (#38750) 2026-09-19 11:13:22 +08:00
Alexandkpham-sgl c3aa09b0db [Kernel] Fuse hc_combine_norm for mid-size verify batches (9-96 rows) (#40208)
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
2026-09-18 20:03:49 -07:00
hmalgewattaandMick Qian c475ac5eaf [diffusion] feat: out of tree platform support (#37547)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-19 09:35:58 +08:00
Xiaoyu Zhang 986959e3c4 [Refactor] Deduplicate kernel helpers and remove unused code (#40197) 2026-09-19 09:27:30 +08:00
Khoa PhamandCursor 10b0bcfd18 [PD] Allow decode radix cache and HiCache L1/L2 with DCP (#40263)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-18 17:54:35 -07:00
Cheng Wan fa7e83fd09 Name the two widths of the WORLD group (#40070) 2026-09-18 17:47:02 -07:00
Cheng Wan 81421b91e9 One read path for every parallel name (#40069) 2026-09-18 17:43:56 -07:00
Cheng Wan afe71f4b9e Read process groups through the runtime context (#40068) 2026-09-18 17:40:32 -07:00
5931fd60ee Support unified memory page-envelope transfers in PD (#39477)
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
2026-09-18 17:39:50 -07:00
Cheng Wan d0730a0e8b Give the attention-DP width and rank one home (#40067) 2026-09-18 17:34:33 -07:00
Lifan Shen 8ea0ee300d perf(sampling): avoid GPU syncs when applying custom logit processors (#39234) 2026-09-18 17:09:30 -07:00
Alison ShaoandXinyuan Tong 21e6c98ccb Fix corrupted chat prompts on mistral_common tokenizers (tool_choice auto never fires) (#39773)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-18 16:12:06 -07:00
2394b231c2 Fix Mistral3 retaining every vision-tower layer to read one (#39185)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-09-18 16:11:54 -07:00
fa521e2758 [MM] Copy placeholder ids to CUDA asynchronously (#40010)
Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-09-18 16:10:37 -07:00
Zhiqiang Xieandcctry ceb1d2e580 [PD] Enable optimistic prefill with buffer-only L3 write-through HiCache (#40043)
Co-authored-by: cctry <csycfl@gmail.com>
2026-09-18 16:04:15 -07:00
5e4b94b134 [MM] Skip VMM error gathers for text-only requests (#40005)
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-09-18 16:00:43 -07:00