Commit Graph
12867 Commits
Author SHA1 Message Date
DarkSharpnessandBBuf d1acbe0746 [DSV4.1] Big fused wo_a quant (#39957)
Co-authored-by: BBuf <1182563586@qq.com>
2026-09-19 19:46:53 +08:00
Xiaoyu Zhang cb22f2451e [Cleanup] Deduplicate kernel tests, diffusion fixtures and benchmark helpers (#40265) 2026-09-19 19:45:38 +08:00
Xiaoyu Zhang 0b0d2c257a [Fix] Repair CI fixtures and ROCm speculative tree device checks (#40325) 2026-09-19 18:16:05 +08:00
Cheng Wan 3a5f52e144 Record a process's placement at publish, not at group build (#40071) 2026-09-18 23:52:37 -07:00
Lianmin Zheng 5d703de9e4 [HiCache] Size MHA host pools from device row width (#40304) 2026-09-18 23:37:58 -07:00
Xiaoyu Zhang 6533223502 [Lint] Fix logits processor formatting on main (#40303) 2026-09-19 14:31:56 +08:00
Xiaozhu Mengandmxz 111aeedd37 [Runtime] Add decode CUDA graph hooks for eager logits processing (#40222)
Co-authored-by: mxz <mxz@fb.com>
2026-09-18 23:27:03 -07:00
6e1338dd1e Fix prefetch attempt cleanup on abort (#40262)
Co-authored-by: cctry <cctry@meta.com>
Co-authored-by: cctry <cctry@fb.com>
2026-09-18 22:46:17 -07:00
677c1cbdc9 [Metrics] Propagate idle gaps across all scheduler loops (#40004)
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
2026-09-18 22:42:00 -07:00
8189e3896b [PD] Skip singleton transfer-status all-reduces (#40003)
Co-authored-by: Pranjal Shankhdhar <pranjalssh@users.noreply.github.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
2026-09-18 22:41:50 -07:00
5b42d10edf fix: restrict SafeUnpickler to explicit globals (#40259)
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Jihui Yang <16509088+jihuiyang@users.noreply.github.com>
2026-09-18 22:38:11 -07:00
5e9342d16f [PP][DeepSeek V4] Overlap communication and optimize SM120 prefill (#38792)
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
Co-authored-by: YAMY <74099316+YAMY1234@users.noreply.github.com>
2026-09-18 21:18:04 -07:00
929230a6f0 Fix GLM-OCR MTP multimodal embeddings and positions (#39088)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-18 21:10:04 -07:00
Mind Lab 111b905bd1 fix(openai): recover logprobs token bytes from token_id (UTF-8 fragments) (#38604) 2026-09-18 20:30:04 -07:00
DefTruthandcopilot-swe-agent[bot] f1fbbd17bb [diffusion] chore: update Cache-DiT to 1.5.1 for DMD Calibrator, SVDQuant DQ, etc (#40104)
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
2026-09-19 11:14:33 +08:00
Vincent Liu 090263eff6 [VLM] avoid CUDA placement on non-CUDA platforms (#38750) 2026-09-19 11:13:22 +08:00
Alexandkpham-sgl c3aa09b0db [Kernel] Fuse hc_combine_norm for mid-size verify batches (9-96 rows) (#40208)
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
2026-09-18 20:03:49 -07:00
hmalgewattaandMick Qian c475ac5eaf [diffusion] feat: out of tree platform support (#37547)
Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
2026-09-19 09:35:58 +08:00
Xiaoyu Zhang 986959e3c4 [Refactor] Deduplicate kernel helpers and remove unused code (#40197) 2026-09-19 09:27:30 +08:00
Khoa PhamandCursor 10b0bcfd18 [PD] Allow decode radix cache and HiCache L1/L2 with DCP (#40263)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-18 17:54:35 -07:00
Cheng Wan fa7e83fd09 Name the two widths of the WORLD group (#40070) 2026-09-18 17:47:02 -07:00
Cheng Wan 81421b91e9 One read path for every parallel name (#40069) 2026-09-18 17:43:56 -07:00
Cheng Wan afe71f4b9e Read process groups through the runtime context (#40068) 2026-09-18 17:40:32 -07:00
5931fd60ee Support unified memory page-envelope transfers in PD (#39477)
Co-authored-by: yhzhuang <yhzhuang@fb.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
Co-authored-by: Yonghao Zhuang <yhzhuang@users.noreply.github.com>
Co-authored-by: Cheng Wan <cheng.wan@radixark.ai>
2026-09-18 17:39:50 -07:00
Cheng Wan d0730a0e8b Give the attention-DP width and rank one home (#40067) 2026-09-18 17:34:33 -07:00
Lifan Shen 8ea0ee300d perf(sampling): avoid GPU syncs when applying custom logit processors (#39234) 2026-09-18 17:09:30 -07:00
Alison ShaoandXinyuan Tong 21e6c98ccb Fix corrupted chat prompts on mistral_common tokenizers (tool_choice auto never fires) (#39773)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-18 16:12:06 -07:00
2394b231c2 Fix Mistral3 retaining every vision-tower layer to read one (#39185)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-09-18 16:11:54 -07:00
fa521e2758 [MM] Copy placeholder ids to CUDA asynchronously (#40010)
Co-authored-by: metamergebot <metamergebot@users.noreply.github.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-09-18 16:10:37 -07:00
Zhiqiang Xieandcctry ceb1d2e580 [PD] Enable optimistic prefill with buffer-only L3 write-through HiCache (#40043)
Co-authored-by: cctry <csycfl@gmail.com>
2026-09-18 16:04:15 -07:00
5e4b94b134 [MM] Skip VMM error gathers for text-only requests (#40005)
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-09-18 16:00:43 -07:00
hunhokimandHun-ho Kim 803f0c93d2 Fix DSA partial DP-TP mode log to use derived attn_tp_size (#39871)
Co-authored-by: Hun-ho Kim <hunho.kim@samsung.com>
2026-09-18 15:54:45 -07:00
Jialin Ouyang f3851486cb [Perf] Fuse SWA page lookup and mapping clear (#38948) 2026-09-18 15:51:07 -07:00
81a199f56a [HiCache] Read the in-flight buffer backup's node id from its snapshot in sanity_check (#40013)
Co-authored-by: Pranjal Shankhdhar <pranjalssh@meta.com>
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
2026-09-18 15:47:56 -07:00
0e5347db82 Support MXFP8 and deferred route weighting in DeepEP v2 (#40030)
Co-authored-by: metamergebot <324680979+metamergebot@users.noreply.github.com>
Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>
Co-authored-by: pranjalssh <14260275+pranjalssh@users.noreply.github.com>
2026-09-18 15:40:38 -07:00
Liangsheng Yin 6cc9090d1f [mem_cache] Release up to owned_kv_len on radix cache insert (#40075) 2026-09-18 15:37:52 -07:00
Zhiqiang Xie a0534f8cca [HiCache] Stop arming a prefetch retry for a too-short storage span (#40042) 2026-09-18 15:33:58 -07:00
Sam (Kesen Li) d346b214fb feat(kv-cache): support SM100 NVFP4 GenMHA and speculative decoding (#36340) 2026-09-18 14:50:46 -07:00
Lianmin ZhengandJialin Ouyang f5a1434700 [HiCache] Document transfer arguments (#40239)
Clarify the legacy host_indices argument and label Mamba test arguments,
including the current staging_tokens parameter. Executable code is unchanged.

Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
2026-09-18 14:08:19 -07:00
amd-danli103 cd4dd81c22 [AMD][DSV4] fix: drop shadowing local get_exec import that breaks model startup on ROCm (#40186) 2026-09-18 13:28:38 -07:00
metamergebotandcctry 6a9c7001d3 [Logprob] Borrow graph-pool memory for input logprob logits construction (#40007)
Co-authored-by: cctry <csycfl@gmail.com>
2026-09-18 13:18:01 -07:00
Jialin Ouyang da2f434951 [Spec] Add explicit prefill shared-read capability for plugins (#39502) 2026-09-18 13:11:01 -07:00
Lianmin Zhengandraghotham 248c202b46 Use runtime token widths for Triton speculative verification (#39859)
Co-authored-by: raghotham <853234+raghotham@users.noreply.github.com>
2026-09-18 11:07:55 -07:00
Lianmin Zheng 6bd1a0af1d Add registration for external model configurations (#39452) 2026-09-18 10:06:24 -07:00
Shuwen Wang 191172fa74 [Unified Tree] fix: exempt host-locked aux nodes from the sanity_check host-LRU check (#39980) 2026-09-18 15:29:29 +00:00
Xiaoyu Zhang 7714b182f2 [Bugfix] Fix top-1 MoE routing with non-unit scaling (#40187) 2026-09-18 22:53:13 +08:00
81363bf8cb [kernel] Share the warp vectorized copy and enforce its alignment (#36176)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: BBuf <1182563586@qq.com>
2026-09-18 22:40:54 +08:00
a6cf05817f dsv4.1: remaining model and runtime integration (#38798)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <xiaoyu.zhang@radixark.ai>
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Zhichen Zeng <zczeng@uw.edu>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
2026-09-18 02:55:30 -07:00
Liangsheng Yin 1b200ffaaa [Quant] Serve 32-wide-K ue8m0 block-FP8 linears through the FlashInfer MXFP8 GEMMs (#40039) 2026-09-18 02:51:12 -07:00
iridiumine 6c7c5e78de [NPU] Run arch35 block-FP8 dense linears on the native MXFP8 GEMM (#39823) 2026-09-18 17:19:31 +08:00