Commit Graph
10735 Commits
Author SHA1 Message Date
8d6549bc40 [Attention Backend] Extend hpc_ops dynamic-scheduled decode to bf16 (#32304)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
2026-07-27 21:31:04 +08:00
inkcherry 5656de2d9a [PD] pool decode bootstrap HTTP sessions (#31543) 2026-07-27 21:07:23 +08:00
Peng Xingchen db9143ee08 [NPU] Fix MTP IndexShare warm-up for attention DP and prefill CP (#32210) 2026-07-27 19:19:38 +08:00
Lianmin Zheng 34454c06b8 [Refactor] Tidy server_args.py section grouping and drop unused alias (#32496) 2026-07-27 04:09:09 -07:00
1d350aaad3 fix(reasoning): let --enable-strict-thinking works for DeepSeek-V4 (#32400)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-07-27 18:15:47 +08:00
Mick 08af5aea57 optimize: optimize EmbeddingGemma prefill performance (#32383) 2026-07-27 17:34:29 +08:00
Jackey HuaandClaude Opus 5 9a0bd24bed model: serve bare Qwen3Model backbone natively as an embedding model (#32457)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 15:49:58 +08:00
McZyWu 169fc1e20c [NPU] Acc fix for afmoe model introduced by topk refactor. (#31280) 2026-07-27 15:40:49 +08:00
McZyWu c0f47a06fc [NPU] Determine the topk norm_type through scoring_func (#31393) 2026-07-27 15:12:38 +08:00
Zheng Wengang 3d3ba4f746 [BugFix][EPD] Fix Mooncake source-MR lifecycle for multi-TP /send (#32071) 2026-07-27 15:06:01 +08:00
Sam Shleifer 5cc273a780 [feat] Opt-in flat response format for prompt top logprobs (#32078) 2026-07-26 23:44:34 -07:00
Yiqi Yangandhzh0425 4ea17169b0 [UnifiedTree] fix: drop prefetched host refill under an un-backed-up parent (#31902)
Co-authored-by: hzh0425 <hzh0425@apache.org>
2026-07-27 13:53:02 +08:00
Ethan (Yusheng) Su ee1736f39a [LoRA] Support LoRA under the breakable/full prefill CUDA graph (#30988) 2026-07-26 22:10:03 -07:00
DAI0818andybyang 2abb1d2c37 fix(hisparse): correct DSA KV memory budget (#31992)
Co-authored-by: ybyang <ybyang7@iflytek.com>
2026-07-27 11:19:32 +08:00
Mick abb8f4b5e3 model: support EmbeddingGemma (#32375) 2026-07-27 10:40:47 +08:00
icarus_zh a358374ae9 [NPU][Fix Issue]: Send expert weights contiguous tensor across cards during EPLB rebalance (#32001) 2026-07-27 09:20:28 +08:00
shadowxz109 e14068d161 [NPU]Add Ascend transfer version compatibility. (#31189) 2026-07-26 21:07:59 +08:00
icarus_zh e8a635a412 Load initial expert location metadata on CPU (#32435) 2026-07-26 20:44:25 +08:00
ming_wang a76b74cbe0 add fill_draft_extend_prepare_buffers_native for NPU (#32427) 2026-07-26 20:42:14 +08:00
Wang, FangYuan 61057bda6c [BugFix] Prevent TBO crash when return_logprob is enabled (#32180) 2026-07-26 00:07:54 -07:00
jacky.cheng 833e1bc601 [Fix][AMD] Qwen3.5 MoE: disable global-slot shared-expert fusion under per-rank EP backends (MoRI + dp-attention init crash) (#31793) 2026-07-25 23:59:47 -07:00
silencejade 72e415dfc8 [Bugfix] Fix prefill suspension caused by delayed negotiate_should_allow_prefill invocation (#32389) 2026-07-26 14:59:26 +08:00
ormandjandMohammad Miadh Angkad 2cbddb842d [DSV4/SM120] Allow fused MHC opt-in with standalone TileLang pre disabled (#30954)
Signed-off-by: David Orman <ormandj@corenode.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
2026-07-26 09:02:54 +08:00
55c4853487 [comm] Enable multi-node custom-AR v2 on a single NVLink clique (#32339)
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-07-25 17:25:18 -07:00
2c63a2f12b Fix --hicache-size allocating ~2x host memory on hybrid SWA (#32373)
Co-authored-by: cctry <cctry@fb.com>
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
2026-07-25 17:19:44 -07:00
Lianmin ZhengandAlec S 9989077f24 Use native batched llguidance mask generation (#32412)
Co-authored-by: Alec S <10566873+alecsolder@users.noreply.github.com>
2026-07-25 16:36:32 -07:00
Lianmin ZhengandXingyu Liu fae84ac0f9 Fix token count localization for replicated attention-TP forwards (#32411)
Co-authored-by: Xingyu Liu <38244988+charlotte12l@users.noreply.github.com>
2026-07-25 16:36:15 -07:00
Liangsheng Yin 3da1071d56 [Spec] Hold the grammar bitmask in one GrammarMask type across all decode paths (#32409) 2026-07-25 15:27:43 -07:00
Jialin Ouyang cd145f840f Radix Cache Split: Spin off TreeCore (#29901) 2026-07-25 14:31:59 -07:00
Kangrui DuandYihao Wang a23f6ea090 [Diffusion] offload rollout weights to pinned host memory (#32032)
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
2026-07-25 13:35:53 -07:00
YAMY 91f386a5b2 fix(disagg): support pipeline-parallel hybrid-linear transfer (#32270) 2026-07-25 13:34:38 -07:00
Cheng WanandShu Wang 659d349b61 [core/loader] Add presharded load format (#24256)
Co-authored-by: Shu Wang <shuwanguc@google.com>
2026-07-25 13:03:39 -07:00
Mohammad Miadh Angkad 9791fc7090 Add configurable FlashInfer autotune skips (#31389) 2026-07-25 11:17:58 -07:00
Ke Bao 69a3c54c70 Fix SWA admission livelock on cached-prefix resumes (#32379) 2026-07-25 22:37:35 +08:00
e943e609dc [DSPARK] Grammar-constrained decoding, incl. tool_choice=auto (#31753)
Co-authored-by: shanemort1982 <shanemort1982@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-07-25 05:20:25 -07:00
a678a42033 [KDA] Add target_verify support for speculative decoding (#26888)
Co-authored-by: yuyanqi <yuyanqi@meituan.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-07-25 19:52:44 +08:00
7c4b22fae5 [Hicache][1/2]Support Mamba branching in Unified Radix Cache with HiCache (#31181)
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-07-25 19:44:40 +08:00
Mick 1054060ef1 perf: speed up marlin moe with occupancy-aware launch specialization (#31552) 2026-07-25 19:38:11 +08:00
Hồ Sỹ Thếandhnyls2002 d021990bf5 [DFLASH] Support grammar-constrained decoding in speculative verify (#30096)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-07-25 04:36:46 -07:00
Liangsheng Yin 17afd8421f [Spec] Share the grammar mask build and verify-tree staging across spec workers (#32393) 2026-07-25 03:40:21 -07:00
Liangsheng Yin 3c5bf1f6d2 [Spec] Derive NGRAM grammar tree links on the host instead of reading back retrive_next_token (#32380) 2026-07-25 02:25:16 -07:00
Yuang Chenand晟海 f5155d9602 [EPD] Fix HTTP dispatch lock blocking cross-request encoder batching (#31275)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
2026-07-25 14:47:58 +08:00
Xun SunandShangming Cai 9eb2dccbb7 [Elastic EP] Fix recovery lifecycle and add manual coverage (#31744)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-07-25 14:32:47 +08:00
f9c14e6bd4 [FEAT] Support fast engine recovery through weight cache (#27139)
Signed-off-by: Michael Qiu <qiudayu.qdy@antgroup.com>
Co-authored-by: liusy58 <liusy58@linux.alibaba.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
2026-07-25 14:31:21 +08:00
SovietPowerandShangming Cai 6a046fad09 [PD] Prevent decode scheduler from blocking on ZMQ sends to a stalled prefill peer (#31144)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-07-25 12:53:41 +08:00
ebcb74abd4 feat(hicache): Add shared memory allocator for host KV cache (#29326)
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-07-25 12:05:37 +08:00
Alex NailsandClaude Opus 4.8 b83041c3cc Migrate CompressedTensorsW4A4Nvfp4MoE TRT-LLM path onto MoeRunner (#32248)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-24 20:57:38 -07:00
Liangsheng Yin cff20a2fbb [Spec] Consolidate the grammar sync decision into ScheduleBatch.grammar_needs_sync (#32353) 2026-07-24 20:42:38 -07:00
Yihao Wang 95865de24f [diffusion] CI: read consistency GT from ci-data-diffusion at per-platform commit (#32297) 2026-07-25 09:55:57 +08:00
cctryandJialin Ouyang a690e5e0b3 Add stream label to TTFT metrics (#32363)
Co-authored-by: Jialin Ouyang <Jialin.Ouyang@gmail.com>
2026-07-24 17:44:04 -07:00