Commit Graph
4586 Commits
Author SHA1 Message Date
Zhanghengand晟海 abed680320 [Unified Cache][5/N]: Integrate external linker mode end to end (#37381)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
2026-09-04 02:02:58 +08:00
cctry 33a22b1b08 [Cache] Forward fast prefix matching capability (#37844) 2026-09-03 10:33:30 -07:00
Vincent Gaoandinkcherry 54cadad151 [Router] Add composable scoring and eligibility policies (#37731)
Co-authored-by: inkcherry <mingzhi.liu@amd.com>
2026-09-04 00:01:41 +08:00
Quanli Liand全力 4e37882a93 [Diffusion][minimax-h3] Add SM120 support for SubBlock sparse attention (#37332)
Co-authored-by: 全力 <liquanli.lql@antgroup.com>
2026-09-03 22:05:22 +08:00
pllimax 3239baef25 [CI][NPU] Fix kimi_k2_6 16p in64k perf test and dsv4-flash testcases (#37760) 2026-09-03 21:57:14 +08:00
2bb25dc18b [Speculative Decoding] Add native UNO serving support (#37667)
Co-authored-by: drproduck <drproduck@MacBook-Air-2.local>
Co-authored-by: BBuf <1182563586@qq.com>
2026-09-03 20:08:41 +08:00
Cheng Wan a11dba1a01 [Feature] Unified memory: support decode context parallelism for the trtllm_mla family (#37693) 2026-09-03 03:37:12 -07:00
27b7a2dc3b [Kimi K3] Rework skipped-think fix as opt-in force_nonempty_content with streaming coverage (#34187)
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
2026-09-03 17:37:28 +08:00
kkandwunhuang a6001478f4 [AMD] Perf Kimi-K3 MoE optimization (#33838)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-09-03 02:28:28 -07:00
Xinyi SongandThomas Wang 7ed29eba80 [AMD] Fix FP4 indexer OOR (#37660)
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
2026-09-03 01:49:46 -07:00
Polisetty V R K Jyothendra Varma e59a576f03 fix test/manual/test_forward_split_prefill.py UT due to many refactors and design changes (#36617)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
2026-09-03 01:48:55 -07:00
3bac084d4e [Model] Add native IFM K2 Horizon serving support (#37654)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-03 16:39:43 +08:00
YC Yen-Ching TsengandPhil Li 1fb85053e7 [AMD][Diffusion] Migrate FlyDSL fused norm kernels to the v0.3.0 stable API (#36349)
Co-authored-by: Phil Li <haicli@amd.com>
2026-09-02 23:03:08 -07:00
xiaobochen-amdandZhang, Jiejing 030d7e7e9b [ROCm] Define the DSA head-gate graph helpers on HIP (#37118)
Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com>
2026-09-02 22:55:45 -07:00
Alex NailsandAlison Shao 28262c20df [CI][RFC] Replace black-jupyter with ruff-format (#37210)
Co-authored-by: Alison Shao <a.shao@wustl.edu>
2026-09-02 19:46:08 -07:00
2641e427be Xpu/weekly simple model enablement 2026 08 30 (#37193)
Co-authored-by: dayanandav <dayananda.vasantha.kumar@intel.com>
Co-authored-by: Girijala, Pavan Sivaram <pavan.sivaram.girijala@intel.com>
Co-authored-by: Cui, Lily <lily.cui@intel.com>
Co-authored-by: Juan Muneton <juan.muneton.gallego@intel.com>
Co-authored-by: Gao, Pengfei <pengfei.gao@intel.com>
2026-09-03 09:35:59 +08:00
Liangsheng Yin a522c8a4b6 [misc] Extract PP dynamic chunk sizing into a DynamicChunkSizer scheduler component (#37674) 2026-09-02 18:35:43 -07:00
Mick 0dd66def7c [chore] harden checkpoint quantization metadata parsing (#36922) 2026-09-03 09:27:22 +08:00
fbf909b460 [Fix] Alpha-channel images and tool-result media ordering (port of #36507) (#37320)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-09-02 18:23:06 -07:00
Lee NauandYangmin Li 6e41f1ad29 [Fix] Preserve FP32 in SM107 MXFP8 fallback (#37489)
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
2026-09-02 17:55:14 -07:00
87d60a2229 Improve CUDA graph and speculative execution output handling (#37329)
Co-authored-by: jiayisuse <jiayisuse@fb.com>
Co-authored-by: Yinghai Lu <yinghai@meta.com>
Co-authored-by: Hao Zhang <zhisbug@users.noreply.github.com>
Co-authored-by: Yichao Fu <yichaofu@meta.com>
2026-09-02 17:25:27 -07:00
Alison Shao db1eb48651 [CI] Graceful teardown for the PD and HiSparse server fixtures (#37485) 2026-09-02 17:22:18 -07:00
ff04a00d73 Reduce tokenizer overhead and offload CUDA VMM publication (#37330)
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>
Co-authored-by: Yinghai Lu <yinghai@meta.com>
2026-09-02 17:21:08 -07:00
Liangsheng Yin 5c46ce37f5 [Fix] Apply the attention-CP broadcast result in PP dynamic-chunk profiling (#37669) 2026-09-02 17:20:07 -07:00
Byron HsuandByronHsu 046cdaabaa [Sampling] Capture masks from sampler support (#36630)
Co-authored-by: ByronHsu <ByronHsu@users.noreply.github.com>
2026-09-02 17:13:28 -07:00
Cheng WanandClaude Opus 5 5ddca6819e Fix unified SWA: size a non-owner's v2p by the id space it must address (#37560)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 16:56:18 -07:00
Cheng WanandClaude Opus 5 d9848b9ecd Build the unified read stream directly, without the page-table rectangle (#37512)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 16:55:16 -07:00
Cheng WanandClaude Opus 5 18d5ffb42a Size the unified read-table grid from bs, and fuse the allocator's tombstone scatters (#37511)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 16:54:26 -07:00
c05f8ae830 [PD] Optimize paged allocator free-list release (#37146)
Co-authored-by: wangwenming.41 <wangwenming.41@jd.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-09-02 16:51:37 -07:00
Mohammad Miadh Angkad 718bd39fe0 [Fix] DP attention: correct the decode->extend prefix off-by-one (#37505) 2026-09-02 16:14:16 -07:00
paulzhang-tmandQiaolin-Yu 3fa6b86504 [Spec] Publish the final multi-layer EAGLE shared-read event (#36752)
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
2026-09-02 15:57:50 -07:00
YAMY 982aa8acfc [Bugfix] Load Qwen3.5 MTP embedding under PP (#37471) 2026-09-02 15:19:40 -07:00
YAMY 3c9cea8f10 [EAGLE] Prune draft-extend logits to selected rows (#35546) 2026-09-02 15:10:08 -07:00
YAMY fe45af1e6f perf(gdn): select ReplaySSM verify loop unrolling by shape (#36970) 2026-09-02 15:07:49 -07:00
cctry 3a855b050a fix(disagg): poll receivers during decode preallocation (#37483) 2026-09-02 14:26:40 -07:00
Liangsheng Yin 19c7679e9e [mem_cache] Make free_swa sync-free on page_size == 1 (#36723) 2026-09-02 14:18:22 -07:00
JoeandBBuf acea43079f Fix native MoE handling of noncontiguous top-k IDs (#36407)
Co-authored-by: BBuf <1182563586@qq.com>
2026-09-02 14:16:19 -07:00
f8cbf000f4 [AMD] Enable FP4 indexer for Deepseek V4 (#37353)
Co-authored-by: 1am9trash <1am9trash@gmail.com>
Co-authored-by: AMD-yanfeiwang <256076023+AMD-yanfeiwang@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
2026-09-02 09:45:08 -07:00
Po-Han Huang (NVIDIA) d585cec4bd Fix nondeterministic FlashInfer GDN alignment test (#37343) 2026-09-01 23:46:58 -07:00
Lianmin Zheng f8f04bafa8 Rust server: align launcher and request validation behavior (#37327) 2026-09-01 23:42:36 -07:00
jasonjk-park 83bd2c473f Allow custom policy for adaptive speculative decoding (#37274) 2026-09-01 23:15:21 -07:00
Liangsheng Yin 01c3a5f54f [misc] Resolve SWA ownership at enqueue time for grouped free() (#36646) 2026-09-01 23:06:55 -07:00
Liangsheng Yin 832d029870 [mem_cache] Split duplicate insert frees at the SWA eviction floor (#37481) 2026-09-01 23:01:40 -07:00
YAMY a6a19f9290 [Bugfix] Skip absent radix lock during cache cleanup (#37494) 2026-09-01 22:50:19 -07:00
Khoa PhamandClaude Opus 5 c66a285c94 [Kernel] GLM 5.3 Flash related kernels (ported from #36507) (#37477)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-01 22:17:09 -07:00
Wes 2d9c64394f Fix reasoning metrics and add TPOT to bench_multiturn (#35443) 2026-09-02 11:28:47 +08:00
Xiaoyu Zhang 403a15c163 [CI] Batch CPU test workers (#37252) 2026-09-02 10:35:14 +08:00
26f760d5c0 [CPU] Support FP8 KV cache (#32733)
Co-authored-by: Zheng, Beilei <beilei.zheng@intel.com>
Co-authored-by: mingfeima <mingfei.ma@intel.com>
2026-09-02 10:20:54 +08:00
ashwini rathiandP V R K Jyothendra Varma e874ae64cd [XPU][CI] Enable nightly-xpu-8-gpu suite: declare + wire runner job (#37230)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
2026-09-02 10:14:40 +08:00
kkandwunhuang dc276264cb [AMD] Perf Kimi-K3 fuse ROCm KDA decode boundary (#34198)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-09-01 18:35:48 -07:00