Commit Graph
1273 Commits
Author SHA1 Message Date
Zhiqiang Xie a480f388b2 [HiCache] L3 storage prefetch lifecycle metrics and cross-tier attribution fixes (#37503) 2026-09-03 16:00:04 -07:00
Yonghao Zhuangandyhzhuang 4dc9cda5f9 [PD] Gate deferred decode KV release on backend capability (#37454)
Co-authored-by: yhzhuang <yhzhuang@fb.com>
2026-09-03 15:42:21 -07:00
Liangsheng Yin 2a980cbf10 [mem_cache] Require page-aligned starts in free_segment and drop the boundary trim (#37729) 2026-09-03 13:28:33 -07:00
Zhiqiang Xie d0c95f6c91 [HiCache] buffer mode: anchor-lock staged prefetches by default (#37464) 2026-09-03 12:08:22 -07:00
Zhanghengand晟海 abed680320 [Unified Cache][5/N]: Integrate external linker mode end to end (#37381)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
2026-09-04 02:02:58 +08:00
cctry 33a22b1b08 [Cache] Forward fast prefix matching capability (#37844) 2026-09-03 10:33:30 -07:00
Vincent Gaoandinkcherry 54cadad151 [Router] Add composable scoring and eligibility policies (#37731)
Co-authored-by: inkcherry <mingzhi.liu@amd.com>
2026-09-04 00:01:41 +08:00
2bb25dc18b [Speculative Decoding] Add native UNO serving support (#37667)
Co-authored-by: drproduck <drproduck@MacBook-Air-2.local>
Co-authored-by: BBuf <1182563586@qq.com>
2026-09-03 20:08:41 +08:00
27b7a2dc3b [Kimi K3] Rework skipped-think fix as opt-in force_nonempty_content with streaming coverage (#34187)
Co-authored-by: yhyang201 <yhyang201@gmail.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
2026-09-03 17:37:28 +08:00
kkandwunhuang a6001478f4 [AMD] Perf Kimi-K3 MoE optimization (#33838)
Co-authored-by: wunhuang <wunhuang@amd.com>
2026-09-03 02:28:28 -07:00
3bac084d4e [Model] Add native IFM K2 Horizon serving support (#37654)
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-09-03 16:39:43 +08:00
xiaobochen-amdandZhang, Jiejing 030d7e7e9b [ROCm] Define the DSA head-gate graph helpers on HIP (#37118)
Co-authored-by: Zhang, Jiejing <jiejing.zhang@amd.com>
2026-09-02 22:55:45 -07:00
Alex NailsandAlison Shao 28262c20df [CI][RFC] Replace black-jupyter with ruff-format (#37210)
Co-authored-by: Alison Shao <a.shao@wustl.edu>
2026-09-02 19:46:08 -07:00
Mick 0dd66def7c [chore] harden checkpoint quantization metadata parsing (#36922) 2026-09-03 09:27:22 +08:00
fbf909b460 [Fix] Alpha-channel images and tool-result media ordering (port of #36507) (#37320)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-09-02 18:23:06 -07:00
Lee NauandYangmin Li 6e41f1ad29 [Fix] Preserve FP32 in SM107 MXFP8 fallback (#37489)
Co-authored-by: Yangmin Li <yangminl@nvidia.com>
2026-09-02 17:55:14 -07:00
87d60a2229 Improve CUDA graph and speculative execution output handling (#37329)
Co-authored-by: jiayisuse <jiayisuse@fb.com>
Co-authored-by: Yinghai Lu <yinghai@meta.com>
Co-authored-by: Hao Zhang <zhisbug@users.noreply.github.com>
Co-authored-by: Yichao Fu <yichaofu@meta.com>
2026-09-02 17:25:27 -07:00
ff04a00d73 Reduce tokenizer overhead and offload CUDA VMM publication (#37330)
Co-authored-by: Shiyan Deng <dsy842974287@meta.com>
Co-authored-by: Yinghai Lu <yinghai@meta.com>
2026-09-02 17:21:08 -07:00
Liangsheng Yin 5c46ce37f5 [Fix] Apply the attention-CP broadcast result in PP dynamic-chunk profiling (#37669) 2026-09-02 17:20:07 -07:00
Cheng WanandClaude Opus 5 5ddca6819e Fix unified SWA: size a non-owner's v2p by the id space it must address (#37560)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 16:56:18 -07:00
Cheng WanandClaude Opus 5 d9848b9ecd Build the unified read stream directly, without the page-table rectangle (#37512)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 16:55:16 -07:00
Cheng WanandClaude Opus 5 18d5ffb42a Size the unified read-table grid from bs, and fuse the allocator's tombstone scatters (#37511)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 16:54:26 -07:00
c05f8ae830 [PD] Optimize paged allocator free-list release (#37146)
Co-authored-by: wangwenming.41 <wangwenming.41@jd.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-09-02 16:51:37 -07:00
Mohammad Miadh Angkad 718bd39fe0 [Fix] DP attention: correct the decode->extend prefix off-by-one (#37505) 2026-09-02 16:14:16 -07:00
paulzhang-tmandQiaolin-Yu 3fa6b86504 [Spec] Publish the final multi-layer EAGLE shared-read event (#36752)
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
2026-09-02 15:57:50 -07:00
YAMY 982aa8acfc [Bugfix] Load Qwen3.5 MTP embedding under PP (#37471) 2026-09-02 15:19:40 -07:00
YAMY 3c9cea8f10 [EAGLE] Prune draft-extend logits to selected rows (#35546) 2026-09-02 15:10:08 -07:00
cctry 3a855b050a fix(disagg): poll receivers during decode preallocation (#37483) 2026-09-02 14:26:40 -07:00
Liangsheng Yin 19c7679e9e [mem_cache] Make free_swa sync-free on page_size == 1 (#36723) 2026-09-02 14:18:22 -07:00
JoeandBBuf acea43079f Fix native MoE handling of noncontiguous top-k IDs (#36407)
Co-authored-by: BBuf <1182563586@qq.com>
2026-09-02 14:16:19 -07:00
f8cbf000f4 [AMD] Enable FP4 indexer for Deepseek V4 (#37353)
Co-authored-by: 1am9trash <1am9trash@gmail.com>
Co-authored-by: AMD-yanfeiwang <256076023+AMD-yanfeiwang@users.noreply.github.com>
Co-authored-by: Thomas Wang <thomawan@amd.com>
2026-09-02 09:45:08 -07:00
Po-Han Huang (NVIDIA) d585cec4bd Fix nondeterministic FlashInfer GDN alignment test (#37343) 2026-09-01 23:46:58 -07:00
Lianmin Zheng f8f04bafa8 Rust server: align launcher and request validation behavior (#37327) 2026-09-01 23:42:36 -07:00
jasonjk-park 83bd2c473f Allow custom policy for adaptive speculative decoding (#37274) 2026-09-01 23:15:21 -07:00
Liangsheng Yin 01c3a5f54f [misc] Resolve SWA ownership at enqueue time for grouped free() (#36646) 2026-09-01 23:06:55 -07:00
Liangsheng Yin 832d029870 [mem_cache] Split duplicate insert frees at the SWA eviction floor (#37481) 2026-09-01 23:01:40 -07:00
YAMY a6a19f9290 [Bugfix] Skip absent radix lock during cache cleanup (#37494) 2026-09-01 22:50:19 -07:00
Wes 2d9c64394f Fix reasoning metrics and add TPOT to bench_multiturn (#35443) 2026-09-02 11:28:47 +08:00
Xiaoyu Zhang 403a15c163 [CI] Batch CPU test workers (#37252) 2026-09-02 10:35:14 +08:00
26f760d5c0 [CPU] Support FP8 KV cache (#32733)
Co-authored-by: Zheng, Beilei <beilei.zheng@intel.com>
Co-authored-by: mingfeima <mingfei.ma@intel.com>
2026-09-02 10:20:54 +08:00
Yikai Zhangandamdpilot-upstream-sync 9978aaec8b [ROCm][Bugfix] Cap the DSA MQA-logits budget at AITER's buffer_store limit (#36960)
Co-authored-by: amdpilot-upstream-sync <amdpilot-upstream-sync@users.noreply.github.com>
2026-09-01 15:40:12 -07:00
Cheng Wan 0b1ce3d140 [Feature] Unified memory: support decode context parallelism for Kimi-Linear (#36890) 2026-09-01 12:44:26 -07:00
Shuwen Wang a2b8681d1d [CI][MLX] Restore the mamba_branching_seqlen attribute the MLX runner reads off a request (#37453) 2026-09-01 11:16:45 -07:00
cctryandcctry 9a05b470fa [Memory] Size the CUDA graph pool from warmup measurements and fix graph-pool borrowing (#36911)
Co-authored-by: cctry <cctry@fb.com>
2026-09-01 09:32:38 -07:00
datdo-msftandShangming Cai 49db27528a fix(test): deflake zmq load-snapshot round-trip tests (#35787)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
2026-09-01 18:14:55 +08:00
Ziang Li 5edcd0a445 [FlashInfer V0.6.18] feat(dsv4): support --dsa-topk-backend flashinfer with fused top-k (#33237) 2026-09-01 01:18:10 -07:00
Liangsheng Yin 3484f7f836 [mem_cache] Add free_kv_row to release a request's kv row by row range (#36721) 2026-09-01 01:14:43 -07:00
Xiaoyu ZhangandCursor 1c3ad92438 [Diffusion] Fuse FLUX.2 ModelOpt FP8 producers and QKV packing (#37162)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 16:14:29 +08:00
Ma Mingfei 5e79110122 [Fix][CPU] fix xeon ci failure by test_qwen35_flashinfer_fusion (#37338) 2026-09-01 14:17:35 +08:00
Liangsheng Yin 959ca033eb refactor(hicache): simplify decode offload state bookkeeping (#37299) 2026-08-31 23:09:41 -07:00