Commit Graph
2870 Commits
Author SHA1 Message Date
Cheng Wan 95d8a75bc9 Bundle set_kv_buffer write targets into KVWriteLoc (loc + swa_loc) (#27695) 2026-06-09 23:09:51 -07:00
Cheng Wan 758fd4bb9a [SWA] Cache full→SWA out_cache_loc per forward across attention backends (#27617) 2026-06-09 22:57:51 -07:00
2495c02c2c [Refactor] Cuda Graph Runner/Backend Refactor (#23906)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-06-09 21:36:57 -07:00
huangtingwei f101b287ef [Unified Tree]fix compatibility with eagle key and l3 hicache (#27655) 2026-06-10 10:54:45 +08:00
Michaelandmichaelzhang-ai f42a093261 [AMD] Migrate 2-GPU kernel allreduce tests into the registered system (#27722)
Co-authored-by: michaelzhang-ai <michaelzhang@example.com>
2026-06-09 19:39:03 -07:00
Mandepudi Rani Chowdary 7e3e616159 Add Arm64 INT8 MoE test coverage (#25007) 2026-06-10 10:36:57 +08:00
Ke Bao 854d232a40 Fix flaky hicache l3 mmlu nightly test (#27688) 2026-06-10 10:01:14 +08:00
Jianhong Zhang 77c4d53f19 [PD] Fix prefill bootstrap registration failure with --host 0.0.0.0 (#27608) 2026-06-10 09:15:26 +08:00
Liangsheng Yin f332e52611 Add UT guarding per-request bookkeeping clock ownership (#27710) 2026-06-09 17:11:49 -07:00
Lianmin Zheng ca716f4734 Add TP server GPU process regression test (#27721) 2026-06-09 16:25:27 -07:00
David Wang 4455abd164 dflash piecewise cuda graphs support (#27468) 2026-06-09 15:44:19 -07:00
decb88e0e3 Support spec v2 for Frozen-KV MTP; remove v1 worker (#27607)
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 15:30:20 -07:00
7f730edfdc fix: correct off-by-one in vocab boundary check for token validation (#22367)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: fzyzcjy <ch271828n@outlook.com>
2026-06-09 15:24:39 -07:00
ziang663 42322947aa [BUG FIX]Fix DSA CPU offload mamba indices signature (#27645) 2026-06-09 14:03:38 -07:00
Liangsheng Yin 186f1e300a [CI] Move JIT kernel tests + benchmarks to test/registered/jit; add in-package guard (#27644) 2026-06-09 12:37:39 -07:00
fzyzcjy 1368717248 Add more testing for chunked prefill (#27506) 2026-06-09 20:19:30 +08:00
fzyzcjy 609f5f549c Add mixed-prefix gsm8k eval and its CPU unit test (#27502) 2026-06-09 20:17:41 +08:00
2218622f50 Fix spec v2 stop output boundary (#25980)
Co-authored-by: gss <2783977641@qq.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-06-09 00:32:21 -07:00
d145a6127a fix: stop-string check misses early matches during speculative decoding (#23802)
Co-authored-by: xythink <xythink@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-06-08 23:25:58 -07:00
c2eae96c56 MSCCL++ Integration (#22734)
Co-authored-by: Caio Rocha <caiorocha@microsof.com>
Co-authored-by: empyreus <rjsouza1995@gmail.com>
2026-06-08 21:13:13 -07:00
Khoa Pham 0f8673851c test: fix gemma GSM8K thresholds in nightly text eval (#27342) 2026-06-08 19:58:42 -07:00
jianan-guandMa Mingfei db143e5212 [Intel GPU][Encoder] Add xpu_attn backend for encoder vision attention (#26460)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-06-09 09:47:44 +08:00
ashwini rathiandClaude Opus 4.7 009a0ceefa [XPU CI] Re-enable stage B with docker-pull flow and split tests (#27526)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-09 09:26:11 +08:00
Zhiy-Zhangandqiufan.zzy a09e85d677 Fix(spec): Fix the crash issue in the FA3 backend when running with top-k > 1 and page_size > 1 (#25077)
Co-authored-by: qiufan.zzy <qiufan.zzy@antgroup.com>
2026-06-08 17:14:29 -07:00
Liangsheng Yin 3fe6bc390b [Spec] Naming cleanup: contiguous draft-loc kernel + accepted->accept (#27599) 2026-06-08 15:04:58 -07:00
Khoa Pham c95179bc85 [Spec] Fuse small kenrels under gather_spec_extras (#27233) 2026-06-08 15:02:04 -07:00
Liangsheng Yin b5c64b94d5 [Spec] Rename token resolver to _resolve_spec_v2_tokens; remove dead V1 helpers (#27552) 2026-06-08 14:42:21 -07:00
YAMYandYuwei An ca66e6fb5e [BCG] Support breakable CUDA graph for DeepSeek V4 DP attention (#25195)
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
2026-06-08 13:54:58 -07:00
Liangsheng Yin 28c1a3cb45 [Spec] Deprecate Spec V1 (#25464) 2026-06-08 13:10:27 -07:00
Lianmin Zheng fca4ef9d69 Fix SWA pool resolution for EAGLE draft workers (#27491) 2026-06-08 11:00:29 -07:00
Elizaveta MartirosianandElizaveta Martirosian 593eb2e0fa [NPU] Fix CI (#27577)
Co-authored-by: Elizaveta Martirosian <elizaveta.martirosian@gmail.com>
2026-06-08 20:44:57 +03:00
Stella-17andxinyue.fan 12de907bc2 [MUSA][23/N] CI: Fix torchada preflight lock cleanup and add LLM server smoke test (#27242)
Co-authored-by: xinyue.fan <xinyue.fan@mthreads.com>
2026-06-08 08:45:16 -07:00
40030d8af8 [NPU] Add GitHub test summary and deduplicate test code. Part 2 (#24689)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Elizaveta Martirosian <elizaveta.martirosian@gmail.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-06-08 18:08:45 +03:00
Niko Ma 18d728967a [PD][MoRI] Drive KV transfers with a sharded synchronous worker pool (#26922) 2026-06-08 00:49:25 -07:00
fzyzcjy 71a0b10462 Fix the _chunked_req_scheduled_last_iter flag with a content-based stash gate (#26938) 2026-06-08 14:55:41 +08:00
fzyzcjy 259a2da3e0 Refactor Req.fill_ids into full_untruncated_fill_ids + fill_len with equivalence (#26637) 2026-06-08 14:52:18 +08:00
fzyzcjy f5fdf9c5d8 Speed up dump comparator percentile computation using numpy (#26874) 2026-06-08 14:50:31 +08:00
fzyzcjy 995e649190 Add parallel-rank dump filenames and pipeline-global layer remapping to dumper (#26850) 2026-06-08 14:49:22 +08:00
Han Yu 3d2165a286 Fix dual-chunk sparse fallback index overflow (#27361) 2026-06-07 23:15:37 -07:00
Liangsheng Yin f68c79675f Support topk > 1 tree drafting for mamba/hybrid-linear models on spec v2 (#27463) 2026-06-07 17:04:09 -07:00
Hubert Lu 10d33bd77e [AMD] Enable Piecewise CUDA Graph for AMD GPUs (#22299) 2026-06-07 16:28:20 -07:00
fzyzcjy 0a190d1c97 Fix PP is_fully_idle missing in-flight microbatches (#27446) 2026-06-07 17:38:26 +08:00
Zhangheng a39c428d3f [UnifiedTree][CI]: Reduce HiCache PP KL test concurrency to avoid decode OOM (#27483) 2026-06-07 17:12:42 +08:00
Liangsheng Yin 0ce3db3c0a [Bug] Fix out-of-range token id crashing tp=1 VocabParallelEmbedding (#27482) 2026-06-06 23:00:17 -07:00
Mohammad Miadh AngkadandLianmin Zheng 52f221cce0 Fix Req array token-id concatenation (#26182)
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-06-06 19:59:51 -07:00
Zhangheng fe548f36b0 [UnifiedTree]: Fix SWA admission budget under-counts HiCache load-back consumption (#27391) 2026-06-07 10:47:08 +08:00
shuwenn e57323cae9 [mem_cache][4/N] refactor: extract MambaTokenToKVPoolAllocator into allocator/ (#27256) 2026-06-07 10:46:29 +08:00
Liangsheng Yin 032c9efb46 Enable async-assert invariant probes by default in CI (#27461) 2026-06-06 16:23:21 -07:00
Qiaolin Yu 4b0f629082 [perf] reduce radix cache match overhead by changing the match algorithm (#27364) 2026-06-06 15:40:28 -07:00
Cheng WanandClaude Opus 4.8 9097647090 Route the eager forward path through the CUDA graph input-buffer registry (#27407)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-06 14:35:53 -07:00