Commit Graph
5092 Commits
Author SHA1 Message Date
luoroger37 54cba63b6c Fix paged SWA free mapping cleanup (#27779) 2026-06-11 18:25:24 +08:00
Cheng Wan 9425ba2478 Remove extra_pytest_path from pr-test.yml (#27856) 2026-06-10 21:16:14 -07:00
Zhangheng 5e0271536a [UnifiedTree]: HybridModel launches HiCache via UnifiedTree by default. (#27759) 2026-06-11 12:03:36 +08:00
Cheng Wan f4f30d7d23 [Fix] Use int64 seq_lens across all CUDA graph runners and backends (#27840) 2026-06-10 19:54:58 -07:00
David Wang 588d1f7bc9 [Feature] Spec V2 DFlash Support (#23000) 2026-06-10 19:27:42 -07:00
Yongji Wuandhzh0425 db061e97c0 [Unified] Fix UnifiedRadixCache write_backup issue in write-back mode(#27108)
Co-authored-by: hzh0425 <hzh0425@apache.org>
2026-06-11 08:28:24 +08:00
740305e1d9 [HiCache] Add opt-in LRU eviction to file storage backend (CP-aware) (#26670)
Co-authored-by: Zhangheng <hzh0425@apache.org>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-11 08:21:59 +08:00
Baizhou Zhang 125ef88892 Disable async assert in Nemotron nightly tests (#27838) 2026-06-10 16:28:14 -07:00
Liangsheng Yin 16124fc9b2 [Metrics] Fix fwd_occupancy reading NaN on every decode log line; probe-free base-a (#27836) 2026-06-10 15:42:19 -07:00
Mohammad Miadh Angkad bdf3ef6421 [CI] Fix registered QK Gemma RMSNorm test location (#27839) 2026-06-10 15:21:59 -07:00
Baizhou Zhang 3c1b0fb226 [1/n] [CP] Simplify prefill context parallel server args (#27312) 2026-06-10 14:11:38 -07:00
shuwennandQiaolin Yu 3600a9ac5f [SPEC] feat: init adaptive spec params from config (#27493)
Co-authored-by: Qiaolin Yu <liin1211@outlook.com>
2026-06-10 18:55:25 +00:00
billishyahaoandHAI 0ae27405d0 [AMD] Support eplb for moriep (#22985)
Co-authored-by: HAI <hixiao@gmail.com>
2026-06-10 10:23:51 -07:00
Kai-Hsun Chen 8c6bbe0658 [deepseek] Enable DP attention + TBO + shared experts fusion (#27510) 2026-06-10 09:42:27 -07:00
Michael 6a16f29af6 [AMD] ci: register 8 framework / unit tests to run on AMD CI (#25939) 2026-06-10 08:41:22 -07:00
Yuan Luoandluoyuan.luo 518e35fae7 [KDA] Add CuteDSL Prefill Kernel on SM100 (#27488)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
2026-06-10 21:25:19 +08:00
Mohammad Miadh Angkad 4faaa9ba92 [CI] Fix stale ngram bookkeeping owner sites (#27803) 2026-06-10 04:51:46 -07:00
111009ea54 [Feature] [Ngram spec] Support ngram spec v2 (#17260)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Ratish P <114130421+Ratish1@users.noreply.github.com>
2026-06-10 02:46:00 -07:00
ChengYao-amdandgithub-actions[bot] 255843d454 Support for Zyphra zaya1 model (#26347)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-06-10 02:44:47 -07:00
DarkSharpnessandClaude b40f365732 [CI] Move misplaced mhc kernel test into test/registered/kernels (#27781)
Co-authored-by: Claude <noreply@anthropic.com>
2026-06-10 01:49:18 -07:00
Ziang Li 01f10acd06 Implement online nvfp4 quantization (#26083) 2026-06-10 00:26:51 -07:00
nbarzilie e76e4959b5 [CI][PD] Add unit tests for nixl backend (#26908) 2026-06-10 14:31:59 +08:00
Cheng Wan 95d8a75bc9 Bundle set_kv_buffer write targets into KVWriteLoc (loc + swa_loc) (#27695) 2026-06-09 23:09:51 -07:00
Cheng Wan 758fd4bb9a [SWA] Cache full→SWA out_cache_loc per forward across attention backends (#27617) 2026-06-09 22:57:51 -07:00
2495c02c2c [Refactor] Cuda Graph Runner/Backend Refactor (#23906)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-06-09 21:36:57 -07:00
huangtingwei f101b287ef [Unified Tree]fix compatibility with eagle key and l3 hicache (#27655) 2026-06-10 10:54:45 +08:00
Michaelandmichaelzhang-ai f42a093261 [AMD] Migrate 2-GPU kernel allreduce tests into the registered system (#27722)
Co-authored-by: michaelzhang-ai <michaelzhang@example.com>
2026-06-09 19:39:03 -07:00
Mandepudi Rani Chowdary 7e3e616159 Add Arm64 INT8 MoE test coverage (#25007) 2026-06-10 10:36:57 +08:00
Ke Bao 854d232a40 Fix flaky hicache l3 mmlu nightly test (#27688) 2026-06-10 10:01:14 +08:00
Jianhong Zhang 77c4d53f19 [PD] Fix prefill bootstrap registration failure with --host 0.0.0.0 (#27608) 2026-06-10 09:15:26 +08:00
Liangsheng Yin f332e52611 Add UT guarding per-request bookkeeping clock ownership (#27710) 2026-06-09 17:11:49 -07:00
Lianmin Zheng ca716f4734 Add TP server GPU process regression test (#27721) 2026-06-09 16:25:27 -07:00
David Wang 4455abd164 dflash piecewise cuda graphs support (#27468) 2026-06-09 15:44:19 -07:00
decb88e0e3 Support spec v2 for Frozen-KV MTP; remove v1 worker (#27607)
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 15:30:20 -07:00
7f730edfdc fix: correct off-by-one in vocab boundary check for token validation (#22367)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: fzyzcjy <ch271828n@outlook.com>
2026-06-09 15:24:39 -07:00
ziang663 42322947aa [BUG FIX]Fix DSA CPU offload mamba indices signature (#27645) 2026-06-09 14:03:38 -07:00
Liangsheng Yin 186f1e300a [CI] Move JIT kernel tests + benchmarks to test/registered/jit; add in-package guard (#27644) 2026-06-09 12:37:39 -07:00
fzyzcjy 1368717248 Add more testing for chunked prefill (#27506) 2026-06-09 20:19:30 +08:00
fzyzcjy 609f5f549c Add mixed-prefix gsm8k eval and its CPU unit test (#27502) 2026-06-09 20:17:41 +08:00
2218622f50 Fix spec v2 stop output boundary (#25980)
Co-authored-by: gss <2783977641@qq.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-06-09 00:32:21 -07:00
d145a6127a fix: stop-string check misses early matches during speculative decoding (#23802)
Co-authored-by: xythink <xythink@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-06-08 23:25:58 -07:00
c2eae96c56 MSCCL++ Integration (#22734)
Co-authored-by: Caio Rocha <caiorocha@microsof.com>
Co-authored-by: empyreus <rjsouza1995@gmail.com>
2026-06-08 21:13:13 -07:00
Khoa Pham 0f8673851c test: fix gemma GSM8K thresholds in nightly text eval (#27342) 2026-06-08 19:58:42 -07:00
jianan-guandMa Mingfei db143e5212 [Intel GPU][Encoder] Add xpu_attn backend for encoder vision attention (#26460)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-06-09 09:47:44 +08:00
ashwini rathiandClaude Opus 4.7 009a0ceefa [XPU CI] Re-enable stage B with docker-pull flow and split tests (#27526)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-09 09:26:11 +08:00
Zhiy-Zhangandqiufan.zzy a09e85d677 Fix(spec): Fix the crash issue in the FA3 backend when running with top-k > 1 and page_size > 1 (#25077)
Co-authored-by: qiufan.zzy <qiufan.zzy@antgroup.com>
2026-06-08 17:14:29 -07:00
Liangsheng Yin 3fe6bc390b [Spec] Naming cleanup: contiguous draft-loc kernel + accepted->accept (#27599) 2026-06-08 15:04:58 -07:00
Khoa Pham c95179bc85 [Spec] Fuse small kenrels under gather_spec_extras (#27233) 2026-06-08 15:02:04 -07:00
Liangsheng Yin b5c64b94d5 [Spec] Rename token resolver to _resolve_spec_v2_tokens; remove dead V1 helpers (#27552) 2026-06-08 14:42:21 -07:00
YAMYandYuwei An ca66e6fb5e [BCG] Support breakable CUDA graph for DeepSeek V4 DP attention (#25195)
Co-authored-by: Yuwei An <ayw.sirius19@gmail.com>
2026-06-08 13:54:58 -07:00