Commit Graph
13818 Commits
Author SHA1 Message Date
DarkSharpnessandClaude b40f365732 [CI] Move misplaced mhc kernel test into test/registered/kernels (#27781)
Co-authored-by: Claude <noreply@anthropic.com>
2026-06-10 01:49:18 -07:00
Cheng Wan 70c71ba183 [NPU] Fix dead patch_model monkey-patch breaking NPU torch.compile capture (#27774) 2026-06-10 01:15:50 -07:00
Ziang Li 01f10acd06 Implement online nvfp4 quantization (#26083) 2026-06-10 00:26:51 -07:00
nbarzilie e76e4959b5 [CI][PD] Add unit tests for nixl backend (#26908) 2026-06-10 14:31:59 +08:00
zijiexiaandClaude Opus 4.8 1a5775a9df [Docs] Remove the legacy release-docs.yml deploy workflow (#27766)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 23:18:41 -07:00
Mick e8a437ef26 [diffusion] doc: update docs architecture (#27767) 2026-06-10 14:18:10 +08:00
Thomas WangandXinyi Song f2bcdb0508 [AMD] Add unified kv attention support in dpsk-v4 (#27380)
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
2026-06-09 23:13:37 -07:00
Cheng Wan 95d8a75bc9 Bundle set_kv_buffer write targets into KVWriteLoc (loc + swa_loc) (#27695) 2026-06-09 23:09:51 -07:00
Cheng Wan 758fd4bb9a [SWA] Cache full→SWA out_cache_loc per forward across attention backends (#27617) 2026-06-09 22:57:51 -07:00
Aleksi Vesanto 08ceb96ea5 [diffusion] fix: remove boolean arithmetic guard to fix compiling (#27065) 2026-06-10 13:55:47 +08:00
Bingxu Chen 4704b10d0d [AMD] Update MoRI to v1.2.0 (#27669) 2026-06-09 22:41:53 -07:00
Liangsheng Yin d1895cb60d [Spec] Extract move_accept_tokens_to_target_kvcache into spec_utils (#27764) 2026-06-09 21:55:26 -07:00
2495c02c2c [Refactor] Cuda Graph Runner/Backend Refactor (#23906)
Co-authored-by: BBuf <1182563586@qq.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-06-09 21:36:57 -07:00
Mick 56f06278c6 [diffusion] refactor: refactor realtime control state and adapters (#27698) 2026-06-10 12:27:16 +08:00
Baizhou Zhang 047e5df3b1 Revert "Share BCG output buffers across capture sizes" (#27758) 2026-06-09 20:51:40 -07:00
Lianmin Zheng 165331a200 Share BCG output buffers across capture sizes (#27659) 2026-06-09 20:33:46 -07:00
fatSheep d21c31f681 fix: forward update_mamba_state_after_mtp_verify in HybridAttnBackend (#25883) 2026-06-09 20:06:50 -07:00
huangtingwei f101b287ef [Unified Tree]fix compatibility with eagle key and l3 hicache (#27655) 2026-06-10 10:54:45 +08:00
Michaelandmichaelzhang-ai f42a093261 [AMD] Migrate 2-GPU kernel allreduce tests into the registered system (#27722)
Co-authored-by: michaelzhang-ai <michaelzhang@example.com>
2026-06-09 19:39:03 -07:00
Mandepudi Rani Chowdary 7e3e616159 Add Arm64 INT8 MoE test coverage (#25007) 2026-06-10 10:36:57 +08:00
Ke Bao 854d232a40 Fix flaky hicache l3 mmlu nightly test (#27688) 2026-06-10 10:01:14 +08:00
zijiexia 6565b7c464 [Docs] Update MegaMoE handling and rerun benchmarks (#27726) 2026-06-09 19:00:43 -07:00
6110ed671f ci(xpu): clean build artifacts in cleanup (#27648)
Co-authored-by: Patil, Jitendra <jitendra.patil@intel.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-10 09:46:54 +08:00
sushil Dubey 5809bbe35d Mistral3 add tensor parallel support for diffusion text encoder (#25950) 2026-06-10 09:43:21 +08:00
Mick af55025644 [diffusion] refactor: refactor realtime and model-specific stage modules (#27697) 2026-06-10 09:39:06 +08:00
bcd9c5a903 update pytorch-xpu to 2.12 (#27133)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: MingxuZh <109504044+MingxuZh@users.noreply.github.com>
2026-06-10 09:25:30 +08:00
Jianhong Zhang 77c4d53f19 [PD] Fix prefill bootstrap registration failure with --host 0.0.0.0 (#27608) 2026-06-10 09:15:26 +08:00
iridiumine 2947781ce6 [NPU] MiMo-V2-Flash Adaptation (#25455) 2026-06-10 09:13:55 +08:00
Jan Bernlöhr f3ecc3688f Fix Gemma3 ModelOpt kv-scale loading (#25794) 2026-06-09 17:53:21 -07:00
Liangsheng Yin f332e52611 Add UT guarding per-request bookkeeping clock ownership (#27710) 2026-06-09 17:11:49 -07:00
Muqi Liandzqlcode 365b7dab9a fix(schema): update tokens_after_end (#27017)
Co-authored-by: zqlcode <1309223143@qq.com>
2026-06-09 16:52:35 -07:00
Mohammad Miadh Angkad bc82086ef8 Remove FlashInfer GB transport workaround (#27453)
Signed-off-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
2026-06-09 16:48:49 -07:00
Wenqiandwenqi 98fe7e326e fix(gemma4): register image/video/audio token_regex for HF-expanded prompts (#26320)
Co-authored-by: wenqi <wenqi@convergence.ai>
2026-06-09 16:30:13 -07:00
Thomas Wang 9ab7a64ee1 [AMD] Update amd qwen3.5 cookbook (#27660) 2026-06-09 16:26:33 -07:00
Lianmin Zheng ca716f4734 Add TP server GPU process regression test (#27721) 2026-06-09 16:25:27 -07:00
Yueming Yuan 53a4b51f8c Fix GLM NextN draft value head dim (#26049) 2026-06-09 16:13:37 -07:00
David Wang 4455abd164 dflash piecewise cuda graphs support (#27468) 2026-06-09 15:44:19 -07:00
decb88e0e3 Support spec v2 for Frozen-KV MTP; remove v1 worker (#27607)
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 15:30:20 -07:00
7f730edfdc fix: correct off-by-one in vocab boundary check for token validation (#22367)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: fzyzcjy <ch271828n@outlook.com>
2026-06-09 15:24:39 -07:00
weizhoublue fde4004429 [Fix] Reset positions tensor in CUDA graph runner when batch size differs from captured size (#24401)
Signed-off-by: weizhoublue <weizhou.lan@daocloud.io>
2026-06-09 15:03:58 -07:00
jacky.cheng 2fef951fe8 [AMD] Replace fp8 mla with fp8 mha kernel for diffusion model aiter backend (#23927) 2026-06-09 14:46:41 -07:00
ziang663 42322947aa [BUG FIX]Fix DSA CPU offload mamba indices signature (#27645) 2026-06-09 14:03:38 -07:00
Lianmin Zhengandlmzheng eb8dceda44 Defer DeepGEMM PDL setup to worker init (#27671)
Co-authored-by: lmzheng <lmzheng@fb.com>
2026-06-09 13:52:30 -07:00
Liangsheng Yin 186f1e300a [CI] Move JIT kernel tests + benchmarks to test/registered/jit; add in-package guard (#27644) 2026-06-09 12:37:39 -07:00
Bi Xue 8ae328e5f0 [sgl] Fix kimi-k2.5 EAGLE3 MLA draft embeds for batched MM prefill (#27647) 2026-06-09 11:26:48 -07:00
Michael 5babb902a9 [AMD] fix: handle per-frame 4D shift in native scale-shift kernel (#27581) 2026-06-09 10:31:00 -07:00
Xiaoyu ZhangandClaude Opus 4.8 aa18a68ac5 [diffusion] Run LTX-2 VAE decode in channels_last_3d (faster decode, lower peak memory) (#27431)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 23:26:40 +08:00
YC Yen-Ching Tseng 17d8c5801d [AMD] Fix test_deepseek_r1_mxfp4_8gpu.py : disable async-assert probes on AMD CI (#27505) 2026-06-09 08:12:35 -07:00
Kangyan-ZhouandClaude Opus 4.8 badab6b136 [router] Add request/TTFT/worker metrics + Grafana dashboard to experimental sgl-router (#27591)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-09 07:25:56 -07:00
fzyzcjy 1368717248 Add more testing for chunked prefill (#27506) 2026-06-09 20:19:30 +08:00