Commit Graph
13189 Commits
Author SHA1 Message Date
Mick 24bcb37efb [diffusion] fix: fix diffusion LoRA consistency cases (#26327) 2026-05-28 06:29:35 +08:00
+2 deaba74745 [AMD][DSV4] DSV4 MTP graph + sparse triton attn optimizations (#26383)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: amd-danli103 <danli103@amd.com>
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
Co-authored-by: yichiche@amd.com <jacky.cheng>
Co-authored-by: yctseng0211 <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
2026-05-27 15:23:35 -07:00
Chunan Zeng e06058ed62 [Kernel] Import flash_mla kernels from sglang kernel for deepseek v4 (#26499) 2026-05-27 14:32:44 -07:00
19663aafcd Support batch size > 1 when enable CP (#23269)
Co-authored-by: Shunkang <182541032+Shunkangz@users.noreply.github.co>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-05-27 14:11:17 -07:00
Kaixi Hou ddf0627254 [NVIDIA] [GDN] Add FlashInfer prefill support for SM100+ (Blackwell) (#22921) 2026-05-27 13:58:05 -07:00
sglang-botandsglang-bot 14f81a67d9 chore: bump sglang-kernel version to 0.4.3 (#26421)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-05-27 11:31:11 -07:00
jasonjk-park d6e1692410 Allow custom speculative algorithm to support disaggregation (#26195) 2026-05-27 09:54:53 -07:00
Sam HandYihao Wang a95b4e2e09 [Feature] WebSocket streaming audio input for ASR (#22848)
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
2026-05-27 22:44:55 +08:00
weireweireandZhangheng 034dd39189 Support KV events for UnifiedRadixCache (#26387)
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-05-27 22:10:40 +08:00
Zhangheng 5f8911183b [UnifiedTree]: Update Unified Radix Cache README (#26485) 2026-05-27 21:50:53 +08:00
gjsheu d9d719b270 [npu] [bugfix] Add contiguous operation during quantized weight loading. (#26309) 2026-05-27 19:55:58 +08:00
zhaozx-cn 83d5f4604c [NPU]add decord2 for npu (#26308)
Signed-off-by: zhaozx-cn <zhaozx2116@163.com>
2026-05-27 19:48:16 +08:00
Cheng Wan 51840ca459 Add xutizhou as code owner for eplb directory (#26479) 2026-05-27 03:32:35 -07:00
loading66 a1ebc4917a [NPU][DOCS]Add faq and feature Compatibilit (#26464) 2026-05-27 17:48:47 +08:00
Makcum888eandronnie_zheng 3afc80d781 [diffusion] Fix multi image input for GLM-Image (#26311)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-05-27 12:33:20 +03:00
Jialin Ouyang 98bc6f3c22 API Perf: Replace pydantic per-element validation with C loop validation (#26355) 2026-05-27 02:04:07 -07:00
Jacob0226 d44584e8d8 [AMD] [CI] Add GLM-5.1 MXFP4 TP2 accuracy gate (#26396) 2026-05-27 01:49:48 -07:00
Jacob0226andCursor bf5bc23431 [AMD] [CI] Add DeepSeek-R1-0528 FP8 HiCache GSM8K test on MI35x (#26395)
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-27 01:48:38 -07:00
Liangsheng Yin 216ed270e5 refresh resolve_seq_lens_cpu comments (#26463) 2026-05-27 00:51:43 -07:00
Mick f70e604101 [diffusion] fix: fix diffusion serve warmup defaults (#26247) 2026-05-27 15:43:47 +08:00
Liangsheng Yin 163b970127 [core] WAR barrier for overlap schedule buffer writes, without fwd occupancy cost (#26380) 2026-05-26 23:58:32 -07:00
ant-yyand得泽 dea85c30f4 Add Ling_2_6 (#23837)
Signed-off-by: vito.yy <vito.yy@antgroup.com>
Co-authored-by: 得泽 <zhangkaihong.zkh@antgroup.com>
2026-05-27 14:57:23 +08:00
Baizhou ZhangandClaude Opus 4.7 d6032c04b6 [docs] Fix V4 Pro balanced recipe (#26451)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 23:50:26 -07:00
Makcum888e 9060509214 [NPU] fix CI (#26390) 2026-05-27 09:45:42 +03:00
YC Yen-Ching Tseng f32ca1e0e5 [AMD] AMD CI - temporarily change to mi325 (#26443) 2026-05-27 13:55:03 +08:00
xdtbynd 21d0e74aff Disable torch.compile for NPU in speculative overlap utils (#26403) 2026-05-27 12:30:53 +08:00
Baizhou Zhang 0c34fc5ace [Misc] Update CI Permission (#26435) 2026-05-26 19:57:17 -07:00
d45ee3f6c5 [HiCache] fix: Mooncake Dummy Client mode for hybrid Mamba models (#25278)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Teng Ma <stmatengss@users.noreply.github.com>
2026-05-27 10:29:52 +08:00
Liangsheng Yin 1051a8456f [core] Maintain req_pool_indices_cpu host mirror (like seq_lens_cpu) (#26425) 2026-05-26 18:59:54 -07:00
Liangsheng Yin 6076066e38 Add mooncake_tcp transfer backend (mooncake over TCP) (#26346) 2026-05-26 18:15:55 -07:00
vikram singh shekhawatandMa Mingfei 737c6cd6d1 [XPU] Add registry mechanism for XPU CI tests (#25405)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-05-27 08:56:59 +08:00
87c3171aaa [CPU] Add support for Qwen3-vl and Qwen3-omni (#12662)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
2026-05-27 08:56:09 +08:00
Dawid Majchrowski c317beda99 [diffusion] model: support a new model (#24994) 2026-05-27 08:51:03 +08:00
Piotr Mazurek 468c565168 Wire YARN rope_parameters through LFM2 and LFM2-MoE attention (#26187) 2026-05-26 23:00:22 +00:00
Joel Schlosser 6989fede3c Purge usage of pytorch named tensors (#25911) 2026-05-26 14:58:57 -07:00
Yilong Zhaoandhappierpig 1a05b511e4 dp: refactor idle batch logic (#25025)
Co-authored-by: happierpig <zhaoyilong217@sjtu.edn.cn>
2026-05-26 14:22:25 -07:00
Qiaolin Yu dd6f073377 Reland "[perf][spec decoding] Skip full-vocab softmax in EAGLE draft when topk == 1 (#26235)" (#26397) 2026-05-26 14:14:48 -07:00
Serge PanevandYihao Wang 499eecce22 [NemotronH] V3 Omni wrapper: WeightsMapper + config round-trip (#25023)
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
2026-05-26 20:34:05 +00:00
Ziang Li 2b1e53c98d [RL] Fix FP8 skip matching for trailing-dot prefixes (#26287) 2026-05-26 20:30:08 +00:00
Zaili Wang 47617cc4df [CPU Doc]Add Xeon CPU info in Qwen3 Cookbook (#25971) 2026-05-26 12:14:07 -07:00
sglang-botandsglang-bot 0753182b50 chore: bump sgl-kernel version to 0.4.3 (#26414)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-05-26 12:10:16 -07:00
zijiexiaandClaude Opus 4.7 6afebc278a [docs] DeepSeek-V4 cookbook: note cu129 image for GB200 Pro DeepEP backend (#26413)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-26 12:08:48 -07:00
Chunan Zeng b66f8e0b96 Sgl flashmla (#26132) 2026-05-26 12:00:23 -07:00
Liangsheng Yin ec6f8d61f7 [Spec] Async-assert probes across EAGLE/MTP; zero tgt_cache_loc (#26335) 2026-05-26 11:34:03 -07:00
Venkatesh GuduruandvguduruTT 6c8128650e [Bugfix] Fix flashinfer_cutlass MoE crash when intermediate_size_per_partition is not 16-aligned (#22627)
Co-authored-by: vguduruTT <venkatesh.guduru@mulitcorewareinc.com>
2026-05-26 16:58:33 +00:00
Zhangheng 38f32c38ab [UnifiedRadixTree]: Support L3 HiStorage framework (#26062) 2026-05-26 22:38:00 +08:00
Jun Liu c47f0e7cdd [PD] Fix top logprobs crash in prefill path (#26299) 2026-05-26 22:01:10 +08:00
Shangming Cai c8c1aed5e9 [PD] Fix cross-rank queue divergence by gating metadata readiness before all-reduce (#26394)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
2026-05-26 21:59:42 +08:00
Yongfei Xu 98eb84497d [PP] Skip PP output communication for pure chunked prefill batches (#26148) 2026-05-26 21:59:18 +08:00
zijiexia a26913158b fix(ci): enforce legacy docs/ gate in Lint workflow (#26322) 2026-05-26 20:06:54 +08:00