 weireweireandZhangheng
|
034dd39189
|
Support KV events for UnifiedRadixCache (#26387)
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-05-27 22:10:40 +08:00 |
|
Zhangheng
|
5f8911183b
|
[UnifiedTree]: Update Unified Radix Cache README (#26485)
|
2026-05-27 21:50:53 +08:00 |
|
gjsheu
|
d9d719b270
|
[npu] [bugfix] Add contiguous operation during quantized weight loading. (#26309)
|
2026-05-27 19:55:58 +08:00 |
|
zhaozx-cn
|
83d5f4604c
|
[NPU]add decord2 for npu (#26308)
Signed-off-by: zhaozx-cn <zhaozx2116@163.com>
|
2026-05-27 19:48:16 +08:00 |
|
Cheng Wan
|
51840ca459
|
Add xutizhou as code owner for eplb directory (#26479)
|
2026-05-27 03:32:35 -07:00 |
|
loading66
|
a1ebc4917a
|
[NPU][DOCS]Add faq and feature Compatibilit (#26464)
|
2026-05-27 17:48:47 +08:00 |
|
 Makcum888eandronnie_zheng
|
3afc80d781
|
[diffusion] Fix multi image input for GLM-Image (#26311)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-05-27 12:33:20 +03:00 |
|
Jialin Ouyang
|
98bc6f3c22
|
API Perf: Replace pydantic per-element validation with C loop validation (#26355)
|
2026-05-27 02:04:07 -07:00 |
|
Jacob0226
|
d44584e8d8
|
[AMD] [CI] Add GLM-5.1 MXFP4 TP2 accuracy gate (#26396)
|
2026-05-27 01:49:48 -07:00 |
|
 Jacob0226andCursor
|
bf5bc23431
|
[AMD] [CI] Add DeepSeek-R1-0528 FP8 HiCache GSM8K test on MI35x (#26395)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-27 01:48:38 -07:00 |
|
Liangsheng Yin
|
216ed270e5
|
refresh resolve_seq_lens_cpu comments (#26463)
|
2026-05-27 00:51:43 -07:00 |
|
Mick
|
f70e604101
|
[diffusion] fix: fix diffusion serve warmup defaults (#26247)
|
2026-05-27 15:43:47 +08:00 |
|
Liangsheng Yin
|
163b970127
|
[core] WAR barrier for overlap schedule buffer writes, without fwd occupancy cost (#26380)
|
2026-05-26 23:58:32 -07:00 |
|
 ant-yyand得泽
|
dea85c30f4
|
Add Ling_2_6 (#23837)
Signed-off-by: vito.yy <vito.yy@antgroup.com>
Co-authored-by: 得泽 <zhangkaihong.zkh@antgroup.com>
|
2026-05-27 14:57:23 +08:00 |
|
 Baizhou ZhangandClaude Opus 4.7
|
d6032c04b6
|
[docs] Fix V4 Pro balanced recipe (#26451)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-26 23:50:26 -07:00 |
|
Makcum888e
|
9060509214
|
[NPU] fix CI (#26390)
|
2026-05-27 09:45:42 +03:00 |
|
YC Yen-Ching Tseng
|
f32ca1e0e5
|
[AMD] AMD CI - temporarily change to mi325 (#26443)
|
2026-05-27 13:55:03 +08:00 |
|
xdtbynd
|
21d0e74aff
|
Disable torch.compile for NPU in speculative overlap utils (#26403)
|
2026-05-27 12:30:53 +08:00 |
|
Baizhou Zhang
|
0c34fc5ace
|
[Misc] Update CI Permission (#26435)
|
2026-05-26 19:57:17 -07:00 |
|
 
|
d45ee3f6c5
|
[HiCache] fix: Mooncake Dummy Client mode for hybrid Mamba models (#25278)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Teng Ma <stmatengss@users.noreply.github.com>
|
2026-05-27 10:29:52 +08:00 |
|
Liangsheng Yin
|
1051a8456f
|
[core] Maintain req_pool_indices_cpu host mirror (like seq_lens_cpu) (#26425)
|
2026-05-26 18:59:54 -07:00 |
|
Liangsheng Yin
|
6076066e38
|
Add mooncake_tcp transfer backend (mooncake over TCP) (#26346)
|
2026-05-26 18:15:55 -07:00 |
|
 vikram singh shekhawatandMa Mingfei
|
737c6cd6d1
|
[XPU] Add registry mechanism for XPU CI tests (#25405)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-05-27 08:56:59 +08:00 |
|
 
|
87c3171aaa
|
[CPU] Add support for Qwen3-vl and Qwen3-omni (#12662)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-05-27 08:56:09 +08:00 |
|
Dawid Majchrowski
|
c317beda99
|
[diffusion] model: support a new model (#24994)
|
2026-05-27 08:51:03 +08:00 |
|
Piotr Mazurek
|
468c565168
|
Wire YARN rope_parameters through LFM2 and LFM2-MoE attention (#26187)
|
2026-05-26 23:00:22 +00:00 |
|
Joel Schlosser
|
6989fede3c
|
Purge usage of pytorch named tensors (#25911)
|
2026-05-26 14:58:57 -07:00 |
|
 Yilong Zhaoandhappierpig
|
1a05b511e4
|
dp: refactor idle batch logic (#25025)
Co-authored-by: happierpig <zhaoyilong217@sjtu.edn.cn>
|
2026-05-26 14:22:25 -07:00 |
|
Qiaolin Yu
|
dd6f073377
|
Reland "[perf][spec decoding] Skip full-vocab softmax in EAGLE draft when topk == 1 (#26235)" (#26397)
|
2026-05-26 14:14:48 -07:00 |
|
 Serge PanevandYihao Wang
|
499eecce22
|
[NemotronH] V3 Omni wrapper: WeightsMapper + config round-trip (#25023)
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
|
2026-05-26 20:34:05 +00:00 |
|
Ziang Li
|
2b1e53c98d
|
[RL] Fix FP8 skip matching for trailing-dot prefixes (#26287)
|
2026-05-26 20:30:08 +00:00 |
|
Zaili Wang
|
47617cc4df
|
[CPU Doc]Add Xeon CPU info in Qwen3 Cookbook (#25971)
|
2026-05-26 12:14:07 -07:00 |
|
 sglang-botandsglang-bot
|
0753182b50
|
chore: bump sgl-kernel version to 0.4.3 (#26414)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-05-26 12:10:16 -07:00 |
|
 zijiexiaandClaude Opus 4.7
|
6afebc278a
|
[docs] DeepSeek-V4 cookbook: note cu129 image for GB200 Pro DeepEP backend (#26413)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-26 12:08:48 -07:00 |
|
Chunan Zeng
|
b66f8e0b96
|
Sgl flashmla (#26132)
|
2026-05-26 12:00:23 -07:00 |
|
Liangsheng Yin
|
ec6f8d61f7
|
[Spec] Async-assert probes across EAGLE/MTP; zero tgt_cache_loc (#26335)
|
2026-05-26 11:34:03 -07:00 |
|
 Venkatesh GuduruandvguduruTT
|
6c8128650e
|
[Bugfix] Fix flashinfer_cutlass MoE crash when intermediate_size_per_partition is not 16-aligned (#22627)
Co-authored-by: vguduruTT <venkatesh.guduru@mulitcorewareinc.com>
|
2026-05-26 16:58:33 +00:00 |
|
Zhangheng
|
38f32c38ab
|
[UnifiedRadixTree]: Support L3 HiStorage framework (#26062)
|
2026-05-26 22:38:00 +08:00 |
|
Jun Liu
|
c47f0e7cdd
|
[PD] Fix top logprobs crash in prefill path (#26299)
|
2026-05-26 22:01:10 +08:00 |
|
Shangming Cai
|
c8c1aed5e9
|
[PD] Fix cross-rank queue divergence by gating metadata readiness before all-reduce (#26394)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-05-26 21:59:42 +08:00 |
|
Yongfei Xu
|
98eb84497d
|
[PP] Skip PP output communication for pure chunked prefill batches (#26148)
|
2026-05-26 21:59:18 +08:00 |
|
zijiexia
|
a26913158b
|
fix(ci): enforce legacy docs/ gate in Lint workflow (#26322)
|
2026-05-26 20:06:54 +08:00 |
|
Michael
|
9409969fd5
|
Revert "[perf][spec decoding] Skip full-vocab softmax in EAGLE draft when topk == 1 (#26235)" (#26358)
|
2026-05-26 02:47:52 -07:00 |
|
fzyzcjy
|
d9c82934c8
|
Extract Scheduler init methods and add skills to enforce the splitting requirements (#26271)
|
2026-05-26 17:45:09 +08:00 |
|
Chao Shi
|
48f3264807
|
[HiCache]: Check return code of cudaHostRegister (#26301)
|
2026-05-26 17:44:15 +08:00 |
|
YC Yen-Ching Tseng
|
d25a220fdb
|
[AMD] Relaxing timeout for AMD CI (#26392)
|
2026-05-26 17:25:05 +08:00 |
|
roikoren755
|
e958f4561f
|
[feat] Support extra_buffer in Mamba2-based models (#15829)
Signed-off-by: Roi Koren <roik@nvidia.com>
|
2026-05-26 16:03:29 +08:00 |
|
Liangsheng Yin
|
7e6e5efe51
|
Revert "fix(tool_call): normalize non-standard JSON Schema types in tool params" (#26379)
|
2026-05-26 00:49:41 -07:00 |
|
Zhonghua Deng
|
dabdd91ef3
|
[EPD] Cross-request batching for image/audio encoder (#25964)
|
2026-05-26 15:38:59 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
d34d4d9f5f
|
[GDN] Support SM100 CuTeDSL GDN Prefill Kernel (#26200)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-05-26 15:38:29 +08:00 |
|