Xinyuan Tong
|
bed20249f1
|
fix(tool_call): reland schema type normalization (#26433)
|
2026-05-28 14:31:18 +08:00 |
|
 Bingxu Chenandyctseng0211
|
6f85957ff2
|
[AMD] Fix aiter checkout (rocm dockerfile) (#26544)
Co-authored-by: yctseng0211 <yctseng@amd.com>
|
2026-05-28 14:25:01 +08:00 |
|
 Mike QiuandMike_Qiu
|
8ff66b707f
|
feat: convert mm_hashes to str in encode_server for Mooncake key compat (#26487)
Co-authored-by: Mike_Qiu <qiudayu.qdy@antgroup.com>
|
2026-05-28 14:16:29 +08:00 |
|
YC Yen-Ching Tseng
|
505f37a63d
|
[AMD] force AITER checkout to bypass CSV CRLF/LF smudge dirty state (#26535)
|
2026-05-28 14:05:53 +08:00 |
|
Xiaoyu Zhang
|
e60f799b40
|
Enable Kimi-K2.5 piecewise CUDA graph (#26382)
|
2026-05-27 22:51:33 -07:00 |
|
Alison Shao
|
5018e5c969
|
[CI] /rerun-test: support glob wildcard patterns (#26422)
|
2026-05-27 21:55:58 -07:00 |
|
 Khoa PhamandClaude Opus 4.7
|
9040feebd8
|
[Gemma4] Add test for MTP models (#24552)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-27 21:36:28 -07:00 |
|
Jay Chun
|
b437a0d066
|
Fix PD decode radix cache double-counting cached_tokens (#25973)
|
2026-05-28 11:50:03 +08:00 |
|
 
|
9fea20a078
|
update XPU Dockerfile (#25174)
Signed-off-by: Matrix Yao <matrix.yao@intel.com>
Co-authored-by: MingxuZh <109504044+MingxuZh@users.noreply.github.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-05-28 10:58:09 +08:00 |
|
Chunyuan WU
|
714fdd9723
|
Fix MiniMax-M2.7 on CPU (#25061)
|
2026-05-28 10:53:13 +08:00 |
|
MingxuZh
|
21ba329dac
|
[Xeon] CPU CI enhancement for Intel Xeon platforms (#24649)
|
2026-05-28 10:49:04 +08:00 |
|
 Bingxu ChenandCursor
|
81663cb5f1
|
[AMD] [CI] Register MI35x GSM8K nightly tests (#26478)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-28 10:31:43 +08:00 |
|
Shaoting
|
14c1bb2721
|
[Feat][LMCache] Support LMCache mp mode (#24089)
Signed-off-by: Shaoting-Feng <stfeng@uw.edu>
|
2026-05-28 10:15:09 +08:00 |
|
jvzibro
|
421bda6d85
|
[Bug Fix] Remove H20 device check for FlashInfer AllReduce Fusion (#26470)
|
2026-05-28 00:45:55 +00:00 |
|
YAMY
|
eae03ce3b2
|
refactor(dsv4): route MHC prenorm through DeepGEMM wrapper (#26238)
|
2026-05-27 17:45:45 -07:00 |
|
 Baizhou ZhangandClaude Opus 4.7
|
68e5b4fdd6
|
[NemotronH] Fix weight-loading unit test broken by Puzzle support (#26522)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-27 17:44:45 -07:00 |
|
Netanel Haber
|
0abe6a85a5
|
Support NemotronHPuzzleForCausalLM (#24429)
Signed-off-by: Netanel Haber <58652339+netanel-haber@users.noreply.github.com>
|
2026-05-27 16:12:44 -07:00 |
|
Alison Shao
|
7c421ed3ec
|
[CI] runner-utilization: count in-flight queue waits + per-job status/links (#26509)
|
2026-05-27 16:06:55 -07:00 |
|
Qiaolin Yu
|
561e54f803
|
Update kimi k25 launch command in cookbook (#26511)
|
2026-05-27 16:04:04 -07:00 |
|
Mick
|
24bcb37efb
|
[diffusion] fix: fix diffusion LoRA consistency cases (#26327)
|
2026-05-28 06:29:35 +08:00 |
|
+2        
|
deaba74745
|
[AMD][DSV4] DSV4 MTP graph + sparse triton attn optimizations (#26383)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: amd-danli103 <danli103@amd.com>
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
Co-authored-by: yichiche@amd.com <jacky.cheng>
Co-authored-by: yctseng0211 <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-05-27 15:23:35 -07:00 |
|
Chunan Zeng
|
e06058ed62
|
[Kernel] Import flash_mla kernels from sglang kernel for deepseek v4 (#26499)
|
2026-05-27 14:32:44 -07:00 |
|
  
|
19663aafcd
|
Support batch size > 1 when enable CP (#23269)
Co-authored-by: Shunkang <182541032+Shunkangz@users.noreply.github.co>
Co-authored-by: Khoa Pham <khoa.pham@radixark.ai>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-05-27 14:11:17 -07:00 |
|
Kaixi Hou
|
ddf0627254
|
[NVIDIA] [GDN] Add FlashInfer prefill support for SM100+ (Blackwell) (#22921)
|
2026-05-27 13:58:05 -07:00 |
|
 sglang-botandsglang-bot
|
14f81a67d9
|
chore: bump sglang-kernel version to 0.4.3 (#26421)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-05-27 11:31:11 -07:00 |
|
jasonjk-park
|
d6e1692410
|
Allow custom speculative algorithm to support disaggregation (#26195)
|
2026-05-27 09:54:53 -07:00 |
|
 Sam HandYihao Wang
|
a95b4e2e09
|
[Feature] WebSocket streaming audio input for ASR (#22848)
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
|
2026-05-27 22:44:55 +08:00 |
|
 weireweireandZhangheng
|
034dd39189
|
Support KV events for UnifiedRadixCache (#26387)
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-05-27 22:10:40 +08:00 |
|
Zhangheng
|
5f8911183b
|
[UnifiedTree]: Update Unified Radix Cache README (#26485)
|
2026-05-27 21:50:53 +08:00 |
|
gjsheu
|
d9d719b270
|
[npu] [bugfix] Add contiguous operation during quantized weight loading. (#26309)
|
2026-05-27 19:55:58 +08:00 |
|
zhaozx-cn
|
83d5f4604c
|
[NPU]add decord2 for npu (#26308)
Signed-off-by: zhaozx-cn <zhaozx2116@163.com>
|
2026-05-27 19:48:16 +08:00 |
|
Cheng Wan
|
51840ca459
|
Add xutizhou as code owner for eplb directory (#26479)
|
2026-05-27 03:32:35 -07:00 |
|
loading66
|
a1ebc4917a
|
[NPU][DOCS]Add faq and feature Compatibilit (#26464)
|
2026-05-27 17:48:47 +08:00 |
|
 Makcum888eandronnie_zheng
|
3afc80d781
|
[diffusion] Fix multi image input for GLM-Image (#26311)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-05-27 12:33:20 +03:00 |
|
Jialin Ouyang
|
98bc6f3c22
|
API Perf: Replace pydantic per-element validation with C loop validation (#26355)
|
2026-05-27 02:04:07 -07:00 |
|
Jacob0226
|
d44584e8d8
|
[AMD] [CI] Add GLM-5.1 MXFP4 TP2 accuracy gate (#26396)
|
2026-05-27 01:49:48 -07:00 |
|
 Jacob0226andCursor
|
bf5bc23431
|
[AMD] [CI] Add DeepSeek-R1-0528 FP8 HiCache GSM8K test on MI35x (#26395)
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-27 01:48:38 -07:00 |
|
Liangsheng Yin
|
216ed270e5
|
refresh resolve_seq_lens_cpu comments (#26463)
|
2026-05-27 00:51:43 -07:00 |
|
Mick
|
f70e604101
|
[diffusion] fix: fix diffusion serve warmup defaults (#26247)
|
2026-05-27 15:43:47 +08:00 |
|
Liangsheng Yin
|
163b970127
|
[core] WAR barrier for overlap schedule buffer writes, without fwd occupancy cost (#26380)
|
2026-05-26 23:58:32 -07:00 |
|
 ant-yyand得泽
|
dea85c30f4
|
Add Ling_2_6 (#23837)
Signed-off-by: vito.yy <vito.yy@antgroup.com>
Co-authored-by: 得泽 <zhangkaihong.zkh@antgroup.com>
|
2026-05-27 14:57:23 +08:00 |
|
 Baizhou ZhangandClaude Opus 4.7
|
d6032c04b6
|
[docs] Fix V4 Pro balanced recipe (#26451)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-26 23:50:26 -07:00 |
|
Makcum888e
|
9060509214
|
[NPU] fix CI (#26390)
|
2026-05-27 09:45:42 +03:00 |
|
YC Yen-Ching Tseng
|
f32ca1e0e5
|
[AMD] AMD CI - temporarily change to mi325 (#26443)
|
2026-05-27 13:55:03 +08:00 |
|
xdtbynd
|
21d0e74aff
|
Disable torch.compile for NPU in speculative overlap utils (#26403)
|
2026-05-27 12:30:53 +08:00 |
|
Baizhou Zhang
|
0c34fc5ace
|
[Misc] Update CI Permission (#26435)
|
2026-05-26 19:57:17 -07:00 |
|
 
|
d45ee3f6c5
|
[HiCache] fix: Mooncake Dummy Client mode for hybrid Mamba models (#25278)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Teng Ma <stmatengss@users.noreply.github.com>
|
2026-05-27 10:29:52 +08:00 |
|
Liangsheng Yin
|
1051a8456f
|
[core] Maintain req_pool_indices_cpu host mirror (like seq_lens_cpu) (#26425)
|
2026-05-26 18:59:54 -07:00 |
|
Liangsheng Yin
|
6076066e38
|
Add mooncake_tcp transfer backend (mooncake over TCP) (#26346)
|
2026-05-26 18:15:55 -07:00 |
|
 vikram singh shekhawatandMa Mingfei
|
737c6cd6d1
|
[XPU] Add registry mechanism for XPU CI tests (#25405)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-05-27 08:56:59 +08:00 |
|