Commit Graph
11824 Commits
Author SHA1 Message Date
ishandhananiandClaude Opus 4.6 2b0f349927 ci: clarify srt-slurm issue filing for incompatible flag combos (#22903)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 15:02:26 -07:00
Xinyu Zhangandxyuzh 13a2cd748d [Ray] Add data parallel (DP) and DP attention support to RayEngine (#21887)
Co-authored-by: xyuzh <xyuzh@users.noreply.github.com>
2026-04-15 15:00:48 -07:00
Sundara Raman Ramachandran 4927975427 [Score API] Add return_pooled_hidden_states to Scoring API for SequenceClassification / RewardModel (#22427) 2026-04-15 14:58:56 -07:00
Lee Nau 4e480d5785 Harden FlashInfer FP4 imports in standard dispatcher (#21776) 2026-04-15 14:54:49 -07:00
ishandhananiandClaude Opus 4.6 9497001b0c ci: add issue filing and suspect PR identification to log analyzer (#22899)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 14:27:14 -07:00
Liangsheng Yin efc267ca29 streaming session: trim spec v2 overshoot in cache_finished_req (#22897) 2026-04-15 14:15:46 -07:00
ishandhanani f61c332cba ci: log analyzer (#22859) 2026-04-15 14:10:00 -07:00
Lianmin Zheng 43925d179d [Speculative] Fix Eagle3/DFLASH aux hidden state capture during CUDA graph init (#22836) 2026-04-15 14:04:54 -07:00
Kurt Shuster 32d9fe5a32 [lora] Speedup triton backend sgemm calls with better grid (#22386) 2026-04-15 13:47:07 -07:00
Baizhou Zhang 113d654152 [Fix] Fix accuracy bug in Flashmla sparse MLA kernel (#22723) 2026-04-15 13:40:04 -07:00
Jimmy Shong 28e915b474 [Bugfix] Preserve auto-detected quant_config for GLM NextN draft model (#22823) 2026-04-15 13:25:36 -07:00
Yuhao Yang 8686f42acb [VLM] Enable per-image ViT cache and avoid TP CUDA context creation for Kimi-K2.5 (#22858) 2026-04-16 01:14:24 +08:00
huangtingweiandhzh0425 7d7fdc1309 [HiCache]Fix CP support for hybrid model (#22782)
Co-authored-by: hzh0425 <hzh0425@apache.org>
2026-04-15 23:50:29 +08:00
ybyang 9e84f53785 [PD] Add a fallback to bypass rust dep for mini_lb (#21982) 2026-04-15 22:34:36 +08:00
Xiaoyu Zhang 695ab705cb [diffusion] quant: update modelopt quantization docs and CI coverage (#22772) 2026-04-15 21:30:28 +08:00
Mick 80718492dd [diffusion] CI: reset thresholds (#22854) 2026-04-15 21:11:00 +08:00
Zhangheng 0a5c9728a1 [HiSparse][BugFix]: Fix the memory leak issue during health checks. (#22882) 2026-04-15 19:49:54 +08:00
Liangsheng Yin ce31934ca8 Streaming session: fix retract tail leak via _free_tail (#22862) 2026-04-15 01:44:27 -07:00
huangtingweiandhzh0425 3511c2deb4 [HiCache] Fix memory host free logic when share_indices_with_anchor enabled (#22767)
Co-authored-by: hzh0425 <hzh0425@apache.org>
2026-04-15 16:31:18 +08:00
Liangsheng Yin aa78564e1a Refactor streaming session abort handling (#22790) 2026-04-15 00:13:05 -07:00
jianzhao-xuandJianzhao Xu 45a83ffbe3 [NPU] Offloading docs update (#22860)
Co-authored-by: Jianzhao Xu <xujianchao@huawei.com>
2026-04-15 15:04:41 +08:00
Hubert LuandHaiShaw b2af34be54 [AMD] Optimize _append_shared_to_topk_output by a single fused Triton kernel for Qwen3.5 (#22844)
Co-authored-by: HaiShaw <hixiao@gmail.com>
2026-04-14 23:50:32 -07:00
Po-Han Huang (NVIDIA)andClaude Opus 4.6 ada52e5972 [Docs] Move ptxas sm_103a workaround into For CUDA 13 section (#22852)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 22:30:21 -07:00
Mick e95c2e73bd [diffusion] CI: refactor diffusion ci and reduce redundancy (#22810) 2026-04-15 10:12:29 +08:00
47ac830c07 [diffusion] rl: support standalone rollout api, denoising environment backpass and sp-aligned log-prob for T2I post-training (#22604)
Co-authored-by: MikukuOvO <mikukuovo@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 10:10:38 +08:00
Michael 39c6bf730c [AMD][CI] Add GLM-5-MXFP4 accuracy and perf nightly tests for MI35x (#21773) 2026-04-14 18:55:36 -07:00
Lianmin Zheng adb310b976 Cleanup server_args.py and minor code tidying (#22820) 2026-04-14 18:52:41 -07:00
ea05ea5abe [AMD] Enable share expert fusion with router experts for Qwen3.5 BF16 & FP8 (#20736)
Co-authored-by: Chen, Todd <zhenchen@amd.com>
Co-authored-by: jacky.cheng <yichiche@amd.com>
2026-04-14 18:52:36 -07:00
Piotr Mazurek 46c8a597ef [VLM] fix LFM2-VL offline inference and GPU JPEG decode (#22448) 2026-04-15 09:13:25 +08:00
ishandhanani 2c9e76d333 ci: skip approval for nightly gb200 runs, keep for manual triggers (#22768) 2026-04-14 16:34:57 -07:00
Alexis MacAskill e15401ee0e Add runai-model-streamer into Python packages installed in Dockerfile and fix NotADirectoryError Docker regression (#22537) 2026-04-14 16:25:41 -07:00
Lianmin Zheng 222eda1598 [Misc] Use cache_once for is_arch_support_pdl in sgl-kernel (#22725) 2026-04-14 15:22:10 -07:00
Jimmy Shong e83560562b Update CI Permissions (#22826) 2026-04-14 15:13:31 -07:00
8092431316 [serving] replace O(n²) stream_buffer string concat with integer offset (#22606)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-14 14:48:44 -07:00
Liangsheng Yin 36891ab514 Rename _alive_streaming_session_count; use _is_streaming helper (#22755) 2026-04-14 13:26:03 -07:00
Liangsheng Yin 0cb7295698 Fix streaming session busy-check double-counting via active_pool_idxs (#22753) 2026-04-14 13:11:06 -07:00
mingyue300 b4616dcbf5 [BugFix] Fix EAGLE speculative decoding missing grammar-based finish … (#21723) 2026-04-14 12:43:50 -07:00
MickandClaude Opus 4.6 d2f479e544 [diffusion] chore: auto-enable best parallel setting if unspecified (#22763)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 00:02:05 +08:00
Ke Bao 3c0a6c6987 Add page_size and SWA coverage to unified radix cache bench test (#22815) 2026-04-14 23:58:05 +08:00
Bi Xue 070c6a2489 [sgl] perf optimization for eplb (#21232) 2026-04-14 22:52:17 +08:00
Ke Bao 9f9e0231bb Refactor unified radix cache UT into parameterized test suite (#22812) 2026-04-14 22:34:33 +08:00
Mick c5e95080d2 [diffusion] model: support Ltx 2.3 two stage ti2v (#22667) 2026-04-14 22:10:08 +08:00
chx96642264 680bd4b429 [NPU] Modify the parameter name and optional values, and add the parameter restrictions. Modify some parameters supported type. (#22804) 2026-04-14 21:34:07 +08:00
McZyWuandroot 1588856e9b [NPU] qwen3next low latency best practice docs. (#22808)
Co-authored-by: root <root@localhost.localdomain>
2026-04-14 21:21:37 +08:00
amote-i ddc7daaf89 [NPU] [DOC] Update NPU docs to match latest code (#22796) 2026-04-14 21:10:28 +08:00
lawtherWu 454228e071 hicache storage backend mooncake support ascend hixl (#20016) 2026-04-14 20:51:06 +08:00
loading66 074c2a476d fix:[NPU]correct the full name of then Kimi model (#22799) 2026-04-14 20:15:22 +08:00
jianzhao-xuandJianzhao Xu 68dfffaaa3 Offloading docs update (#22795)
Co-authored-by: Jianzhao Xu <xujianchao@huawei.com>
2026-04-14 20:03:29 +08:00
xdtbynd 88253c39b0 [Docs] Fix formatting of tool-call-parser options (#22793) 2026-04-14 19:21:31 +08:00
amote-i 368cdfbe2f [NPU] [DOC] Fix outdated descriptions in the NPU documentation (#22707) 2026-04-14 19:21:15 +08:00