Commit Graph
11824 Commits
Author SHA1 Message Date
Byron HsuandByron Hsu cf9845f8e3 [Bug Fix] Ensure prefill_info_table is populated before honoring disagg_prefill_dp_rank (#22990)
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
2026-04-17 11:10:31 +08:00
Jan Bernlöhr 04a53955b9 feat: add coordinated checkpoint prefetch for network filesystem loading (#20843) 2026-04-16 20:08:19 -07:00
Yuhao Yang a77abbe005 [VLM] Reduce GPU memory footprint of CUDA IPC MM feature transport (#22662) 2026-04-17 10:38:36 +08:00
Khoa Pham 5a0eea8ac5 [CI] Adding Gemma 4 to Nightly CI (#22408) 2026-04-16 19:30:16 -07:00
Yuxuan Zhang 16d11c2a10 Fix for the low-probability garbled output issue in the GLM-5 series models. (#22811) 2026-04-17 09:52:13 +08:00
Alison Shao 0052093178 test(4-gpu-b200): split test_qwen35_models.py + bump partitions 5→6 (#22913) 2026-04-16 18:51:59 -07:00
Makcum888e e353630b57 [Diffusion] [NPU] Fix multimodal gen CI (#22879) 2026-04-17 04:09:44 +03:00
Egor Filimonov ba850d3a9d [Bugfix] [NPU] Fix check_env on Ascend for CANN 8.5 (#22888) 2026-04-17 04:05:20 +03:00
Mick 3d2d57c6cc [diffusion] refactor: extract LTX2 image encoding from denoising stage (#22976) 2026-04-17 08:35:15 +08:00
Daifeng Li 2cc52d8326 feat: Support MXFP4 quantized dense models on AMD CDNA2/CDNA3 GPUs (#19143) 2026-04-16 16:51:32 -07:00
f639425ff0 add check for none status code in FinishAbort (#22535)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-04-16 16:21:07 -07:00
Tarushii Goel 2211b4d9c6 [sgl] improve accuracy of additional page requirement during spec decode (#22406) 2026-04-16 15:50:51 -07:00
Liangsheng Yin db7a751d48 refactor: extract FanOutCommunicator and use declarative spec table (#22967) 2026-04-16 15:37:19 -07:00
mqhc2020andHubert Lu 52f0b86f5d [AMD] Qwen3.5 MXFP4 breaks after shared expert fusion is enabled (#22948)
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
2026-04-16 15:25:33 -07:00
Liangsheng Yin c83ef4fdb6 use envs in server_args (#22994) 2026-04-16 15:01:33 -07:00
Xinyu Zhangandxyuzh c0172aef6e [Ray] Bind scheduler actors to GPU-local NUMA node (#22989)
Co-authored-by: xyuzh <xyuzh@users.noreply.github.com>
2026-04-16 14:52:15 -07:00
Xinyu Zhang d430034bde [Ray] Support multi-replica serving by making scheduler actor names unique (#22917) 2026-04-16 14:51:01 -07:00
Qiaolin Yu a87806a65f [misc] refine outdated comments for chain-style multi-layer MTP (#22996) 2026-04-16 14:49:43 -07:00
Qiaolin Yu 12266cf953 [misc] update .github/CODEOWNERS (#22993) 2026-04-16 14:19:41 -07:00
ybyang 41258f874d [PD]feat(bench): add --fake-prefill flag for decode-only stress testing (#22973) 2026-04-16 13:57:55 -07:00
Mick 29f56cb230 CI: fix lint (#22991) 2026-04-17 02:09:04 +08:00
Yujun Dong 0882f5c132 [Doc] correct the HTTP endpoint for stopping profiling in benchmark_and_profiling.md (#22523) 2026-04-16 12:54:28 -04:00
Zaire 71377deda7 [Docs] fix profiling endpoint (#22982)
Signed-off-by: Zaire404 <3147879462@qq.com>
2026-04-16 12:51:39 -04:00
Xinyuan Tong 082eaed0a4 test: fix flaky required function calling assertion (#22890) 2026-04-16 09:44:26 -07:00
Yuhao Yang 9da998a882 [diffusion] feat: disaggregated diffusion (#21701) 2026-04-16 23:51:32 +08:00
Zhangheng 14bcdfca21 [HiSparse]: Adding e2e ut for hisparse (#22979) 2026-04-16 23:20:07 +08:00
amote-i 78147306b7 [NPU] [DOC] Update npu best practice docs to match latest code (#22975) 2026-04-16 20:45:22 +08:00
Liangsheng Yin bbd8f9ba09 migrate CPU-only unit tests from openai_server to unit/ (#22965) 2026-04-16 03:53:33 -07:00
Liangsheng Yin 62309f09db fix(loads): preserve include filtering after watching mode switch (#22959) 2026-04-16 03:04:53 -07:00
ybyang 03fef357a6 fix(loads): switch get_loads_communicator to watching mode (#22919) 2026-04-16 02:12:22 -07:00
ybyang fbd6dc3565 fix: normalize tool message content for GLM5.1 chat template (#22595) 2026-04-16 16:48:38 +08:00
Aleksi Vesanto aaa682346e [diffusion] model: Properly validate device for Mistral 3 attention (#22690) 2026-04-16 00:29:23 -07:00
jhchouuu 1412e287bf [AMD][MoRI] bump MoRI to v1.1.0 (#22870) 2026-04-16 00:11:55 -07:00
Lianmin Zheng 35da90cb76 [misc] Configure logging before ServerArgs.__post_init__ (#22926) 2026-04-15 23:53:15 -07:00
yuefeng Wuandgemini-code-assist[bot] 65bc839a5f [Fix] eagle/eagle3 speculative decoding conflicts with xgrammar in NPU (#20989)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-15 23:34:23 -07:00
Bi Xue c43716a357 [sgl] provide an option to send control req to all dp ranks rank0 (#22758) 2026-04-16 14:24:26 +08:00
Byron Hsu 3600465e81 [Bug Fix] Remove follow_bootstrap_room fast path in PD disaggregation DP rank resolution (#22901) 2026-04-15 22:53:29 -07:00
hhwxw 2480cc2a16 docs: fix incorrect default max-payload-size in gateway config reference (#22923) 2026-04-16 13:25:27 +08:00
LHXuuu e7ad7c587a [EPD][VLM] Support Kimi VL EPD (#22490)
Signed-off-by: LHXuuu <xulianhao.xlh@antgroup.com>
2026-04-16 12:40:02 +08:00
CYYYC0310andcyy 58c6b871b2 Remove compatibility restriction between Pipeline Parallelism and Mixed Chunked Prefill (#22920)
Co-authored-by: cyy <cy02433585@alibaba-inc.com>
2026-04-16 11:25:31 +08:00
Xinyuan Tong 34fef07a15 Upgrade transformers to 5.5.3 and refactor hf_transformers_utils into subpackage (#21569) 2026-04-15 20:03:44 -07:00
JINZandZhangheng 14e122cdee [BugFix][RadixTree]:Fix stale eviction assertion in HiMambaRadixCache host eviction path (#22592)
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-04-16 10:49:30 +08:00
Vladimir (Vova) Vagaytsev 5ef67cee16 [AMD] Fix aiter import failure in ROCm Docker images (#22363) 2026-04-15 18:40:05 -07:00
Yuhao Yang b8794baa6d [Step3p5] Optimize allreduce in MoE layers (#22773) 2026-04-16 09:33:12 +08:00
Liangsheng Yin a4cf2ea128 streaming session: spec v2 bonus accounting + comprehensive test matrix (#22651) 2026-04-15 17:12:41 -07:00
Lianmin Zheng ccff59254c Update .codespellrc (#22912) 2026-04-15 16:25:55 -07:00
ishandhananiandClaude Opus 4.6 761259448d ci: re-enable fp8 nightly benchmark configs (#22910)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-15 15:57:49 -07:00
Xinyu Zhang e8c6e5466c [Ray] Auto-create placement group in RayEngine when none is detected (#22898) 2026-04-15 15:17:52 -07:00
Qiaolin Yu 0b1b07db72 [misc] fix ray folder lint (#22905) 2026-04-15 15:08:18 -07:00
Liangsheng Yin f9792166c3 trim_overshoot: cap swa_evicted_seqlen + unit test (#22900) 2026-04-15 15:05:35 -07:00