Commit Graph
11742 Commits
Author SHA1 Message Date
YC Yen-Ching Tsengandbingxche f399997d2f [AMD] mirror nightly images to local registry and prefer LAN pulls (#23073)
Co-authored-by: bingxche <bingxche@amd.com>
2026-04-17 19:49:26 +08:00
YC Yen-Ching Tsengandbingxche 8c13295842 [AMD] fix AMD CI gate (#22974)
Co-authored-by: bingxche <bingxche@amd.com>
2026-04-17 18:32:26 +08:00
Opher Lieber 6e3bbef568 expose num_embeddings in VocabParallelEmbeddingWithLoRA (#22547) 2026-04-17 02:35:13 -07:00
CYYYC0310andcyy a12ea979d4 [test] Add GSM8K accuracy test for PP with mixed chunk prefill (#23029)
Co-authored-by: cyy <cy02433585@alibaba-inc.com>
2026-04-17 17:09:53 +08:00
ybyang 271c177443 [NPU]chore(docker): use editable install for sglang in npu.Dockerfile (#23040) 2026-04-17 17:08:39 +08:00
Mick 0b2058853d [diffusion] doc: update doc (#23052) 2026-04-17 16:23:46 +08:00
Jonah Bernard 0d031335ed [Pipeline Parallelism][Bug] Fix scheduler hang in pipeline parallelism setup (#23006) 2026-04-17 14:50:47 +08:00
Duyi-Wang 8c190f6b91 [AMD] Add SGLANG_MORI_MOE_MAX_INPUT_TOKENS to truncate dispatch before MoE. (#22952) 2026-04-16 23:40:15 -07:00
xdtbynd 53f87c463d [Docs] [npu] change the feature support status (#23041) 2026-04-17 14:34:54 +08:00
Alex Nails 43eb66028f ci: install rust toolchain in ci_install_dependency.sh (#23017) 2026-04-16 23:18:22 -07:00
RichardoMuandMu Huai 7390eddf28 feat(observability): add OpenTelemetry tracing for speculative decoding (#19545)
Co-authored-by: Mu Huai <tianbowen.tbw@antgroup.com>
2026-04-17 14:01:58 +08:00
5fa0c6a52e Allow piecewise CUDA graph with speculative decoding (#22128)
Co-authored-by: luhongyu.4869 <luhongyu.4869@bytedance.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-17 13:39:30 +08:00
Xiaoyu Zhang 91679d935d [codex] Update diffusion skills (#23028) 2026-04-17 13:29:26 +08:00
Bingxu Chenandbingxche 7ac337df94 [AMD] CI Job Monitor: fix queue time, utilization, and summary metrics (#22274)
Co-authored-by: bingxche <binxche@amd.com>
2026-04-16 22:03:37 -07:00
0dcfae5553 [CPU] Add gemma4_rmsnorm_cpu kernel (#22842)
Co-authored-by: Copilot <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-04-17 13:03:16 +08:00
Chunyuan WUandMa Mingfei 6c89214584 [CPU][sgl-kernel] extend_attention_cpu and flash_attn_varlen_func: fix nan for large seq (#22434)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-04-17 13:01:01 +08:00
YC Yen-Ching Tseng f0f0148167 Revert "feat: Support MXFP4 quantized dense models on AMD CDNA2/CDNA3 GPUs (#19143)" (#23031) 2026-04-16 21:53:25 -07:00
Zhangheng 7d47f40a96 [UnifiedRadixTree]: Add HiCache hook interface for TreeComponent (#22924) 2026-04-17 12:09:41 +08:00
Byron HsuandByron Hsu cf9845f8e3 [Bug Fix] Ensure prefill_info_table is populated before honoring disagg_prefill_dp_rank (#22990)
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
2026-04-17 11:10:31 +08:00
Jan Bernlöhr 04a53955b9 feat: add coordinated checkpoint prefetch for network filesystem loading (#20843) 2026-04-16 20:08:19 -07:00
Yuhao Yang a77abbe005 [VLM] Reduce GPU memory footprint of CUDA IPC MM feature transport (#22662) 2026-04-17 10:38:36 +08:00
Khoa Pham 5a0eea8ac5 [CI] Adding Gemma 4 to Nightly CI (#22408) 2026-04-16 19:30:16 -07:00
Yuxuan Zhang 16d11c2a10 Fix for the low-probability garbled output issue in the GLM-5 series models. (#22811) 2026-04-17 09:52:13 +08:00
Alison Shao 0052093178 test(4-gpu-b200): split test_qwen35_models.py + bump partitions 5→6 (#22913) 2026-04-16 18:51:59 -07:00
Makcum888e e353630b57 [Diffusion] [NPU] Fix multimodal gen CI (#22879) 2026-04-17 04:09:44 +03:00
Egor Filimonov ba850d3a9d [Bugfix] [NPU] Fix check_env on Ascend for CANN 8.5 (#22888) 2026-04-17 04:05:20 +03:00
Mick 3d2d57c6cc [diffusion] refactor: extract LTX2 image encoding from denoising stage (#22976) 2026-04-17 08:35:15 +08:00
Daifeng Li 2cc52d8326 feat: Support MXFP4 quantized dense models on AMD CDNA2/CDNA3 GPUs (#19143) 2026-04-16 16:51:32 -07:00
f639425ff0 add check for none status code in FinishAbort (#22535)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
2026-04-16 16:21:07 -07:00
Tarushii Goel 2211b4d9c6 [sgl] improve accuracy of additional page requirement during spec decode (#22406) 2026-04-16 15:50:51 -07:00
Liangsheng Yin db7a751d48 refactor: extract FanOutCommunicator and use declarative spec table (#22967) 2026-04-16 15:37:19 -07:00
mqhc2020andHubert Lu 52f0b86f5d [AMD] Qwen3.5 MXFP4 breaks after shared expert fusion is enabled (#22948)
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
2026-04-16 15:25:33 -07:00
Liangsheng Yin c83ef4fdb6 use envs in server_args (#22994) 2026-04-16 15:01:33 -07:00
Xinyu Zhangandxyuzh c0172aef6e [Ray] Bind scheduler actors to GPU-local NUMA node (#22989)
Co-authored-by: xyuzh <xyuzh@users.noreply.github.com>
2026-04-16 14:52:15 -07:00
Xinyu Zhang d430034bde [Ray] Support multi-replica serving by making scheduler actor names unique (#22917) 2026-04-16 14:51:01 -07:00
Qiaolin Yu a87806a65f [misc] refine outdated comments for chain-style multi-layer MTP (#22996) 2026-04-16 14:49:43 -07:00
Qiaolin Yu 12266cf953 [misc] update .github/CODEOWNERS (#22993) 2026-04-16 14:19:41 -07:00
ybyang 41258f874d [PD]feat(bench): add --fake-prefill flag for decode-only stress testing (#22973) 2026-04-16 13:57:55 -07:00
Mick 29f56cb230 CI: fix lint (#22991) 2026-04-17 02:09:04 +08:00
Yujun Dong 0882f5c132 [Doc] correct the HTTP endpoint for stopping profiling in benchmark_and_profiling.md (#22523) 2026-04-16 12:54:28 -04:00
Zaire 71377deda7 [Docs] fix profiling endpoint (#22982)
Signed-off-by: Zaire404 <3147879462@qq.com>
2026-04-16 12:51:39 -04:00
Xinyuan Tong 082eaed0a4 test: fix flaky required function calling assertion (#22890) 2026-04-16 09:44:26 -07:00
Yuhao Yang 9da998a882 [diffusion] feat: disaggregated diffusion (#21701) 2026-04-16 23:51:32 +08:00
Zhangheng 14bcdfca21 [HiSparse]: Adding e2e ut for hisparse (#22979) 2026-04-16 23:20:07 +08:00
amote-i 78147306b7 [NPU] [DOC] Update npu best practice docs to match latest code (#22975) 2026-04-16 20:45:22 +08:00
Liangsheng Yin bbd8f9ba09 migrate CPU-only unit tests from openai_server to unit/ (#22965) 2026-04-16 03:53:33 -07:00
Liangsheng Yin 62309f09db fix(loads): preserve include filtering after watching mode switch (#22959) 2026-04-16 03:04:53 -07:00
ybyang 03fef357a6 fix(loads): switch get_loads_communicator to watching mode (#22919) 2026-04-16 02:12:22 -07:00
ybyang fbd6dc3565 fix: normalize tool message content for GLM5.1 chat template (#22595) 2026-04-16 16:48:38 +08:00
Aleksi Vesanto aaa682346e [diffusion] model: Properly validate device for Mistral 3 attention (#22690) 2026-04-16 00:29:23 -07:00