 YC Yen-Ching Tsengandbingxche
|
f399997d2f
|
[AMD] mirror nightly images to local registry and prefer LAN pulls (#23073)
Co-authored-by: bingxche <bingxche@amd.com>
|
2026-04-17 19:49:26 +08:00 |
|
 YC Yen-Ching Tsengandbingxche
|
8c13295842
|
[AMD] fix AMD CI gate (#22974)
Co-authored-by: bingxche <bingxche@amd.com>
|
2026-04-17 18:32:26 +08:00 |
|
Opher Lieber
|
6e3bbef568
|
expose num_embeddings in VocabParallelEmbeddingWithLoRA (#22547)
|
2026-04-17 02:35:13 -07:00 |
|
 CYYYC0310andcyy
|
a12ea979d4
|
[test] Add GSM8K accuracy test for PP with mixed chunk prefill (#23029)
Co-authored-by: cyy <cy02433585@alibaba-inc.com>
|
2026-04-17 17:09:53 +08:00 |
|
ybyang
|
271c177443
|
[NPU]chore(docker): use editable install for sglang in npu.Dockerfile (#23040)
|
2026-04-17 17:08:39 +08:00 |
|
Mick
|
0b2058853d
|
[diffusion] doc: update doc (#23052)
|
2026-04-17 16:23:46 +08:00 |
|
Jonah Bernard
|
0d031335ed
|
[Pipeline Parallelism][Bug] Fix scheduler hang in pipeline parallelism setup (#23006)
|
2026-04-17 14:50:47 +08:00 |
|
Duyi-Wang
|
8c190f6b91
|
[AMD] Add SGLANG_MORI_MOE_MAX_INPUT_TOKENS to truncate dispatch before MoE. (#22952)
|
2026-04-16 23:40:15 -07:00 |
|
xdtbynd
|
53f87c463d
|
[Docs] [npu] change the feature support status (#23041)
|
2026-04-17 14:34:54 +08:00 |
|
Alex Nails
|
43eb66028f
|
ci: install rust toolchain in ci_install_dependency.sh (#23017)
|
2026-04-16 23:18:22 -07:00 |
|
 RichardoMuandMu Huai
|
7390eddf28
|
feat(observability): add OpenTelemetry tracing for speculative decoding (#19545)
Co-authored-by: Mu Huai <tianbowen.tbw@antgroup.com>
|
2026-04-17 14:01:58 +08:00 |
|
 
|
5fa0c6a52e
|
Allow piecewise CUDA graph with speculative decoding (#22128)
Co-authored-by: luhongyu.4869 <luhongyu.4869@bytedance.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-17 13:39:30 +08:00 |
|
Xiaoyu Zhang
|
91679d935d
|
[codex] Update diffusion skills (#23028)
|
2026-04-17 13:29:26 +08:00 |
|
 Bingxu Chenandbingxche
|
7ac337df94
|
[AMD] CI Job Monitor: fix queue time, utilization, and summary metrics (#22274)
Co-authored-by: bingxche <binxche@amd.com>
|
2026-04-16 22:03:37 -07:00 |
|
 
|
0dcfae5553
|
[CPU] Add gemma4_rmsnorm_cpu kernel (#22842)
Co-authored-by: Copilot <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-17 13:03:16 +08:00 |
|
 Chunyuan WUandMa Mingfei
|
6c89214584
|
[CPU][sgl-kernel] extend_attention_cpu and flash_attn_varlen_func: fix nan for large seq (#22434)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-04-17 13:01:01 +08:00 |
|
YC Yen-Ching Tseng
|
f0f0148167
|
Revert "feat: Support MXFP4 quantized dense models on AMD CDNA2/CDNA3 GPUs (#19143)" (#23031)
|
2026-04-16 21:53:25 -07:00 |
|
Zhangheng
|
7d47f40a96
|
[UnifiedRadixTree]: Add HiCache hook interface for TreeComponent (#22924)
|
2026-04-17 12:09:41 +08:00 |
|
 Byron HsuandByron Hsu
|
cf9845f8e3
|
[Bug Fix] Ensure prefill_info_table is populated before honoring disagg_prefill_dp_rank (#22990)
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
|
2026-04-17 11:10:31 +08:00 |
|
Jan Bernlöhr
|
04a53955b9
|
feat: add coordinated checkpoint prefetch for network filesystem loading (#20843)
|
2026-04-16 20:08:19 -07:00 |
|
Yuhao Yang
|
a77abbe005
|
[VLM] Reduce GPU memory footprint of CUDA IPC MM feature transport (#22662)
|
2026-04-17 10:38:36 +08:00 |
|
Khoa Pham
|
5a0eea8ac5
|
[CI] Adding Gemma 4 to Nightly CI (#22408)
|
2026-04-16 19:30:16 -07:00 |
|
Yuxuan Zhang
|
16d11c2a10
|
Fix for the low-probability garbled output issue in the GLM-5 series models. (#22811)
|
2026-04-17 09:52:13 +08:00 |
|
Alison Shao
|
0052093178
|
test(4-gpu-b200): split test_qwen35_models.py + bump partitions 5→6 (#22913)
|
2026-04-16 18:51:59 -07:00 |
|
Makcum888e
|
e353630b57
|
[Diffusion] [NPU] Fix multimodal gen CI (#22879)
|
2026-04-17 04:09:44 +03:00 |
|
Egor Filimonov
|
ba850d3a9d
|
[Bugfix] [NPU] Fix check_env on Ascend for CANN 8.5 (#22888)
|
2026-04-17 04:05:20 +03:00 |
|
Mick
|
3d2d57c6cc
|
[diffusion] refactor: extract LTX2 image encoding from denoising stage (#22976)
|
2026-04-17 08:35:15 +08:00 |
|
Daifeng Li
|
2cc52d8326
|
feat: Support MXFP4 quantized dense models on AMD CDNA2/CDNA3 GPUs (#19143)
|
2026-04-16 16:51:32 -07:00 |
|
 
|
f639425ff0
|
add check for none status code in FinishAbort (#22535)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
|
2026-04-16 16:21:07 -07:00 |
|
Tarushii Goel
|
2211b4d9c6
|
[sgl] improve accuracy of additional page requirement during spec decode (#22406)
|
2026-04-16 15:50:51 -07:00 |
|
Liangsheng Yin
|
db7a751d48
|
refactor: extract FanOutCommunicator and use declarative spec table (#22967)
|
2026-04-16 15:37:19 -07:00 |
|
 mqhc2020andHubert Lu
|
52f0b86f5d
|
[AMD] Qwen3.5 MXFP4 breaks after shared expert fusion is enabled (#22948)
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
|
2026-04-16 15:25:33 -07:00 |
|
Liangsheng Yin
|
c83ef4fdb6
|
use envs in server_args (#22994)
|
2026-04-16 15:01:33 -07:00 |
|
 Xinyu Zhangandxyuzh
|
c0172aef6e
|
[Ray] Bind scheduler actors to GPU-local NUMA node (#22989)
Co-authored-by: xyuzh <xyuzh@users.noreply.github.com>
|
2026-04-16 14:52:15 -07:00 |
|
Xinyu Zhang
|
d430034bde
|
[Ray] Support multi-replica serving by making scheduler actor names unique (#22917)
|
2026-04-16 14:51:01 -07:00 |
|
Qiaolin Yu
|
a87806a65f
|
[misc] refine outdated comments for chain-style multi-layer MTP (#22996)
|
2026-04-16 14:49:43 -07:00 |
|
Qiaolin Yu
|
12266cf953
|
[misc] update .github/CODEOWNERS (#22993)
|
2026-04-16 14:19:41 -07:00 |
|
ybyang
|
41258f874d
|
[PD]feat(bench): add --fake-prefill flag for decode-only stress testing (#22973)
|
2026-04-16 13:57:55 -07:00 |
|
Mick
|
29f56cb230
|
CI: fix lint (#22991)
|
2026-04-17 02:09:04 +08:00 |
|
Yujun Dong
|
0882f5c132
|
[Doc] correct the HTTP endpoint for stopping profiling in benchmark_and_profiling.md (#22523)
|
2026-04-16 12:54:28 -04:00 |
|
Zaire
|
71377deda7
|
[Docs] fix profiling endpoint (#22982)
Signed-off-by: Zaire404 <3147879462@qq.com>
|
2026-04-16 12:51:39 -04:00 |
|
Xinyuan Tong
|
082eaed0a4
|
test: fix flaky required function calling assertion (#22890)
|
2026-04-16 09:44:26 -07:00 |
|
Yuhao Yang
|
9da998a882
|
[diffusion] feat: disaggregated diffusion (#21701)
|
2026-04-16 23:51:32 +08:00 |
|
Zhangheng
|
14bcdfca21
|
[HiSparse]: Adding e2e ut for hisparse (#22979)
|
2026-04-16 23:20:07 +08:00 |
|
amote-i
|
78147306b7
|
[NPU] [DOC] Update npu best practice docs to match latest code (#22975)
|
2026-04-16 20:45:22 +08:00 |
|
Liangsheng Yin
|
bbd8f9ba09
|
migrate CPU-only unit tests from openai_server to unit/ (#22965)
|
2026-04-16 03:53:33 -07:00 |
|
Liangsheng Yin
|
62309f09db
|
fix(loads): preserve include filtering after watching mode switch (#22959)
|
2026-04-16 03:04:53 -07:00 |
|
ybyang
|
03fef357a6
|
fix(loads): switch get_loads_communicator to watching mode (#22919)
|
2026-04-16 02:12:22 -07:00 |
|
ybyang
|
fbd6dc3565
|
fix: normalize tool message content for GLM5.1 chat template (#22595)
|
2026-04-16 16:48:38 +08:00 |
|
Aleksi Vesanto
|
aaa682346e
|
[diffusion] model: Properly validate device for Mistral 3 attention (#22690)
|
2026-04-16 00:29:23 -07:00 |
|