Michael
39c6bf730c
[AMD][CI] Add GLM-5-MXFP4 accuracy and perf nightly tests for MI35x ( #21773 )
2026-04-14 18:55:36 -07:00
Lianmin Zheng
adb310b976
Cleanup server_args.py and minor code tidying ( #22820 )
2026-04-14 18:52:41 -07:00
ea05ea5abe
[AMD] Enable share expert fusion with router experts for Qwen3.5 BF16 & FP8 ( #20736 )
...
Co-authored-by: Chen, Todd <zhenchen@amd.com >
Co-authored-by: jacky.cheng <yichiche@amd.com >
2026-04-14 18:52:36 -07:00
Piotr Mazurek
46c8a597ef
[VLM] fix LFM2-VL offline inference and GPU JPEG decode ( #22448 )
2026-04-15 09:13:25 +08:00
ishandhanani
2c9e76d333
ci: skip approval for nightly gb200 runs, keep for manual triggers ( #22768 )
2026-04-14 16:34:57 -07:00
Alexis MacAskill
e15401ee0e
Add runai-model-streamer into Python packages installed in Dockerfile and fix NotADirectoryError Docker regression ( #22537 )
2026-04-14 16:25:41 -07:00
Lianmin Zheng
222eda1598
[Misc] Use cache_once for is_arch_support_pdl in sgl-kernel ( #22725 )
2026-04-14 15:22:10 -07:00
Jimmy Shong
e83560562b
Update CI Permissions ( #22826 )
2026-04-14 15:13:31 -07:00
8092431316
[serving] replace O(n²) stream_buffer string concat with integer offset ( #22606 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-04-14 14:48:44 -07:00
Liangsheng Yin
36891ab514
Rename _alive_streaming_session_count; use _is_streaming helper ( #22755 )
2026-04-14 13:26:03 -07:00
Liangsheng Yin
0cb7295698
Fix streaming session busy-check double-counting via active_pool_idxs ( #22753 )
2026-04-14 13:11:06 -07:00
mingyue300
b4616dcbf5
[BugFix] Fix EAGLE speculative decoding missing grammar-based finish … ( #21723 )
2026-04-14 12:43:50 -07:00
Mick and Claude Opus 4.6
d2f479e544
[diffusion] chore: auto-enable best parallel setting if unspecified ( #22763 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-15 00:02:05 +08:00
Ke Bao
3c0a6c6987
Add page_size and SWA coverage to unified radix cache bench test ( #22815 )
2026-04-14 23:58:05 +08:00
Bi Xue
070c6a2489
[sgl] perf optimization for eplb ( #21232 )
2026-04-14 22:52:17 +08:00
Ke Bao
9f9e0231bb
Refactor unified radix cache UT into parameterized test suite ( #22812 )
2026-04-14 22:34:33 +08:00
Mick
c5e95080d2
[diffusion] model: support Ltx 2.3 two stage ti2v ( #22667 )
2026-04-14 22:10:08 +08:00
chx96642264
680bd4b429
[NPU] Modify the parameter name and optional values, and add the parameter restrictions. Modify some parameters supported type. ( #22804 )
2026-04-14 21:34:07 +08:00
McZyWu and root
1588856e9b
[NPU] qwen3next low latency best practice docs. ( #22808 )
...
Co-authored-by: root <root@localhost.localdomain >
2026-04-14 21:21:37 +08:00
amote-i
ddc7daaf89
[NPU] [DOC] Update NPU docs to match latest code ( #22796 )
2026-04-14 21:10:28 +08:00
lawtherWu
454228e071
hicache storage backend mooncake support ascend hixl ( #20016 )
2026-04-14 20:51:06 +08:00
loading66
074c2a476d
fix:[NPU]correct the full name of then Kimi model ( #22799 )
2026-04-14 20:15:22 +08:00
jianzhao-xu and Jianzhao Xu
68dfffaaa3
Offloading docs update ( #22795 )
...
Co-authored-by: Jianzhao Xu <xujianchao@huawei.com >
2026-04-14 20:03:29 +08:00
xdtbynd
88253c39b0
[Docs] Fix formatting of tool-call-parser options ( #22793 )
2026-04-14 19:21:31 +08:00
amote-i
368cdfbe2f
[NPU] [DOC] Fix outdated descriptions in the NPU documentation ( #22707 )
2026-04-14 19:21:15 +08:00
Jia Guo and Claude Opus 4.6
6da3aba6a5
perf: optimize PCG inductor path for FP8 models ( #21734 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-14 17:51:27 +08:00
xutizhou
3cb3f7c018
fix: EPLB dispatch OOB when shared experts fusion enabled under DeepEP ( #22525 )
2026-04-14 02:33:27 -07:00
Jincong Chen
6760c790bd
[bugfix] avoid attention padding tokens computation in pcg ( #17706 )
2026-04-14 16:08:23 +08:00
Michael and HaiShaw
eab045b2b7
[AMD] Add MiniMax-M2.7 accuracy and performance nightly tests ( #22722 )
...
Co-authored-by: HaiShaw <hixiao@gmail.com >
2026-04-14 00:30:11 -07:00
xiaobochen-amd and kk
d7ecab5113
[ROCm]fix(aiter): cast fp8 prefill output back to model dtype ( #22626 )
...
Co-authored-by: kk <43161300+kkHuang-amd@users.noreply.github.com >
2026-04-14 00:25:09 -07:00
Xiaoyu Zhang
f97c608caa
[diffusion] quant: add FLUX.1-dev modelopt nvfp4 support ( #22672 )
2026-04-14 15:00:59 +08:00
Sahithi Chigurupati and ishandhanani
7c1bde2e38
[CI] Add optional image input to GB200 nightly workflow_dispatch ( #22745 )
...
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com >
2026-04-13 23:57:15 -07:00
Colin Z and HAI
b10f852118
GLM-5/5.1 MXFP4 Checkpoint Inference Compatibility Fix ( #22543 )
...
Co-authored-by: HAI <hixiao@gmail.com >
2026-04-13 23:56:48 -07:00
Baizhou Zhang and Claude Opus 4.6
8fe9bbffb6
[CI] Reinstall flashinfer-jit-cache on CUDA version mismatch ( #22741 )
...
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-13 23:04:23 -07:00
YC Yen-Ching Tseng and bingxche
d44eb16ac6
[AMD] Replace push trigger with scheduled runs and enable parallel stage execution ( #22489 )
...
Co-authored-by: bingxche <bingxche@amd.com >
2026-04-14 13:33:29 +08:00
YAMY
657945c338
Replace all-reduce + dp_scatter with reduce_scatterv for DP attention ( #22642 )
2026-04-13 21:51:10 -07:00
ishandhanani
520ce526b9
Restore Qwen3 rope config fallback ( #22739 )
2026-04-13 21:47:37 -07:00
Xuwei
a9a2ae4a68
[Anthropic] Fix clock mismatch in received_time causing negative Prometheus metrics ( #22247 )
...
Signed-off-by: Xuwei Li <lixuwei.xy@gmail.com >
2026-04-13 21:22:00 -07:00
Jia Guo and Claude Opus 4.6
bc16130a17
ci: skip full rerun when sgl-kernel wheel already built ( #22534 )
...
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-04-13 20:32:55 -07:00
e9d6b9eb2d
[HiCache & HybridModel] mooncake backend support DSA & mamba model ( #21259 )
...
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com >
Co-authored-by: hzh0425 <hzh0425@apache.org >
Co-authored-by: pansicheng <sicheng.pan.chn@gmail.com >
Co-authored-by: ispobock <ispobaoke@gmail.com >
Co-authored-by: Vladislav Nosivskoy <vladnosiv@gmail.com >
2026-04-13 18:47:36 -07:00
ishandhanani
cc449ac4e5
feat(metrics): expose raw KV cache pool token counts as prometheus gauges ( #22726 )
2026-04-13 18:30:36 -07:00
huangtingwei
945d73824f
[HiSparse] Clarify decode token usage logs ( #22331 )
2026-04-13 18:03:25 -07:00
Zhai Feiyue
c456cba7fd
[gateway] Support SGLANG_LOG_MS for millisecond precision in router logs ( #22506 )
2026-04-13 17:28:00 -07:00
yuki-brook
1ec018f27a
[Feature] Add SiMM as sglang HiCache Storage backend ( #18016 )
2026-04-13 17:12:37 -07:00
Sahithi Chigurupati
ff61b2e470
[CI] Add workflow_dispatch and environment gate to GB200 nightly pipeline ( #22733 )
2026-04-13 17:08:18 -07:00
Liangsheng Yin
33a3ba256f
Delete dead rematch path in SessionAwareCache.release_session ( #22735 )
2026-04-13 17:02:40 -07:00
Lianmin Zheng
9fb00ede15
Clean up TokenizerManager and req_time_stats: reduce overhead and simplify ( #21646 )
2026-04-13 16:47:32 -07:00
Jia Guo and Claude Opus 4.6
a2b5111962
perf: skip KV cache in FA backend for embedding mode ( #21971 )
...
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
2026-04-13 16:27:52 -07:00
Lianmin Zheng
8f9553bccb
[Misc] Migrate SGLANG_SET_CPU_AFFINITY to envs and refactor model config building ( #22730 )
2026-04-13 16:10:31 -07:00
mqhc2020 and HAI
f4f9e68189
[AMD] Add MoE weights and scales padding ( #21097 )
...
Co-authored-by: HAI <hixiao@gmail.com >
2026-04-13 15:50:15 -07:00