Commit Graph
16764 Commits
Author SHA1 Message Date
Siyuan Chen fcdaaf8a5d [Feature] Optimize TP LMHead with All-to-All (#32313) 2026-08-17 19:55:27 -07:00
jiayisunxandMa Mingfei d6c837489a [XPU] Enable fused GDN QKV split Triton kernel on XPU (#30144)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-08-18 10:42:58 +08:00
Ma Mingfei 7f51e6bba0 [CPU] Explicitly import sgl_kernel in CPU kernel tests (#35119) 2026-08-18 10:20:28 +08:00
MickandYiqi Yang d55f1c28e2 [diffusion] feat: load quantized H3 text encoder checkpoints (#34986)
Co-authored-by: Yiqi Yang <yangyiqi8787@gmail.com>
2026-08-18 09:10:54 +08:00
Zaili Wang 0ea262e6e5 [XPU] xpu kernel release workflow (#33679) 2026-08-17 18:07:48 -07:00
Baizhou Zhang c2c1c4d3eb [CI] Move DSA PD+MTP+CP Layersplit test to basic B300 test suite (#35220) 2026-08-17 17:44:31 -07:00
sglang-botandsglang-bot 9ffc2856fb docs: sync LMSYS SGLang blog cards (#35218)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-08-17 17:30:28 -07:00
Faradawn Yang 91144797c5 Update Qwen3.5 H200 FP8 for AgentX HiCache MTP (#35194) 2026-08-17 17:12:07 -07:00
Liangsheng Yin c0b6474b43 [Spec] Reduce host-side overhead in ngram draft prep (#35207) 2026-08-17 16:40:06 -07:00
Cheng Wan b3c8f0d923 config: one control-plane log for the process (#35028) 2026-08-17 16:19:00 -07:00
Cheng Wan c70c7d72a8 config: the readback and the resolving view say what they are (#35027) 2026-08-17 16:18:19 -07:00
Cheng Wan cba3c5d5ac config: the per-instance families read the bags (#35026) 2026-08-17 16:17:53 -07:00
Cheng Wan a97bc8db32 config: the DP/EP topology reads come from the parallel bag (#35025) 2026-08-17 16:17:26 -07:00
Cheng Wan d2bc697396 spec: size the speculative buffers from the bags, not the startup record (#35024) 2026-08-17 16:16:56 -07:00
Cheng Wan 3d7ec00179 config: publish before a process reads configuration (#35023) 2026-08-17 16:16:20 -07:00
Cheng Wan 2b278b4ac4 config: retire the multi-engine accommodation in the runtime context (#35022) 2026-08-17 16:15:33 -07:00
Baizhou Zhang bc312d185d Clean deprecated DeepSeek V4 Environs (#34926) 2026-08-17 16:07:00 -07:00
Lennox Fu b42abbb1ba [metrics] Fix prefill FLOPs estimate to count prefix and per-request causal pairs (#34316) 2026-08-17 15:18:54 -07:00
Jimmy ShongandClaude Opus 5 b956e916ae docs(cookbook): add Qwen3.8-27B DGX Spark configs (#35121)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 15:04:17 -07:00
816ea65058 [AMD] Add Kimi-K3 8-GPU MI35x nightly accuracy CI (#32568)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Michael <michaelzhang-ai@users.noreply.github.com>
2026-08-17 14:48:29 -07:00
Lianmin Zheng 198a7b2fc9 [Misc] Clean up python/sglang package structure (#35062) 2026-08-17 14:24:35 -07:00
770e7b47a2 [Spec] Support output logprobs with DSpark (#34478)
Co-authored-by: zhisbug <1654062+zhisbug@users.noreply.github.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-08-17 14:23:24 -07:00
Liangsheng Yin 032fe9c891 [Spec] Relay ngram accept tokens through the FutureMap (#35198) 2026-08-17 14:21:07 -07:00
Yuhao Yang 861eca8e25 docs: add NVFP4 quantization option to Kimi-K3 deploy panel (#35168) 2026-08-17 11:01:47 -07:00
cctryandcctry 2e7c85da68 [PD] Preserve decode KV across retraction in HiCache (#34801)
Co-authored-by: cctry <cctry@fb.com>
2026-08-17 08:49:11 -07:00
Lianmin Zheng af743371cc Clean up environ.py: remove dead env vars, unify deprecation handling, move examples to a unit test (#35060) 2026-08-17 06:53:34 -07:00
Mick d97b796c16 [diffusion] chore: reuse SRT CLIP encoder blocks (#35004) 2026-08-17 19:51:51 +08:00
Ke Bao 6e8a4abb57 Add bit-exact class for MTP (#35143) 2026-08-17 19:51:39 +08:00
Mick e9ad8102a2 [diffusion] chore: reuse SRT SigLIP in Pi0.5 (#34992) 2026-08-17 19:33:36 +08:00
WenhaoZhang f33b83b4cc [diffusion] fix: fix h3 swap peft SwiGLU lora_B halves when loading FFN Lora (#34940) 2026-08-17 18:47:31 +08:00
Mohammad Miadh Angkad 82995a001b Stabilize GB300 nightly tests (#35044) 2026-08-17 03:30:16 -07:00
744740dbea [XPU] upgrade sglang xpu backend to PyTorch 2.13 (#31751)
Co-authored-by: MingxuZh <109504044+MingxuZh@users.noreply.github.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-08-17 18:29:15 +08:00
kangwangamd c82e928fe5 [AMD] diffusion: normalize ModelOpt-FP8 weights to e4m3fnuz on gfx942 (#35111) 2026-08-17 02:35:10 -07:00
Alan Kao 056808723e [AMD] Guard ROCm 7.0 build from using hipMemcpyBatchAsync (#35128) 2026-08-17 02:22:58 -07:00
92bce3d7bb [AMD] [GLM5] fp8 MLA absorbed bmm for GLM-5.2 on gfx950 (#30519)
Co-authored-by: Thomas Wang <thomawan@amd.com>
Co-authored-by: sogalin_codegen <39478626+sogalin@users.noreply.github.com>
2026-08-17 02:15:16 -07:00
8cc112d486 [DSA] Skip indexer KV cache for skip-topk layers (#30531)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
Co-authored-by: mmangkad <mohammad.angkad@radixark.ai>
2026-08-17 02:02:23 -07:00
YAMY 7c423cfd41 [PD] Avoid unused PREBUILT prompt tensor transfer (#35070) 2026-08-17 16:48:12 +08:00
Liangsheng Yin 711bdacb82 [Spec] Resolve shared-read ends from the backend declaration alone (#35059) 2026-08-17 01:35:29 -07:00
b83d507cd7 [NPU] Support DeepSeek-V4 DSpark and refactor DSV4 cache management (#33676)
Co-authored-by: JiaruiChang5268 <jc5268@columbia.edu>
Co-authored-by: Kelon <kelonlu@163.com>
Co-authored-by: unknown <z8ruev42yk@gmail.com>
Co-authored-by: Talantan1102 <545811257@qq.com>
Co-authored-by: Talantan1102 <44429302+Talantan1102@users.noreply.github.com>
2026-08-17 16:27:44 +08:00
Jimmy ShongandClaude Fable 5 e03c53fc13 docs(cookbook): Qwen3.8-27B deployment grid rework (#35065)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-17 00:56:08 -07:00
Ke Bao 4cad864361 Fix sconv track refresh on graph capture (#35042) 2026-08-17 15:51:45 +08:00
Mohammad Miadh AngkadandMohammad Angkad 5769b6d637 [JIT Kernel] Migrate causal_conv1d_fwd and causal_conv1d_update from AOT to JIT (#35031)
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
2026-08-17 15:22:37 +08:00
Liangsheng Yin 0d8c850a35 [Fix] Read the DSA prefill CP flag from the parallel config bag in bootstrap (#35110) 2026-08-17 00:21:41 -07:00
Mohammad Miadh Angkad 721e359ca7 Suppress expected FlashInfer TRT-LLM workspace warnings (#34921) 2026-08-17 15:16:16 +08:00
Lianmin Zhengandwangwenchen0407 12a455a910 Fix world-size-one aliasing in MLP batch sync (#34997)
Co-authored-by: wangwenchen0407 <wangwenchen@meta.com>
2026-08-17 00:04:09 -07:00
vikram singh shekhawatandClaude Sonnet 4.5 f7a404e9c3 Fix rope config compatibility and VL/transformers-fallback weight loading (#31575)
Co-authored-by: Claude Sonnet 4.5 (1M context) <noreply@anthropic.com>
2026-08-17 14:53:59 +08:00
Liangsheng Yin 0099107e8b Revert "[AMD] [GLM5] Fuse shared-expert append into aiter grouped-topk (skip per-layer append kernel)" (#35105) 2026-08-16 23:49:49 -07:00
ziang663andZhangheng 43226af812 fix(hicache): limit load-back pending to write-back (#34519)
Co-authored-by: Zhangheng <hzh0425@apache.org>
2026-08-17 14:23:54 +08:00
MingxuZh eafbe2cb6f [CI] Install sgl-eval in xeon (CPU) Docker image (#34818) 2026-08-17 14:00:44 +08:00
Jimmy ShongandClaude Fable 5 07a28ec5cf docs: fix Qwen3.8-27B mamba ratio calculator for speculative decoding (#35064)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 22:54:57 -07:00