Commit Graph
16469 Commits
Author SHA1 Message Date
Juan MunetonandMa Mingfei cac3269305 [XPU] Fix NemotronH (hybrid mamba2) launch on --device xpu (#32227)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-08-12 13:23:23 +08:00
Liangsheng Yin 256981ce16 [CI] Align rerun-test environment with the test stages (#34195) 2026-08-11 21:37:10 -07:00
Xiaoyu Zhang 22e4b3a81f [Diffusion] Avoid slow cuBLASLt GELU epilogue on SM120 (#34350) 2026-08-12 12:12:00 +08:00
Mick a9a355774a [diffusion] feat: support dynamically cpu offload components (#34391) 2026-08-12 11:38:27 +08:00
+3 5899674504 [Fix] Make DeepSeek-V4 reasoning and tool-call streaming parsing chunk-invariant (#34458)
Co-authored-by: hao-cyber <89575785+hao-cyber@users.noreply.github.com>
Co-authored-by: Enrico Falco <enrico9034@gmail.com>
Co-authored-by: Svyatoslav <85786374+slivanovich@users.noreply.github.com>
Co-authored-by: Andreas Hassellof <andreas@ombori.com>
Co-authored-by: Leoyzen <leoyzen@gmail.com>
Co-authored-by: Chenglun Hu <chenglunhu@gmail.com>
Co-authored-by: robellliu-dev <robell.liu@huawei.com>
Co-authored-by: Gavin.Zhu <gavin.z@gmicloud.ai>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: tancheng33 <garrytancheng@gmail.com>
Co-authored-by: dineshx29 <dinesh.b.offl@gmail.com>
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
2026-08-11 20:03:28 -07:00
Mick 2be9773a21 [diffusion] doc: update cosmos3 edge and distilled cookbook (#34497) 2026-08-12 10:51:11 +08:00
Xiaoyu Zhang 4aff4b1822 [Diffusion] Improve bit-exact fusion fallback diagnostics (#34412) 2026-08-12 10:42:48 +08:00
Xiaoyu Zhang 3f9d184833 [Diffusion] Tune QK head LayerNorm for SM120 (#34349) 2026-08-12 10:41:54 +08:00
f7b6800b22 [diffusion] model: support cosmos3 edge and distilled (#31590)
Co-authored-by: Kedi Wu <31940276+kediwu0331@users.noreply.github.com>
Co-authored-by: Kedi Wu <kediw@nvidia.com>
2026-08-12 10:29:40 +08:00
MickandClaude Opus 5 81c88da1ab [diffusion] fix: nightly diffusion benchmark passes the retired --warmup flag (#34423)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 10:25:14 +08:00
NVShreyasandClaude Opus 4.8 30cb848d4b [Bugfix][DSA] Fix num_splits "(b+1)" crash on prefill-CP speculative decode (#34443)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-11 19:13:49 -07:00
Xiaoyu Zhang a53d3636ce [diffusion][model] Add native SANA-Video T2V support (#32921) 2026-08-12 10:07:24 +08:00
zijiec 93e9db5eb8 [Fix][Qwen]: fused shared-expert detection PP-safe protection (#34447) 2026-08-11 18:50:05 -07:00
Xiaoyu ZhangandClaude Fable 5 37c631ef23 [diffusion] Ideogram-4: fuse Qwen3-style RoPE and SwiGLU silu-mul (denoise -5.1% H100 / -4.7% H200, bit-exact) (#34314)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-12 09:19:35 +08:00
Xiaoyu Zhang b1b8ce715b [Diffusion][MiniMax H3] Fix SM120 QKNorm+RoPE rounding (#34347) 2026-08-12 09:18:16 +08:00
Mick a2d723820e [diffusion] fix: fix model-driven dit layerwise offload auto policy (#34401) 2026-08-12 09:08:31 +08:00
Yanbin Jiang 1ce515a53d Refocus LoRA tests on regression coverage (#34464) 2026-08-11 17:37:52 -07:00
Shu Wang c7c03ec53b [NVIDIA] Add flashinfer MNNVL backend for allreduce only (#30700) 2026-08-11 16:46:46 -07:00
Baizhou Zhang 983dfd6a9a [Misc] Sanitize the structure of environ.py (#34472) 2026-08-11 16:31:33 -07:00
Faradawn Yang 857910bd35 docs(semianalysis): Update Qwen3.5 B200 NVFP4 MTP config (#34357) 2026-08-11 16:12:17 -07:00
d82a1d4802 [XPU] Pad MoE expert weight row stride to avoid L3 aliasing (#33905)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
2026-08-11 15:54:13 -07:00
fde9ad2531 [Feature] Add Muse Glimmer model support (#34262)
Co-authored-by: sglang-bot <232288953+sglang-bot@users.noreply.github.com>
Co-authored-by: Brayden Zhong <brayden.zhong@radixark.ai>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-08-11 15:41:52 -07:00
Hanming Lu 9c1517df4a Raise PD zmq per-context socket cap via SGLANG_DISAGGREGATION_ZMQ_MAX_SOCKETS (#34450) 2026-08-11 15:27:32 -07:00
Mohammad Miadh Angkad 7fb6e61b95 Fix CUDA 13.0 VMM handle type compatibility (#34431) 2026-08-11 15:16:34 -07:00
ilyasher-harmonic 9ced8d0981 Optimize FP32 LM head for bf16/fp16 (#32370) 2026-08-11 15:14:50 -07:00
gongwei1027 2c07ca5e8d [Fix] Allow flashinfer_sparse_mla DSA backend for HiSparse on SM120 FP8 KV (#33075) 2026-08-11 15:05:04 -07:00
Douglas YangandClaude Opus 5 59450c4f18 docs(cookbook): Kimi-K3 — drop --enable-symm-mem from the GB cells (#34444)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 19:23:41 +00:00
Xinyuan Tong d5d41d07ed [Docs] Add Ling-3.0-tiny INT4 recipes (#34395) 2026-08-12 03:16:03 +08:00
93c1bff1d4 Fix default dtype restoration after model loader errors (#34440)
Co-authored-by: zhisbug <1654062+zhisbug@users.noreply.github.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2026-08-11 11:33:04 -07:00
Dmitrii SergeevandZhiqiang Xie c58953d90a O(1) slot allocation in ReqToTokenPool.alloc() (#32208)
Co-authored-by: Zhiqiang Xie <xiezhq@stanford.edu>
2026-08-11 11:26:05 -07:00
Jeremy Zhang aadb9720fe fix: route scheduler aborts to multi-tokenizer workers (#33940) 2026-08-11 11:15:16 -07:00
Ke Bao f9153df62e Add bit-exact hicache logprob-consistency test (#34356) 2026-08-12 01:54:45 +08:00
huangtingweiandHanming Lu 8f3d3a31f4 [HiCache] Fix Mamba track-boundary bookkeeping under overlap scheduling (#29792)
Co-authored-by: Hanming Lu <hanminglu@meta.com>
2026-08-12 00:36:48 +08:00
Mohammad Miadh Angkad a3bd7d9401 Bump CuTeDSL to 4.6.2 (#34372) 2026-08-12 00:35:35 +08:00
Mick 8267d76c2c [VLM] replace deprecated image processor use_fast (#34175) 2026-08-12 00:14:07 +08:00
Ke Bao b20c375c10 Fix flaky decode cache-hit check in Inkling test (#34405) 2026-08-11 21:47:01 +08:00
Jinyan Yi f148eb6e6e Add Hunyuan3 On Ascend Doc (#30223) 2026-08-11 21:20:30 +08:00
gjsheu a50ab9cec7 [npu] [bugfix] Fix HiCache MHA backup for NPU (#34341) 2026-08-11 21:09:45 +08:00
Faradawn YangandRyan Stewart 3add7e19ff Add NVIDIA Nemotron 3.5 Lightning cookbook (#33481)
Signed-off-by: Faradawn Yang <73060648+faradawn@users.noreply.github.com>
Signed-off-by: Ryan Stewart <rystewart@nvidia.com>
Co-authored-by: Ryan Stewart <rystewart@nvidia.com>
2026-08-11 06:00:57 -07:00
Mohammad Miadh Angkad 2d193077f7 [JIT Kernel] Migrate per-token FP8 quantization from AOT to JIT (#34257) 2026-08-11 20:40:40 +08:00
a0a76e4485 [DSV4] perf: Enable alt stream during BCG prefill (#29070)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
2026-08-11 05:10:19 -07:00
Xiaoyu ZhangandClaude Fable 5 546965fc72 [diffusion] LTX-2: mount the bit-exact fused modulate at the 8 bare adaLN sites (ltx23-one-stage denoise -2.8% H100 / -2.6% H200) (#34315)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 18:23:08 +08:00
Mick dd8c5849af [diffusion] refactor: move dit execution capabilities to runtime models (#34249) 2026-08-11 18:20:43 +08:00
Xiaoyu ZhangandClaude Fable 5 071f0f1e9d [diffusion] ERNIE-Image: fuse rotate-half RoPE + GELU-mul and hoist rope cos/sin (denoise -16.2% H100 / -12.7% H200, bit-exact) (#34306)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 18:18:37 +08:00
Xiaoyu ZhangandClaude Fable 5 ba3dc16401 [diffusion] weight-only FP8: dequantize linear weights once at first use (Ideogram-4 denoise -18.8% H200 / -7.8% H100, bit-exact) (#34305)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 18:15:26 +08:00
Zhiqiang XieandTingwei Huang 5469faec45 HiSparse: shared-index (IndexShare) plan-then-IO swap-in prefetch (#34329)
Co-authored-by: Tingwei Huang <huangtingwei9988@gmail.com>
2026-08-11 01:58:28 -07:00
396722e490 [NPU] Disable failed test cases (#34377)
Co-authored-by: Even Zhou <even.y.zhou@outlook.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
2026-08-11 16:38:15 +08:00
Bingxu ChenandYC Yen-Ching Tseng 6f3fe13a9c [AMD] Install AITER's pinned Triton wheel in the ROCm 7.2 image (#34364)
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
2026-08-11 16:36:45 +08:00
Liangsheng Yin d8a61c26a6 [CI] Add a scheduled workflow to close stale PRs (#34380) 2026-08-11 00:54:02 -07:00
e74ea5b1d7 [ROCm/gfx95] Fix fp8 per-channel attention for Kimi-K2.7-code-mxfp4 o… (#31105)
Co-authored-by: Hung <Emmanuel0612@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
2026-08-11 00:51:57 -07:00