Commit Graph
464 Commits
Author SHA1 Message Date
Liangsheng Yin 9b853e6832 [Scheduler] Enable decode retraction ordering under speculative decoding (#32023) 2026-07-23 00:42:56 -07:00
sglang-botandsglang-bot 9bda6fdb9a docs: sync LMSYS SGLang blog cards (#32127)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-07-23 00:57:01 +00:00
Xiaoyu ZhangandClaude Opus 4.8 99f636a86f [Kernel] RFC #29630 finale: retire sglang.jit_kernel into sglang.kernels (#32072)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-23 08:35:09 +08:00
Brayden Zhong e7511141ea Support CuteDSL GEMM BF16 on SM100 on by default when allowed by heuristic (#30567) 2026-07-22 14:13:12 -07:00
0a6d1930c3 [Attention Backend] Add HPC-Ops attention backend (#30540)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
2026-07-22 22:06:22 +08:00
axx-ty911 7dcf3f7cbf [Doc] Rename A5 product name (#31998) 2026-07-22 15:09:05 +08:00
4a55fdba0b docs(cookbook): re-benchmark DeepSeek-V4 on sglang 0.5.15 (#31363)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
2026-07-21 22:55:36 +00:00
9a6c96083f [Cookbook] Add Laguna-S-2.1 (#31918)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-07-21 17:01:50 +00:00
dcd9014f15 [AMD][MXFP4] Reland "Online MXFP4 quantization 2/N - FP8 to MXFP4 requantization on AMD GPUs" (#28291)
Co-authored-by: Bowen Bao <bowenbao@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
2026-07-21 02:50:23 -07:00
Rahul Vijayaraghavan fa0ced195e [XPU] Enable breakable prefill CUDA graph on XPU (#30273) 2026-07-21 09:09:40 +08:00
zijiexiaandClaude Opus 4.8 0a2d3ca071 [cookbook] Inkling: add measured accuracy numbers to benchmark cards (#31823)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 22:43:35 +00:00
Douglas YangandClaude Opus 4.8 e149cdb337 docs(cookbook): revert MiniMax-M3 to dev image (model not yet in a release) (#31819)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 12:16:47 -07:00
Ke Bao 7fc545b649 Align reasoning_effort schema across chat, tokenize, and responses (#31784) 2026-07-20 23:22:43 +08:00
amote-i 97e0647bdc [NPU] [DOC] Fix issues about npu docs found by aidd (#31302) 2026-07-20 16:31:27 +08:00
jacky.cheng 17fdd8487f [AMD] Update qwen3.5 cookbook (#31737) 2026-07-20 14:14:14 +08:00
b8ec544946 [DSA] Integrate Q8KV8 FP8 Sparse MLA Prefill into the DSA Backend (DeepSeek-V3.2) (#30514)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
2026-07-19 11:58:16 +08:00
Gabriel Wu faf6894093 Implement SM120 DeepSeek V4 flashinfer_mxfp4 moe runner backend + TP2 (#30272) 2026-07-18 03:01:06 -07:00
44e4999ab2 [Diffusion] Use SGLang server for ERNIE-Image prompt enhancement (#31354)
Co-authored-by: Elizaveta Martirosian <you@example.com>
Co-authored-by: Elizaveta Martirosian <elizaveta.martirosian@gmail.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-07-18 06:55:28 +03:00
67e7f8d13a [JIT] Refactor dtype traits into DTypeTrait and unify warp reductions (#30838)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: BBuf <xiaoyu.zhang@radixark.ai>
Co-authored-by: jessiewei7 <jessiewei747@gmail.com>
Co-authored-by: root <root@GPUC5A6.maas>
2026-07-18 10:07:18 +08:00
Brayden Zhong 238b2b2c9c Remove QServe and FBGEMM FP8 quantization (#31109) 2026-07-17 17:10:34 -07:00
Douglas YangandClaude Opus 4.8 a01a8e1ed9 docs(cookbook): replace pinned nightly/dev images with :latest (#31610)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 23:33:19 +00:00
Sam (Kesen Li) ec6a3163b7 [Feature] Add FP4 KV Cache Design and support SM120 GPUs (#21601) 2026-07-17 14:49:43 -07:00
Brayden ZhongandBrayden Zhong 7fc3fb9657 Remove deprecated Mamba flags from doc, wrong FP8 GEMM docstrings and change Nemotron image to 0.5.15 (#31094)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
2026-07-17 14:34:29 -07:00
Douglas YangandClaude Opus 4.8 ae3f62613a docs(cookbook): fix stale/pruned Docker image tags (#31508)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 13:24:46 -07:00
Mick 85ac56c823 docs: simplify diffusion new model guide (#30109) 2026-07-17 19:39:38 +08:00
Baizhou Zhang eaeb779ea4 [Doc] Update GLM5.2 Cookbook with LayerSplit usage (#31577) 2026-07-17 02:25:50 -07:00
zijiexiaandClaude Fable 5 8f765bc1c9 [Docs] Inkling cookbook: mark B300/GB300 recipes verified, tune B300 MTP mem fractions (#31550)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-17 16:07:09 +08:00
Junlin Wu bbd2a3fe4a [llm][npu][quant] Add W4A4 MXFP4 quantization support for Qwen3 Dense on Ascend NPU (#23795) 2026-07-17 09:06:30 +03:00
AuFlowandAuFlow bf417440e9 [Scheduler] Add SGLANG_MAX_NEW_TOKENS_LIMIT to cap per-request max_new_tokens (#22591)
Co-authored-by: AuFlow <AuFlow@users.noreply.github.com>
2026-07-16 21:34:10 -07:00
Duyi-Wang 27a52d2530 [Docs] Tune DeepSeek-V4 HiCache for MI355X PD (#31452)
Signed-off-by: Duyi-Wang <duyi.wang@amd.com>
2026-07-17 11:50:58 +08:00
zijiexiaandClaude Fable 5 40a3bd7659 [Docs] Mistral Medium 3.5 cookbook: replace stale day-0 dev images with latest (#31507)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 20:19:01 -07:00
zijiexiaandClaude Fable 5 68f4de162d [Docs] Remove Inkling H200 LoRA BF16 cookbook command (#31489)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-16 19:57:37 +00:00
ChangLiu0709 8c9833f9a9 cookbook(qwen3.5): bump AMD ROCm docker images to v0.5.15.post1 (#31454) 2026-07-16 12:41:24 -07:00
seungrokjandClaude Opus 4.6 3264477a07 [Docs] Add AMD-specific HiCache config for DeepSeek V4 playground (#31122)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-07-16 01:12:29 -07:00
zijiexia 1b9f228838 [Docs] Playground: migrate CP knob to canonical prefill-CP flags, align gating with runtime semantics (#31411) 2026-07-16 01:07:34 -07:00
Cameron Quilici a614821341 [Docs] Align B200 DeepSeek-V4-Pro balanced recipe with MegaMoE (#31373) 2026-07-16 00:24:02 -07:00
amote-i 6275114548 [NPU] [DOC] Update model names supported on Ascend NPU (#31316) 2026-07-16 14:45:59 +08:00
Yanbin Jiang 40517b593b [docs] Inkling cookbook: LoRA cells require --disable-prefill-cuda-graph (#31418) 2026-07-16 05:19:49 +00:00
Zaili Wangandzijiexia bc525dcf90 [Cookbook][CPU]Update CPU model support info in Cookbook (#30520)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-07-16 03:42:09 +00:00
sglang-botandsglang-bot 34f5691ea1 docs: sync LMSYS SGLang blog cards (#31386)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-07-16 00:53:19 +00:00
Yuhao Yang dd2e4cdc99 Add Inkling cookbook (#31360) 2026-07-16 02:24:31 +08:00
Jun LiuandXinyuan Tong d2b1243be0 docs: document CUDA crash dump output (#31333)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
2026-07-15 22:18:36 +08:00
Артем Савкинandronnie_zheng 8ed82afcc8 [MoE Refactor] [NPU] Refactor Ascend MoE implementation to reduce code duplication and align with community design (#25663)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-07-15 14:59:42 +03:00
fbcbe0a986 cookbook(deepseek-v4): add MORI disagg backend for AMD + bump MI355X image (#30651)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
2026-07-15 07:33:05 +00:00
Thomas Wang aafa706f8f [AMD] Update qwen3.5 cookbook (#31258) 2026-07-15 10:59:12 +08:00
sglang-botandsglang-bot b8a00e2ec8 docs: sync LMSYS SGLang blog cards (#31242)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
2026-07-15 00:49:30 +00:00
bdc9848c25 [Doc]Standardize the names of PyTorch NPU-related software throughout the documentation by replacing them all with TorchNPU. (#29886)
Signed-off-by: a60124901 <anyuxin4@h-partners.com>
Signed-off-by: axx-ty911 <anyuxin4@h-partners.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
2026-07-14 21:56:08 +03:00
Mick 43241b7f3f [diffusion] model: support fal Ideogram V4 Fast and Instant (#31177) 2026-07-14 19:51:35 +08:00
zijiexiaandClaude Fable 5 7e0b29ac03 docs: add VLA card image to cookbook overview (#31132)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 07:08:49 +00:00
zijiexiaandClaude Fable 5 96c2ebc58b [docs] Note the default dsa-topk-backend on all DSA-model cookbook pages (#31124)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-14 00:04:47 -07:00