Rahul Vijayaraghavan
|
fa0ced195e
|
[XPU] Enable breakable prefill CUDA graph on XPU (#30273)
|
2026-07-21 09:09:40 +08:00 |
|
 zijiexiaandClaude Opus 4.8
|
0a2d3ca071
|
[cookbook] Inkling: add measured accuracy numbers to benchmark cards (#31823)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-20 22:43:35 +00:00 |
|
 Douglas YangandClaude Opus 4.8
|
e149cdb337
|
docs(cookbook): revert MiniMax-M3 to dev image (model not yet in a release) (#31819)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-20 12:16:47 -07:00 |
|
Ke Bao
|
7fc545b649
|
Align reasoning_effort schema across chat, tokenize, and responses (#31784)
|
2026-07-20 23:22:43 +08:00 |
|
amote-i
|
97e0647bdc
|
[NPU] [DOC] Fix issues about npu docs found by aidd (#31302)
|
2026-07-20 16:31:27 +08:00 |
|
jacky.cheng
|
17fdd8487f
|
[AMD] Update qwen3.5 cookbook (#31737)
|
2026-07-20 14:14:14 +08:00 |
|
 
|
b8ec544946
|
[DSA] Integrate Q8KV8 FP8 Sparse MLA Prefill into the DSA Backend (DeepSeek-V3.2) (#30514)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-07-19 11:58:16 +08:00 |
|
Gabriel Wu
|
faf6894093
|
Implement SM120 DeepSeek V4 flashinfer_mxfp4 moe runner backend + TP2 (#30272)
|
2026-07-18 03:01:06 -07:00 |
|
  
|
44e4999ab2
|
[Diffusion] Use SGLang server for ERNIE-Image prompt enhancement (#31354)
Co-authored-by: Elizaveta Martirosian <you@example.com>
Co-authored-by: Elizaveta Martirosian <elizaveta.martirosian@gmail.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-18 06:55:28 +03:00 |
|
   
|
67e7f8d13a
|
[JIT] Refactor dtype traits into DTypeTrait and unify warp reductions (#30838)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: BBuf <xiaoyu.zhang@radixark.ai>
Co-authored-by: jessiewei7 <jessiewei747@gmail.com>
Co-authored-by: root <root@GPUC5A6.maas>
|
2026-07-18 10:07:18 +08:00 |
|
Brayden Zhong
|
238b2b2c9c
|
Remove QServe and FBGEMM FP8 quantization (#31109)
|
2026-07-17 17:10:34 -07:00 |
|
 Douglas YangandClaude Opus 4.8
|
a01a8e1ed9
|
docs(cookbook): replace pinned nightly/dev images with :latest (#31610)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-17 23:33:19 +00:00 |
|
Sam (Kesen Li)
|
ec6a3163b7
|
[Feature] Add FP4 KV Cache Design and support SM120 GPUs (#21601)
|
2026-07-17 14:49:43 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
7fc3fb9657
|
Remove deprecated Mamba flags from doc, wrong FP8 GEMM docstrings and change Nemotron image to 0.5.15 (#31094)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-07-17 14:34:29 -07:00 |
|
 Douglas YangandClaude Opus 4.8
|
ae3f62613a
|
docs(cookbook): fix stale/pruned Docker image tags (#31508)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-17 13:24:46 -07:00 |
|
Mick
|
85ac56c823
|
docs: simplify diffusion new model guide (#30109)
|
2026-07-17 19:39:38 +08:00 |
|
Baizhou Zhang
|
eaeb779ea4
|
[Doc] Update GLM5.2 Cookbook with LayerSplit usage (#31577)
|
2026-07-17 02:25:50 -07:00 |
|
 zijiexiaandClaude Fable 5
|
8f765bc1c9
|
[Docs] Inkling cookbook: mark B300/GB300 recipes verified, tune B300 MTP mem fractions (#31550)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-17 16:07:09 +08:00 |
|
Junlin Wu
|
bbd2a3fe4a
|
✨ [llm][npu][quant] Add W4A4 MXFP4 quantization support for Qwen3 Dense on Ascend NPU (#23795)
|
2026-07-17 09:06:30 +03:00 |
|
 AuFlowandAuFlow
|
bf417440e9
|
[Scheduler] Add SGLANG_MAX_NEW_TOKENS_LIMIT to cap per-request max_new_tokens (#22591)
Co-authored-by: AuFlow <AuFlow@users.noreply.github.com>
|
2026-07-16 21:34:10 -07:00 |
|
Duyi-Wang
|
27a52d2530
|
[Docs] Tune DeepSeek-V4 HiCache for MI355X PD (#31452)
Signed-off-by: Duyi-Wang <duyi.wang@amd.com>
|
2026-07-17 11:50:58 +08:00 |
|
 zijiexiaandClaude Fable 5
|
40a3bd7659
|
[Docs] Mistral Medium 3.5 cookbook: replace stale day-0 dev images with latest (#31507)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-16 20:19:01 -07:00 |
|
 zijiexiaandClaude Fable 5
|
68f4de162d
|
[Docs] Remove Inkling H200 LoRA BF16 cookbook command (#31489)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-16 19:57:37 +00:00 |
|
ChangLiu0709
|
8c9833f9a9
|
cookbook(qwen3.5): bump AMD ROCm docker images to v0.5.15.post1 (#31454)
|
2026-07-16 12:41:24 -07:00 |
|
 seungrokjandClaude Opus 4.6
|
3264477a07
|
[Docs] Add AMD-specific HiCache config for DeepSeek V4 playground (#31122)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-07-16 01:12:29 -07:00 |
|
zijiexia
|
1b9f228838
|
[Docs] Playground: migrate CP knob to canonical prefill-CP flags, align gating with runtime semantics (#31411)
|
2026-07-16 01:07:34 -07:00 |
|
Cameron Quilici
|
a614821341
|
[Docs] Align B200 DeepSeek-V4-Pro balanced recipe with MegaMoE (#31373)
|
2026-07-16 00:24:02 -07:00 |
|
amote-i
|
6275114548
|
[NPU] [DOC] Update model names supported on Ascend NPU (#31316)
|
2026-07-16 14:45:59 +08:00 |
|
Yanbin Jiang
|
40517b593b
|
[docs] Inkling cookbook: LoRA cells require --disable-prefill-cuda-graph (#31418)
|
2026-07-16 05:19:49 +00:00 |
|
 Zaili Wangandzijiexia
|
bc525dcf90
|
[Cookbook][CPU]Update CPU model support info in Cookbook (#30520)
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
|
2026-07-16 03:42:09 +00:00 |
|
 sglang-botandsglang-bot
|
34f5691ea1
|
docs: sync LMSYS SGLang blog cards (#31386)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-07-16 00:53:19 +00:00 |
|
Yuhao Yang
|
dd2e4cdc99
|
Add Inkling cookbook (#31360)
|
2026-07-16 02:24:31 +08:00 |
|
 Jun LiuandXinyuan Tong
|
d2b1243be0
|
docs: document CUDA crash dump output (#31333)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-07-15 22:18:36 +08:00 |
|
 Артем Савкинandronnie_zheng
|
8ed82afcc8
|
[MoE Refactor] [NPU] Refactor Ascend MoE implementation to reduce code duplication and align with community design (#25663)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-15 14:59:42 +03:00 |
|
  
|
fbcbe0a986
|
cookbook(deepseek-v4): add MORI disagg backend for AMD + bump MI355X image (#30651)
Co-authored-by: Zijie Xia <zijie.xia@radixark.ai>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
|
2026-07-15 07:33:05 +00:00 |
|
Thomas Wang
|
aafa706f8f
|
[AMD] Update qwen3.5 cookbook (#31258)
|
2026-07-15 10:59:12 +08:00 |
|
 sglang-botandsglang-bot
|
b8a00e2ec8
|
docs: sync LMSYS SGLang blog cards (#31242)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-07-15 00:49:30 +00:00 |
|
 
|
bdc9848c25
|
[Doc]Standardize the names of PyTorch NPU-related software throughout the documentation by replacing them all with TorchNPU. (#29886)
Signed-off-by: a60124901 <anyuxin4@h-partners.com>
Signed-off-by: axx-ty911 <anyuxin4@h-partners.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-14 21:56:08 +03:00 |
|
Mick
|
43241b7f3f
|
[diffusion] model: support fal Ideogram V4 Fast and Instant (#31177)
|
2026-07-14 19:51:35 +08:00 |
|
 zijiexiaandClaude Fable 5
|
7e0b29ac03
|
docs: add VLA card image to cookbook overview (#31132)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-14 07:08:49 +00:00 |
|
 zijiexiaandClaude Fable 5
|
96c2ebc58b
|
[docs] Note the default dsa-topk-backend on all DSA-model cookbook pages (#31124)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-14 00:04:47 -07:00 |
|
  ![gemini-code-assist[bot]](/assets/img/avatar_default.png)
|
702bddcee8
|
[Model] Add support for JetBrains' Mellum v2 code generation model (#27375)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Jiminator <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-07-13 22:54:38 -07:00 |
|
Mick
|
23b2c6f1ce
|
docs: fix diffusion cookbook overview cards (#31101)
|
2026-07-14 10:40:42 +08:00 |
|
 Brayden Zhongandroot
|
9756f768a6
|
Refactor FP4 quantization and remove deprecated JIT kernels (#30448)
Co-authored-by: root <root@sgl-b300-inference.datacrunch.io>
|
2026-07-14 09:22:07 +08:00 |
|
     
|
423b8485fb
|
[Quantization] add humming quantization kernel (#23754)
Co-authored-by: guzekai01 <zekai01@antgroup.com>
Co-authored-by: Julian Huang <huangzhilin.hzl@gmail.com>
Co-authored-by: 墨楼 <huangzhilin.hzl@antgroup.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Peng Zhang <aniz1905@gmail.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-07-14 08:42:56 +08:00 |
|
 
|
cfe4eefabb
|
[diffusion] model: support LongLive 2.0 T2V and I2V inference (#27639)
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-07-13 17:39:30 -07:00 |
|
zijiexia
|
50ed4c011f
|
Remove legacy Sphinx docs/ and finish the Mintlify cutover (#28964)
|
2026-07-13 15:06:08 -07:00 |
|
Liangsheng Yin
|
e2728ac504
|
[Spec] Remove dead padded_static_len and stale SGLANG_ENABLE_SPEC_V2 references (#30998)
|
2026-07-13 15:30:31 -05:00 |
|
 sglang-botandsglang-bot
|
805385414e
|
chore: bump docs install version to 0.5.15 (#31058)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-07-13 18:58:48 +00:00 |
|
 sglang-botandsglang-bot
|
b677babc62
|
docs: sync LMSYS SGLang blog cards (#30571)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-07-13 18:44:39 +00:00 |
|