Nan Jiang
|
89f4a80c1f
|
Support fastsafetensors no-GDS loading and page-cache release (#31859)
|
2026-07-31 23:12:32 +08:00 |
|
amote-i
|
0d6bef6b6d
|
[NPU] [DOC] renew triton-ascend installation guide location (#32986)
|
2026-07-31 14:46:04 +08:00 |
|
amote-i
|
c039e1a7ee
|
[NPU][DOC] Restructure ascend-npus docs into layered navigation (#32857)
|
2026-07-31 09:49:46 +08:00 |
|
 Trang DoandCheng Wan
|
a1c30701aa
|
Integrate pplx a2a backend (#30756)
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-07-30 15:33:19 -07:00 |
|
Xiaoyu Zhang
|
7784ac8f91
|
[diffusion][docs] Fix Cosmos3 model sizes (#32916)
|
2026-07-30 22:10:23 +08:00 |
|
Mick
|
b129e8a299
|
[diffusion] docs: surface diffusion AR and PE guides (#32932)
|
2026-07-30 21:05:54 +08:00 |
|
Mick
|
db3da62333
|
[diffusion] feat: unify encoder folding and batch data-parallel encoding (#30211)
|
2026-07-30 20:15:22 +08:00 |
|
Mick
|
22faf9fef8
|
embedding: centralize capabilities and complete OpenAI compatibility (#32481)
|
2026-07-30 10:28:52 +08:00 |
|
amote-i
|
20d5b91e5b
|
[NPU] [DOC] update feature name to follow the code changement (#32749)
|
2026-07-30 09:54:08 +08:00 |
|
YAMY
|
fddfc1fb5e
|
[GDN] Support FlashInfer GDN prefill with extra-buffer radix cache (#29735)
|
2026-07-30 00:47:35 +08:00 |
|
Xiaoyu Zhang
|
c32c4ef79c
|
[Kernel] Move sgl-kernel under sglang.kernels.aot (#32648)
|
2026-07-29 17:25:00 +08:00 |
|
 
|
f05c92fb6d
|
✨ [llm][npu][quant] Add W8A8 MXFP8 quantization for Qwen3 MoE on Ascend NPU (#30768)
Co-authored-by: Артем Савкин <58187114+OrangeRedeng@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-29 10:39:36 +03:00 |
|
amote-i
|
cb12a1547b
|
[NPU] [DOC] update supported features on ascend npu (#32647)
|
2026-07-29 09:05:21 +08:00 |
|
Mick
|
161fffedfc
|
docs: clarify diffusion stage reuse guidance (#32639)
|
2026-07-28 19:52:52 +08:00 |
|
 
|
8d6549bc40
|
[Attention Backend] Extend hpc_ops dynamic-scheduled decode to bf16 (#32304)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
|
2026-07-27 21:31:04 +08:00 |
|
 Jackey HuaandClaude Opus 5
|
9a0bd24bed
|
model: serve bare Qwen3Model backbone natively as an embedding model (#32457)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
2026-07-27 15:49:58 +08:00 |
|
Mohammad Miadh Angkad
|
9791fc7090
|
Add configurable FlashInfer autotune skips (#31389)
|
2026-07-25 11:17:58 -07:00 |
|
 SovietPowerandShangming Cai
|
6a046fad09
|
[PD] Prevent decode scheduler from blocking on ZMQ sends to a stalled prefill peer (#31144)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-07-25 12:53:41 +08:00 |
|
 sglang-botandsglang-bot
|
1ef2b85ef7
|
chore: bump docs install version to 0.5.16 (#32347)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
|
2026-07-24 17:19:39 -07:00 |
|
Zhihao Wang
|
0a212c6119
|
[RL] DSV4: add env to quantize SWA KV cache from bf16-rounded values (#31086)
Signed-off-by: zhihaow6 <zhihaow6@illinois.edu>
|
2026-07-24 15:52:00 -07:00 |
|
YAMY
|
de816e1eb5
|
[Disagg][StagingBuffer][1/2] Robustness and failure handling (#31217)
|
2026-07-24 17:22:55 +08:00 |
|
 Polisetty V R K Jyothendra VarmaandMa Mingfei
|
319055c191
|
[Intel GPU] Add XPU Platform support (#31949)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-07-24 12:49:29 +08:00 |
|
Jinyan Yi
|
b98a577fbe
|
Doc/update ascend quickstart (#32205)
|
2026-07-23 20:25:20 +08:00 |
|
Liangsheng Yin
|
9b853e6832
|
[Scheduler] Enable decode retraction ordering under speculative decoding (#32023)
|
2026-07-23 00:42:56 -07:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
99f636a86f
|
[Kernel] RFC #29630 finale: retire sglang.jit_kernel into sglang.kernels (#32072)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-23 08:35:09 +08:00 |
|
Brayden Zhong
|
e7511141ea
|
Support CuteDSL GEMM BF16 on SM100 on by default when allowed by heuristic (#30567)
|
2026-07-22 14:13:12 -07:00 |
|
 
|
0a6d1930c3
|
[Attention Backend] Add HPC-Ops attention backend (#30540)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Halcyon <56064364+VAthree@users.noreply.github.com>
|
2026-07-22 22:06:22 +08:00 |
|
axx-ty911
|
7dcf3f7cbf
|
[Doc] Rename A5 product name (#31998)
|
2026-07-22 15:09:05 +08:00 |
|
  
|
dcd9014f15
|
[AMD][MXFP4] Reland "Online MXFP4 quantization 2/N - FP8 to MXFP4 requantization on AMD GPUs" (#28291)
Co-authored-by: Bowen Bao <bowenbao@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
Co-authored-by: Claude <noreply@anthropic.com>
|
2026-07-21 02:50:23 -07:00 |
|
Rahul Vijayaraghavan
|
fa0ced195e
|
[XPU] Enable breakable prefill CUDA graph on XPU (#30273)
|
2026-07-21 09:09:40 +08:00 |
|
amote-i
|
97e0647bdc
|
[NPU] [DOC] Fix issues about npu docs found by aidd (#31302)
|
2026-07-20 16:31:27 +08:00 |
|
 
|
b8ec544946
|
[DSA] Integrate Q8KV8 FP8 Sparse MLA Prefill into the DSA Backend (DeepSeek-V3.2) (#30514)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-07-19 11:58:16 +08:00 |
|
  
|
44e4999ab2
|
[Diffusion] Use SGLang server for ERNIE-Image prompt enhancement (#31354)
Co-authored-by: Elizaveta Martirosian <you@example.com>
Co-authored-by: Elizaveta Martirosian <elizaveta.martirosian@gmail.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-18 06:55:28 +03:00 |
|
   
|
67e7f8d13a
|
[JIT] Refactor dtype traits into DTypeTrait and unify warp reductions (#30838)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: BBuf <xiaoyu.zhang@radixark.ai>
Co-authored-by: jessiewei7 <jessiewei747@gmail.com>
Co-authored-by: root <root@GPUC5A6.maas>
|
2026-07-18 10:07:18 +08:00 |
|
Brayden Zhong
|
238b2b2c9c
|
Remove QServe and FBGEMM FP8 quantization (#31109)
|
2026-07-17 17:10:34 -07:00 |
|
Sam (Kesen Li)
|
ec6a3163b7
|
[Feature] Add FP4 KV Cache Design and support SM120 GPUs (#21601)
|
2026-07-17 14:49:43 -07:00 |
|
 Brayden ZhongandBrayden Zhong
|
7fc3fb9657
|
Remove deprecated Mamba flags from doc, wrong FP8 GEMM docstrings and change Nemotron image to 0.5.15 (#31094)
Co-authored-by: Brayden Zhong <brayden@radixark.ai>
|
2026-07-17 14:34:29 -07:00 |
|
Mick
|
85ac56c823
|
docs: simplify diffusion new model guide (#30109)
|
2026-07-17 19:39:38 +08:00 |
|
Junlin Wu
|
bbd2a3fe4a
|
✨ [llm][npu][quant] Add W4A4 MXFP4 quantization support for Qwen3 Dense on Ascend NPU (#23795)
|
2026-07-17 09:06:30 +03:00 |
|
 AuFlowandAuFlow
|
bf417440e9
|
[Scheduler] Add SGLANG_MAX_NEW_TOKENS_LIMIT to cap per-request max_new_tokens (#22591)
Co-authored-by: AuFlow <AuFlow@users.noreply.github.com>
|
2026-07-16 21:34:10 -07:00 |
|
amote-i
|
6275114548
|
[NPU] [DOC] Update model names supported on Ascend NPU (#31316)
|
2026-07-16 14:45:59 +08:00 |
|
 Jun LiuandXinyuan Tong
|
d2b1243be0
|
docs: document CUDA crash dump output (#31333)
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-07-15 22:18:36 +08:00 |
|
 Артем Савкинandronnie_zheng
|
8ed82afcc8
|
[MoE Refactor] [NPU] Refactor Ascend MoE implementation to reduce code duplication and align with community design (#25663)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-15 14:59:42 +03:00 |
|
 
|
bdc9848c25
|
[Doc]Standardize the names of PyTorch NPU-related software throughout the documentation by replacing them all with TorchNPU. (#29886)
Signed-off-by: a60124901 <anyuxin4@h-partners.com>
Signed-off-by: axx-ty911 <anyuxin4@h-partners.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-14 21:56:08 +03:00 |
|
Mick
|
43241b7f3f
|
[diffusion] model: support fal Ideogram V4 Fast and Instant (#31177)
|
2026-07-14 19:51:35 +08:00 |
|
  ![gemini-code-assist[bot]](/assets/img/avatar_default.png)
|
702bddcee8
|
[Model] Add support for JetBrains' Mellum v2 code generation model (#27375)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Jiminator <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-07-13 22:54:38 -07:00 |
|
 Brayden Zhongandroot
|
9756f768a6
|
Refactor FP4 quantization and remove deprecated JIT kernels (#30448)
Co-authored-by: root <root@sgl-b300-inference.datacrunch.io>
|
2026-07-14 09:22:07 +08:00 |
|
     
|
423b8485fb
|
[Quantization] add humming quantization kernel (#23754)
Co-authored-by: guzekai01 <zekai01@antgroup.com>
Co-authored-by: Julian Huang <huangzhilin.hzl@gmail.com>
Co-authored-by: 墨楼 <huangzhilin.hzl@antgroup.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Peng Zhang <aniz1905@gmail.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
|
2026-07-14 08:42:56 +08:00 |
|
 
|
cfe4eefabb
|
[diffusion] model: support LongLive 2.0 T2V and I2V inference (#27639)
Co-authored-by: Yihao Wang <42559837+AgainstEntropy@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-07-13 17:39:30 -07:00 |
|
zijiexia
|
50ed4c011f
|
Remove legacy Sphinx docs/ and finish the Mintlify cutover (#28964)
|
2026-07-13 15:06:08 -07:00 |
|