Dawid Majchrowski
|
c317beda99
|
[diffusion] model: support a new model (#24994)
|
2026-05-27 08:51:03 +08:00 |
|
roikoren755
|
e958f4561f
|
[feat] Support extra_buffer in Mamba2-based models (#15829)
Signed-off-by: Roi Koren <roik@nvidia.com>
|
2026-05-26 16:03:29 +08:00 |
|
+2        
|
3f5e2c7688
|
[AMD] Dsv4/pr2 compressor opt (#26208)
Co-authored-by: wunhuang <wunhuang@amd.com>
Co-authored-by: Thomas Wang <1am9trash@gmail.com>
Co-authored-by: Xinyi Song <86638975+RolaoDenthu@users.noreply.github.com>
Co-authored-by: HaiShaw <hixiao@gmail.com>
Co-authored-by: amd-danli103 <danli103@amd.com>
Co-authored-by: Lin, Soga <soga.lin@amd.com>
Co-authored-by: Raiden-Makoto <Raiden-Makoto@users.noreply.github.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
Co-authored-by: yichiche@amd.com <jacky.cheng>
Co-authored-by: yctseng0211 <yctseng@amd.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
|
2026-05-25 23:54:40 -07:00 |
|
Ziang Li
|
2b9dd9c8b3
|
[FlashInfer v0.6.10] [RL] [DSv32] [GLM-5] Add --dsa-topk-backend and integrate FlashInfer and pytorch topk (#22851)
|
2026-05-25 13:08:03 -07:00 |
|
Xiaoyu Zhang
|
533ef41112
|
[Diffusion] Default NVFP4 backend to FlashInfer TRTLLM (#25523)
|
2026-05-25 18:14:06 +08:00 |
|
 Zhanghengand晟海
|
a4db563c87
|
[hisparse]: update user guide (#26249)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
|
2026-05-25 17:54:55 +08:00 |
|
Alex O. P.
|
ae7c4226eb
|
[diffusion] model: support FLUX.2-klein-base (#25661)
|
2026-05-22 11:24:46 +08:00 |
|
silencejade
|
e603beab55
|
[NPU] Add Qwen3.5-397B-A17B best practice doc (#25594)
|
2026-05-21 10:02:12 +08:00 |
|
 Cheng WanandClaude Sonnet 4.6
|
8131641bc6
|
[Refactor] Rename NSA → DSA: user-facing aliases, file/class/import rename (#25821)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-20 00:18:04 -07:00 |
|
 Arseniy MironovandNapkin-AI
|
45a85efc3a
|
[Diffusion][NPU]Add attention backends for diffusion models for Ascend NPU (#23482)
Co-authored-by: Napkin-AI <arseniy.mironov.dev@gmail.com>
|
2026-05-19 12:46:55 +03:00 |
|
Ziang Li
|
78cb38ed5e
|
[FlashInfer v0.6.11] [RL] Support FlashInfer per-token NVFP4 MoE (#22918)
|
2026-05-19 01:04:48 -07:00 |
|
 ![gemini-code-assist[bot]](/assets/img/avatar_default.png)
|
a080358cac
|
[Refactor] Refactor DeepEP dispatcher (#22822)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-05-18 04:36:42 +03:00 |
|
Mick
|
416fdbbb3d
|
[diffusion] feat: generalize layerwise offload residency mixin to all components (#24593)
|
2026-05-16 11:44:46 +08:00 |
|
 Lewisand百麒
|
0680f1b3d1
|
Add IntraNode NVLink configration in PD disaggregation docs (#23329)
Co-authored-by: 百麒 <yaozhong.lyz@alibaba-inc.com>
|
2026-05-13 23:23:57 -07:00 |
|
  
|
c701a08765
|
feat: [2/2][DeepEP] Add waterfill load balancing for shared expert dispatch (#19290)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Co-authored-by: root <aichenf@nvidia.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-05-13 19:23:41 -07:00 |
|
 
|
34c0029f0a
|
[diffusion] [AMD] feat: support online MXFP4 and fp8 quantization (#21431)
Co-authored-by: Bowen Bao <bowenbao@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-05-14 08:52:01 +08:00 |
|
Mick
|
ff70aeac30
|
[diffusion] feat: add performance mode server args (#24491)
|
2026-05-14 00:57:46 +08:00 |
|
Khoa Pham
|
c665edec6e
|
[env] Make max KV chunk capacity configurable via SGLANG_MAX_KV_CHUNK_CAPACITY (#25120)
|
2026-05-12 22:37:45 -07:00 |
|
 RulinJuiceandRulinJuice
|
3f048c80b8
|
Reject repetition_penalty=0 in SamplingParams.verify() (#24874)
Co-authored-by: RulinJuice <265952454+RulinJuice@users.noreply.github.com>
|
2026-05-12 21:25:23 -07:00 |
|
 TianheandClaude Sonnet 4.6
|
95985f983d
|
feat(trace): support SGLANG_TRACE_LEVEL env var for startup trace level (#24716)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
2026-05-11 10:31:03 -07:00 |
|
 
|
f9c315e85d
|
docs: clarify how /tag-and-rerun-ci kicks off CI on the current commit (#24774)
Co-authored-by: Byron Hsu <byron@periodiclabs.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-10 11:28:41 +08:00 |
|
YAMY
|
560829a171
|
feat(scheduler): add adaptive queue-based prefill delayer trigger (#23189)
|
2026-05-08 16:54:30 -07:00 |
|
cen121212
|
461bc8af49
|
[NPU][Doc] Update GLM-5 docs, enabling deepep by default (#23708)
|
2026-05-08 11:12:35 +08:00 |
|
Revanth Reddy Airre
|
be088f8076
|
fix(router): configure HTTP client connection settings (#24330)
Signed-off-by: Revanth Reddy Airre <revanthreddy@hippocraticai.com>
|
2026-05-07 11:42:45 -07:00 |
|
![github-actions[bot]](/assets/img/avatar_default.png) 
|
80a6014243
|
✨ [diffusion][npu][quant] Add MXFP8 quantization support for Wan2.2 Diffusion on Ascend NPU (#20922)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2026-05-07 21:30:56 +03:00 |
|
inkcherry
|
3b2c730320
|
[AMD] Enable dual-stream MoE on ROCm (#24005)
Signed-off-by: inkcherry <mingzhi.liu@amd.com>
|
2026-05-07 02:27:24 -07:00 |
|
Revanth Reddy Airre
|
d363315de9
|
fix(router): make HTTP pool idle timeout configurable (#24329)
Signed-off-by: Revanth Reddy Airre <revanthreddy@hippocraticai.com>
|
2026-05-06 22:11:11 -07:00 |
|
Jianhong Zhang
|
c7019ff33d
|
[NIXL][XPU] Use np.uint64 for pointer/length arrays in disaggregation KV transfer (#24188)
|
2026-05-06 10:09:03 +08:00 |
|
Xiaoyu Zhang
|
8c703f215e
|
Add HunyuanVideo ModelOpt FP8 diffusion support (#23199)
|
2026-05-05 19:27:28 +08:00 |
|
Mick
|
2f7d99b7f7
|
[diffusion] cli: support component attention backend overrides (#24320)
|
2026-05-05 08:39:27 +08:00 |
|
 Xiaoyu ZhangandMick
|
f2d1390909
|
[Diffusion] Add Qwen Image ModelOpt FP8 support (#23155)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-05-04 00:24:22 +08:00 |
|
Glen Liu
|
76b9c8de6f
|
[Feature] add LoRADrainer to address high P99 TTFT (#17913)
|
2026-05-02 16:13:43 -07:00 |
|
 egvenediktovandronnie_zheng
|
83bf5d6869
|
[NPU]TP Communications compression For Qwen3 models for NPU (#20520)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-05-02 14:29:11 +03:00 |
|
Xiaoyu Zhang
|
589f90b368
|
[diffusion] chore: use lmsys as org for modelopt checkpoints (#23924)
|
2026-05-02 17:18:58 +08:00 |
|
Lianmin Zheng
|
ece8a1a788
|
Refactor device timer, clean up metrics collector, and add fwd occupancy metric (#24197)
|
2026-05-01 10:25:25 -07:00 |
|
 
|
3272af2f00
|
[Apple Silicon] [MLX] MLX decode partial overlap scheduling for generation (async eval) (#22416)
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-04-29 12:21:14 -07:00 |
|
AndyLi429
|
4c1eefca4f
|
[NPU] ascend backend support qwen3 moe attention cp (#21685)
|
2026-04-29 19:25:17 +08:00 |
|
Lianmin Zheng
|
d66eb3a91b
|
docs: update contribution guide with coding style guidelines (#23977)
|
2026-04-28 19:51:58 -07:00 |
|
Xun Sun
|
9a53ab3d6d
|
[6/N] (Elastic EP) Recover failed ranks (#15771)
|
2026-04-28 00:44:26 -07:00 |
|
Pai Liu
|
7b9ff79f93
|
docs: update Python prerequisite to 3.10 (#23801)
|
2026-04-27 15:36:38 -07:00 |
|
 1874.andronnie_zheng
|
046c14a3ed
|
[NPU] Support GGUF quantization for Ascend NPU (dense + MoE) (#17883)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-04-25 17:16:47 +03:00 |
|
Yujing
|
6175946db7
|
[Feature]Add MSProbe dump support in SGLang (#18349)
|
2026-04-25 10:12:50 +03:00 |
|
   
|
6d03861476
|
support Hy3 preview (#23533)
Co-authored-by: pengmeng <pengmeng@tencent.com>
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
Co-authored-by: chengvjiang <chengvjiang@tencent.com>
Co-authored-by: russellfeng <russellfeng@tencent.com>
|
2026-04-24 12:03:24 -07:00 |
|
Mick
|
cd1fa7506a
|
[diffusion] model: support LTX2.3 high quality pipeline (#23366)
|
2026-04-24 14:18:20 +08:00 |
|
zijiexia
|
6490afe36e
|
[docs] add deprecation notice banner to legacy documentation site (#23516)
|
2026-04-22 20:17:53 -07:00 |
|
 jianzhao-xuandJianzhao Xu
|
2f3e6a3143
|
[NPU] offloading docs update (#23378)
Co-authored-by: Jianzhao Xu <xujianchao@huawei.com>
|
2026-04-22 11:01:55 +08:00 |
|
Yanbin Jiang
|
4f764dfbb8
|
[Lora] Support LoRA and multi-batch in bench_one_batch_server (#23047)
|
2026-04-21 14:20:11 -07:00 |
|
amote-i
|
301604f953
|
[NPU] [DOC] Quick start doc for Ascend NPU (#23238)
|
2026-04-21 11:19:09 +08:00 |
|
 shuwennandQiaolin-Yu
|
b65799cf83
|
[SPEC][1/N] feat: add adaptive speculative_num_steps for EAGLE topk=1 (#21599)
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
|
2026-04-20 14:25:04 -07:00 |
|
    
|
7ca3566130
|
Multi platform Plugin (#21388)
Co-authored-by: root <root@tjzj-inf-sci-k8s-bzz2-0183.tjzj.baidu.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
Co-authored-by: Alex Nails <alexj.nails@gmail.com>
Co-authored-by: root <root@tjzj-inf-sci-k8s-bzz2-0000.tjzj.baidu.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-04-19 17:23:51 -07:00 |
|