Alex O. P.
|
ae7c4226eb
|
[diffusion] model: support FLUX.2-klein-base (#25661)
|
2026-05-22 11:24:46 +08:00 |
|
silencejade
|
e603beab55
|
[NPU] Add Qwen3.5-397B-A17B best practice doc (#25594)
|
2026-05-21 10:02:12 +08:00 |
|
 Cheng WanandClaude Sonnet 4.6
|
8131641bc6
|
[Refactor] Rename NSA → DSA: user-facing aliases, file/class/import rename (#25821)
Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
|
2026-05-20 00:18:04 -07:00 |
|
 Arseniy MironovandNapkin-AI
|
45a85efc3a
|
[Diffusion][NPU]Add attention backends for diffusion models for Ascend NPU (#23482)
Co-authored-by: Napkin-AI <arseniy.mironov.dev@gmail.com>
|
2026-05-19 12:46:55 +03:00 |
|
Ziang Li
|
78cb38ed5e
|
[FlashInfer v0.6.11] [RL] Support FlashInfer per-token NVFP4 MoE (#22918)
|
2026-05-19 01:04:48 -07:00 |
|
 ![gemini-code-assist[bot]](/assets/img/avatar_default.png)
|
a080358cac
|
[Refactor] Refactor DeepEP dispatcher (#22822)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-05-18 04:36:42 +03:00 |
|
Mick
|
416fdbbb3d
|
[diffusion] feat: generalize layerwise offload residency mixin to all components (#24593)
|
2026-05-16 11:44:46 +08:00 |
|
 Lewisand百麒
|
0680f1b3d1
|
Add IntraNode NVLink configration in PD disaggregation docs (#23329)
Co-authored-by: 百麒 <yaozhong.lyz@alibaba-inc.com>
|
2026-05-13 23:23:57 -07:00 |
|
  
|
c701a08765
|
feat: [2/2][DeepEP] Add waterfill load balancing for shared expert dispatch (#19290)
Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Co-authored-by: root <aichenf@nvidia.com>
Co-authored-by: Cheng Wan <54331508+ch-wan@users.noreply.github.com>
|
2026-05-13 19:23:41 -07:00 |
|
 
|
34c0029f0a
|
[diffusion] [AMD] feat: support online MXFP4 and fp8 quantization (#21431)
Co-authored-by: Bowen Bao <bowenbao@amd.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-05-14 08:52:01 +08:00 |
|
Mick
|
ff70aeac30
|
[diffusion] feat: add performance mode server args (#24491)
|
2026-05-14 00:57:46 +08:00 |
|
Khoa Pham
|
c665edec6e
|
[env] Make max KV chunk capacity configurable via SGLANG_MAX_KV_CHUNK_CAPACITY (#25120)
|
2026-05-12 22:37:45 -07:00 |
|
 RulinJuiceandRulinJuice
|
3f048c80b8
|
Reject repetition_penalty=0 in SamplingParams.verify() (#24874)
Co-authored-by: RulinJuice <265952454+RulinJuice@users.noreply.github.com>
|
2026-05-12 21:25:23 -07:00 |
|
 TianheandClaude Sonnet 4.6
|
95985f983d
|
feat(trace): support SGLANG_TRACE_LEVEL env var for startup trace level (#24716)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
|
2026-05-11 10:31:03 -07:00 |
|
 
|
f9c315e85d
|
docs: clarify how /tag-and-rerun-ci kicks off CI on the current commit (#24774)
Co-authored-by: Byron Hsu <byron@periodiclabs.ai>
Co-authored-by: Cursor <cursoragent@cursor.com>
|
2026-05-10 11:28:41 +08:00 |
|
YAMY
|
560829a171
|
feat(scheduler): add adaptive queue-based prefill delayer trigger (#23189)
|
2026-05-08 16:54:30 -07:00 |
|
cen121212
|
461bc8af49
|
[NPU][Doc] Update GLM-5 docs, enabling deepep by default (#23708)
|
2026-05-08 11:12:35 +08:00 |
|
Revanth Reddy Airre
|
be088f8076
|
fix(router): configure HTTP client connection settings (#24330)
Signed-off-by: Revanth Reddy Airre <revanthreddy@hippocraticai.com>
|
2026-05-07 11:42:45 -07:00 |
|
![github-actions[bot]](/assets/img/avatar_default.png) 
|
80a6014243
|
✨ [diffusion][npu][quant] Add MXFP8 quantization support for Wan2.2 Diffusion on Ascend NPU (#20922)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
|
2026-05-07 21:30:56 +03:00 |
|
inkcherry
|
3b2c730320
|
[AMD] Enable dual-stream MoE on ROCm (#24005)
Signed-off-by: inkcherry <mingzhi.liu@amd.com>
|
2026-05-07 02:27:24 -07:00 |
|
Revanth Reddy Airre
|
d363315de9
|
fix(router): make HTTP pool idle timeout configurable (#24329)
Signed-off-by: Revanth Reddy Airre <revanthreddy@hippocraticai.com>
|
2026-05-06 22:11:11 -07:00 |
|
Jianhong Zhang
|
c7019ff33d
|
[NIXL][XPU] Use np.uint64 for pointer/length arrays in disaggregation KV transfer (#24188)
|
2026-05-06 10:09:03 +08:00 |
|
Xiaoyu Zhang
|
8c703f215e
|
Add HunyuanVideo ModelOpt FP8 diffusion support (#23199)
|
2026-05-05 19:27:28 +08:00 |
|
Mick
|
2f7d99b7f7
|
[diffusion] cli: support component attention backend overrides (#24320)
|
2026-05-05 08:39:27 +08:00 |
|
 Xiaoyu ZhangandMick
|
f2d1390909
|
[Diffusion] Add Qwen Image ModelOpt FP8 support (#23155)
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-05-04 00:24:22 +08:00 |
|
Glen Liu
|
76b9c8de6f
|
[Feature] add LoRADrainer to address high P99 TTFT (#17913)
|
2026-05-02 16:13:43 -07:00 |
|
 egvenediktovandronnie_zheng
|
83bf5d6869
|
[NPU]TP Communications compression For Qwen3 models for NPU (#20520)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-05-02 14:29:11 +03:00 |
|
Xiaoyu Zhang
|
589f90b368
|
[diffusion] chore: use lmsys as org for modelopt checkpoints (#23924)
|
2026-05-02 17:18:58 +08:00 |
|
Lianmin Zheng
|
ece8a1a788
|
Refactor device timer, clean up metrics collector, and add fwd occupancy metric (#24197)
|
2026-05-01 10:25:25 -07:00 |
|
 
|
3272af2f00
|
[Apple Silicon] [MLX] MLX decode partial overlap scheduling for generation (async eval) (#22416)
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-04-29 12:21:14 -07:00 |
|
AndyLi429
|
4c1eefca4f
|
[NPU] ascend backend support qwen3 moe attention cp (#21685)
|
2026-04-29 19:25:17 +08:00 |
|
Lianmin Zheng
|
d66eb3a91b
|
docs: update contribution guide with coding style guidelines (#23977)
|
2026-04-28 19:51:58 -07:00 |
|
Xun Sun
|
9a53ab3d6d
|
[6/N] (Elastic EP) Recover failed ranks (#15771)
|
2026-04-28 00:44:26 -07:00 |
|
Pai Liu
|
7b9ff79f93
|
docs: update Python prerequisite to 3.10 (#23801)
|
2026-04-27 15:36:38 -07:00 |
|
 1874.andronnie_zheng
|
046c14a3ed
|
[NPU] Support GGUF quantization for Ascend NPU (dense + MoE) (#17883)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-04-25 17:16:47 +03:00 |
|
Yujing
|
6175946db7
|
[Feature]Add MSProbe dump support in SGLang (#18349)
|
2026-04-25 10:12:50 +03:00 |
|
   
|
6d03861476
|
support Hy3 preview (#23533)
Co-authored-by: pengmeng <pengmeng@tencent.com>
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
Co-authored-by: chengvjiang <chengvjiang@tencent.com>
Co-authored-by: russellfeng <russellfeng@tencent.com>
|
2026-04-24 12:03:24 -07:00 |
|
Mick
|
cd1fa7506a
|
[diffusion] model: support LTX2.3 high quality pipeline (#23366)
|
2026-04-24 14:18:20 +08:00 |
|
zijiexia
|
6490afe36e
|
[docs] add deprecation notice banner to legacy documentation site (#23516)
|
2026-04-22 20:17:53 -07:00 |
|
 jianzhao-xuandJianzhao Xu
|
2f3e6a3143
|
[NPU] offloading docs update (#23378)
Co-authored-by: Jianzhao Xu <xujianchao@huawei.com>
|
2026-04-22 11:01:55 +08:00 |
|
Yanbin Jiang
|
4f764dfbb8
|
[Lora] Support LoRA and multi-batch in bench_one_batch_server (#23047)
|
2026-04-21 14:20:11 -07:00 |
|
amote-i
|
301604f953
|
[NPU] [DOC] Quick start doc for Ascend NPU (#23238)
|
2026-04-21 11:19:09 +08:00 |
|
 shuwennandQiaolin-Yu
|
b65799cf83
|
[SPEC][1/N] feat: add adaptive speculative_num_steps for EAGLE topk=1 (#21599)
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
|
2026-04-20 14:25:04 -07:00 |
|
    
|
7ca3566130
|
Multi platform Plugin (#21388)
Co-authored-by: root <root@tjzj-inf-sci-k8s-bzz2-0183.tjzj.baidu.com>
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
Co-authored-by: Alex Nails <alexj.nails@gmail.com>
Co-authored-by: root <root@tjzj-inf-sci-k8s-bzz2-0000.tjzj.baidu.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-04-19 17:23:51 -07:00 |
|
amote-i
|
ea20f1baa4
|
[NPU] [DOC] Update npu best practice docs to match latest code (#23077)
|
2026-04-18 14:17:00 +08:00 |
|
Mick
|
0d94c3366a
|
[diffusion] feat: introduce ltx-2-two-stage device manager (#22869)
|
2026-04-18 11:04:33 +08:00 |
|
Xiaoyu Zhang
|
615d6c93b2
|
[codex] Add flashinfer TRTLLM backend for diffusion NVFP4 (#22717)
|
2026-04-18 09:06:28 +08:00 |
|
Lianmin Zheng
|
44e67c6835
|
Remove deprecated double sparsity feature (#23009)
|
2026-04-17 13:33:12 -07:00 |
|
Mick
|
0b2058853d
|
[diffusion] doc: update doc (#23052)
|
2026-04-17 16:23:46 +08:00 |
|
Duyi-Wang
|
8c190f6b91
|
[AMD] Add SGLANG_MORI_MOE_MAX_INPUT_TOKENS to truncate dispatch before MoE. (#22952)
|
2026-04-16 23:40:15 -07:00 |
|