jacky.cheng
|
e78051a419
|
[AMD] [Fix] Fix --attention-backend triton work for DeepSeek MLA on MI355 (null-K + decode dispatch + RoPE) (#30355)
|
2026-07-15 14:19:23 -07:00 |
|
Yuhao Yang
|
5abec3fbf8
|
Fix MiMo-V2 on Blackwell: FA3 fallback and TP-aware audio weight loading (#31343)
|
2026-07-15 13:26:27 -07:00 |
|
Shaun Kotek
|
ab627e5d75
|
fix: load the right mtp lm head quantization (#30976)
|
2026-07-15 12:12:04 -07:00 |
|
Xuwei
|
7e7129acd7
|
[Bugfix] Release Mamba cache after PP dynamic chunk profiling (#31321)
|
2026-07-16 02:32:07 +08:00 |
|
 YAMYandXuwei Li
|
2d00e20a52
|
[Disagg][Qwen3.5] Fix heterogeneous attn-TP scatter transfer: GDN conv sub-block slice + GQA replicated-KV head map (#30997)
Co-authored-by: Xuwei Li <lixuwei.xy@gmail.com>
|
2026-07-16 02:31:37 +08:00 |
|
 AuFlowandAuFlow
|
9d147fdca1
|
[Multimodal] Support n>1 outputs for GLM-Image generation (#31027)
Co-authored-by: AuFlow <AuFlow@users.noreply.github.com>
|
2026-07-15 20:55:19 +03:00 |
|
Liangsheng Yin
|
dc078ddd2a
|
[Spec] Extract stateless draft prepare helpers into eagle_worker_common (#31257)
|
2026-07-15 10:54:31 -07:00 |
|
 Bingxu Chenandntgiang71096
|
d36e96ce23
|
[AMD] Enable mamba-extra-buffer for Qwen3.5 on ROCm (#30359)
Co-authored-by: ntgiang71096 <nguyentruonggiang71096@gmail.com>
|
2026-07-15 10:18:12 -07:00 |
|
 YC Yen-Ching Tsengandkangwangamd
|
a8b60433c2
|
[AMD] Fix DSV4 JIT build on rocm (#31131)
Co-authored-by: kangwangamd <kangwang@amd.com>
|
2026-07-15 09:58:12 -07:00 |
|
 Andy YeandClaude Fable 5
|
c879f3da5c
|
[diffusion] rl: support rl rollout for the wan pipeline via a per-request scheduler switch (#30036)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-15 22:30:51 +08:00 |
|
 
|
495ae9aaa6
|
Fix Ministral3 accuracy issue by aligning YaRN RoPE scaling with Transformers implementation (#31232)
Co-authored-by: Elizaveta Martirosian <elizaveta.martirosian@gmail.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-15 15:23:04 +03:00 |
|
 Артем Савкинandronnie_zheng
|
8ed82afcc8
|
[MoE Refactor] [NPU] Refactor Ascend MoE implementation to reduce code duplication and align with community design (#25663)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-15 14:59:42 +03:00 |
|
 
|
c9b17403e7
|
Fix image URL response for multiple outputs (#30621)
Co-authored-by: AuFlow <AuFlow@users.noreply.github.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-15 14:49:04 +03:00 |
|
Mick
|
947a14d617
|
feat: unify multimodal feature transport (#30904)
|
2026-07-15 17:42:38 +08:00 |
|
 weireweireandweireweire
|
f2c875d1c8
|
[PD] Route PD server warmup to every DP rank (#30748)
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
|
2026-07-15 15:59:41 +08:00 |
|
  
|
241937af87
|
[NPU] Determine the topk norm_type through scoring_func (#31107)
Co-authored-by: iridiumine <42236072+iridiumine@users.noreply.github.com>
Co-authored-by: zhaozx-cn <59479021+zhaozx-cn@users.noreply.github.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
|
2026-07-15 10:26:44 +03:00 |
|
 Sam ShleiferandAlex Nails
|
dec0836302
|
Fix processor config loading for object-storage model paths (#31211)
Co-authored-by: Alex Nails <alex.nails@radixark.ai>
|
2026-07-15 00:21:31 -07:00 |
|
 skyandBrayden Zhong
|
980acd6eca
|
Fix MoE TP allreduce to use NCCL symmetric memory via in-pool output allocation (#29007)
Signed-off-by: wangfakang <fakangwang@gmail.com>
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-07-15 15:06:37 +08:00 |
|
Ma Mingfei
|
41e0b4b369
|
[CPU] add fused input proj for qwen3.5 (#31171)
|
2026-07-15 15:06:24 +08:00 |
|
 Yuan Luoandluoyuan.luo
|
a649b5a9db
|
[KDA] Add FlashInfer SM100 KDA decode + MTP (target_verify) backend (#30113)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-07-15 15:04:20 +08:00 |
|
fzyzcjy
|
1afab30577
|
Fix bookkeeping fields not encapsulated with real allocations in normal alloc, PD pre-alloc, DFlash and EAGLE (#29432)
|
2026-07-15 14:52:21 +08:00 |
|
fzyzcjy
|
e789ca24a7
|
Lightweight extract allocation logic from mem_cache/common.py to more clearly show nearly parallel variants (#29431)
|
2026-07-15 14:49:37 +08:00 |
|
fzyzcjy
|
c315df49bb
|
Fix abusing presence of req.req_pool_idx to indicate the presence of req.kv resources (#29430)
|
2026-07-15 14:48:39 +08:00 |
|
fzyzcjy
|
2d979f1d8c
|
Let the presence of req.kv indicate the existence of owned kv resources (#29429)
|
2026-07-15 14:47:16 +08:00 |
|
fzyzcjy
|
27256aee5b
|
Let cache backend do not couple with owned committed kv details and avoid kv_committed_freed/kv_overallocated_freed fields (#29428)
|
2026-07-15 14:43:10 +08:00 |
|
fzyzcjy
|
d8d76c4d12
|
Introduce req.kv container for coupled owned kv field lifecycle (#29427)
|
2026-07-15 14:40:38 +08:00 |
|
fzyzcjy
|
201ddeaba1
|
Avoid relaying per-step outputs through ScheduleBatch fields in disagg prefill and PP (#30677)
|
2026-07-15 14:33:37 +08:00 |
|
fzyzcjy
|
01343d2759
|
Avoid implicit running_batch access in dllm and pdmux scheduling (#30676)
|
2026-07-15 14:33:00 +08:00 |
|
fzyzcjy
|
21c62b9830
|
Rewrite pause_generation retract path as req-level release and requeue for clarity (#30675)
|
2026-07-15 14:32:12 +08:00 |
|
fzyzcjy
|
1967b9ec99
|
Fix missed hisparse release and stale field cleanup in pause retract (#30674)
|
2026-07-15 14:31:11 +08:00 |
|
fzyzcjy
|
b6cc897fea
|
Fix non-existent abort mode in Scheduler.pause_generation and inline retract_all (#30673)
|
2026-07-15 14:27:48 +08:00 |
|
fzyzcjy
|
52a88fb212
|
Avoid mutating ScheduleBatch fields in place (#30672)
|
2026-07-15 14:27:08 +08:00 |
|
fzyzcjy
|
e77d95c3d5
|
Pass per-forward overrides to ForwardBatch.init_new as explicit arguments (#30670)
|
2026-07-15 14:25:59 +08:00 |
|
fzyzcjy
|
861d97d24d
|
Remove dead ScheduleBatch fields and avoid inplace seq_lens bump (#30669)
|
2026-07-15 14:23:31 +08:00 |
|
AMD-yanfeiwang
|
a3194d3585
|
[AMD] Remove ROCm page_first+kernel -> layer_first HiCache fallback (follow-up to #28534) (#30622)
|
2026-07-14 22:48:16 -07:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
4aadf94146
|
[Kernel] Relocate vendored fla and mamba kernel trees to sglang.kernels (RFC #29630, Phase 2.5, 7/7) (#30795)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-15 12:52:15 +08:00 |
|
ANSHUMAN TRIPATHY
|
23f2b77d82
|
Make UTs compatible for XPU (#27106)
|
2026-07-15 12:35:56 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
c00131ebaa
|
[Kernel] Migrate linear-attention, MiniMax-sparse and diffusion kernels to sglang.kernels (RFC #29630, Phase 2.5, 6/7) (#30793)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-15 11:21:36 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
ba5be86d42
|
[Kernel] Migrate DSA + DSV4 attention kernels to sglang.kernels (RFC #29630, Phase 2.5, 5/7) (#30792)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-15 11:11:22 +08:00 |
|
 
|
4ae9cc3c81
|
Fix gate stride for 4D decode layouts (#31231)
Co-authored-by: lmzheng <lmzheng@fb.com>
Co-authored-by: michael604work <michael604@meta.com>
|
2026-07-14 20:06:50 -07:00 |
|
Lianmin Zheng
|
b4fdce3b63
|
Fix post-capture KV sizing for SWA pools (#31092)
|
2026-07-14 20:06:15 -07:00 |
|
Polisetty V R K Jyothendra Varma
|
532cd337ed
|
[Intel GPU] DeepSeek V4 12/N: use sgl-kernel implementation of silu_and_mul_clamp to run on XPU (#28428)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
|
2026-07-15 11:03:53 +08:00 |
|
 Polisetty V R K Jyothendra VarmaandMa Mingfei
|
46b675ce70
|
[Intel GPU] DeepSeek V4 11/N: support fp8_paged_mqa_logits_triton from sgl-kernel to run on XPU (#28059)
Signed-off-by: P V R K Jyothendra Varma <polisettyvarma@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-07-15 11:03:35 +08:00 |
|
Mick
|
43124cdd90
|
fix: fix image benchmark backend parity (#30867)
|
2026-07-15 10:11:22 +08:00 |
|
Liangsheng Yin
|
76dc427806
|
[Spec] Single-source num_tokens_per_req derivation and access (#31013)
|
2026-07-14 18:41:08 -07:00 |
|
Liangsheng Yin
|
ca0ee3f1a8
|
[Spec] Consolidate spec-worker weight updates into BaseSpecWorker via draft_runners (#31078)
|
2026-07-14 18:39:46 -07:00 |
|
 weireweireandweireweire
|
a9cf5e68e6
|
[DSV4] Remove per-step seqlen D2H from speculative to make overlap scheduler work (#30365)
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
|
2026-07-14 18:01:22 -07:00 |
|
WenhaoZhang
|
90f10cbe26
|
[diffusion] post_training: Add LoRA IPC weight sync via lora_merge mode (#31029)
|
2026-07-15 08:42:54 +08:00 |
|
Liangsheng Yin
|
50d1edaa7f
|
[misc] Move SchedulerRecvSkipper into scheduler_components (#31222)
|
2026-07-14 16:35:04 -07:00 |
|
paulzhang-tm
|
463a3f4248
|
[Mamba] Support configurable conv-window layouts (#31059)
|
2026-07-14 14:41:10 -07:00 |
|