This website requires JavaScript.
6275114548
[NPU] [DOC] Update model names supported on Ascend NPU (#31316 )
amote-i
2026-07-16 14:45:59 +08:00
e73f323464
[JIT] Reduce MoE fused gate CI test sweep (#31400 )
Xiaoyu Zhang
2026-07-16 14:42:45 +08:00
8c5e0cee18
[AMD] Bump MoRI to f7e6ac6 to fix ROCm install_dependency build break (#31406 )
YC Yen-Ching Tseng
2026-07-16 14:40:18 +08:00
0a64139c94
Fix --moe-a2a-backend silently ignored for LongCat-2.0 (moe_topk missing from gate) (#30975 )
2026-07-16 14:40:01 +08:00
5af65d8542
[DeepSeek-V4] Support BF16 Compress State for Online C128 (#29609 )
Ryan Zzz and zhujunyu
2026-07-16 14:17:29 +08:00
dc60f65661
chore: bump tokenspeed_mla to 0.1.8 (#31385 )
Khoa Pham and Claude Fable 5
2026-07-15 22:37:54 -07:00
ac4fa6496c
Skip no-op EAGLE sampling renormalization (#31294 )
weireweire and weireweire
2026-07-16 13:37:10 +08:00
40517b593b
[docs] Inkling cookbook: LoRA cells require --disable-prefill-cuda-graph (#31418 )
Yanbin Jiang
2026-07-15 22:19:49 -07:00
0b372d03de
[CI] Guard partition consumers against failed check-changes and degenerate fits (#31416 )
Liangsheng Yin
2026-07-15 22:09:40 -07:00
238448f4c8
ci: fix runner utilization report undercounting busy time ~25x (#31396 )
Alison Shao
2026-07-15 22:02:38 -07:00
b871a509e6
[Fix] Wire the detokenizer soft watchdog into the multi-http-worker event loop (#31392 )
Liangsheng Yin
2026-07-15 21:55:46 -07:00
bc525dcf90
[Cookbook][CPU]Update CPU model support info in Cookbook (#30520 )
Zaili Wang and zijiexia
2026-07-16 11:42:09 +08:00
dee91c51cf
perf(deepseek_v4): enable SGLANG_OPT_FP8_WO_A_GEMM on sm90 (Hopper) (#28983 )
guzekai01
2026-07-16 11:16:32 +08:00
7f9a902cd9
fix(humming): handle missing quant_method (#31185 )
guzekai01
2026-07-16 11:15:52 +08:00
22453ca63c
docker: build HPC-Ops into the GPU image (#31390 )
Xiaoyu Zhang and Claude Fable 5
2026-07-16 11:07:01 +08:00
f28ce5c420
[XPU]REPO cache dtype xpu align with cuda (#31140 )
Swift.Sun
2026-07-16 10:30:01 +08:00
3a8ddd3fd0
[XPU] Route topk_sigmoid and topk_softmax to AOT sgl-kernel-xpu symbols (#31038 )
2026-07-16 07:59:19 +05:30
14095ef78a
[AMD] Disable CUDA IPC multimodal transport on ROCm in MMMU VLM tests (#31342 )
YC Yen-Ching Tseng
2026-07-16 09:42:30 +08:00
d9003dd452
fix: skip unsafe automatic prefill graph capture (#31204 )
Mick
2026-07-16 09:27:38 +08:00
1c4892d7bb
[Mamba] Fix spec-v2 + extra_buffer crash (guard None mamba_next_track_idx) (#27998 )
Serge Panev
2026-07-15 18:23:14 -07:00
7647a9d260
Fuse the preprocess kernels of trtllm-gen attention (#29690 )
Brayden Zhong and Brayden Zhong
2026-07-15 18:21:34 -07:00
871c648203
[NPU]revert add scoring func for GLM 4.7 Flash (#31388 )
McZyWu
2026-07-16 09:16:57 +08:00
edb2059139
Support Flashinfer one-sided A2A + CuteDSL MoE for Nemotron Ultra (#28309 )
Brayden Zhong and Brayden Zhong
2026-07-15 18:14:20 -07:00
ac23be8d09
Skip MXFP8 autotune on dense GEMM, which causes IMA (#29669 )
Brayden Zhong and Brayden Zhong
2026-07-15 18:06:55 -07:00
0b04e9da83
Use fused A GEMM for fc1_latent_proj in NemotronH (#29692 )
Brayden Zhong and Brayden Zhong
2026-07-15 18:05:22 -07:00
1cc9493747
[Spec] Converge DP-attention spec width scaling onto num_tokens_per_req (#31244 )
Liangsheng Yin
2026-07-15 17:57:08 -07:00
5d004a20c5
Fix FlashInfer A2A top-k ID dtype (#29929 )
Po-Han Huang (NVIDIA)
2026-07-16 08:56:11 +08:00
34f5691ea1
docs: sync LMSYS SGLang blog cards (#31386 )
sglang-bot and sglang-bot
2026-07-15 17:53:19 -07:00
b0b2dfbda1
[Spec] Extract the shared draft() tail into build_eagle_verify_input (#31375 )
Liangsheng Yin
2026-07-15 15:59:31 -07:00
7a973c03a0
[Bugfix] Stamp capture-time num_tokens_per_req in multi-layer EAGLE; close jit_kernel CI filter gaps (#31367 )
2026-07-15 15:24:32 -07:00
3101c1258c
[DSv4] Use BF16 instead of FP32 for indexer score computation (#30012 )
Lewis and 百麒
2026-07-16 06:06:55 +08:00
26cb0fcdda
Empty _REQ_TYPES_WITH_OPAQUE_FIELDS on the msgpack IPC path (#29465 Task 4) (#30182 )
Jorge António
2026-07-15 22:55:06 +01:00
67148447a6
[AMD] Register 3 CPU-bound / triton unit + light-integration tests for AMD 1-GPU PR CI (#31088 )
Michael
2026-07-15 14:36:05 -07:00
ec32590025
feat(moriep): add fp4 combine dtype (SGLANG_MORI_COMBINE_DTYPE=fp4) (#30706 )
karverma-amd
2026-07-15 16:21:41 -05:00
e78051a419
[AMD] [Fix] Fix --attention-backend triton work for DeepSeek MLA on MI355 (null-K + decode dispatch + RoPE) (#30355 )
jacky.cheng
2026-07-16 05:19:23 +08:00
e76cc75cfa
[CI] Remove nightly registrations redundant with scheduled stage runs (#31371 )
Liangsheng Yin
2026-07-15 14:12:07 -07:00
5abec3fbf8
Fix MiMo-V2 on Blackwell: FA3 fallback and TP-aware audio weight loading (#31343 )
Yuhao Yang
2026-07-16 04:26:27 +08:00
18043aec20
[CI] Fix TRTLLM MHA graph metadata test fixture (#31332 )
Mohammad Miadh Angkad
2026-07-16 03:44:51 +08:00
ab627e5d75
fix: load the right mtp lm head quantization (#30976 )
Shaun Kotek
2026-07-15 22:12:04 +03:00
7e7129acd7
[Bugfix] Release Mamba cache after PP dynamic chunk profiling (#31321 )
Xuwei
2026-07-16 02:32:07 +08:00
2d00e20a52
[Disagg][Qwen3.5] Fix heterogeneous attn-TP scatter transfer: GDN conv sub-block slice + GQA replicated-KV head map (#30997 )
YAMY and Xuwei Li
2026-07-15 11:31:37 -07:00
dd2e4cdc99
Add Inkling cookbook (#31360 )
Yuhao Yang
2026-07-16 02:24:31 +08:00
9d147fdca1
[Multimodal] Support n>1 outputs for GLM-Image generation (#31027 )
AuFlow and AuFlow
2026-07-16 01:55:19 +08:00
dc078ddd2a
[Spec] Extract stateless draft prepare helpers into eagle_worker_common (#31257 )
Liangsheng Yin
2026-07-15 10:54:31 -07:00
d36e96ce23
[AMD] Enable mamba-extra-buffer for Qwen3.5 on ROCm (#30359 )
Bingxu Chen and ntgiang71096
2026-07-16 01:18:12 +08:00
a8b60433c2
[AMD] Fix DSV4 JIT build on rocm (#31131 )
YC Yen-Ching Tseng and kangwangamd
2026-07-16 00:58:12 +08:00
c879f3da5c
[diffusion] rl: support rl rollout for the wan pipeline via a per-request scheduler switch (#30036 )
Andy Ye and Claude Fable 5
2026-07-15 07:30:51 -07:00
d2b1243be0
docs: document CUDA crash dump output (#31333 )
Jun Liu and Xinyuan Tong
2026-07-15 23:18:36 +09:00
495ae9aaa6
Fix Ministral3 accuracy issue by aligning YaRN RoPE scaling with Transformers implementation (#31232 )
2026-07-15 15:23:04 +03:00
8ed82afcc8
[MoE Refactor] [NPU] Refactor Ascend MoE implementation to reduce code duplication and align with community design (#25663 )
Артем Савкин and ronnie_zheng
2026-07-15 14:59:42 +03:00
c9b17403e7
Fix image URL response for multiple outputs (#30621 )
2026-07-15 19:49:04 +08:00
947a14d617
feat: unify multimodal feature transport (#30904 )
Mick
2026-07-15 17:42:38 +08:00
f2c875d1c8
[PD] Route PD server warmup to every DP rank (#30748 )
weireweire and weireweire
2026-07-15 15:59:41 +08:00
5af670284e
[CI] Lower GLM-5.2 NVFP4 MTP speed threshold (#31289 )
Mohammad Miadh Angkad
2026-07-15 15:39:58 +08:00
fbcbe0a986
cookbook(deepseek-v4): add MORI disagg backend for AMD + bump MI355X image (#30651 )
2026-07-15 08:33:05 +01:00
241937af87
[NPU] Determine the topk norm_type through scoring_func (#31107 )
2026-07-15 15:26:44 +08:00
dec0836302
Fix processor config loading for object-storage model paths (#31211 )
Sam Shleifer and Alex Nails
2026-07-15 00:21:31 -07:00
980acd6eca
Fix MoE TP allreduce to use NCCL symmetric memory via in-pool output allocation (#29007 )
sky and Brayden Zhong
2026-07-15 15:06:37 +08:00
41e0b4b369
[CPU] add fused input proj for qwen3.5 (#31171 )
Ma Mingfei
2026-07-15 15:06:24 +08:00
a649b5a9db
[KDA] Add FlashInfer SM100 KDA decode + MTP (target_verify) backend (#30113 )
Yuan Luo and luoyuan.luo
2026-07-15 15:04:20 +08:00
1afab30577
Fix bookkeeping fields not encapsulated with real allocations in normal alloc, PD pre-alloc, DFlash and EAGLE (#29432 )
fzyzcjy
2026-07-15 14:52:21 +08:00
e789ca24a7
Lightweight extract allocation logic from mem_cache/common.py to more clearly show nearly parallel variants (#29431 )
fzyzcjy
2026-07-15 14:49:37 +08:00
c315df49bb
Fix abusing presence of req.req_pool_idx to indicate the presence of req.kv resources (#29430 )
fzyzcjy
2026-07-15 14:48:39 +08:00
2d979f1d8c
Let the presence of req.kv indicate the existence of owned kv resources (#29429 )
fzyzcjy
2026-07-15 14:47:16 +08:00
27256aee5b
Let cache backend do not couple with owned committed kv details and avoid kv_committed_freed/kv_overallocated_freed fields (#29428 )
fzyzcjy
2026-07-15 14:43:10 +08:00
d8d76c4d12
Introduce req.kv container for coupled owned kv field lifecycle (#29427 )
fzyzcjy
2026-07-15 14:40:38 +08:00
201ddeaba1
Avoid relaying per-step outputs through ScheduleBatch fields in disagg prefill and PP (#30677 )
fzyzcjy
2026-07-15 14:33:37 +08:00
01343d2759
Avoid implicit running_batch access in dllm and pdmux scheduling (#30676 )
fzyzcjy
2026-07-15 14:33:00 +08:00
21c62b9830
Rewrite pause_generation retract path as req-level release and requeue for clarity (#30675 )
fzyzcjy
2026-07-15 14:32:12 +08:00
1967b9ec99
Fix missed hisparse release and stale field cleanup in pause retract (#30674 )
fzyzcjy
2026-07-15 14:31:11 +08:00
b6cc897fea
Fix non-existent abort mode in Scheduler.pause_generation and inline retract_all (#30673 )
fzyzcjy
2026-07-15 14:27:48 +08:00
52a88fb212
Avoid mutating ScheduleBatch fields in place (#30672 )
fzyzcjy
2026-07-15 14:27:08 +08:00
e77d95c3d5
Pass per-forward overrides to ForwardBatch.init_new as explicit arguments (#30670 )
fzyzcjy
2026-07-15 14:25:59 +08:00
861d97d24d
Remove dead ScheduleBatch fields and avoid inplace seq_lens bump (#30669 )
fzyzcjy
2026-07-15 14:23:31 +08:00
a3194d3585
[AMD] Remove ROCm page_first+kernel -> layer_first HiCache fallback (follow-up to #28534 ) (#30622 )
AMD-yanfeiwang
2026-07-15 13:48:16 +08:00
0832d856ca
[Bugfix] fix quickreduce acc error in cudagraph mode (#29508 )
haoyangli0109
2026-07-15 13:16:02 +08:00
4aadf94146
[Kernel] Relocate vendored fla and mamba kernel trees to sglang.kernels (RFC #29630 , Phase 2.5, 7/7) (#30795 )
Xiaoyu Zhang and Claude Fable 5
2026-07-15 12:52:15 +08:00
23f2b77d82
Make UTs compatible for XPU (#27106 )
ANSHUMAN TRIPATHY
2026-07-15 10:05:56 +05:30
c00131ebaa
[Kernel] Migrate linear-attention, MiniMax-sparse and diffusion kernels to sglang.kernels (RFC #29630 , Phase 2.5, 6/7) (#30793 )
Xiaoyu Zhang and Claude Fable 5
2026-07-15 11:21:36 +08:00
ba5be86d42
[Kernel] Migrate DSA + DSV4 attention kernels to sglang.kernels (RFC #29630 , Phase 2.5, 5/7) (#30792 )
Xiaoyu Zhang and Claude Fable 5
2026-07-15 11:11:22 +08:00
4ae9cc3c81
Fix gate stride for 4D decode layouts (#31231 )
2026-07-14 20:06:50 -07:00
b4fdce3b63
Fix post-capture KV sizing for SWA pools (#31092 )
Lianmin Zheng
2026-07-14 20:06:15 -07:00
532cd337ed
[Intel GPU] DeepSeek V4 12/N: use sgl-kernel implementation of silu_and_mul_clamp to run on XPU (#28428 )
Polisetty V R K Jyothendra Varma
2026-07-15 08:33:53 +05:30
46b675ce70
[Intel GPU] DeepSeek V4 11/N: support fp8_paged_mqa_logits_triton from sgl-kernel to run on XPU (#28059 )
Polisetty V R K Jyothendra Varma and Ma Mingfei
2026-07-15 08:33:35 +05:30
aafa706f8f
[AMD] Update qwen3.5 cookbook (#31258 )
Thomas Wang
2026-07-15 10:59:12 +08:00
43124cdd90
fix: fix image benchmark backend parity (#30867 )
Mick
2026-07-15 10:11:22 +08:00
b22f20b660
[CI] Fix CUDA 12 NVIDIA wheel cleanup (#31035 )
Hank Han
2026-07-15 10:06:11 +08:00
76dc427806
[Spec] Single-source num_tokens_per_req derivation and access (#31013 )
Liangsheng Yin
2026-07-14 18:41:08 -07:00
ca0ee3f1a8
[Spec] Consolidate spec-worker weight updates into BaseSpecWorker via draft_runners (#31078 )
Liangsheng Yin
2026-07-14 18:39:46 -07:00
a9cf5e68e6
[DSV4] Remove per-step seqlen D2H from speculative to make overlap scheduler work (#30365 )
weireweire and weireweire
2026-07-15 09:01:22 +08:00
b8a00e2ec8
docs: sync LMSYS SGLang blog cards (#31242 )
sglang-bot and sglang-bot
2026-07-14 17:49:30 -07:00
90f10cbe26
[diffusion] post_training: Add LoRA IPC weight sync via lora_merge mode (#31029 )
WenhaoZhang
2026-07-15 08:42:54 +08:00
50d1edaa7f
[misc] Move SchedulerRecvSkipper into scheduler_components (#31222 )
Liangsheng Yin
2026-07-14 16:35:04 -07:00
21a6d08557
ci: strip invisible Unicode format chars from slash-command input (#31234 )
Alison Shao
2026-07-14 16:33:45 -07:00
771e386332
Disable flaky DSV4-Flash FP4 BCG determinism test (nondeterminism from #30898 idle-rank dummy extend) (#31125 )
Alison Shao
2026-07-14 16:21:10 -07:00
463a3f4248
[Mamba] Support configurable conv-window layouts (#31059 )
paulzhang-tm
2026-07-14 14:41:10 -07:00
08c46e1f1a
Add dummy forward batch preparation hook (#31070 )
paulzhang-tm
2026-07-14 14:30:31 -07:00
0d89564d27
Support scheduler_recv_interval (recv skipper) under DP-attention (#30457 )
2026-07-14 16:02:58 -05:00
bdc9848c25
[Doc]Standardize the names of PyTorch NPU-related software throughout the documentation by replacing them all with TorchNPU. (#29886 )
2026-07-15 02:56:08 +08:00
cb47a68717
[PD] Stride KV token->page indices on device before D2H copy (#31173 )
cctry and cctry
2026-07-14 10:08:30 -07:00