JoyFuture
|
a309f1f8f4
|
fix(cuda_graph): zero out_cache_loc_swa on pad and use int32 (hybrid-SWA accuracy fix) (#24743)
|
2026-05-09 18:22:12 +08:00 |
|
Liangsheng Yin
|
ba625d5290
|
slash command rerun UX: emoji semantics + result writeback (#24802)
|
2026-05-09 03:19:24 -07:00 |
|
 Brayden Zhongandb8zhong
|
f4b7e73699
|
Enable trtllm-gen BF16 MoE for MTP (#24260)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
|
2026-05-09 03:14:17 -07:00 |
|
sglang-npu-bot
|
f1a9a455e0
|
Revert "[NPU] fix profiler on npu" (#24815)
|
2026-05-09 17:53:02 +08:00 |
|
zhaozx-cn
|
e2527df8b6
|
[NPU] fix profiler on npu (#24685)
Signed-off-by: zhaozx-cn <zhaozx2116@163.com>
|
2026-05-09 17:48:24 +08:00 |
|
Jia Guo
|
fd636410a2
|
Restrict fa_skip_kv_cache to non-MLA backends (#24097)
|
2026-05-09 09:25:02 +00:00 |
|
 Brayden Zhongandb8zhong
|
8f33bee31b
|
Reland Cute-DSL FP4 dense GEMM (#23590)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
|
2026-05-09 02:20:58 -07:00 |
|
Yuxuan Zhang
|
d49fc092cb
|
[Bug Fix] GLM-5.1: drop constexpr on page_indice_batch_offset, skip offloader post_init on draft worker, support N=32 in copy_to_gpu_no_ce (#23550)
|
2026-05-09 15:43:45 +08:00 |
|
 shuwennandClaude Opus 4.7
|
9d12f9e6fa
|
[HiCache] ci: lower est_time for test_hicache_spec_file_storage (#24713)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-09 00:33:18 -07:00 |
|
Liangsheng Yin
|
78da0d3106
|
[Spec] Move accept_tokens off EagleDraftInput; pass via method arg (#24735)
|
2026-05-08 23:24:18 -07:00 |
|
Khoa Pham
|
1610aa77ab
|
Reduce gemma4 moe deterministic test runtime (#24754)
|
2026-05-08 20:46:56 -07:00 |
|
Chi McIsaac
|
8e534e8f15
|
[diffusion] fix: fix diffusers executor crash when component residency manager is absent (#24573)
|
2026-05-09 11:45:06 +08:00 |
|
Liangsheng Yin
|
44a527f6f4
|
fix patch_torch test queue race (#24739)
|
2026-05-08 20:25:59 -07:00 |
|
storyicon
|
590b13b513
|
[diffusion] fix: fix NCCL deadlock in ulysses sp when sequence length has remainder (#24694)
Signed-off-by: storyicon <storyicon@foxmail.com>
|
2026-05-09 11:05:37 +08:00 |
|
 Polisetty V R K Jyothendra VarmaandMa Mingfei
|
50ed01674e
|
fix is_arch_support_pdl function usage (#24600)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-05-09 09:39:34 +08:00 |
|
Liangsheng Yin
|
1613bae412
|
[Spec] Disambiguate verified_id into bonus_token(s) / accept_tokens (#24724)
|
2026-05-08 18:24:33 -07:00 |
|
 Yuan Luoandluoyuan.luo
|
a61a14f416
|
[KDA] Optimize prefill kernels with diagonal and recompute fuse (#24271)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-05-09 08:52:51 +08:00 |
|
 Brayden Zhongandb8zhong
|
9ee830346f
|
Disable Custom AR V2 when in multi-node (#24729)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
|
2026-05-08 17:50:05 -07:00 |
|
 
|
d1c5937428
|
env: add SGLANG_RADIX_FORCE_MISS to force radix prefix-cache miss (#24726)
Co-authored-by: sihan-zzz <228612289+sihan-zzz@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-08 17:46:38 -07:00 |
|
YAMY
|
560829a171
|
feat(scheduler): add adaptive queue-based prefill delayer trigger (#23189)
|
2026-05-08 16:54:30 -07:00 |
|
YAMY
|
6971a03fe6
|
fix(fa3): skip scheduler_metadata precompute under DP attention (#24632)
|
2026-05-08 16:19:20 -07:00 |
|
Niko Ma
|
62c2e091f6
|
[PD] MORI-IO: Add state transfer, inline transfer model, and high-concurrency fixes (#22665)
|
2026-05-08 16:07:22 -07:00 |
|
Michael
|
190b15c8fe
|
[AMD] Register 8 CPU-bound unit tests for AMD 1-GPU PR CI (#24569)
|
2026-05-08 16:01:58 -07:00 |
|
Alison Shao
|
5fbec0e445
|
ci: prune per-commit CUDA tests — move 25 files + 13 testcases to test/manual/ (#24721)
|
2026-05-08 15:53:23 -07:00 |
|
Alison Shao
|
aefd8e257f
|
Re-land #23109: rebase-required mode + fix for grep-no-match abort (#24180)
|
2026-05-08 15:28:57 -07:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Jimmy Shongandgemini-code-assist[bot]
|
fa8985486e
|
[test/fix]: isolate VLM MMMU eval output dirs to fix nightly-4-gpu cross-test pollution (#24623)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-05-08 15:01:53 -07:00 |
|
Liangsheng Yin
|
5dc4c7bef1
|
Add speculative decoding naming convention rule (#24094)
|
2026-05-08 14:52:31 -07:00 |
|
Jimmy Shong
|
096ad02b06
|
[Model] Laguna-XS.2 Model Support (#24204)
|
2026-05-09 05:43:13 +08:00 |
|
Cheng Wan
|
7b707c9222
|
disable the combination of --enable-two-batch-overlap and --enforce-s… (#24720)
|
2026-05-08 14:27:35 -07:00 |
|
Yuhao Yang
|
09912fd89d
|
Remove unnecessary bf16 assert in rotate_activation (#24686)
|
2026-05-09 05:00:52 +08:00 |
|
Yilong Zhao
|
f30d1d0b0a
|
logits: remove blocking H2D copy (#24627)
|
2026-05-08 13:22:13 -07:00 |
|
Ethan Feng
|
672f778512
|
[NemotronH] Fix expert scale weight loading (#24434)
|
2026-05-08 12:37:06 -07:00 |
|
 zhongdaor-nvandzhongdaor-nv
|
2cf1a4ab38
|
feat: Add KV events for Mamba radix cache (#23678)
Signed-off-by: zhongdaor-nv <220807034+zhongdaor-nv@users.noreply.github.com>
Co-authored-by: zhongdaor-nv <220807034+zhongdaor-nv@users.noreply.github.com>
|
2026-05-08 11:53:36 -07:00 |
|
 Xu Zouandxz-keg
|
ca7a8cc61d
|
[Bugfix] Fix a bug causing NVFP4 to be tested on all gpus like SM90 devices. (#24604)
Co-authored-by: xz-keg <xuzou_keg@outlook.com>
|
2026-05-08 11:51:30 -07:00 |
|
Lianmin Zheng
|
e40e339c72
|
Filter non-int token ids in benchmark and observe decode-side bootstrap/alloc metrics (#24684)
|
2026-05-08 11:45:37 -07:00 |
|
Mick
|
73b8eda103
|
[diffusion] fix: fix FA3 varlen out argument handling (#24688)
|
2026-05-08 19:01:49 +08:00 |
|
Mick
|
17888fa92a
|
[diffusion] doc: update ltx2 multi-gpu deployment guide (#24682)
|
2026-05-08 18:38:05 +08:00 |
|
 fanxingranandfanxingran
|
7f8e7a9130
|
fix(aiter): drop FP8 KV upcast; use native FP8 path in paged_attentio… (#24129)
Co-authored-by: fanxingran <fanxingran@amd.com>
|
2026-05-08 02:47:48 -07:00 |
|
jacky.cheng
|
f21d4868dc
|
[AMD] Replace naive triton RMSNorm with aiter RMSNorm for diffusion model (#24360)
|
2026-05-08 02:44:13 -07:00 |
|
YC Yen-Ching Tseng
|
e1150f66db
|
[AMD][diffusion] Temporal-unfolded batched Conv2D for ROCm VAE decode (#22971)
|
2026-05-08 02:32:14 -07:00 |
|
amote-i
|
d32e283947
|
[NPU] [DOC] refresh npu supported model list (#24676)
|
2026-05-08 17:08:15 +08:00 |
|
 Brayden Zhongandb8zhong
|
80d0226b68
|
Turn on JIT custom AR implementation by default (#24363)
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
|
2026-05-08 02:05:31 -07:00 |
|
HAI
|
73792629d4
|
[AMD] Intro SGLANG_DIFFUSION_AITER_FP8_ATTN (#24677)
|
2026-05-08 01:31:00 -07:00 |
|
jacky.cheng
|
76a1f169b3
|
[AMD] Add AMD FP8 MLA attention test for Wan2.2-T2V-A14B (#23955)
|
2026-05-08 01:03:51 -07:00 |
|
jacky.cheng
|
b22d3cd606
|
[AMD] Support fp8 MLA for diffusion model (#20319)
|
2026-05-08 00:56:24 -07:00 |
|
Thomas Wang
|
19afe73e03
|
[AMD] Cherry-pick aiter commit for mhc_pre fix (#24665)
|
2026-05-07 23:49:39 -07:00 |
|
Yibo Cai
|
55d8223c2b
|
[sgl-kernel/cpu] support w8a8 int8 model for arm cpu (#16045)
skip gpu test as this one is not related to gpu backend.
|
2026-05-08 14:47:06 +08:00 |
|
amote-i
|
47e9ec11ad
|
[NPU] [DOC] fix ascend_npu_support_new_models TOC (#24658)
|
2026-05-08 14:07:00 +08:00 |
|
JoyFuture
|
e1bc001872
|
fix(mimo_v2): auto-disable multimodal when vision/audio configs are absent (#24652)
|
2026-05-08 13:40:08 +08:00 |
|
 maocheng23andlawrence-harmonic
|
7deed98e1b
|
[fix] /pause_generation and /continue_generation wrong for --tokenizer-worker-num > 1 (#24462)
Co-authored-by: lawrence-harmonic <185285563+lawrence-harmonic@users.noreply.github.com>
|
2026-05-07 21:32:21 -07:00 |
|