Xinyuan Tong
|
db12bfcdc8
|
[JIT] Add kpool_topk_transform JIT kernel (#28670)
|
2026-06-22 01:04:21 -07:00 |
|
Liangsheng Yin
|
64e455d4bf
|
Fix lint break on main (#28886)
|
2026-06-21 22:46:08 -07:00 |
|
Bingxu Chen
|
e2540188ce
|
[AMD] Clean up DeepSeek-R1-MXFP4 TP2/TP4 MLA GSM8K tests (#27243)
|
2026-06-21 21:41:19 -07:00 |
|
Lianmin Zheng
|
886b96621d
|
Migrate more server args to annotated style (#28830)
|
2026-06-21 20:50:17 -07:00 |
|
 cctryandcctry
|
0c065671c9
|
[Spec] Redo: split init_backends; account draft weights in --mem-fraction-static (#28855)
Co-authored-by: cctry <cctry@fb.com>
|
2026-06-21 20:45:16 -07:00 |
|
Bingxu Chen
|
fd7874d11b
|
[AMD] Register DP attention test (#28495)
|
2026-06-21 20:21:41 -07:00 |
|
Liangsheng Yin
|
e6722c751b
|
[Feature] Add graceful scheduler shutdown; free hisparse host buffer on exit (#28779)
|
2026-06-21 15:08:10 -07:00 |
|
 Lianmin Zhengandhnyls2002
|
a4d0ff3def
|
[misc] Make NaN-logit sanitization opt-in (default off) (#28829)
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
|
2026-06-21 14:31:35 -07:00 |
|
cctry
|
6d4ca9bc54
|
Cap SWA pool sizing with chunk cache (#28755)
|
2026-06-21 01:06:59 -07:00 |
|
 Jairo David Campaña RoseroandXinyuan Tong
|
b4dda8b3ce
|
fix(anthropic): handle mid-conversation system messages (#26773)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-06-21 04:34:47 +00:00 |
|
karverma-amd
|
2552b860a3
|
[AMD][bugfix] Place TBO cuda-graph num_token_non_padded buffer on model devices (#28337)
|
2026-06-20 18:06:22 -07:00 |
|
    
|
f42ec350b4
|
[mtp] add rejection sampling for speculative decoding (#26312)
Co-authored-by: lyc508653 <lyc508653@alibaba-inc.com>
Co-authored-by: Qiaolin-Yu <liin1211@outlook.com>
Co-authored-by: Huiqiang Jiang <30883354+iofu728@users.noreply.github.com>
Co-authored-by: Yi Zhang <25844240+yizhang2077@users.noreply.github.com>
Co-authored-by: Yizhong Cao <114661107+cao1zhg@users.noreply.github.com>
|
2026-06-20 15:10:42 -07:00 |
|
Lianmin Zheng
|
95fb1ef697
|
[CI] Remove deprecated test/srt legacy CI setup (#28810)
|
2026-06-20 15:09:33 -07:00 |
|
 Rita BrugarolasandClaude Opus 4.6
|
d6d06cdc17
|
[AMD] Fix no-op dtype cast in _topk_ids_logical_to_physical_dynamic on HIP (#28074)
Signed-off-by: Rita Brugarolas Brufau <rita.brugarolasbrufau@amd.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-06-20 11:07:22 -07:00 |
|
shuwenn
|
ff1fc1fbdf
|
[mem_cache][5/N] refactor: extract host KV cache base layer into pool_host package (#27273)
|
2026-06-20 20:44:08 +08:00 |
|
 Zhiyao JiangandXinyu Jiang
|
1115373668
|
[AMD] Fix garbled unquantized Qwen3-30B-A3B output on ROCm/aiter where the aiter CK fused-MoE falls back to Triton with pre-shuffled weights (#28244)
Co-authored-by: Xinyu Jiang <xinyuj2@andrew.cmu.edu>
|
2026-06-20 01:25:01 -07:00 |
|
 Lianmin ZhengandYinghai Lu
|
45d203fb08
|
Fix tokenizer state cleanup on dispatch failure (#28694)
Co-authored-by: Yinghai Lu <yinghai@meta.com>
|
2026-06-19 21:55:39 -07:00 |
|
Bi Xue
|
28e2096d1c
|
[sgl] wire SGLANG_OPT_SWA_RELEASE_LEAF_LOCK_AFTER_WINDOW on Unified Cache. (#28161)
|
2026-06-20 12:06:52 +08:00 |
|
Michael
|
653f6735a0
|
[AMD] add nightly kv_canary JIT benchmark suite (#28611)
|
2026-06-19 19:46:34 -07:00 |
|
 Jae B.andR0CKSTAR
|
2cbe1e6404
|
[Apple Silicon] [MLX] Fix MlxModelRunnerStub.initialize() signature desync with base (#28660)
Co-authored-by: R0CKSTAR <yeahdongcn@gmail.com>
|
2026-06-19 18:05:22 -07:00 |
|
Michael
|
871ed0dc0c
|
Revert "ci: add 4-GPU mi35x runner and rebalance off the saturated 8-GPU pool" (#28751)
|
2026-06-19 16:23:18 -07:00 |
|
Michael
|
420004827c
|
[AMD] register 3 tests to stage-b-test-1-gpu-large-amd (batch-6) (#28736)
|
2026-06-19 15:45:16 -07:00 |
|
 Michaelandmichaelzhang-ai
|
13aab2fc06
|
ci: add 4-GPU mi35x runner and rebalance off the saturated 8-GPU pool (#28745)
Co-authored-by: michaelzhang-ai <michaelzhang-ai@users.noreply.github.com>
|
2026-06-19 15:41:19 -07:00 |
|
  
|
3ed46f599f
|
[core] Don't force seq_lens_cpu publication under piecewise CUDA graph (#28633)
Co-authored-by: jonnykong <jonnykong@fb.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
|
2026-06-19 15:12:07 -07:00 |
|
Liangsheng Yin
|
d271de64fe
|
[misc] Move bench_one_batch_server into sglang/benchmark/ with a back-compat shim (#28625)
|
2026-06-19 14:19:29 -07:00 |
|
 
|
ab0714d0ee
|
ci: run GB300 nightly suite in the standard Nvidia nightly workflow (#28536)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Mohammad Miadh Angkad <176301910+mmangkad@users.noreply.github.com>
|
2026-06-19 12:24:38 -07:00 |
|
Michael
|
7c505c2927
|
[AMD] fix(jit): port kv_canary write/verify/plan kernels to ROCm (#28357)
|
2026-06-19 12:19:37 -07:00 |
|
 
|
c436a8161a
|
[AMD] Enable HiSparse on ROCm (#26639)
Co-authored-by: clintg6 <7388379+clintg6@users.noreply.github.com>
Co-authored-by: HAI <hixiao@gmail.com>
|
2026-06-19 11:59:45 -07:00 |
|
Jaybe
|
ca88b7f1d2
|
fix: remove manual rope parameters injection in PretrainedConfig (#23910)
|
2026-06-19 17:41:51 +00:00 |
|
Vladislav Nosivskoy
|
2ad9a5b576
|
[UnifiedTree] Use dense model for HiCache+CP KL tests (#28726)
|
2026-06-20 00:18:01 +08:00 |
|
Oguz Ulgen
|
3af991fb3e
|
[AMD] Make breakable CUDA graph run on ROCm/HIP (#28173)
|
2026-06-19 07:16:00 -07:00 |
|
Liangsheng Yin
|
9bb9d17e1a
|
[Spec] Unify speculative grammar token-accept path in decode processing (#28682)
|
2026-06-19 02:49:40 -07:00 |
|
  
|
4d94e9471a
|
[AMD] Relax allreduce-fusion residual accuracy tolerance to 1 bf16 ULP (#28226)
Co-authored-by: kangwangamd <kangwangamd@users.noreply.github.com>
Co-authored-by: Bingxu Chen <bingxche@amd.com>
Co-authored-by: Hubert Lu <55214931+hubertlu-tw@users.noreply.github.com>
|
2026-06-18 19:18:53 -07:00 |
|
giang_ng_tr
|
b36360dc5b
|
[AMD][Perf] Split-KV flash-decode attention for EAGLE target-verify (Triton backend) (#27382)
|
2026-06-18 19:11:10 -07:00 |
|
 huangtingweiandhzh0425
|
6b7ecca663
|
[HiCache]Support hybrid pool staged H2D kernel (#28434)
Co-authored-by: hzh0425 <hzh0425@apache.org>
|
2026-06-19 09:48:03 +08:00 |
|
Rain Jiang
|
ef01618dfb
|
support MPServer and embedded server for granian to enable muti tokenizer worker (#28573)
|
2026-06-18 17:59:12 -07:00 |
|
Michael
|
62ab09a478
|
[AMD] register 2 spec tests to stage-b-test-1-gpu-large-amd (batch-5) (#28558)
|
2026-06-18 16:33:23 -07:00 |
|
giang_ng_tr
|
c1067f88d6
|
[AMD][Perf] Tune extend attention block sizes for gfx950 (head_dim > 128) (#27793)
|
2026-06-18 15:16:52 -07:00 |
|
Baizhou Zhang
|
e3026ef016
|
[3/N][CP] Implement zigzag CP strategy (#28421)
|
2026-06-18 15:10:30 -07:00 |
|
    
|
9fc9d37f6d
|
Fix spec decoding with grammar in disagg (#24082)
Co-authored-by: jimmy.shong <jimmy.shong@radixark.ai>
Co-authored-by: Jimmy Shong <69131491+Jiminator@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Codex <codex@example.com>
|
2026-06-18 14:58:32 -07:00 |
|
 
|
bea282cede
|
[DeepSeek-V4] Fuse UE8M0 scale rounding into FP8 group quantization (#26766)
Co-authored-by: liqichao <liqichao@baidu.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
|
2026-06-18 14:41:56 -07:00 |
|
Zhaoyi Li
|
27a374eaef
|
[AMD][CI] Use non-gated Qwen3-8B for MI35x disaggregation tests (#28678)
|
2026-06-18 14:29:35 -07:00 |
|
Liangsheng Yin
|
8f6d9ef9a5
|
[misc] Drop redundant req_pool_indices_cpu guards; fold hisparse into GLM-5.1 e2e (#28607)
|
2026-06-18 12:51:23 -07:00 |
|
 Zhanghengandispobock
|
867707f1f2
|
[UnifiedTree]: move some kl test from base to extra stage (#28627)
Co-authored-by: ispobock <ispobaoke@gmail.com>
|
2026-06-18 01:39:46 -07:00 |
|
 Vladislav NosivskoyandZhangheng
|
b7ae7149e8
|
[HiCache] Fix SWA L3 cache miss due to a prefetch/hit len mismatch (#27291)
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-06-18 14:55:00 +08:00 |
|
DovLin
|
0188c54fbe
|
[FA3] Add unit test for only_qv (NoPE) KV decode path (#28595)
Signed-off-by: Shijin Zhang <dovis.zhang02@gmail.com>
|
2026-06-17 23:41:44 -07:00 |
|
Michael
|
5d1949152d
|
[AMD] ci: add extra-a 1-gpu-large tier (fp8kv-triton, streaming-session, spec-standalone) (#28458)
|
2026-06-17 23:31:32 -07:00 |
|
 Michaelandmichaelzhang-ai
|
0e5a66dca4
|
[AMD] Register 3 JIT kernel unit tests for AMD CI (#27837)
Co-authored-by: michaelzhang-ai <michaelzhang@example.com>
|
2026-06-17 23:28:56 -07:00 |
|
cctry
|
7976928c57
|
Abort during chunked prefill + PD peer-liveness abort (#28086)
|
2026-06-17 23:13:28 -07:00 |
|
 Yuan Luoandluoyuan.luo
|
3340f4e3da
|
[GDN][KDA][mem_cache] int8 checkpoint pool for the linear-attn prefix cache (#28185)
Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
|
2026-06-17 20:41:46 -07:00 |
|