 
|
f8e62a9224
|
[NPU] Add PR test cases (#32392)
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
|
2026-08-02 23:00:02 +08:00 |
|
Liangsheng Yin
|
558c9bdcc2
|
[misc] Improve benchmark determinism and dataset API coverage (#33255)
|
2026-08-02 01:39:50 -07:00 |
|
Alison Shao
|
43be25b2b7
|
[CI] Graceful teardown for kv_canary and EAGLE spec fixtures (#32829)
|
2026-08-02 01:09:59 -07:00 |
|
  
|
33ecf4bcd8
|
Add pr tests (#31952)
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
Co-authored-by: Cherry_ming <136634645@qq.com>
|
2026-08-01 15:03:21 +08:00 |
|
Sugar920
|
fd96a35fb0
|
add NPU GSM8K accuracy tests for 7 models (#32649)
|
2026-08-01 14:22:57 +08:00 |
|
Cheng Wan
|
55b6769b0e
|
config: read resolved config via namespace accessors (#33013)
|
2026-07-31 15:06:59 -07:00 |
|
Rain Jiang
|
4af8ddb576
|
support rust sglang server (#29799)
|
2026-07-31 11:56:31 -07:00 |
|
 weireweireandweireweire
|
55c1963df4
|
Remove unused draft-extend CUDA graph top-k (#31430)
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
|
2026-07-30 17:45:01 -07:00 |
|
Jimmy Shong
|
ed361ae7f0
|
Fix attention backends for models with per-layer head counts (num_attention_heads_per_layer) (#32625)
|
2026-07-29 20:03:00 -07:00 |
|
YAMY
|
fddfc1fb5e
|
[GDN] Support FlashInfer GDN prefill with extra-buffer radix cache (#29735)
|
2026-07-30 00:47:35 +08:00 |
|
pllimax
|
d004a15a3e
|
Fix GLM4-7B-Flash accuracy test configuration, tune Qwen3.6-27B/35B performance test parameters, and harden Ascend NPU multi-node E2E test utilities against pod name format errors. (#32371)
|
2026-07-29 21:44:17 +08:00 |
|
YAMY
|
2428f56145
|
[Bugfix] Fix Kimi-Linear state transfer across heterogeneous TP (#32262)
|
2026-07-24 10:31:17 -07:00 |
|
ANSHUMAN TRIPATHY
|
f0f78a6c93
|
Add deterministic inference for eagle parity test (#30026)
|
2026-07-24 12:48:21 +08:00 |
|
Liangsheng Yin
|
99b29bf188
|
[Fix] Support ENCODER_ONLY target-verify in the trtllm_mha backend (#32178)
|
2026-07-23 20:33:39 -07:00 |
|
Cheng Wan
|
eac7c7d7cd
|
fix(attention): read per-runner kv cache dtype off model_runner (#32251)
|
2026-07-23 20:08:57 -07:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
2d1a7be8c4
|
[Kernel] Reclassify kernel tests by ops group + move helpers out of the package (RFC #29630) (#32128)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
2026-07-23 12:18:27 +08:00 |
|
 Mohammad Miadh AngkadandBrayden Zhong
|
0c29c8fece
|
Bump FlashInfer to 0.6.15.post1 (#31927)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-07-22 14:21:59 -07:00 |
|
 Xiaoyu ZhangandClaude Opus 4.8
|
74338e94f1
|
[Kernel] Phase 4 batch-3: migrate tangled JIT subsystems + new groups into kernels.ops (RFC #29630) (#32045)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-22 21:15:03 +08:00 |
|
 
|
3217b7e3ce
|
fix(hicache): support staged write-back for asymmetric MHA (#30981)
Co-authored-by: 晟海 <huangtingwei.htw@antgroup.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-07-22 09:58:49 +08:00 |
|
 ![gemini-code-assist[bot]](/assets/img/avatar_default.png)
|
df39a7b0b6
|
[CI][PD] Add NIXL disaggregation functional tests (#27894)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-07-21 18:17:06 +08:00 |
|
iridiumine
|
d6ef68881e
|
[NPU] Adapt MiMo-V2.5-W8A8 (#29131)
|
2026-07-21 09:16:48 +08:00 |
|
 
|
fafa302e41
|
[XPU][NIGHTLY] Add 8 XPU nightly tests, enable 1-gpu suite (#30246)
Co-authored-by: arathi-hlab <arathi-hlab@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
2026-07-20 16:15:50 +08:00 |
|
       
|
02236fa38c
|
Add Inkling model support (#31681)
Co-authored-by: Chunan Zeng <zcnrex@gmail.com>
Co-authored-by: Ke Bao <ispobaoke@gmail.com>
Co-authored-by: Yanbin Jiang <jybsuper@gmail.com>
Co-authored-by: Yuhao Yang <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: Qiaolin Yu <qiaolin.yu@radixark.ai>
Co-authored-by: Zhichen Zeng <zczeng@uw.edu>
Co-authored-by: Aurick Qiao <aurick@thinkingmachines.ai>
Co-authored-by: Joseph <jk@thinkingmachines.ai>
|
2026-07-19 22:57:37 -07:00 |
|
Yuhao Yang
|
a03ca46a28
|
Fix KDA prefix caching under mamba extra_buffer and enable it for kimi_linear (#31474)
|
2026-07-19 20:09:03 +08:00 |
|
Alison Shao
|
a5c0b94034
|
Let CI server launches wait longer for ports held by a dying predecessor (#31281)
|
2026-07-17 21:45:24 -07:00 |
|
Lianmin Zheng
|
c95026aed3
|
Upgrade llguidance to 1.7.6 (#31484)
|
2026-07-17 16:31:44 -07:00 |
|
Liangsheng Yin
|
19f4859b30
|
[CI] Exclude current process memory from GPU idle check (#31571)
|
2026-07-17 02:11:41 -07:00 |
|
Liangsheng Yin
|
27ad9d11b1
|
[CI] Wait for GPU memory release before each test class setUpClass (#31509)
|
2026-07-16 21:06:29 -07:00 |
|
pllimax
|
4ad418d2c3
|
Push test case scripts from test repo to main upstream community repository (#31114)
|
2026-07-16 21:24:01 +08:00 |
|
YC Yen-Ching Tseng
|
14095ef78a
|
[AMD] Disable CUDA IPC multimodal transport on ROCm in MMMU VLM tests (#31342)
|
2026-07-15 18:42:30 -07:00 |
|
 Артем Савкинandronnie_zheng
|
8ed82afcc8
|
[MoE Refactor] [NPU] Refactor Ascend MoE implementation to reduce code duplication and align with community design (#25663)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-07-15 14:59:42 +03:00 |
|
fzyzcjy
|
c315df49bb
|
Fix abusing presence of req.req_pool_idx to indicate the presence of req.kv resources (#29430)
|
2026-07-15 14:48:39 +08:00 |
|
fzyzcjy
|
d8d76c4d12
|
Introduce req.kv container for coupled owned kv field lifecycle (#29427)
|
2026-07-15 14:40:38 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
4aadf94146
|
[Kernel] Relocate vendored fla and mamba kernel trees to sglang.kernels (RFC #29630, Phase 2.5, 7/7) (#30795)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-15 12:52:15 +08:00 |
|
 Xiaoyu ZhangandClaude Fable 5
|
ba5be86d42
|
[Kernel] Migrate DSA + DSV4 attention kernels to sglang.kernels (RFC #29630, Phase 2.5, 5/7) (#30792)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-15 11:11:22 +08:00 |
|
 weireweireandweireweire
|
a9cf5e68e6
|
[DSV4] Remove per-step seqlen D2H from speculative to make overlap scheduler work (#30365)
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
|
2026-07-14 18:01:22 -07:00 |
|
fzyzcjy
|
725920915f
|
Introduce ModelRunner.ps ParallelState (#31161)
|
2026-07-14 16:01:14 +08:00 |
|
 Zhanghengandispobock
|
afa3c06d1f
|
Using UnifiedRadixTree by default for SWA, Mamba, and DSA models (#30468)
Co-authored-by: ispobock <ispobaoke@gmail.com>
|
2026-07-14 15:17:08 +08:00 |
|
Liangsheng Yin
|
e2728ac504
|
[Spec] Remove dead padded_static_len and stale SGLANG_ENABLE_SPEC_V2 references (#30998)
|
2026-07-13 15:30:31 -05:00 |
|
Jialin Ouyang
|
47030b28be
|
Fix MockDSV4ModelRunner missing spec_algorithm (#31056)
|
2026-07-13 15:24:28 -05:00 |
|
Liangsheng Yin
|
c0f1f7e062
|
[Spec] Rename num_tokens_per_bs to num_tokens_per_req (#30977)
|
2026-07-13 13:47:53 -05:00 |
|
Jae B.
|
a91c2e6596
|
[Apple Silicon] [CI] Move the MLX lane to the check-changes + pr-gate composite (#30121)
|
2026-07-10 22:00:15 -07:00 |
|
Alison Shao
|
90688366d9
|
test(disagg): set MC_GID_INDEX on RoCE hosts so mooncake KV transfer works (#30737)
|
2026-07-11 11:29:22 +08:00 |
|
Xinyuan Tong
|
b76dd0be69
|
Fix Mistral GSM8K chat eval (#27757)
|
2026-07-09 21:08:48 -07:00 |
|
Cheng Wan
|
1f15308dca
|
[refactor] Retire the legacy config accessor and the remaining process singletons (#30493)
|
2026-07-09 02:10:47 -07:00 |
|
Cheng Wan
|
e703f9e566
|
[refactor] Adopt get_parallel() everywhere and close out the parallel wrapper surface (#30492)
|
2026-07-09 02:09:39 -07:00 |
|
Liangsheng Yin
|
bc5d376c2c
|
[Bench] Add fixed-prompt mode and per-request spec accept length metrics (#30615)
|
2026-07-09 02:06:04 -07:00 |
|
Cheng Wan
|
b14f7b4f75
|
[refactor] Move model-capability adjustments into the resolution pipeline (#30299)
|
2026-07-07 21:26:55 -07:00 |
|
   
|
16372b4c5f
|
[Spec] Anchor GLM-5.2 MTP IndexShare topk on the draft-extend step (#29787)
Co-authored-by: kpham-sgl <264503018+kpham-sgl@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
2026-07-06 20:36:48 -07:00 |
|
hhhh1252023
|
1b481deade
|
feat: sync npu nightly test improvements from Ascend testcases (#29403)
|
2026-07-06 22:41:15 +08:00 |
|