  
|
08d4c2072b
|
move topk capturers to srt/state_capturer/ (#24450)
Co-authored-by: Yueming Yuan <yym022502@gmail.com>
Co-authored-by: fzyzcjy <5236035+fzyzcjy@users.noreply.github.com>
Co-authored-by: Ziang Li <ziangli@umich.edu>
|
2026-05-05 15:54:01 -07:00 |
|
Liangsheng Yin
|
47a416fc62
|
add indexer-topk capture (V3.2 NSA + infra) (#24392)
|
2026-05-05 15:05:15 -07:00 |
|
Zhangheng
|
e299ec1bff
|
[UnifiedRadixTree]: Fix flaky ci (#24421)
|
2026-05-05 20:22:19 +08:00 |
|
Khoa Pham
|
d22853480d
|
Fix deterministic inference on models with SWAKVPool (#24395)
|
2026-05-05 20:20:46 +08:00 |
|
Michael
|
244531bc4f
|
[AMD] Add Kimi-K2.6 in nightly tests for MI30x and MI35x (#23848)
|
2026-05-04 23:37:14 -07:00 |
|
Vladislav Nosivskoy
|
60a1dacd89
|
[HiCache] return cached_tokens_details in sglext for streaming responses (#22055)
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
|
2026-05-04 12:30:17 -07:00 |
|
Zhangheng
|
05aed5e1d5
|
[UnifiedRadixTree]: Add KL accuracy CI for UnifiedTree with HiCache (#24346)
|
2026-05-04 20:18:10 +08:00 |
|
  
|
952b3caf18
|
feat: use structural tags to enable strict tool calling and reasoning for more models (#21722)
Signed-off-by: Yuchuan <yuchuan.7streams@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Ubospica <ubospica@gmail.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
|
2026-05-04 02:30:28 -07:00 |
|
Khoa Pham
|
ef2b1b6d89
|
Fix flashinfer workspace OOM (#24172)
|
2026-05-04 01:26:35 -07:00 |
|
 Bingxu ChenandYC Yen-Ching Tseng
|
5eff3c489a
|
[AMD] Deepseek v4 Flash / Pro nightly tests for MI35x ROCm 7.2 (#24203)
Co-authored-by: YC Yen-Ching Tseng <yctseng@amd.com>
|
2026-05-04 00:02:33 -07:00 |
|
Ke Bao
|
aea527afdc
|
Fix swa chunk req deferred (#24318)
|
2026-05-04 14:52:15 +08:00 |
|
 ZhanghengandShangming Cai
|
9a5450ad73
|
[PD]: Support incremental transfer for mooncake transfer engine (#24257)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-05-04 00:57:59 +08:00 |
|
 
|
c0f5950636
|
[UnifiedRadixTree]: Support HiCache Framework for UnifiedRadixTree (#23316)
Co-authored-by: JINZ <1023553676@qq.com>
Co-authored-by: diemchai <diemchai@tencent.com>
|
2026-05-03 22:13:22 +08:00 |
|
 ZhanghengandShangming Cai
|
44ca2d01fc
|
[pd]: (Bug Fix) Incorrect out_cache_loc slicing in prepare_for_prebuilt (#24230)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-05-03 18:35:16 +08:00 |
|
Glen Liu
|
76b9c8de6f
|
[Feature] add LoRADrainer to address high P99 TTFT (#17913)
|
2026-05-02 16:13:43 -07:00 |
|
    
|
88bb5dffe4
|
[Dependency] Upgrade to Torch 2.11.0 (#21247)
Co-authored-by: Kangyan Zhou <zky314343421@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: b8zhong <b8zhong@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
|
2026-05-02 12:25:36 -07:00 |
|
 egvenediktovandronnie_zheng
|
83bf5d6869
|
[NPU]TP Communications compression For Qwen3 models for NPU (#20520)
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-05-02 14:29:11 +03:00 |
|
   
|
ebbaab5597
|
[NPU] Add GitHub test summary and deduplicate test code. Part 1 (#23835)
Co-authored-by: Elizaveta Martirosian <elizaveta.martirosian@gmail.com>
Co-authored-by: root <root@localhost.localdomain>
Co-authored-by: Elizaveta Martirosian <you@example.com>
Co-authored-by: ronnie_zheng <zl19940307@163.com>
|
2026-05-02 14:18:18 +03:00 |
|
billishyahao
|
b939d5410f
|
[AMD] enable sdma for moriep unittest (#24259)
|
2026-05-01 23:34:53 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.7
|
cd27baaffd
|
[ci][cu13] Bump torch_memory_saver to 0.0.9.post1; restore manual tests (#23182)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
|
2026-05-01 22:50:38 -07:00 |
|
 
|
d41e8c459d
|
Support RunAI loading for quantized checkpoints (#23850)
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Sam Shleifer <sam@thinkingmachines.ai>
|
2026-05-02 11:11:40 +08:00 |
|
Cheng Wan
|
b47fab6f5d
|
[bugfix] Support MIXED forward mode in TBO splitter for DP attention (#24241)
|
2026-05-01 16:01:23 -07:00 |
|
Lianmin Zheng
|
ece8a1a788
|
Refactor device timer, clean up metrics collector, and add fwd occupancy metric (#24197)
|
2026-05-01 10:25:25 -07:00 |
|
 
|
4a50cd781e
|
[BugFix][HiMamba] Fix host-protected node deletion in HiMamba tombstone del (#23696)
Co-authored-by: diemchai <diemchai@tencent.com>
Co-authored-by: Zhangheng <hzh0425@apache.org>
|
2026-05-01 21:57:47 +08:00 |
|
ishandhanani
|
5b7ce417d0
|
[P/D disagg] - support decode side radix cache (#19746)
|
2026-05-01 21:55:34 +08:00 |
|
Qiaolin Yu
|
4197c55968
|
[spec decoding] add tests for chain-style multi layer eagle + return_logprob (#24192)
|
2026-05-01 01:48:48 -07:00 |
|
billishyahao
|
a578bf814c
|
[AMD] fix moriep unittest failure (#24205)
|
2026-04-30 23:36:20 -07:00 |
|
Yanbin Jiang
|
8975479f87
|
[LoRA][MOE] Fix EP correctness in MoE LoRA slicing and virtual-experts kernels (#24171)
|
2026-04-30 22:42:10 -07:00 |
|
 AlecandKangyan-Zhou
|
9d95783603
|
Add Docker image provenance metadata (#24090)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
|
2026-04-30 21:40:42 -07:00 |
|
Yihao Wang
|
0acc569edd
|
[Bench] extend MMMU answer extractor with explicit-commit patterns (#24084)
|
2026-04-30 19:08:08 -07:00 |
|
  ![gemini-code-assist[bot]](/assets/img/avatar_default.png)
|
da7f890788
|
[Intel GPU] Integrate flash_mla_decode in Intel XPU attention backend (#23557)
Signed-off-by: P V R K Jyothendra Varma <polisetty.v.r.k.jyothendra.varma@intel.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-05-01 07:21:28 +08:00 |
|
 
|
e35ac95cdc
|
[Test] Add XPU device support to unit tests (#22236)
Co-authored-by: vshekhawat-hlab <vshekhawat@habana.ai>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-05-01 07:18:51 +08:00 |
|
 
|
9c5cad3914
|
Use device-agnostic helpers for Mamba tests and core ops (#20234)
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-05-01 07:14:53 +08:00 |
|
 Kalyan KumarandMa Mingfei
|
8a9e424faa
|
Replace hardcoded CUDA device with get_device() for XPU support (#13599)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-05-01 07:13:46 +08:00 |
|
 Alison ShaoandAlison Shao
|
694ef516cb
|
Revert "[ci] split stage-c-test-4-gpu-b200 to enable a low-disk runner pool" (#24163)
Co-authored-by: Alison Shao <alisonshao@radixark.ai>
|
2026-04-30 15:57:19 -07:00 |
|
Xinyuan Tong
|
989a16187d
|
[Bench] Fix bench_serving missing reasoning_content stream chunks (#23954)
|
2026-04-30 15:00:27 -07:00 |
|
    
|
651af06a0b
|
[Feature] Xiaomi MiMo-V2.5 day0 support (#23811)
Co-authored-by: 张袁 <zhangyuan36@xiaomi.com>
Co-authored-by: 刘安岐 <liuanqi6@xiaomi.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Shangming Cai <csmthu@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
|
2026-05-01 00:02:26 +08:00 |
|
 
|
aa74911448
|
[NPU] fix some npu error with OffloaderV2 (#19541)
Co-authored-by: Jianzhao Xu <xujianchao@huawei.com>
Co-authored-by: sglang-npu-bot <sglangnpu@163.com>
|
2026-04-30 15:05:35 +03:00 |
|
Yaochen Han
|
577dbc4ab9
|
[4/N] Quantization Refactor: AWQ schemes and Kernel call and weight init split (#21126)
|
2026-04-30 14:51:01 +03:00 |
|
Lianmin Zheng
|
b1ef99f65f
|
[CI] Remove orphaned test/srt/ascend and test/srt/configs (#24145)
|
2026-04-30 04:43:11 -07:00 |
|
kkyyxhll
|
936c9c2355
|
fix(qwen3_5): broadcast per-tensor scale in _make_packed_weight_loader for FP8 models (#23062)
|
2026-04-30 14:16:57 +08:00 |
|
 Opher LieberandYanbin Jiang
|
c8c1c9261d
|
LoRA support for qwen3.5 and nemotron3 (#23594)
Co-authored-by: Yanbin Jiang <jybsuper@gmail.com>
|
2026-04-29 21:51:53 -07:00 |
|
Liangsheng Yin
|
c54ada994b
|
fix: rename mimo spec threshold attr to num_accepted_drafts_thres (#24118)
|
2026-04-29 21:00:55 -07:00 |
|
Qiaolin Yu
|
2bbd30a27a
|
relax the threshold in test_step3p5_flash_chain_mtp (#24105)
|
2026-04-29 16:53:35 -07:00 |
|
Jimmy Shong
|
3d31ac2672
|
[Fix] FP8 Qwen3-Next quant error by removing fallback fused shards (#23973)
|
2026-04-29 17:33:47 -04:00 |
|
Qiaolin Yu
|
79dbfe4505
|
Use spec v2 by default (#21062)
|
2026-04-29 13:40:42 -07:00 |
|
AndyLi429
|
4c1eefca4f
|
[NPU] ascend backend support qwen3 moe attention cp (#21685)
|
2026-04-29 19:25:17 +08:00 |
|
Sam (Kesen Li)
|
73e93bebd6
|
[1/4] NVFP4 KV cache: quantization strategy abstraction and kernel (#21954)
|
2026-04-29 01:45:48 -07:00 |
|
Xinyuan Tong
|
832b4f59ed
|
[Bench] fix MMMU answer-extraction regex dropping multi-line responses (#23864)
|
2026-04-29 14:48:49 +08:00 |
|
shuwenn
|
2c41ef4c93
|
[HiCache] feat: add draft KV cache backing for L2/L3 (#21125)
|
2026-04-28 23:47:31 -07:00 |
|