Ziang Li
|
a19ef3a615
|
[FlashInver v0.6.7] Integrate flashinfer_trtllm mxfp8 gemm (#21576)
|
2026-04-01 15:55:06 -04:00 |
|
Cherry_ming
|
e67b95d66b
|
[NPU]Add a full test pipeline on NPU, resolve issues in the NPU test architecture (#20751)
|
2026-04-01 19:56:31 +08:00 |
|
Liangsheng Yin
|
ac039bd04e
|
Use CustomTestCase for TestSessionControl to enable CI retry (#21830)
|
2026-04-01 04:26:11 -07:00 |
|
Yuhao Yang
|
1aabe44b64
|
[VLM] remove AsyncMMDataProcessor wrapper (#21651)
|
2026-04-01 17:39:50 +08:00 |
|
wduan-hai
|
95b881452e
|
Fix in-place mode in pause generation (#21705)
|
2026-04-01 01:36:28 -07:00 |
|
   
|
72d3d8f4cf
|
[Feature Restoration] repetition_penalty is essential for GLM-V models (#21258)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
|
2026-03-31 23:29:49 -07:00 |
|
 Ethan (Yusheng) SuandBaizhou Zhang
|
cffc95edf4
|
[3/n] lora moe - Support Qwen3-VL-30B-A3B-Instruct (#21469)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-03-31 23:15:16 -07:00 |
|
 
|
ca3ba05a7a
|
chore: bump flashinfer version to 0.6.7 (#21422)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-03-31 21:18:16 -07:00 |
|
KnightLTC
|
2488233ad5
|
[bugfix]GLM-4V model (#17122)
|
2026-04-01 10:37:40 +08:00 |
|
Qiaolin Yu
|
5f6250769a
|
Reduce redundant speculative decoding CI tests (#21779)
|
2026-03-31 17:40:20 -07:00 |
|
Liangsheng Yin
|
b6fe0cca99
|
Switch MooncakeSpec to EAGLE3 + Llama-3.1 (#21794)
|
2026-03-31 17:12:20 -07:00 |
|
Liangsheng Yin
|
d047d41bad
|
Increase hicache eval to 200 examples (#21791)
|
2026-03-31 16:58:44 -07:00 |
|
Liangsheng Yin
|
7932e4c3e6
|
Remove redundant test_moe_eval_accuracy_large (#21787)
|
2026-03-31 16:45:03 -07:00 |
|
Mohammad Miadh Angkad
|
883ba640b2
|
[CI] Remove more redundant PCG tests (#21554)
|
2026-03-31 16:25:30 -07:00 |
|
 Ethan (Yusheng) SuandBaizhou Zhang
|
3c91ebdf55
|
[2/n] lora - Shared outer experts and support qwen3_30b_a3b_instruct (#21466)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-03-31 14:06:23 -07:00 |
|
JD
|
20d07c4384
|
Fix remote weight info nnode>1 and dp>1 (#17389)
|
2026-03-31 21:17:18 +08:00 |
|
Ke Bao
|
91ab75b505
|
[CI] Fix ring test timeout (#21751)
|
2026-03-31 18:26:08 +08:00 |
|
 weireweireandWeiliangl User
|
4455d17619
|
[PD] Refactor Disagg Conn and Fix Hang with total_request/total_tokens Balancing (#21299)
Co-authored-by: Weiliangl User <weiliangl@login-node.hosted.internal>
|
2026-03-31 18:01:50 +08:00 |
|
Ke Bao
|
08d3c1a134
|
Fix disaggregation hybrid attention ci (#21745)
|
2026-03-31 16:22:54 +08:00 |
|
Baizhou Zhang
|
d52757fe97
|
[CI]Remove msgm-en and mmlu tests which cause timeout (#21733)
|
2026-03-31 01:10:05 -07:00 |
|
  
|
1d6424d5ad
|
fix: Mistral Small 4 fails to start due to config/weight format mismatch (#21620)
Co-authored-by: mengxiancheng03 <mengxiancheng03@kuaishou.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-30 01:57:35 -07:00 |
|
 blzhengandMa Mingfei
|
ed01e1d5d6
|
[CPU] add kernel apply_rotary_pos_emb_cpu for Qwen3-VL and Qwen3-Omni (#13121)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-03-29 23:43:46 -07:00 |
|
Aishwarya Ramasethu
|
c32ee48886
|
MFU metrics in Prometheus (#19395)
|
2026-03-29 23:40:06 -07:00 |
|
Ziang Li
|
1a4b383fac
|
[CI] [FlashInfer v0.6.7] Use offline quantized checkpoint for MXFP8 Gemm tests (#21625)
|
2026-03-29 22:47:46 -07:00 |
|
 Feng Suandzhangxiaolei123456
|
9b4dd27478
|
[Fix] Fix Qwen3.5 MoE model loading and Mamba cache sharding in PP mode (#21448)
Co-authored-by: zhangxiaolei123456 <zhangxiaolei.666@bytedance.com>
|
2026-03-30 11:57:26 +08:00 |
|
 Lianmin ZhengandClaude Opus 4.6
|
1d9c8e8c9e
|
Simplify routed experts test and move base64 encoding to tokenizer manager (#21634)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-29 12:44:01 -07:00 |
|
 
|
22e4733ab9
|
Add subprocess liveness monitor to detect scheduler crashes (#18582)
Co-authored-by: 继优 <jiyou.ljy@alibaba-inc.com>
Co-authored-by: shuwenn <47200617+alphabetc1@users.noreply.github.com>
|
2026-03-29 00:09:13 -07:00 |
|
Junrong Lin
|
35f5a0ff35
|
[CI] Lossen test_return_routed_experts threshold (#21270)
|
2026-03-28 22:04:53 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.6
|
9d64a82173
|
feat(ci): add GB300 nightly benchmark test suites (#21487)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-28 21:54:03 -07:00 |
|
Shangming Cai
|
166e9090ee
|
[CI] Skip flaky elastic EP test (#21619)
|
2026-03-29 12:50:40 +08:00 |
|
 Lianmin ZhengandClaude Opus 4.6
|
ba6b501f3a
|
Clean up detokenizer and remove dead multimodal_gen code (#21588)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-28 21:44:40 -07:00 |
|
 eigenandAvery Huang
|
3ab9afd653
|
fix: piecewise_cuda_graph get correct qo_indptr (#21452)
Co-authored-by: Avery Huang <averyh@nvidia.com>
|
2026-03-28 15:57:29 -07:00 |
|
Xinyuan Tong
|
ced69c9f84
|
feat: enable CUDA graph and timestamp for the whisper model(#21190)
|
2026-03-29 01:46:03 +08:00 |
|
Yuan Luo
|
ee15c104ef
|
[CI] hot-fix ci lint (#21608)
|
2026-03-28 21:32:39 +08:00 |
|
 Jacob0226andClaude Opus 4.6
|
7078e385ea
|
[AMD] Add GLM-4.7-FP8 accuracy CI test for MI35x (#21534)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
|
2026-03-28 00:28:56 -07:00 |
|
Baizhou Zhang
|
6ef4318ec0
|
[CI] Move v32 cp test to deepep running suite (#21585)
|
2026-03-27 22:49:06 -07:00 |
|
Trevor Morris
|
7160b6cb76
|
[NVIDIA] Enable automatic NUMA configuration (#19452)
|
2026-03-27 18:44:13 -07:00 |
|
 Vladislav NosivskoyandLianmin Zheng
|
c37200f5e4
|
Scope streaming backlog coalescing to incremental_streaming_output mode (#21037)
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
|
2026-03-27 17:29:54 -07:00 |
|
 Ethan (Yusheng) SuandBaizhou Zhang
|
6d48719e31
|
[1/n] lora support - Auto detect lora target modules (#21439)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-03-27 16:08:36 -07:00 |
|
Qiaolin Yu
|
4a41aec844
|
Fix flaky test_pp_single_node (#21564)
|
2026-03-27 14:33:46 -07:00 |
|
Baizhou Zhang
|
4e905febd2
|
[CI] Relax several thresholds in flaky CIs (#21562)
|
2026-03-27 13:16:49 -07:00 |
|
Lianmin Zheng
|
9a91323c9f
|
test: point DSV3 int8 MLA CI models to lmsys Hugging Face org (#21561)
|
2026-03-27 13:04:01 -07:00 |
|
Bi Xue
|
30397e0a1e
|
[rl][sgl] fix tensor mismatch after pause (#21514)
|
2026-03-27 23:02:30 +08:00 |
|
Baizhou Zhang
|
0138129d3c
|
[CI] Fix nemotron nvfp4 test estimated time (#21516)
|
2026-03-26 21:53:09 -07:00 |
|
Mohammad Miadh Angkad
|
eaf392b9cc
|
Remove redundant DeepSeek V3 FP4 PCG test (#21485)
|
2026-03-26 21:52:47 -07:00 |
|
Shangming Cai
|
1487f80158
|
chore: bump mooncake version to 0.3.10 (#20942)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
|
2026-03-27 10:35:31 +08:00 |
|
Aurick Qiao
|
c2b3e42ad6
|
Fix sessions with mm inputs (#21269)
|
2026-03-26 17:38:23 -07:00 |
|
Liangsheng Yin
|
8a4cdcd538
|
Simplify flush_cache: reject concurrent requests, remove client-side retry (#21490)
|
2026-03-26 16:31:04 -07:00 |
|
SevenJ
|
2e65c27b29
|
Api add flush cache timeout (#21413)
Signed-off-by: root <wenjun7j@gmail.com>
|
2026-03-26 14:44:37 -07:00 |
|
Qiaolin Yu
|
8c3ccef2d9
|
Fix Kimi K2.5 dp attention+ spec decoding launch crash (#21391)
|
2026-03-26 14:40:26 -07:00 |
|