Commit Graph
2164 Commits
Author SHA1 Message Date
Ziang Li a19ef3a615 [FlashInver v0.6.7] Integrate flashinfer_trtllm mxfp8 gemm (#21576) 2026-04-01 15:55:06 -04:00
Cherry_ming e67b95d66b [NPU]Add a full test pipeline on NPU, resolve issues in the NPU test architecture (#20751) 2026-04-01 19:56:31 +08:00
Liangsheng Yin ac039bd04e Use CustomTestCase for TestSessionControl to enable CI retry (#21830) 2026-04-01 04:26:11 -07:00
Yuhao Yang 1aabe44b64 [VLM] remove AsyncMMDataProcessor wrapper (#21651) 2026-04-01 17:39:50 +08:00
wduan-hai 95b881452e Fix in-place mode in pause generation (#21705) 2026-04-01 01:36:28 -07:00
72d3d8f4cf [Feature Restoration] repetition_penalty is essential for GLM-V models (#21258)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
2026-03-31 23:29:49 -07:00
Ethan (Yusheng) SuandBaizhou Zhang cffc95edf4 [3/n] lora moe - Support Qwen3-VL-30B-A3B-Instruct (#21469)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-03-31 23:15:16 -07:00
ca3ba05a7a chore: bump flashinfer version to 0.6.7 (#21422)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-03-31 21:18:16 -07:00
KnightLTC 2488233ad5 [bugfix]GLM-4V model (#17122) 2026-04-01 10:37:40 +08:00
Qiaolin Yu 5f6250769a Reduce redundant speculative decoding CI tests (#21779) 2026-03-31 17:40:20 -07:00
Liangsheng Yin b6fe0cca99 Switch MooncakeSpec to EAGLE3 + Llama-3.1 (#21794) 2026-03-31 17:12:20 -07:00
Liangsheng Yin d047d41bad Increase hicache eval to 200 examples (#21791) 2026-03-31 16:58:44 -07:00
Liangsheng Yin 7932e4c3e6 Remove redundant test_moe_eval_accuracy_large (#21787) 2026-03-31 16:45:03 -07:00
Mohammad Miadh Angkad 883ba640b2 [CI] Remove more redundant PCG tests (#21554) 2026-03-31 16:25:30 -07:00
Ethan (Yusheng) SuandBaizhou Zhang 3c91ebdf55 [2/n] lora - Shared outer experts and support qwen3_30b_a3b_instruct (#21466)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-03-31 14:06:23 -07:00
JD 20d07c4384 Fix remote weight info nnode>1 and dp>1 (#17389) 2026-03-31 21:17:18 +08:00
Ke Bao 91ab75b505 [CI] Fix ring test timeout (#21751) 2026-03-31 18:26:08 +08:00
weireweireandWeiliangl User 4455d17619 [PD] Refactor Disagg Conn and Fix Hang with total_request/total_tokens Balancing (#21299)
Co-authored-by: Weiliangl User <weiliangl@login-node.hosted.internal>
2026-03-31 18:01:50 +08:00
Ke Bao 08d3c1a134 Fix disaggregation hybrid attention ci (#21745) 2026-03-31 16:22:54 +08:00
Baizhou Zhang d52757fe97 [CI]Remove msgm-en and mmlu tests which cause timeout (#21733) 2026-03-31 01:10:05 -07:00
1d6424d5ad fix: Mistral Small 4 fails to start due to config/weight format mismatch (#21620)
Co-authored-by: mengxiancheng03 <mengxiancheng03@kuaishou.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-30 01:57:35 -07:00
blzhengandMa Mingfei ed01e1d5d6 [CPU] add kernel apply_rotary_pos_emb_cpu for Qwen3-VL and Qwen3-Omni (#13121)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
2026-03-29 23:43:46 -07:00
Aishwarya Ramasethu c32ee48886 MFU metrics in Prometheus (#19395) 2026-03-29 23:40:06 -07:00
Ziang Li 1a4b383fac [CI] [FlashInfer v0.6.7] Use offline quantized checkpoint for MXFP8 Gemm tests (#21625) 2026-03-29 22:47:46 -07:00
Feng Suandzhangxiaolei123456 9b4dd27478 [Fix] Fix Qwen3.5 MoE model loading and Mamba cache sharding in PP mode (#21448)
Co-authored-by: zhangxiaolei123456 <zhangxiaolei.666@bytedance.com>
2026-03-30 11:57:26 +08:00
Lianmin ZhengandClaude Opus 4.6 1d9c8e8c9e Simplify routed experts test and move base64 encoding to tokenizer manager (#21634)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-29 12:44:01 -07:00
22e4733ab9 Add subprocess liveness monitor to detect scheduler crashes (#18582)
Co-authored-by: 继优 <jiyou.ljy@alibaba-inc.com>
Co-authored-by: shuwenn <47200617+alphabetc1@users.noreply.github.com>
2026-03-29 00:09:13 -07:00
Junrong Lin 35f5a0ff35 [CI] Lossen test_return_routed_experts threshold (#21270) 2026-03-28 22:04:53 -07:00
Kangyan-ZhouandClaude Opus 4.6 9d64a82173 feat(ci): add GB300 nightly benchmark test suites (#21487)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-28 21:54:03 -07:00
Shangming Cai 166e9090ee [CI] Skip flaky elastic EP test (#21619) 2026-03-29 12:50:40 +08:00
Lianmin ZhengandClaude Opus 4.6 ba6b501f3a Clean up detokenizer and remove dead multimodal_gen code (#21588)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-28 21:44:40 -07:00
eigenandAvery Huang 3ab9afd653 fix: piecewise_cuda_graph get correct qo_indptr (#21452)
Co-authored-by: Avery Huang <averyh@nvidia.com>
2026-03-28 15:57:29 -07:00
Xinyuan Tong ced69c9f84 feat: enable CUDA graph and timestamp for the whisper model(#21190) 2026-03-29 01:46:03 +08:00
Yuan Luo ee15c104ef [CI] hot-fix ci lint (#21608) 2026-03-28 21:32:39 +08:00
Jacob0226andClaude Opus 4.6 7078e385ea [AMD] Add GLM-4.7-FP8 accuracy CI test for MI35x (#21534)
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-28 00:28:56 -07:00
Baizhou Zhang 6ef4318ec0 [CI] Move v32 cp test to deepep running suite (#21585) 2026-03-27 22:49:06 -07:00
Trevor Morris 7160b6cb76 [NVIDIA] Enable automatic NUMA configuration (#19452) 2026-03-27 18:44:13 -07:00
Vladislav NosivskoyandLianmin Zheng c37200f5e4 Scope streaming backlog coalescing to incremental_streaming_output mode (#21037)
Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
2026-03-27 17:29:54 -07:00
Ethan (Yusheng) SuandBaizhou Zhang 6d48719e31 [1/n] lora support - Auto detect lora target modules (#21439)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
2026-03-27 16:08:36 -07:00
Qiaolin Yu 4a41aec844 Fix flaky test_pp_single_node (#21564) 2026-03-27 14:33:46 -07:00
Baizhou Zhang 4e905febd2 [CI] Relax several thresholds in flaky CIs (#21562) 2026-03-27 13:16:49 -07:00
Lianmin Zheng 9a91323c9f test: point DSV3 int8 MLA CI models to lmsys Hugging Face org (#21561) 2026-03-27 13:04:01 -07:00
Bi Xue 30397e0a1e [rl][sgl] fix tensor mismatch after pause (#21514) 2026-03-27 23:02:30 +08:00
Baizhou Zhang 0138129d3c [CI] Fix nemotron nvfp4 test estimated time (#21516) 2026-03-26 21:53:09 -07:00
Mohammad Miadh Angkad eaf392b9cc Remove redundant DeepSeek V3 FP4 PCG test (#21485) 2026-03-26 21:52:47 -07:00
Shangming Cai 1487f80158 chore: bump mooncake version to 0.3.10 (#20942)
Signed-off-by: Shangming Cai <csmthu@gmail.com>
2026-03-27 10:35:31 +08:00
Aurick Qiao c2b3e42ad6 Fix sessions with mm inputs (#21269) 2026-03-26 17:38:23 -07:00
Liangsheng Yin 8a4cdcd538 Simplify flush_cache: reject concurrent requests, remove client-side retry (#21490) 2026-03-26 16:31:04 -07:00
SevenJ 2e65c27b29 Api add flush cache timeout (#21413)
Signed-off-by: root <wenjun7j@gmail.com>
2026-03-26 14:44:37 -07:00
Qiaolin Yu 8c3ccef2d9 Fix Kimi K2.5 dp attention+ spec decoding launch crash (#21391) 2026-03-26 14:40:26 -07:00