Ke Bao
|
b21db86e2f
|
[CI] Fix gpu deps import in cpu test (#21950)
|
2026-04-03 00:06:31 +08:00 |
|
 Liangsheng YinandDarkSharpness
|
9d9537fbd3
|
Migrate ngram corpus from torch cpp_extension to TVM FFI jit_kernel (#21920)
Co-authored-by: DarkSharpness <2040703891@qq.com>
|
2026-04-02 02:18:11 -07:00 |
|
 foraxeandyunzhi
|
e55a35fbcd
|
test: add manual init test for mooncake transfer engine (#21842)
Co-authored-by: yunzhi <ningyunxiao.nyx@antgroup.com>
|
2026-04-02 16:01:10 +08:00 |
|
Khoa Pham
|
f836658077
|
[Spec][Ngram] 4/N: Remove max_match_window_size and min_match_window_size, matching all suffixes of the Trie (#21225)
|
2026-04-01 22:09:46 -07:00 |
|
Liangsheng Yin
|
269589ad71
|
Return HTTP 400 for streaming validation errors (#21900)
|
2026-04-01 21:58:12 -07:00 |
|
Khoa Pham
|
153359b4dd
|
Multi tool streaming fix (#20004)
|
2026-04-01 21:53:05 -07:00 |
|
Mook
|
7a59e05dd1
|
[Kernel] Fuse temperature + softmax in sampling for decode speedup (#20501)
|
2026-04-02 12:46:36 +08:00 |
|
David Cheung
|
ed427e1299
|
Migrate all callers from /get_server_info to /server_info (#21463)
|
2026-04-01 21:17:50 -07:00 |
|
Kangyan-Zhou
|
648632b6c4
|
[CI] Remove crashing Kimi K2.5 EAGLE3/MTP variants, keep TP8 and TP8+DP8 (#21898)
|
2026-04-01 20:27:24 -07:00 |
|
 Liangsheng YinandClaude Opus 4.6
|
875a615993
|
fix(ci): update est_time for 57 tests based on runtime analysis (#21896)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-01 20:16:13 -07:00 |
|
Yuhao Yang
|
2ef12073f4
|
[VLM] Add VLM TP=4 per-commit CI test and improve MMMU eval prompt/parser (#21841)
|
2026-04-01 20:09:47 -07:00 |
|
Noa Neria
|
8d9145d97e
|
Direct model loading from object storage with Runai Model Streamer (#17948)
Signed-off-by: Noa Neria <noa@run.ai>
|
2026-04-01 18:41:22 -07:00 |
|
 Derek YuandBrayden Zhong
|
51ad717089
|
[CI] Add Per-Tensor, Blockwise FP8 Tests on SM120 (#20717)
Co-authored-by: Brayden Zhong <b8zhong@uwaterloo.ca>
|
2026-04-02 01:20:14 +00:00 |
|
Derek Yu
|
83c3158014
|
[CI] Add Llama 3.1 8B Instruct FP4 CI test on SM120 (#20648)
|
2026-04-02 01:17:38 +00:00 |
|
Liangsheng Yin
|
d7256eb69a
|
Unify GSM8K eval path to Chat API for regression CI readiness (#21667)
|
2026-04-01 17:12:19 -07:00 |
|
 Alison ShaoandAlison Shao
|
1ac74e652e
|
[Misc] Fix comparator e2e tests: add polars dep + fix dp-attention test (#21804)
Co-authored-by: Alison Shao <alison.shao@mac.lan>
|
2026-04-01 15:44:35 -07:00 |
|
Ziang Li
|
a19ef3a615
|
[FlashInver v0.6.7] Integrate flashinfer_trtllm mxfp8 gemm (#21576)
|
2026-04-01 15:55:06 -04:00 |
|
Cherry_ming
|
e67b95d66b
|
[NPU]Add a full test pipeline on NPU, resolve issues in the NPU test architecture (#20751)
|
2026-04-01 19:56:31 +08:00 |
|
Liangsheng Yin
|
ac039bd04e
|
Use CustomTestCase for TestSessionControl to enable CI retry (#21830)
|
2026-04-01 04:26:11 -07:00 |
|
Yuhao Yang
|
1aabe44b64
|
[VLM] remove AsyncMMDataProcessor wrapper (#21651)
|
2026-04-01 17:39:50 +08:00 |
|
wduan-hai
|
95b881452e
|
Fix in-place mode in pause generation (#21705)
|
2026-04-01 01:36:28 -07:00 |
|
   
|
72d3d8f4cf
|
[Feature Restoration] repetition_penalty is essential for GLM-V models (#21258)
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
|
2026-03-31 23:29:49 -07:00 |
|
 Ethan (Yusheng) SuandBaizhou Zhang
|
cffc95edf4
|
[3/n] lora moe - Support Qwen3-VL-30B-A3B-Instruct (#21469)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-03-31 23:15:16 -07:00 |
|
 
|
ca3ba05a7a
|
chore: bump flashinfer version to 0.6.7 (#21422)
Co-authored-by: sglang-bot <sglang-bot@users.noreply.github.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-03-31 21:18:16 -07:00 |
|
KnightLTC
|
2488233ad5
|
[bugfix]GLM-4V model (#17122)
|
2026-04-01 10:37:40 +08:00 |
|
Qiaolin Yu
|
5f6250769a
|
Reduce redundant speculative decoding CI tests (#21779)
|
2026-03-31 17:40:20 -07:00 |
|
Liangsheng Yin
|
b6fe0cca99
|
Switch MooncakeSpec to EAGLE3 + Llama-3.1 (#21794)
|
2026-03-31 17:12:20 -07:00 |
|
Liangsheng Yin
|
d047d41bad
|
Increase hicache eval to 200 examples (#21791)
|
2026-03-31 16:58:44 -07:00 |
|
Liangsheng Yin
|
7932e4c3e6
|
Remove redundant test_moe_eval_accuracy_large (#21787)
|
2026-03-31 16:45:03 -07:00 |
|
Mohammad Miadh Angkad
|
883ba640b2
|
[CI] Remove more redundant PCG tests (#21554)
|
2026-03-31 16:25:30 -07:00 |
|
 Ethan (Yusheng) SuandBaizhou Zhang
|
3c91ebdf55
|
[2/n] lora - Shared outer experts and support qwen3_30b_a3b_instruct (#21466)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-03-31 14:06:23 -07:00 |
|
JD
|
20d07c4384
|
Fix remote weight info nnode>1 and dp>1 (#17389)
|
2026-03-31 21:17:18 +08:00 |
|
Ke Bao
|
91ab75b505
|
[CI] Fix ring test timeout (#21751)
|
2026-03-31 18:26:08 +08:00 |
|
 weireweireandWeiliangl User
|
4455d17619
|
[PD] Refactor Disagg Conn and Fix Hang with total_request/total_tokens Balancing (#21299)
Co-authored-by: Weiliangl User <weiliangl@login-node.hosted.internal>
|
2026-03-31 18:01:50 +08:00 |
|
Ke Bao
|
08d3c1a134
|
Fix disaggregation hybrid attention ci (#21745)
|
2026-03-31 16:22:54 +08:00 |
|
Baizhou Zhang
|
d52757fe97
|
[CI]Remove msgm-en and mmlu tests which cause timeout (#21733)
|
2026-03-31 01:10:05 -07:00 |
|
  
|
1d6424d5ad
|
fix: Mistral Small 4 fails to start due to config/weight format mismatch (#21620)
Co-authored-by: mengxiancheng03 <mengxiancheng03@kuaishou.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-30 01:57:35 -07:00 |
|
 blzhengandMa Mingfei
|
ed01e1d5d6
|
[CPU] add kernel apply_rotary_pos_emb_cpu for Qwen3-VL and Qwen3-Omni (#13121)
Co-authored-by: Ma Mingfei <mingfei.ma@intel.com>
|
2026-03-29 23:43:46 -07:00 |
|
Aishwarya Ramasethu
|
c32ee48886
|
MFU metrics in Prometheus (#19395)
|
2026-03-29 23:40:06 -07:00 |
|
Ziang Li
|
1a4b383fac
|
[CI] [FlashInfer v0.6.7] Use offline quantized checkpoint for MXFP8 Gemm tests (#21625)
|
2026-03-29 22:47:46 -07:00 |
|
 Feng Suandzhangxiaolei123456
|
9b4dd27478
|
[Fix] Fix Qwen3.5 MoE model loading and Mamba cache sharding in PP mode (#21448)
Co-authored-by: zhangxiaolei123456 <zhangxiaolei.666@bytedance.com>
|
2026-03-30 11:57:26 +08:00 |
|
 Lianmin ZhengandClaude Opus 4.6
|
1d9c8e8c9e
|
Simplify routed experts test and move base64 encoding to tokenizer manager (#21634)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-29 12:44:01 -07:00 |
|
 
|
22e4733ab9
|
Add subprocess liveness monitor to detect scheduler crashes (#18582)
Co-authored-by: 继优 <jiyou.ljy@alibaba-inc.com>
Co-authored-by: shuwenn <47200617+alphabetc1@users.noreply.github.com>
|
2026-03-29 00:09:13 -07:00 |
|
Junrong Lin
|
35f5a0ff35
|
[CI] Lossen test_return_routed_experts threshold (#21270)
|
2026-03-28 22:04:53 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.6
|
9d64a82173
|
feat(ci): add GB300 nightly benchmark test suites (#21487)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-28 21:54:03 -07:00 |
|
Shangming Cai
|
166e9090ee
|
[CI] Skip flaky elastic EP test (#21619)
|
2026-03-29 12:50:40 +08:00 |
|
 Lianmin ZhengandClaude Opus 4.6
|
ba6b501f3a
|
Clean up detokenizer and remove dead multimodal_gen code (#21588)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-03-28 21:44:40 -07:00 |
|
 eigenandAvery Huang
|
3ab9afd653
|
fix: piecewise_cuda_graph get correct qo_indptr (#21452)
Co-authored-by: Avery Huang <averyh@nvidia.com>
|
2026-03-28 15:57:29 -07:00 |
|
Xinyuan Tong
|
ced69c9f84
|
feat: enable CUDA graph and timestamp for the whisper model(#21190)
|
2026-03-29 01:46:03 +08:00 |
|
Yuan Luo
|
ee15c104ef
|
[CI] hot-fix ci lint (#21608)
|
2026-03-28 21:32:39 +08:00 |
|