 Alison ShaoandAlison Shao
|
2ad5e6df12
|
[CI] Relax gpt-oss 4GPU accuracy threshold from 0.60 to 0.58 (#22237)
Co-authored-by: Alison Shao <alison.shao@Mac.attlocal.net>
|
2026-04-08 02:20:23 -07:00 |
|
Sundara Raman Ramachandran
|
712c8c5051
|
[Score API] Add SequenceClassification Model support (#22118)
|
2026-04-08 01:30:58 -07:00 |
|
Baizhou Zhang
|
213af1d4f7
|
Add CI tests for GLM-5 (#22285)
|
2026-04-08 01:05:36 -07:00 |
|
Michael
|
db60a620db
|
[AMD] Add GLM-5-FP8 nightly performance benchmarks for MI30x and MI35x (#21710)
|
2026-04-07 22:43:14 -07:00 |
|
 Alison ShaoandAlison Shao
|
36f05810c9
|
[CI] Move manual-only nightly tests out of test/registered/ (#22298)
Co-authored-by: Alison Shao <alison.shao@Mac.attlocal.net>
|
2026-04-07 21:03:52 -07:00 |
|
 Alex NailsandClaude Opus 4.6
|
493ec91cbe
|
[CI] Fix stage-b-test-1-gpu-large (0) timeout by reordering LoRA tests and using tokenizer from cache (#22292)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-07 20:00:44 -07:00 |
|
 Qiaolin YuandLiangsheng Yin
|
117508dcd7
|
Switch eagle_infer_beta to EAGLE3 (#22303)
Co-authored-by: Liangsheng Yin <hnyls2002@users.noreply.github.com>
|
2026-04-07 18:43:48 -07:00 |
|
Kangyan-Zhou
|
dd73e9a62e
|
Revert "[CI] Update nightly test models for H200/B200 (#22288)" (#22297)
|
2026-04-07 17:04:06 -07:00 |
|
 
|
f6fc39569a
|
[CI] Migrate mgsm_en eval to gsm8k to remove openaipublic dependency (#21931)
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Kangyan-Zhou <zky314343421@gmail.com>
|
2026-04-07 16:29:20 -07:00 |
|
 Kangyan-ZhouandClaude Opus 4.6
|
e6652309c4
|
[CI] Update nightly test models for H200/B200 (#22288)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-07 15:44:52 -07:00 |
|
    
|
f08726fd56
|
[Feature] Add DFLASH speculative decoding support (#22077)
Co-authored-by: Jian Chen <141193260+jianc99@users.noreply.github.com>
Co-authored-by: Zhijian Liu <5782437+zhijian-liu@users.noreply.github.com>
Co-authored-by: Richard Gong <8001209+gongy@users.noreply.github.com>
Co-authored-by: David Wang <21328423+dcw02@users.noreply.github.com>
Co-authored-by: yilian49 <43861414+yilian49@users.noreply.github.com>
Co-authored-by: xm:D <38322020+xiaomin-d@users.noreply.github.com>
|
2026-04-07 14:48:51 -07:00 |
|
YC Yen-Ching Tseng
|
e14876742a
|
[AMD] Fix test_kimi_k25_mxfp4.py : stage-c-test-large-8-gpu-amd-mi35x (linux-mi35x-gpu-8, 1) (#22188)
|
2026-04-07 13:48:37 -07:00 |
|
Liangsheng Yin
|
cc35714b03
|
[tiny] migrate /get_server_info; print accept length in accuracy tests (#22282)
|
2026-04-07 13:08:35 -07:00 |
|
Rain Jiang
|
1a8eb890f6
|
Kernels community fa3 (#20796)
|
2026-04-07 12:48:44 -07:00 |
|
Ke Bao
|
be42fbbbd7
|
Support HTTP2 server (#21700)
|
2026-04-08 00:42:52 +08:00 |
|
Ke Bao
|
fae90abf6e
|
Move ring test to nightly (#22267)
|
2026-04-07 21:56:39 +08:00 |
|
Xingyu Liu
|
98f38b14df
|
Add registration API for external linear attention backend (#21983)
Signed-off-by: Xingyu Liu <charlotteliu12x@gmail.com>
|
2026-04-07 02:47:40 -07:00 |
|
Michael
|
ba78f6e0ef
|
[AMD] Add Qwen3.5-397B FP8 nightly perf benchmarks for MI30x and MI35x (#21669)
|
2026-04-06 23:46:00 -07:00 |
|
 ![gemini-code-assist[bot]](/assets/img/avatar_default.png)    
|
2813cb6d9a
|
[New Model] Gemma 4 (#21952)
Signed-off-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
Co-authored-by: Pengyu Chen <pychen96@gmail.com>
Co-authored-by: kpham-sgl <khoa.pham@radixark.ai>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Andy Luo <andy.luo@amd.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: adarshxs <adarsh.shirawalmath@gmail.com>
|
2026-04-06 20:24:44 -07:00 |
|
Liangsheng Yin
|
be0277f9a0
|
[Spec][Ngram] Add output-as-corpus accept length benchmark for external SAM (#22199)
|
2026-04-06 19:09:52 -07:00 |
|
 Lianmin ZhengandClaude Opus 4.6
|
494bb86169
|
Cache sub-objects in __getitem__ to ensure identity stability (#22184)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
2026-04-06 18:53:38 -07:00 |
|
Liangsheng Yin
|
e4b1366a46
|
[Spec][Ngram] Support multiple SAMs with dynamic HTTP API (#22203)
|
2026-04-06 18:49:22 -07:00 |
|
Trevor Morris
|
56266de624
|
[CI] Add basic unit test for Minimax-M2.5 (#21792)
|
2026-04-06 15:48:33 -07:00 |
|
 Alison ShaoandAlison Shao
|
6f1412f4f5
|
[CI] Relax transformers MMLU threshold from 0.65 to 0.64 (#22210)
Co-authored-by: Alison Shao <alison.shao@Mac.attlocal.net>
|
2026-04-06 15:32:09 -07:00 |
|
Ratish P
|
7f2fcc0b08
|
[VLM]: allow Qwen3.5 models for encoder disaggregation (#21849)
|
2026-04-07 02:07:24 +08:00 |
|
Zhangheng
|
d72f58d1c1
|
[Qwen3-Specv2]: Fix flaky ci (#22194)
|
2026-04-07 00:40:44 +08:00 |
|
Ke Bao
|
9ca2ae1c6c
|
Update test skills and guide (#22189)
|
2026-04-06 20:30:25 +08:00 |
|
Aurick Qiao
|
3178f3959f
|
Align incremental streaming logprobs with streamed output tokens (#21583)
|
2026-04-06 00:30:02 -07:00 |
|
Khoa Pham
|
12272b6791
|
[Spec][Ngram] 6/N: Load an external corpus and construct a Suffix Automaton (#21425)
|
2026-04-06 00:11:14 -07:00 |
|
Liangsheng Yin
|
6de2ff2a80
|
[Spec][Ngram] Followup fixes for MatchState incremental advance (#22180)
|
2026-04-05 23:04:28 -07:00 |
|
 YAMYandShangming Cai
|
dc125afffb
|
Add staging buffer CI test and documentation for heterogeneous TP (#21921)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-04-06 14:00:20 +08:00 |
|
Khoa Pham
|
b2008bf9e0
|
[Spec][Ngram] 5/N: Store and advance anchor match state across decode steps (#21243)
|
2026-04-05 22:21:05 -07:00 |
|
Ke Bao
|
c4240218cb
|
Fix ut module importing (#22176)
|
2026-04-06 11:53:58 +08:00 |
|
Lidang Jiang
|
682ddf6416
|
[CI] Add unit tests for srt/utils/auth.py (#21400)
|
2026-04-06 10:30:46 +08:00 |
|
Lidang Jiang
|
30f5b87608
|
[CI] Add unit tests for function_call detectors (hermes, llama32, mistral) (#21399)
|
2026-04-06 10:29:55 +08:00 |
|
ROSINE HE
|
8d8aca8f26
|
[Test] Add unit tests for srt/tokenizer/tiktoken_tokenizer (#21107)
|
2026-04-06 10:23:19 +08:00 |
|
Shangming Cai
|
75ab75d027
|
Fix create_grammar_backend test calls with think_end_id (#22158)
|
2026-04-06 00:54:35 +08:00 |
|
Zhangheng
|
f6c9072e42
|
[SpecV2]: Reopen kl accuracy test for qwen3 + SpecV2 (#22104)
|
2026-04-05 23:26:40 +08:00 |
|
 Liangsheng YinandShangming Cai
|
3a4f4cbc52
|
DEBUG: reproduce flaky test_load_weights_from_remote_instance (#22150)
Co-authored-by: Shangming Cai <csmthu@gmail.com>
|
2026-04-05 22:46:36 +08:00 |
|
 iLeGendandBaizhou Zhang
|
5a35316417
|
Enable IndexCache for DeepSeek V3.2 (#21405)
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2026-04-05 02:45:58 -07:00 |
|
Baizhou Zhang
|
088203454b
|
[Fix] Fix nightly tests (#22140)
|
2026-04-05 02:26:42 -07:00 |
|
Liangsheng Yin
|
bd6a585605
|
Consolidate reasoning tests into test/registered/reasoning/ (#22139)
|
2026-04-05 01:09:11 -07:00 |
|
Liangsheng Yin
|
904bb476d8
|
Migrate reasoning_tokens tests to existing server fixtures (#22102)
|
2026-04-05 00:30:57 -07:00 |
|
Baizhou Zhang
|
723ed6c3c4
|
[CI]Temporary ban auto benchmark tool test (#22138)
|
2026-04-04 23:18:19 -07:00 |
|
Liangsheng Yin
|
675100b53a
|
Remove flaky TestToolChoiceLfm2Moe from test_tool_choice (#22137)
|
2026-04-04 22:53:11 -07:00 |
|
 RoyWangandRoyWang
|
dd49127fe6
|
[AMD]: Support MLA with nhead<16 and FP8 KV cache for TP=8 (Kimi K2.5… (#21213)
Co-authored-by: RoyWang <RoyWang@amd.com>
|
2026-04-04 22:13:29 -07:00 |
|
Xiaoyu Zhang
|
0f0f004f1f
|
[Benchmark] Add auto benchmark tool with YAML-driven server flag search and canonical dataset format (#21736)
|
2026-04-04 21:46:58 +08:00 |
|
Liangsheng Yin
|
e9d92b0e33
|
Relax spec decoding accuracy threshold to fix flaky test (#22100)
|
2026-04-04 02:38:35 -07:00 |
|
    
|
1ad6839659
|
[Feature] Add Reasoning Tokens Usage (#15562)
Signed-off-by: Muqi Li <muqi1029@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: Mufeez Amjad <mufeez.amjad@outlook.com>
Co-authored-by: cklxx <1293822641@qq.com>
Co-authored-by: hnyls2002 <lsyincs@gmail.com>
Co-authored-by: Liangsheng Yin <hnyls2002@gmail.com>
|
2026-04-04 02:18:10 -07:00 |
|
 narutolhyandluhongyu.4869
|
24763256b9
|
[Speculative Decoding] Add FA4-based Spec Support (#21080)
Co-authored-by: luhongyu.4869 <luhongyu.4869@bytedance.com>
|
2026-04-04 02:09:45 -07:00 |
|