+3        
|
52fecfdf09
|
support qwen 3.8 flash next (#37500)
Co-authored-by: ch-wan <54331508+ch-wan@users.noreply.github.com>
Co-authored-by: ispobock <26454835+ispobock@users.noreply.github.com>
Co-authored-by: JustinTong0323 <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: samuellees <26428561+samuellees@users.noreply.github.com>
Co-authored-by: YAMY1234 <74099316+YAMY1234@users.noreply.github.com>
Co-authored-by: yhyang201 <47235274+yhyang201@users.noreply.github.com>
Co-authored-by: yizhang2077 <25844240+yizhang2077@users.noreply.github.com>
Co-authored-by: zijiexia <37504505+zijiexia@users.noreply.github.com>
Co-authored-by: Shinto C V <cshintov@gmail.com>
Co-authored-by: Julian Huang <huangzhilin.hzl@antgroup.com>
Co-authored-by: Xiaoyu Zhang <1182563586@qq.com>
Co-authored-by: yhyang201 <yhyang201@gmail.com>
|
2026-09-08 13:56:21 -07:00 |
|
Liangsheng Yin
|
9a489f8d2f
|
[Test] Move gpqa and aime25 onto sgl-eval, drop unused eval paths (#36979)
|
2026-08-29 17:36:13 -07:00 |
|
Baizhou Zhang
|
2d88c79b3e
|
[CI] Add Kimi-K3 MMMU-Pro accuracy coverage (#36284)
|
2026-08-25 16:33:46 -07:00 |
|
Liangsheng Yin
|
50cc1aa241
|
[CI] Route mmlu and GB300 MMMU-Pro evals through sgl-eval (#34477)
|
2026-08-12 19:33:21 -07:00 |
|
Rain Jiang
|
4af8ddb576
|
support rust sglang server (#29799)
|
2026-07-31 11:56:31 -07:00 |
|
Xinyuan Tong
|
b76dd0be69
|
Fix Mistral GSM8K chat eval (#27757)
|
2026-07-09 21:08:48 -07:00 |
|
Liangsheng Yin
|
bc5d376c2c
|
[Bench] Add fixed-prompt mode and per-request spec accept length metrics (#30615)
|
2026-07-09 02:06:04 -07:00 |
|
fzyzcjy
|
609f5f549c
|
Add mixed-prefix gsm8k eval and its CPU unit test (#27502)
|
2026-06-09 20:17:41 +08:00 |
|
Liangsheng Yin
|
d7256eb69a
|
Unify GSM8K eval path to Chat API for regression CI readiness (#21667)
|
2026-04-01 17:12:19 -07:00 |
|
Baizhou Zhang
|
5e12c4e08e
|
[DSA] Support trtllm sparse mla kernel for prefill batches (#21783)
|
2026-04-01 13:55:05 -07:00 |
|
Liangsheng Yin
|
09907795e1
|
Add latency and throughput metrics to run_eval (#21793)
|
2026-03-31 18:36:14 -07:00 |
|
Liangsheng Yin
|
7581d814ae
|
Add CompletionSampler for non-chat eval in run_eval (#21785)
|
2026-03-31 16:33:07 -07:00 |
|
Mohammad Miadh Angkad
|
4fbb311234
|
[Fix][Eval] Keep --dataset-path scoped to longbench_v2 (#21156)
|
2026-03-24 02:25:11 -07:00 |
|
Jia Guo
|
87549f8f0b
|
perf(mamba): use Triton conv1d for non-contiguous input to avoid .contiguous() copy (#20469)
|
2026-03-19 19:38:46 -07:00 |
|
Mohammad Miadh Angkad
|
ca997b7ba9
|
Add min_p and chat-template kwargs support to run_eval (#19571)
|
2026-03-09 14:53:09 -07:00 |
|
 Kaixi HouandClaude Opus 4.5
|
4181290efd
|
[NVIDIA] Add --top-k argument to run_eval.py (#18025)
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
|
2026-02-02 22:17:53 -08:00 |
|
YAMY
|
2740ed1ae7
|
[eval] GSM8k support for run_eval (#17041)
|
2026-01-16 11:10:17 +08:00 |
|
hlu1
|
0e86de7c0b
|
Remove deepseek-r1 from THINKING_MODE_CHOICES in run_eval.py (#17178)
|
2026-01-15 16:53:06 -08:00 |
|
hlu1
|
aeb480c11f
|
Add top-p to run_eval.py (#16844)
|
2026-01-10 17:10:37 +08:00 |
|
![gemini-code-assist[bot]](/assets/img/avatar_default.png) Hudson Xingandgemini-code-assist[bot]
|
f4ab2ec5be
|
Add unified metrics collection framework (v1) (#16064)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
|
2026-01-03 16:30:38 -08:00 |
|
 ![gemini-code-assist[bot]](/assets/img/avatar_default.png)     
|
1f1f05a85e
|
vlm: refactor engine vlm params and support processor output as input (#14091)
Co-authored-by: Mick <mickjagger19@icloud.com>
Co-authored-by: zhaochenyang20 <zhaochenyang20@gmail.com>
Co-authored-by: Xinyuan Tong <115166877+JustinTong0323@users.noreply.github.com>
Co-authored-by: BenYao21 <cyao22@asu.edu>
Co-authored-by: minleminzui <minleminzui@gmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: 赵晨阳 <zhaochen20@outlook.com>
|
2025-12-20 18:31:24 +08:00 |
|
 
|
3e4d431a44
|
[Feature] Add AIME25 dataset support for SGLang simple_eval (#14990)
Co-authored-by: zkexorability <zkexorability@gmail.com>
Co-authored-by: Baizhou Zhang <sobereddiezhang@gmail.com>
|
2025-12-15 21:59:40 -08:00 |
|
Baizhou Zhang
|
ab3ffd1c8e
|
Add nightly accuracy test for DeepSeek V3.2 (#14935)
|
2025-12-13 12:11:16 -08:00 |
|
Binyao Jiang
|
312df1d6c0
|
Fix TestGLM41VPPAccuracy test flakiness (#14848)
|
2025-12-10 16:59:58 -08:00 |
|
Binyao Jiang
|
0aa65f94f1
|
[Fix] Improve longbench prompt and other logics (#11474)
|
2025-10-12 15:04:28 -07:00 |
|
Al-Ekram Elahee Hridoy
|
533e58a15d
|
Feature/longbench v2 evaluation utils (#10949)
|
2025-10-07 14:17:31 +08:00 |
|
Liangsheng Yin
|
04b86b3c5c
|
[hot-fix] Fix CI break which caused by adding thinking_mode in eval (#11192)
|
2025-10-03 18:29:27 +08:00 |
|
hlu1
|
d6777a706d
|
Add --thinking-mode to run_eval (#11189)
Signed-off-by: Hao Lu <14827759+hlu1@users.noreply.github.com>
|
2025-10-03 16:49:39 +08:00 |
|
Liangsheng Yin
|
9710f718fb
|
[Eval] Add --repeat in run_eval (#11101)
|
2025-09-30 23:35:54 +08:00 |
|
Mick
|
777eb53897
|
ci: refactor nightly test (#10495)
|
2025-09-26 15:24:30 -07:00 |
|
fzyzcjy
|
442534aa44
|
Add CI for gpt-oss model on hopper (#8851)
|
2025-08-09 00:34:23 -07:00 |
|
Lifu Huang
|
6e2da51561
|
Replace time.time() to time.perf_counter() for benchmarking. (#6178)
Signed-off-by: Lifu Huang <lifu.hlf@gmail.com>
|
2025-05-11 14:32:49 -07:00 |
|
Lianmin Zheng
|
ad4125d1a9
|
Fuse more ops & Simplify token mapping (#1758)
|
2024-10-22 23:20:43 -07:00 |
|
Lianmin Zheng
|
0c1c72a0b4
|
Fix accuracy test (#1051)
|
2024-08-12 19:48:40 +10:00 |
|
Lianmin Zheng
|
41598e0d8e
|
Add longer accuracy test on CI (#1049)
|
2024-08-12 09:21:38 +00:00 |
|
Ying Sheng
|
995af5a54b
|
Improve the structure of CI (#911)
|
2024-08-03 23:09:21 -07:00 |
|
Ying Sheng
|
e90e3a50d4
|
Add benchmark: HumanEval (#889)
|
2024-08-02 00:46:41 -07:00 |
|
Ying Sheng
|
ae7ee01a8e
|
Add accuracy test to CI: MMLU (#882)
|
2024-08-01 21:20:17 -07:00 |
|