[Test] Move gpqa and aime25 onto sgl-eval, drop unused eval paths (#36979)
This commit is contained in:
@@ -122,7 +122,6 @@ Also, do not rely on the "Latency/Output throughput" from this script, as it is
|
||||
GSM8K is too easy for state-of-the-art models nowadays. Please try your own more challenging accuracy tests.
|
||||
You can find additional accuracy eval examples in:
|
||||
- [test_eval_accuracy_large.py](https://github.com/sgl-project/sglang/blob/main/test/manual/eval/test_eval_accuracy_large.py)
|
||||
- [test_gpt_oss_1gpu.py](https://github.com/sgl-project/sglang/blob/main/test/manual/core/test_gpt_oss_1gpu.py)
|
||||
|
||||
## Benchmark the speed
|
||||
|
||||
|
||||
Reference in New Issue
Block a user