diff --git a/docs/cookbook/autoregressive/DeepSeek/DeepSeek-V3_2.mdx b/docs/cookbook/autoregressive/DeepSeek/DeepSeek-V3_2.mdx
index 14166f925..ddd4baf70 100644
--- a/docs/cookbook/autoregressive/DeepSeek/DeepSeek-V3_2.mdx
+++ b/docs/cookbook/autoregressive/DeepSeek/DeepSeek-V3_2.mdx
@@ -813,7 +813,7 @@ python3 -m sglang.test.run_eval --port 30000 --eval-name gpqa --num-examples 198
#### 5.2.4 AIME 2025 Benchmark
-Results on AIME 2025 (8×B200), evaluated with [NeMo-Skills](https://github.com/NVIDIA/NeMo-Skills.git):
+Results on AIME 2025 (8×B200), evaluated with [sgl-eval](https://github.com/sgl-project/sgl-eval):
@@ -846,25 +846,19 @@ Results on AIME 2025 (8×B200), evaluated with [NeMo-Skills](https://github.com/
-**Reproduction.** Install [NeMo-Skills](https://github.com/NVIDIA/NeMo-Skills), launch the server with the tool-call and reasoning parsers, then run `ns eval`:
+**Reproduction.** Install [sgl-eval](https://github.com/sgl-project/sgl-eval), launch the server with the tool-call and reasoning parsers, then run `sgl-eval run`. The AIME 2025 dataset ships with sgl-eval, and thinking is on by default for it:
```bash Command
-pip install git+https://github.com/NVIDIA/NeMo-Skills.git --ignore-installed blinker
+pip install git+https://github.com/sgl-project/sgl-eval.git
-export NEMO_SKILLS_DISABLE_UNCOMMITTED_CHANGES_CHECK=1
-ns prepare_data aime25
-
-ns eval \
- --benchmarks=aime25:4 \
- --server_type=sglang \
- --model=deepseek-ai/DeepSeek-V3.2-Exp \
- --server_address=http://localhost:30000/v1 \
- --output_dir=nemo_skills_aime25_output \
- ++chat_template_kwargs.thinking=true \
- ++inference.temperature=1.0 \
- ++inference.top_p=0.95 \
- ++inference.tokens_to_generate=64000
- # Use ++inference.tokens_to_generate=120000 for the DeepSeek-V3.2-Speciale model
+sgl-eval run aime25 \
+ --base-url http://localhost:30000/v1 \
+ --model deepseek-ai/DeepSeek-V3.2-Exp \
+ --n-repeats 4 \
+ --temperature 1.0 \
+ --top-p 0.95 \
+ --max-tokens 64000
+ # Use --max-tokens 120000 for the DeepSeek-V3.2-Speciale model
```
### 5.3 Speed Benchmark on Hopper
diff --git a/docs/cookbook/autoregressive/MiniMax/MiniMax-M2.7.mdx b/docs/cookbook/autoregressive/MiniMax/MiniMax-M2.7.mdx
index aa357327a..7cbf6666f 100644
--- a/docs/cookbook/autoregressive/MiniMax/MiniMax-M2.7.mdx
+++ b/docs/cookbook/autoregressive/MiniMax/MiniMax-M2.7.mdx
@@ -487,30 +487,25 @@ This section uses **industry-standard configurations** for comparable benchmark
### 5.1 Accuracy Benchmark
-**Evaluation Tool**: [NVIDIA NeMo-Skills](https://github.com/NVIDIA-NeMo/Skills)
+**Evaluation Tool**: [sgl-eval](https://github.com/sgl-project/sgl-eval)
-**Evaluation Settings**: temperature=0.6, top_p=0.95, 8 seeds, max_tokens=120,000, `parse_reasoning=True`
+**Evaluation Settings**: temperature=0.6, top_p=0.95, 8 repeats, max_tokens=120,000 (GPQA and AIME; MMLU-Pro is greedy and single-pass, see below)
#### 5.1.1 GPQA Diamond
- Dataset: [GPQA Diamond](https://huggingface.co/datasets/Idavidrein/gpqa) (198 questions)
-- Prompt: `eval/aai/mcq-4choices` (4-choice multiple choice, matching [Artificial Analysis methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking))
+- Prompt: `eval/aai/mcq-4choices` (4-choice multiple choice, matching [Artificial Analysis methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking)) -- sgl-eval's default for this benchmark
- Evaluation command:
```bash Command
-ns prepare_data gpqa
+pip install git+https://github.com/sgl-project/sgl-eval.git
-ns eval \
- --cluster=local \
- --server_type=openai \
- --model=MiniMaxAI/MiniMax-M2.7 \
- --server_address=http://localhost:30000/v1 \
- --output_dir=./m2.7-eval/ \
- --benchmarks=gpqa:8 \
- ++prompt_config=eval/aai/mcq-4choices \
- ++inference.tokens_to_generate=120000 \
- ++inference.temperature=0.6 \
- ++inference.top_p=0.95 \
- ++parse_reasoning=True
+sgl-eval run gpqa \
+ --base-url http://localhost:30000/v1 \
+ --model MiniMaxAI/MiniMax-M2.7 \
+ --n-repeats 8 \
+ --max-tokens 120000 \
+ --temperature 0.6 \
+ --top-p 0.95
```
- Test Results:
@@ -544,22 +539,16 @@ ns eval \
#### 5.1.2 AIME 2025
- Dataset: AIME 2025 (30 problems)
-- Prompt: `generic/math` (boxed answer format)
+- Prompt: `generic/math` (boxed answer format) -- sgl-eval's default for this benchmark
- Evaluation command:
```bash Command
-ns prepare_data aime25
-
-ns eval \
- --cluster=local \
- --server_type=openai \
- --model=MiniMaxAI/MiniMax-M2.7 \
- --server_address=http://localhost:30000/v1 \
- --output_dir=./m2.7-eval/ \
- --benchmarks=aime25:8 \
- ++inference.tokens_to_generate=120000 \
- ++inference.temperature=0.6 \
- ++inference.top_p=0.95 \
- ++parse_reasoning=True
+sgl-eval run aime25 \
+ --base-url http://localhost:30000/v1 \
+ --model MiniMaxAI/MiniMax-M2.7 \
+ --n-repeats 8 \
+ --max-tokens 120000 \
+ --temperature 0.6 \
+ --top-p 0.95
```
- Test Results:
@@ -593,22 +582,14 @@ ns eval \
#### 5.1.3 MMLU-Pro
- Dataset: [MMLU-Pro](https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro) (12,032 questions, 10-choice)
-- Prompt: `eval/aai/mcq-10choices` (10-choice multiple choice)
+- Prompt: `eval/aai/mcq-10choices` (10-choice multiple choice) -- sgl-eval's default for this benchmark
- Evaluation command:
```bash Command
-ns prepare_data mmlu-pro
-
-ns eval \
- --cluster=local \
- --server_type=openai \
- --model=MiniMaxAI/MiniMax-M2.7 \
- --server_address=http://localhost:30000/v1 \
- --output_dir=./m2.7-eval/ \
- --benchmarks=mmlu-pro \
- ++prompt_config=eval/aai/mcq-10choices \
- ++inference.tokens_to_generate=32768 \
- ++inference.temperature=0.0 \
- ++parse_reasoning=True
+sgl-eval run mmlu_pro \
+ --base-url http://localhost:30000/v1 \
+ --model MiniMaxAI/MiniMax-M2.7 \
+ --max-tokens 32768 \
+ --temperature 0.0
```
- Test Results:
diff --git a/docs/cookbook/autoregressive/Moonshotai/Kimi-K2.6.mdx b/docs/cookbook/autoregressive/Moonshotai/Kimi-K2.6.mdx
index 178a7c1f7..368c014b8 100644
--- a/docs/cookbook/autoregressive/Moonshotai/Kimi-K2.6.mdx
+++ b/docs/cookbook/autoregressive/Moonshotai/Kimi-K2.6.mdx
@@ -621,32 +621,25 @@ python tool_calls_eval.py tool-calls/samples.jsonl \
#### 5.1.2 AIME 2025
- Dataset: [AIME 2025](https://huggingface.co/datasets/nvidia/aime25) (30 problems)
-- Evaluation Tool: [NVIDIA NeMo-Skills](https://github.com/NVIDIA/NeMo-Skills)
+- Evaluation Tool: [sgl-eval](https://github.com/sgl-project/sgl-eval)
- Prompt: `eval/matharena/aime` (MathArena format with `\boxed{}` answers)
-- Settings: temperature=1.0, top_p=0.95, max_tokens=131,072, 32 seeds
+- Settings: temperature=1.0, top_p=0.95, max_tokens=131,072, 32 repeats
**Evaluation Command:**
```shell Command
-# Prepare dataset
-python3 nemo_skills/dataset/aime25/prepare.py
+pip install git+https://github.com/sgl-project/sgl-eval.git
-# Run 32 seeds in parallel
-for RS in $(seq 0 31); do
- python3 nemo_skills/inference/generate.py \
- input_file=nemo_skills/dataset/aime25/test.jsonl \
- output_file=results/kimi-k26/aime25/output-rs${RS}.jsonl \
- prompt_config=eval/matharena/aime \
- prompt_format=openai \
- +server.server_type=openai \
- +server.model=moonshotai/Kimi-K2.6 \
- +server.base_url=http://localhost:30000/v1 \
- ++inference.temperature=1.0 \
- ++inference.top_p=0.95 \
- ++inference.tokens_to_generate=131072 \
- ++inference.random_seed=${RS} \
- max_concurrent_requests=512 &
-done
+# The dataset ships with sgl-eval; --n-repeats runs all 32 samples per problem.
+sgl-eval run aime25 \
+ --base-url http://localhost:30000/v1 \
+ --model moonshotai/Kimi-K2.6 \
+ --n-repeats 32 \
+ --temperature 1.0 \
+ --top-p 0.95 \
+ --max-tokens 131072 \
+ --prompt matharena-aime \
+ --num-threads 512
```
**Results:**