[Benchmark] Add auto benchmark tool with YAML-driven server flag search and canonical dataset format (#21736)

This commit is contained in:
Xiaoyu Zhang
2026-04-04 21:46:58 +08:00
committed by GitHub
parent da25b471e3
commit 0f0f004f1f
48 changed files with 6688 additions and 0 deletions
@@ -0,0 +1,496 @@
---
name: sglang-auto-benchmark
description: Run SGLang auto benchmark searches with tiered server-flag sweeps, canonical dataset preparation, ShareGPT auto-download, custom-data conversion/validation, SLA or fixed-QPS benchmarking, CSV export, and optional second-stage speculative/EAGLE tuning. Use when the user wants an AI-operated benchmark workflow rather than a one-off bench_serving command.
---
# SGLang Auto Benchmark
This skill is for repeatable, AI-driven SGLang performance tuning.
The preferred workflow is:
- start from a mostly pure-TP baseline command,
- move the rest of the performance knobs into `search_space`,
- let auto benchmark search and compare candidates under the target SLA.
The implementation lives in:
- `python -m sglang.auto_benchmark`
- canonical dataset loader in `python -m sglang.bench_serving --dataset-name autobench`
- cookbook-derived LLM reference configs in `.claude/skills/sglang-auto-benchmark/references/cookbook-llm/`
## Preconditions
- SGLang can already launch and serve the target model in this environment.
- The model path exists, or the model is otherwise launchable.
- The goal is clear:
- benchmark a fixed QPS list, or
- search the maximum QPS that satisfies `max_ttft_ms` / `max_tpot_ms`.
If those are not true yet, fix them before running a large search.
Environment consistency check:
- if the benchmark will run from a remote repo copy or ad-hoc synced workspace,
verify that the remote `python/sglang/bench_serving.py` matches the local
feature level needed by auto benchmark before launching a long run
- at minimum, run a preflight such as `PYTHONPATH=<repo>/python python3 -m
sglang.bench_serving --help` and confirm that the dataset choices include
`autobench`
- if `autobench` is missing remotely, do not start the benchmark; sync
`python/sglang/bench_serving.py` and any required dataset modules first
## Remote Run Logging
If the benchmark is executed on a remote machine, the progress bar output must be
mirrored back to a local file for humans to watch.
Scope note:
- use the remote-log mirroring workflow only when the benchmark is running in a
different machine or a different remote container than the one the agent is
actively operating in
- if the agent itself is already running inside the target container where auto
benchmark is executing, do not add a separate log-return loop just for parity;
inspect the live log files and result files directly in the current container
- in other words, "remote container" needs mirrored local logs, while "current
container" should use direct local inspection
Required behavior:
- start the remote run with a persistent terminal/session log, for example with
`script -q -f <log> -c "<cmd>"`; on Linux containers that use util-linux
`script`, prefer the explicit `-c` form instead of BSD-style positional
command arguments
- continuously sync a cleaned version of that remote session log back to a local
`progress.log`; this local `progress.log` should already have terminal control
sequences removed, because `script` + `tqdm` progress bars will otherwise leave
ANSI cursor-control bytes and carriage-return redraws that look like garbled text
- if the benchmark itself is executed inside a remote container, the cleaned local
`progress.log` must be refreshed automatically at least once every 30
seconds while the run is active; do not rely on one-off manual polling
- implement the sync loop as a dedicated local script file checked into neither
git nor the benchmark config; avoid fragile one-line `nohup zsh -lc '...'`
command strings with heavy nested quoting
- prefer running the sync loop inside a long-lived local session such as a
dedicated `tmux` pane, `screen`, or the agent's own persistent PTY session;
detached child processes started from short-lived command runners can be
reaped unexpectedly, so plain `nohup ... &` is not the most stable default
- immediately after starting the sync loop, verify that `progress.log` is
actually updating by checking its timestamp or size twice across a short wait;
if it is not changing, treat that as a broken sync setup and fix it before
telling the user that live log mirroring is working
- tell the user the local log path up front
- keep final result files synced back locally after the run ends
- when scenario-level or top-level markdown summaries are produced, sync those
`summary.md` / `SUMMARY.md` files back locally as first-class result artifacts
rather than leaving them only on the remote machine
This is important because long searches can run for hours, and people need a
stable local file they can tail without logging into the remote box. The final
local run folder should also be self-contained enough for someone to review the
benchmark outcome without re-entering the remote environment.
Recommended cleanup pipeline for the local mirrored log:
```bash
perl -pe 's/\e\[[0-9;?]*[ -\/]*[@-~]//g; s/\r/\n/g; s/\x08//g;' raw_progress.log \
> progress.log
```
Recommended remote-container sync pattern:
```bash
cat > sync_progress.sh <<'EOF'
#!/bin/zsh
set -euo pipefail
while true; do
ssh <remote-host> "tail -n 200 <remote-progress-log>" > raw_progress.log
perl -pe 's/\e\[[0-9;?]*[ -\/]*[@-~]//g; s/\r/\n/g; s/\x08//g;' raw_progress.log \
> progress.log
sleep 15
done
EOF
chmod +x sync_progress.sh
```
Run that script from a long-lived local session, for example:
```bash
tmux new-session -d -s autobench-sync './sync_progress.sh'
```
Use a persistent local background job, `tmux` pane, `screen`, or equivalent
long-lived sync process so that humans can watch the cleaned local log in real
time. Use `sleep 15` by default for long runs unless there is a specific need
for tighter polling, and keep the cleaned local `progress.log` within the
required 30-second refresh window while the run is active.
At the end of the run, make sure the local artifact set includes any generated:
- `results.jsonl`
- `results.csv`
- `summary.md`
- `SUMMARY.md`
- `scenario_summary.jsonl`
- `scenario_summary.csv`
Required health check after starting the sync script:
```bash
stat -f '%m %z' progress.log
sleep 5
stat -f '%m %z' progress.log
```
If the timestamp and size both stay unchanged while the remote benchmark is known
to be producing new output, the sync loop is broken. Fix the script before
continuing.
Do not make the cleaned log optional. The default local progress artifact should
be the cleaned `progress.log` that humans actually read.
## Most Important Rule
If the user wants the best command for a **real production or real workload scenario**, the benchmark must use **their real request distribution**.
That means:
- real prompt lengths,
- real output lengths,
- real multi-turn patterns,
- real tool / reasoning / sampling settings,
- real prefix-sharing behavior if it exists.
`sharegpt`, `random`, and `generated-shared-prefix` are useful for sanity checks and broad tuning, but they are not a substitute for the user’s real traffic.
The cookbook reference configs now default to `random` because it is portable and immediately runnable, but that should still be treated as a fallback benchmark shape rather than the final answer for a real deployment.
## Supported Dataset Kinds
The current implementation intentionally keeps the dataset surface small:
- `sharegpt`
- Supports auto-download when no file path is provided.
- Will be prepared into canonical autobench JSONL on disk before benchmarking.
- `custom`
- Supports two cases:
- old `bench_serving` custom conversation JSONL,
- already-converted canonical autobench JSONL.
- `random`
- Uses SGLang’s existing synthetic/random benchmark path.
- This is the default dataset mode in the cookbook reference configs.
- `input_len` and `output_len` can be lists of equal length.
- Each aligned pair becomes one full benchmark scenario, not a cartesian product.
- Example:
```yaml
dataset:
kind: random
scenario_names: [chat, summarization]
input_len: [1000, 8000]
output_len: [1000, 1000]
```
- The workflow will run one full search for `1000 -> 1000` and one full search for `8000 -> 1000`.
- `generated-shared-prefix`
- Uses SGLang’s existing shared-prefix synthetic generator.
Everything is normalized into one canonical autobench JSONL file before the benchmark loop starts.
## Canonical Dataset Format
Canonical format is JSONL, one request per line.
Minimal rows:
```json
{"prompt": "Write a summary of this document.", "output_len": 256}
{"prompt": [{"role": "user", "content": "Summarize this document."}], "output_len": 256}
{"prompt": ["first turn", "follow-up turn"], "output_len": 128}
```
Optional fields:
```json
{
"prompt": [{"role": "user", "content": "Use the weather tool."}],
"output_len": 256,
"extra_request_body": {"temperature": 0.0, "top_p": 0.95},
"image_data": ["file:///tmp/example.png"],
"timestamp": 1710000000,
"routing_key": "group-a",
"metadata": {"source": "custom-upload"}
}
```
Compatibility:
- legacy `messages`
- legacy `prompt_origin`
- legacy `param_send`
- legacy `system + content`
## ShareGPT Auto-Prepare
`sharegpt` does not need a full path.
Example:
```bash
python3 -m sglang.auto_benchmark convert \
--kind sharegpt \
--tokenizer /path/to/tokenizer \
--num-prompts 200 \
--output /tmp/sharegpt.autobench.jsonl
```
This will:
- auto-download ShareGPT through the existing SGLang cache path when needed,
- convert it into canonical autobench JSONL,
- save it to the requested output path.
## Custom User Data Workflow
When the user uploads custom data:
1. Inspect a few raw rows first.
2. Decide whether the file is:
- already canonical autobench JSONL,
- old `bench_serving` custom format,
- or an unsupported custom schema that must be transformed manually.
3. If manual transformation is needed:
- map it into canonical JSONL,
- never hallucinate missing turns or answers,
- never keep the final assistant answer as part of the benchmark prompt if that answer is the target completion,
- preserve per-request generation settings in `extra_request_body`.
4. Run:
```bash
python3 -m sglang.auto_benchmark validate \
--dataset-path /path/to/converted.autobench.jsonl \
--tokenizer /path/to/tokenizer
```
5. Manually inspect at least 3 converted rows and confirm:
- prompt shape is correct,
- final assistant answer was not accidentally left in the prompt,
- `output_len` is sensible,
- request extras were preserved.
## Search Tiers
`search.tier` controls search breadth.
- Tier 1
- Fastest and smallest sweep.
- Best for smoke tests, config validation, and quickly checking whether a model can run at all.
- Uses a very small subset of the search space and mainly does one-at-a-time changes on top of the baseline.
- Lowest search cost, but also the easiest to miss a better configuration.
- Tier 2
- Recommended default.
- Good balance between coverage and runtime.
- Runs a small cartesian search on the first few high-priority keys, then expands the rest one at a time.
- Usually the right choice for everyday tuning when you want meaningful search without waiting too long.
- Tier 3
- Largest search space.
- Runs the full cartesian product of the provided search space.
- Search time is the longest by far.
- Only use it when the search space is already tightly bounded and you intentionally want the most exhaustive sweep.
- This is the best chance of finding the strongest config, but it is also the easiest way to turn a benchmark into a multi-hour or multi-day run.
`search.max_candidates` still applies at all tiers, including tier 3.
When it is set together with tier 3, the workflow still enumerates the full cartesian order conceptually, but only keeps the first `max_candidates` unique candidates after deduplication.
That makes it useful as a safety valve, but it also means tier 3 is no longer truly exhaustive unless you remove the cap or raise it high enough.
If `search.max_candidates` is omitted, the workflow now defaults to `8`.
Set it to `null` only when you intentionally want an unbounded sweep.
The reference configs now default to tier 2 with `search.max_candidates: 8`.
## Interrupt And Resume
Long searches may need to be stopped and resumed later.
Use:
```yaml
search:
tier: 2
resume: true
```
Behavior:
- every completed trial is appended to `live_results.jsonl`
- if the process receives `SIGINT` or `SIGTERM`, it will first save partial
`results.jsonl`, `results.csv`, and `summary.md`
- on the next run with the same config and `search.resume: true`, completed
trials are reused and only unfinished trials are executed
- resume works per scenario directory, so it is safest to keep the same
`benchmark.output_dir`
Notes:
- resume assumes the candidate order and dataset are unchanged
- for maximum safety, reuse the same prepared dataset or keep the same dataset
seed/config
- `SIGKILL` cannot be handled gracefully, so only the already-written
`live_results.jsonl` can be reused after a hard kill
YAML key order matters. Put the most important search keys first.
## What Is Tunable
This workflow is not limited to attention backend tuning.
`server.base_flags` and `server.search_space` are passed directly to `sglang.launch_server`, so in practice any valid server CLI flag can be set or searched.
There is also a small convenience layer for parallel search:
- `server.parallel.tp`
- `server.parallel.pp_size`
When `server.parallel` is used and `dp_size` is not set explicitly, the workflow auto-derives:
`dp_size = visible_gpus / (tp_size * pp_size)`
Visible GPU count is inferred from `server.env.CUDA_VISIBLE_DEVICES` by default, or from `server.parallel.gpu_count` if you set it explicitly.
The most important performance-related groups are:
- Kernel / backend
- `attention_backend`
- `prefill_attention_backend`
- `decode_attention_backend`
- `sampling_backend`
- `grammar_backend`
- Batching / scheduling
- `max_running_requests`
- `max_queued_requests`
- `chunked_prefill_size`
- `prefill_max_requests`
- `max_prefill_tokens`
- `schedule_conservativeness`
- `num_continuous_decode_steps`
- `stream_interval`
- Memory / cache
- `max_total_tokens`
- `page_size`
- `disable_radix_cache`
- Parallel / distributed execution
- `tp_size`
- `pp_size`
- `dp_size`
- `ep_size`
- `load_balance_method`
- `enable_dp_attention`
- `enable_mixed_chunk`
- `disable_overlap_schedule`
- Runtime / CUDA graph
- keep CUDA graph enabled by default for performance benchmarking
- `cuda_graph_max_bs`
- `disable_cuda_graph_padding`
- `enable_cudagraph_gc`
- Optional speculative / EAGLE stage
- `speculative_num_steps`
- `speculative_eagle_topk`
- `speculative_num_draft_tokens`
- `speculative_attention_mode`
- `speculative_draft_attention_backend`
- `speculative_accept_threshold_single`
- `speculative_accept_threshold_acc`
For cookbook-derived reference configs, keep `mem_fraction_static` and
`schedule_policy` pinned to the cookbook baseline unless the user explicitly
asks to search them. They are useful knobs, but they add a lot of search width
for relatively low validation value in the default workflow.
Do not put these into the default search space:
- `mem_fraction_static`
- `schedule_policy`
- `enable_hierarchical_cache`
- `hicache_ratio`
- `hicache_size`
- `enable_lmcache`
Those features are not treated as standard auto-benchmark sweep knobs in this workflow.
Budget guardrails for the default workflow:
- use `dataset.num_prompts: 80` unless the user asks for a heavier study
- prefer a coarse QPS search tolerance
- keep `benchmark.qps.max_rounds <= 5`
- keep `search.max_duration_hours <= 12`
## Base Tuning Before EAGLE
Never start by tuning EAGLE first.
Use this order:
1. Tune the non-speculative base server first.
2. Find the best normal config for the target dataset and SLA.
3. Only if the user explicitly asks for speculative/EAGLE tuning, and provides the required draft model or equivalent assets, run the second-stage speculative search.
Do not put `disable_cuda_graph` into the default search space. For normal performance tuning, CUDA graph should stay enabled unless the user is debugging compatibility issues.
When a candidate OOMs, keep it in the final result table as a failed row and add a hint such as:
- increase GPU count, or
- use GPUs with larger memory.
## Running The Workflow
Prepare a dataset explicitly:
```bash
python3 -m sglang.auto_benchmark convert \
--kind custom \
--path /path/to/data.jsonl \
--tokenizer /path/to/tokenizer \
--output /tmp/data.autobench.jsonl
```
Run from config:
```bash
python3 -m sglang.auto_benchmark run --config /path/to/config.yaml
```
Outputs:
- prepared canonical dataset JSONL
- per-run `results.jsonl`
- summary `results.csv`
- per-candidate server logs
## Config Template
Standalone example (uses ShareGPT as dataset, a good starting point for non-cookbook models):
- `references/qwen3-32b.yaml`
Cookbook-derived configs live in `references/cookbook-llm/`.
They default to synthetic `random` traffic and are runnable out of the box.
See `references/cookbook-llm/README.md` for the full list.
Representative picks from that folder:
- `references/cookbook-llm/llama-3.1-70b-instruct.yaml`
- `references/cookbook-llm/llama-3.3-70b-instruct.yaml`
- `references/cookbook-llm/llama-4-scout-17b-16e-instruct.yaml`
- `references/cookbook-llm/llama-4-maverick-17b-128e-instruct-fp8.yaml`
- `references/cookbook-llm/minimax-m2.5.yaml`
- `references/cookbook-llm/minimax-m2.1.yaml`
- `references/cookbook-llm/deepseek-v3.yaml`
- `references/cookbook-llm/deepseek-v3.1.yaml`
- `references/cookbook-llm/deepseek-v3.2.yaml`
- `references/cookbook-llm/deepseek-r1-0528.yaml`
- `references/cookbook-llm/qwen3-235b-a22b.yaml`
- `references/cookbook-llm/qwen35-397b-a17b-fp8.yaml`
- `references/cookbook-llm/mistral-small-4-119b-2603.yaml`
- `references/cookbook-llm/kimi-k2-instruct.yaml`
All reference configs use Hugging Face repo IDs by default.
Replace `model_path` and `tokenizer` with local paths when the weights are already on disk.
## What To Report Back
After a run, summarize:
- which tier was used,
- which dataset kind was used,
- whether the dataset was synthetic or real user traffic,
- best base config,
- best QPS that satisfied SLA,
- whether speculative tuning was skipped or run,
- paths to:
- prepared dataset JSONL
- `results.jsonl`
- `results.csv`
- key server logs
@@ -0,0 +1,76 @@
# Cookbook LLM References
These configs are derived from `sgl-cookbook` autoregressive text-model pages and normalized into the auto-benchmark config format.
Rules used here:
- Keep the baseline as close as possible to a pure-TP launch command.
- Keep `mem_fraction_static` and `schedule_policy` at the cookbook baseline by
default; search higher-ROI knobs first.
- Move the remaining common performance knobs into `search_space`.
- Add `ep_size` search for relevant MoE pages.
- Keep CUDA graph enabled by default.
- Prefer cookbook H200 defaults first, then H100 defaults when H200 is not available; if neither exists, fall back to the cookbook's published baseline for that model and say so in the config comments.
- Default to synthetic `random` data so every config is runnable out of the box.
- Default to `dataset.num_prompts: 80` so the reference sweep stays cheap enough
for interactive validation.
- Default to a coarse QPS search with `benchmark.qps.max_rounds <= 5`.
- Default to `search.tier: 2` so the shipped configs stay reasonably practical to run.
- Default to `search.max_candidates: 8` so the candidate sweep stays bounded by
default.
- Default to `search.max_duration_hours: 12` because longer searches do not fit
the intended workflow budget.
- Treat `dataset.input_len` and `dataset.output_len` as aligned scenario lists, not a cartesian product.
- If a candidate OOMs, the result table should recommend increasing GPU count or using GPUs with larger memory.
Default random scenarios in these configs:
- `1000 -> 1000` for a chat-like shape
- `8000 -> 1000` for a summarization-like shape
Each scenario should run a full search independently, and each scenario should have its own best launch command and summary table.
Excluded from this folder because they are OCR/VL-oriented rather than text-serving benchmark configs:
- DeepSeekOCR / DeepSeekOCR2
- GLMOCR
- GLM45V / GLM46V
- Qwen2.5-VL / Qwen3-VL
- Step3-VL-10B
Configs in this folder:
- `deepseek-v3.2.yaml`
- `deepseek-math-v2.yaml`
- `deepseek-r1-0528.yaml`
- `deepseek-v3.1.yaml`
- `deepseek-v3.yaml`
- `devstral-small-2-24b-instruct-2512.yaml`
- `ernie-4.5-21b-a3b-pt.yaml`
- `glm-4.5.yaml`
- `glm-4.6.yaml`
- `glm-4.7.yaml`
- `glm-4.7-flash.yaml`
- `glm-5-fp8.yaml`
- `gpt-oss-120b.yaml`
- `glyph.yaml`
- `intern-s1.yaml`
- `kimi-k2.5.yaml`
- `kimi-k2-instruct.yaml`
- `kimi-linear-48b-a3b-instruct.yaml`
- `llada2-1-mini.yaml`
- `ling-2.5-1t.yaml`
- `llama-3.1-70b-instruct.yaml`
- `llama-3.3-70b-instruct.yaml`
- `llama-4-scout-17b-16e-instruct.yaml`
- `llama-4-maverick-17b-128e-instruct-fp8.yaml`
- `mimo-v2-flash.yaml`
- `minimax-m2.5.yaml`
- `minimax-m2.1.yaml`
- `ministral-3-8b-instruct-2512.yaml`
- `mistral-small-4-119b-2603.yaml`
- `nemotron-3-nano-30b-a3b-bf16.yaml`
- `nemotron-3-super-120b-a12b-bf16.yaml`
- `qwen35-397b-a17b-fp8.yaml`
- `qwen3-coder-480b-a35b-instruct.yaml`
- `qwen3-coder-next.yaml`
- `qwen3-235b-a22b.yaml`
- `qwen3-next-80b-a3b-instruct.yaml`
- `ring-2.5-1t.yaml`
- `step-3.5-flash.yaml`
@@ -0,0 +1,81 @@
# Cookbook auto benchmark config for deepseek math v2.
# Cookbook does not expose an H100/H200 default for this model; this config follows the cookbook B200 baseline instead.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.8/deepseek-math-v2.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
model_path: deepseek-ai/DeepSeek-Math-V2
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- flashinfer
decode_attention_backend:
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
ep_size:
- 1
- 4
- 8
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: deepseek-ai/DeepSeek-Math-V2
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/deepseek-math-v2
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,84 @@
# Cookbook auto benchmark config for deepseek r1 0528.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.6/deepseek-r1.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
enable_symm_mem: true
model_path: deepseek-ai/DeepSeek-R1-0528
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
ep_size:
- 1
- 4
- 8
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: deepseek-ai/DeepSeek-R1-0528
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/deepseek-r1-0528
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,83 @@
# Cookbook auto benchmark config for deepseek v3.1.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: src/components/autoregressive/DeepSeekV31ConfigGenerator/index.js
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
model_path: deepseek-ai/DeepSeek-V3.1
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
ep_size:
- 1
- 4
- 8
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: deepseek-ai/DeepSeek-V3.1
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/deepseek-v3.1
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,83 @@
# Cookbook auto benchmark config for deepseek v3.2.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.6/deepseek.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
model_path: deepseek-ai/DeepSeek-V3.2
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
ep_size:
- 1
- 4
- 8
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: deepseek-ai/DeepSeek-V3.2
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/deepseek-v3.2
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,84 @@
# Cookbook auto benchmark config for deepseek v3.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: src/components/autoregressive/DeepSeekV3ConfigGenerator/index.js
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
enable_symm_mem: true
model_path: deepseek-ai/DeepSeek-V3
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
ep_size:
- 1
- 4
- 8
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: deepseek-ai/DeepSeek-V3
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/deepseek-v3
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,78 @@
# Cookbook auto benchmark config for devstral small 2 24b instruct 2512.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: src/components/autoregressive/Devstral2ConfigGenerator/index.js
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: '0'
base_flags:
model_path: mistralai/Devstral-Small-2-24B-Instruct-2512
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: mistralai/Devstral-Small-2-24B-Instruct-2512
max_concurrency:
- null
- 16
- 32
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 16.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/devstral-small-2-24b-instruct-2512
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,76 @@
# Cookbook auto benchmark config for ernie 4.5 21b a3b pt.
# Cookbook does not expose an H100/H200 default for this model; this config follows the cookbook MI300X baseline instead.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: src/components/autoregressive/Ernie45ConfigGenerator/index.js
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: '0'
base_flags:
model_path: baidu/ERNIE-4.5-21B-A3B-PT
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
command_prefix:
- python3
- -m
- sglang.launch_server
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: baidu/ERNIE-4.5-21B-A3B-PT
max_concurrency:
- null
- 16
- 32
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 16.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/ernie-4.5-21b-a3b-pt
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,74 @@
# Cookbook auto benchmark config for glm 4.5.
# Cookbook does not expose an H100/H200 default for this model; this config follows the cookbook MI300X baseline instead.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: src/components/autoregressive/GLM45ConfigGenerator/index.js
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3
base_flags:
tp_size: 4
context_length: 8192
model_path: zai-org/GLM-4.5
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: zai-org/GLM-4.5
max_concurrency:
- null
- 8
- 16
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 8.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/glm-4.5
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,83 @@
# Cookbook auto benchmark config for glm 4.6.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.6/glm46.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
model_path: zai-org/GLM-4.6
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
ep_size:
- 1
- 4
- 8
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: zai-org/GLM-4.6
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/glm-4.6
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,78 @@
# Cookbook auto benchmark config for glm 4.7 flash.
# Baseline follows the cookbook H100 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: src/components/autoregressive/GLM47FlashConfigGenerator/index.js
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: '0'
base_flags:
model_path: zai-org/GLM-4.7-Flash
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: zai-org/GLM-4.7-Flash
max_concurrency:
- null
- 16
- 32
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 16.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/glm-4.7-flash
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,78 @@
# Cookbook auto benchmark config for glm 4.7.
# Cookbook does not expose an H100/H200 default for this model; this config follows the cookbook MI300X baseline instead.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: src/components/autoregressive/GLM47ConfigGenerator/index.js
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3
base_flags:
tp_size: 4
context_length: 8192
model_path: zai-org/GLM-4.7
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
ep_size:
- 1
- 2
- 4
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: zai-org/GLM-4.7
max_concurrency:
- null
- 8
- 16
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 8.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/glm-4.7
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,83 @@
# Cookbook auto benchmark config for glm 5 fp8.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.8/glm5.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
model_path: zai-org/GLM-5-FP8
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
ep_size:
- 1
- 4
- 8
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: zai-org/GLM-5-FP8
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/glm-5-fp8
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,81 @@
# Cookbook auto benchmark config for glyph.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: src/components/autoregressive/GlyphConfigGenerator/index.js
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3
base_flags:
tp_size: 4
reasoning_parser: glm45
tool_call_parser: glm45
model_path: zai-org/Glyph
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: zai-org/Glyph
max_concurrency:
- null
- 8
- 16
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 8.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/glyph
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,83 @@
# Cookbook auto benchmark config for gpt oss 120b.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.6/gpt-oss.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
model_path: openai/gpt-oss-120b
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
ep_size:
- 1
- 4
- 8
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: openai/gpt-oss-120b
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/gpt-oss-120b
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,84 @@
# Cookbook auto benchmark config for intern s1.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.6/intern-s1.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
trust_remote_code: true
model_path: internlm/Intern-S1
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
ep_size:
- 1
- 4
- 8
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: internlm/Intern-S1
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/intern-s1
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,83 @@
# Cookbook auto benchmark config for kimi k2 instruct.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.6/kimi-k2.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
trust_remote_code: true
model_path: moonshotai/Kimi-K2-Instruct
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
ep_size:
- 1
- 4
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: moonshotai/Kimi-K2-Instruct
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/kimi-k2-instruct
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,80 @@
# Cookbook auto benchmark config for kimi k2.5.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.8/kimi-k25.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
trust_remote_code: true
model_path: moonshotai/Kimi-K2.5
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: moonshotai/Kimi-K2.5
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/kimi-k2.5
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,79 @@
# Cookbook auto benchmark config for kimi linear 48b a3b instruct.
# Cookbook does not expose an H100/H200 default for this model; this config follows the cookbook MI300X baseline instead.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: src/components/autoregressive/KimiK2linearConfigGenerator/index.js
server:
launch: true
host: 127.0.0.1
port: 30000
env:
SGLANG_ROCM_FUSED_DECODE_MLA: '0'
CUDA_VISIBLE_DEVICES: 0,1,2,3
base_flags:
tp_size: 4
trust_remote_code: true
model_path: moonshotai/Kimi-Linear-48B-A3B-Instruct
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
command_prefix:
- python3
- -m
- sglang.launch_server
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: moonshotai/Kimi-Linear-48B-A3B-Instruct
max_concurrency:
- null
- 8
- 16
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 8.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/kimi-linear-48b-a3b-instruct
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,91 @@
# Cookbook auto benchmark config for ling 2.5 1t.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# This cookbook baseline is multi-node or external-service oriented. Keep launch: false and benchmark an already running deployment.
# Source: src/components/autoregressive/Ling25ConfigGenerator/index.js
server:
launch: false
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
pp_size: 2
nnodes: 2
trust_remote_code: true
tool_call_parser: qwen
model_path: inclusionAI/Ling-2.5-1T
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
pp_size:
- 1
- 2
command_prefix:
- python3
- -m
- sglang.launch_server
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: inclusionAI/Ling-2.5-1T
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 2.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/ling-2.5-1t
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,83 @@
# Cookbook auto benchmark config for llada2 1 mini.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.6/llada21.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: '0'
base_flags:
tp_size: 1
dllm_algorithm: JointThreshold
trust_remote_code: true
max_running_requests: 1
attention_backend: flashinfer
model_path: inclusionAI/LLaDA2.1-mini
mem_fraction_static: 0.77
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 1
- 2
- 4
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: inclusionAI/LLaDA2.1-mini
max_concurrency:
- 1
- 2
- 4
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/llada2-1-mini
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,79 @@
# Cookbook auto benchmark config for llama 3.1 70b instruct.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.6/llama31.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3
base_flags:
tp_size: 4
model_path: meta-llama/Llama-3.1-70B-Instruct
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: meta-llama/Llama-3.1-70B-Instruct
max_concurrency:
- null
- 8
- 16
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 12.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/llama-3.1-70b-instruct
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,73 @@
# Cookbook auto benchmark config for llama 3.3 70b instruct.
# Cookbook does not expose an H100/H200 default for this model; this config follows the cookbook MI300X baseline instead.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: src/components/autoregressive/Llama33ConfigGenerator/index.js
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: '0'
base_flags:
tool_call_parser: llama3
model_path: meta-llama/Llama-3.3-70B-Instruct
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: meta-llama/Llama-3.3-70B-Instruct
max_concurrency:
- null
- 16
- 32
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 16.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/llama-3.3-70b-instruct
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,79 @@
# Cookbook auto benchmark config for llama 4 maverick 17b 128e instruct fp8.
# Cookbook does not expose an H100/H200 default for this model; this config follows the cookbook GENERIC-8GPU baseline instead.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: docs/autoregressive/Llama/Llama4.md
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
context_length: 1000000
trust_remote_code: true
enable_multimodal: true
model_path: meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 4
- 8
- 12
command_prefix:
- sglang
- serve
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8
max_concurrency:
- null
- 2
- 4
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 2.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/llama-4-maverick-17b-128e-instruct-fp8
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE3
draft_model_path: lmsys/sglang-EAGLE3-Llama-4-Maverick-17B-128E-Instruct-v1
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,83 @@
# Cookbook auto benchmark config for llama 4 scout 17b 16e instruct.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.6/llama4scout.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
enable_multimodal: true
context_length: 65536
dtype: bfloat16
trust_remote_code: true
model_path: meta-llama/Llama-4-Scout-17B-16E-Instruct
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 8
- 16
- 24
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: meta-llama/Llama-4-Scout-17B-16E-Instruct
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/llama-4-scout-17b-16e-instruct
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE3
draft_model_path: lmsys/sglang-EAGLE3-Llama-4-Scout-17B-16E-Instruct-v1
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,91 @@
# Cookbook auto benchmark config for mimo v2 flash.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: src/components/autoregressive/MiMoConfigGenerator/index.js
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
trust_remote_code: true
max_running_requests: 128
chunked_prefill_size: 16384
model_loader_extra_config: '{"enable_multithread_load": "true","num_threads":
64}'
attention_backend: fa3
reasoning_parser: qwen3
tool_call_parser: mimo
model_path: XiaomiMiMo/MiMo-V2-Flash
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
command_prefix:
- python3
- -m
- sglang.launch_server
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: XiaomiMiMo/MiMo-V2-Flash
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/mimo-v2-flash
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,77 @@
# Cookbook auto benchmark config for minimax m2.1.
# Cookbook does not expose an H100/H200 default for this model; this config follows the cookbook MI300X baseline instead.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: src/components/autoregressive/MiniMaxM2ConfigGenerator/index.js
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3
base_flags:
tp_size: 4
trust_remote_code: true
model_path: MiniMaxAI/MiniMax-M2.1
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
command_prefix:
- sglang
- serve
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: MiniMaxAI/MiniMax-M2.1
max_concurrency:
- null
- 8
- 16
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 8.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/minimax-m2.1
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,83 @@
# Cookbook auto benchmark config for minimax m2.5.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: src/components/autoregressive/MiniMaxM25ConfigGenerator/index.js
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3
base_flags:
tp_size: 4
trust_remote_code: true
model_path: MiniMaxAI/MiniMax-M2.5
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
ep_size:
- 1
- 4
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: MiniMaxAI/MiniMax-M2.5
max_concurrency:
- null
- 8
- 16
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/minimax-m2.5
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,77 @@
# Cookbook auto benchmark config for ministral 3 8b instruct 2512.
# Cookbook does not expose an H100/H200 default for this model; this config follows the cookbook MI300X baseline instead.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: src/components/autoregressive/Ministral3ConfigGenerator/index.js
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: '0'
base_flags:
trust_remote_code: true
tool_call_parser: mistral
model_path: mistralai/Ministral-3-8B-Instruct-2512
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
command_prefix:
- sglang
- serve
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: mistralai/Ministral-3-8B-Instruct-2512
max_concurrency:
- null
- 16
- 32
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 16.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/ministral-3-8b-instruct-2512
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,79 @@
# Cookbook auto benchmark config for mistral small 4 119b 2603.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.8/mistral-small-4.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1
base_flags:
tp_size: 2
model_path: mistralai/Mistral-Small-4-119B-2603
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: mistralai/Mistral-Small-4-119B-2603
max_concurrency:
- null
- 8
- 16
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 6.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/mistral-small-4-119b-2603
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,81 @@
# Cookbook auto benchmark config for nemotron 3 nano 30b a3b bf16.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.6/nemotron.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: '0'
base_flags:
tp_size: 1
trust_remote_code: true
kv_cache_dtype: fp8_e4m3
model_path: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
max_concurrency:
- null
- 16
- 32
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 16.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/nemotron-3-nano-30b-a3b-bf16
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,81 @@
# Cookbook auto benchmark config for nemotron 3 super 120b a12b bf16.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.8/nemotron-super.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3
base_flags:
tp_size: 4
trust_remote_code: true
kv_cache_dtype: fp8_e4m3
model_path: nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
max_concurrency:
- null
- 8
- 16
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 6.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/nemotron-3-super-120b-a12b-bf16
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,83 @@
# Cookbook auto benchmark config for qwen3 235b a22b.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.6/qwen.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
model_path: Qwen/Qwen3-235B-A22B
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
ep_size:
- 1
- 4
- 8
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: Qwen/Qwen3-235B-A22B
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/qwen3-235b-a22b
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,82 @@
# Cookbook auto benchmark config for qwen3 coder 480b a35b instruct.
# Cookbook does not expose an H100/H200 default for this model; this config follows the cookbook B200 baseline instead.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: src/components/autoregressive/Qwen3CoderConfigGenerator/index.js
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
ep_size: 2
moe_runner_backend: triton
model_path: Qwen/Qwen3-Coder-480B-A35B-Instruct
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- flashinfer
decode_attention_backend:
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
ep_size:
- 1
- 2
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: Qwen/Qwen3-Coder-480B-A35B-Instruct
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/qwen3-coder-480b-a35b-instruct
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,79 @@
# Cookbook auto benchmark config for qwen3 coder next.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.8/qwen3codernext.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1
base_flags:
tp_size: 2
model_path: Qwen/Qwen3-Coder-Next
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: Qwen/Qwen3-Coder-Next
max_concurrency:
- null
- 8
- 16
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 12.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/qwen3-coder-next
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,82 @@
# Cookbook auto benchmark config for qwen3 next 80b a3b instruct.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.6/qwen3next.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1
base_flags:
tp_size: 2
model_path: Qwen/Qwen3-Next-80B-A3B-Instruct
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
ep_size:
- 1
- 2
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: Qwen/Qwen3-Next-80B-A3B-Instruct
max_concurrency:
- null
- 8
- 16
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 12.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/qwen3-next-80b-a3b-instruct
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,83 @@
# Cookbook auto benchmark config for qwen35 397b a17b fp8.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.8/qwen35.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3
base_flags:
tp_size: 4
model_path: Qwen/Qwen3.5-397B-A17B-FP8
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
ep_size:
- 1
- 4
- 8
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: Qwen/Qwen3.5-397B-A17B-FP8
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 4.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/qwen35-397b-a17b-fp8
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,79 @@
# Cookbook auto benchmark config for ring 2.5 1t.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.8/ring25.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3,4,5,6,7
base_flags:
tp_size: 8
model_path: inclusionAI/Ring-2.5-1T
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 32
- 48
- 64
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: inclusionAI/Ring-2.5-1T
max_concurrency:
- null
- 4
- 8
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 2.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/ring-2.5-1t
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,83 @@
# Cookbook auto benchmark config for step 3.5 flash.
# Baseline follows the cookbook H200 default command.
# If you run on larger-memory or stronger GPUs, you can often reduce tp/ep/pp scale or total GPU count.
# Replace server.base_flags.model_path with a local model path if the weights are already present on disk.
# These configs default to synthetic random traffic. Use your real workload data if you want production-faithful tuning.
# Each input_len/output_len pair defines one benchmark scenario, for example [1000, 1000] and [8000, 1000].
# Source: data/models/generated/v0.5.8/step35.yaml
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: 0,1,2,3
base_flags:
tp_size: 4
trust_remote_code: true
model_path: stepfun-ai/Step-3.5-Flash
mem_fraction_static: 0.82
schedule_policy: lpm
search_space:
prefill_attention_backend:
- fa3
- flashinfer
decode_attention_backend:
- fa3
- flashinfer
chunked_prefill_size:
- 4096
- 8192
max_running_requests:
- 64
- 96
- 128
ep_size:
- 1
- 4
dataset:
kind: random
num_prompts: 80
scenario_names:
- chat
- summarization
input_len:
- 1000
- 8000
output_len:
- 1000
- 1000
benchmark:
backend: auto
tokenizer: stepfun-ai/Step-3.5-Flash
max_concurrency:
- null
- 8
- 16
extra_request_body:
temperature: 0.0
qps:
lower: 0.25
upper: 8.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/cookbook-llm/step-3.5-flash
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
draft_model_path: ''
search_space:
speculative_num_steps:
- 3
- 5
speculative_eagle_topk:
- 1
- 2
speculative_num_draft_tokens:
- 4
- 8
@@ -0,0 +1,62 @@
server:
launch: true
host: 127.0.0.1
port: 30000
env:
CUDA_VISIBLE_DEVICES: "0"
# Optional parallel search sugar:
# parallel:
# # dp_size is auto-derived as visible_gpus / (tp_size * pp_size).
# tp: [4, 2]
# # pp_size: [1, 2]
base_flags:
# Replace with a local model path if you have one.
model_path: Qwen/Qwen3-32B
tp_size: 1
trust_remote_code: true
# Keep the baseline close to a pure-TP launch command.
mem_fraction_static: 0.82
schedule_policy: lpm
# Keep CUDA graph enabled for normal performance tuning.
search_space:
prefill_attention_backend: [fa3, flashinfer]
decode_attention_backend: [fa3, flashinfer]
chunked_prefill_size: [4096, 8192]
max_running_requests: [64, 96, 128]
dataset:
kind: random
num_prompts: 80
scenario_names: [chat, summarization]
input_len: [1000, 8000]
output_len: [1000, 1000]
benchmark:
backend: auto
# Replace with a local tokenizer/model path if you have one.
tokenizer: Qwen/Qwen3-32B
max_concurrency: [null, 16, 32]
extra_request_body:
temperature: 0.0
qps:
lower: 1.0
upper: 12.0
tolerance: 0.1
sla:
max_ttft_ms: 1500
max_tpot_ms: 30
output_dir: ./auto_benchmark_results/qwen3-32b
search:
tier: 2
max_candidates: 8 # Cap search breadth so full searches stay tractable.
resume: true
speculative:
enabled: false
algorithm: EAGLE
# Fill this only if you explicitly want a second-stage speculative search.
draft_model_path: ""
search_space:
speculative_num_steps: [3, 5]
speculative_eagle_topk: [1, 4]
speculative_num_draft_tokens: [4, 8]