[Simulator] Add high-fidelity CPU-based inference simulator (#33824)
Co-authored-by: zhouhaizhu.zhz <zhouhaizhu.zhz@alibaba-inc.com> Co-authored-by: LinSiyuan814 <linsiyuan.lsy@alibaba-inc.com> Co-authored-by: hzh0425 <hzh0425@apache.org>
This commit is contained in:
co-authored by
zhouhaizhu.zhz
LinSiyuan814
hzh0425
parent
a5f07b1241
commit
59799a3687
@@ -0,0 +1,258 @@
|
||||
# SGLang Simulator
|
||||
|
||||
SGLang Simulator reuses SGLang's scheduler and cache implementation while
|
||||
replacing model forward execution with a latency predictor. It supports
|
||||
timestamped trace replay, synthetic workloads, hierarchical cache simulation,
|
||||
and serving-compatible metrics without loading model weights.
|
||||
|
||||
See the [SGLang Simulator advanced-feature guide](../../docs/docs/advanced_features/sglang_simulator.mdx)
|
||||
for the user-facing setup and serving workflow.
|
||||
|
||||
## Compatibility
|
||||
|
||||
SGLang Simulator tracks the current SGLang `main` branch and maintains compatibility
|
||||
with recent SGLang releases. The current integration is validated with `v0.5.16`,
|
||||
`v0.5.17`, `v0.5.18`, and `main`. Compatibility code uses API and capability
|
||||
checks instead of branching on version numbers.
|
||||
|
||||
## Requirements
|
||||
|
||||
- A compatible SGLang checkout. The simulator uses the SGLang source from the
|
||||
same monorepo checkout.
|
||||
- A local model directory containing model configuration files. Tokenizer files
|
||||
are also required unless tokenizer initialization is disabled.
|
||||
- Predictor data for AIConfigurator, ML, or replay mode.
|
||||
|
||||
Use an official SGLang image matching the checkout when validating GPU and
|
||||
runtime compatibility.
|
||||
|
||||
## Installation
|
||||
|
||||
From the SGLang repository:
|
||||
|
||||
```bash
|
||||
pip install -e tools/sglang-simulator
|
||||
```
|
||||
|
||||
The simulator does not install or pin a second `sglang` package. Run it from a
|
||||
checkout whose `python/sglang` package is available on `PYTHONPATH`, or from a
|
||||
matching official SGLang image.
|
||||
|
||||
AIConfigurator is optional. Install it separately when using the
|
||||
`aiconfigurator` predictor. The `aic` extra pins AIConfigurator to the exact release
|
||||
validated with the simulator so upstream API changes cannot silently alter an
|
||||
installation. In a clean virtual environment, install the extra with:
|
||||
|
||||
```bash
|
||||
pip install -e "tools/sglang-simulator[aic]"
|
||||
```
|
||||
|
||||
In an existing SGLang image, install the same pin without dependency resolution
|
||||
to avoid replacing its NumPy/CUDA stack:
|
||||
|
||||
```bash
|
||||
pip install --no-deps "aiconfigurator==0.10.0"
|
||||
```
|
||||
|
||||
Upgrade this pin only after rerunning the AIC predictor and compatibility tests.
|
||||
|
||||
## Quick start
|
||||
|
||||
The maintained tests define the supported first-version scope:
|
||||
|
||||
- [`test/test_simulation_sglang_runner.py`](test/test_simulation_sglang_runner.py):
|
||||
direct Python use of the repository-level
|
||||
[`SGLangBenchmarkRunner`](../../benchmark/simulator/bench_runner.py);
|
||||
- [`test/test_simulation_sglang_serving.py`](test/test_simulation_sglang_serving.py):
|
||||
server plus benchmark-client use through the HTTP serving path with AIC, ML,
|
||||
and replay predictors and ShareGPT or timestamped traffic;
|
||||
- [`test/test_simulation_offline_blocking.py`](test/test_simulation_offline_blocking.py):
|
||||
equivalent logical results in `OFFLINE` and `BLOCKING` modes;
|
||||
- [`test/test_simulation_cache_hit_ratio.py`](test/test_simulation_cache_hit_ratio.py):
|
||||
reusable-prefix accounting and cache-tier hit metrics across repeated runs.
|
||||
|
||||
From `tools/sglang-simulator`:
|
||||
|
||||
```bash
|
||||
python3 -m pytest -q test/test_simulation_sglang_runner.py
|
||||
python3 -m pytest -q test/test_simulation_sglang_serving.py
|
||||
```
|
||||
|
||||
Read these tests as the minimal maintained examples for constructing a dataset,
|
||||
running a benchmark, starting a simulator server, sending programmatic, ShareGPT,
|
||||
or timestamped traffic, comparing execution modes, and collecting request,
|
||||
latency, throughput, and prefix-cache metrics.
|
||||
|
||||
## Serving mode
|
||||
|
||||
Choose a fresh output directory and export it in the server terminal before
|
||||
starting the server:
|
||||
|
||||
```bash
|
||||
export SGLANG_USE_CPU_ENGINE=1
|
||||
export CUDA_VISIBLE_DEVICES=""
|
||||
export SGLANG_SIMULATOR_OUTPUT_MODE=OFFLINE
|
||||
export SIMULATOR_OUTPUT_DIR=/tmp/sglang-simulator-serving-001
|
||||
test ! -e "$SIMULATOR_OUTPUT_DIR"
|
||||
export SGLANG_SIMULATOR_OUTPUT_DIR="$SIMULATOR_OUTPUT_DIR"
|
||||
|
||||
python3 -m sglang_simulator.simulation.sglang.launch_server \
|
||||
--model-path /absolute/path/to/model \
|
||||
--sim-config-path /absolute/path/to/simulator.json \
|
||||
--port 30000
|
||||
```
|
||||
|
||||
In the benchmark terminal, export the same output directory before sending
|
||||
timestamped traffic with the simulator-aware benchmark adapter:
|
||||
|
||||
```bash
|
||||
cd /path/to/sglang
|
||||
export SIMULATOR_OUTPUT_DIR=/tmp/sglang-simulator-serving-001
|
||||
export SGLANG_SIMULATOR_OUTPUT_DIR="$SIMULATOR_OUTPUT_DIR"
|
||||
|
||||
python3 benchmark/simulator/bench_serving.py \
|
||||
--simulator-mode offline \
|
||||
--backend sglang \
|
||||
--base-url http://127.0.0.1:30000 \
|
||||
--model /absolute/path/to/model \
|
||||
--dataset-name autobench \
|
||||
--dataset-path /absolute/path/to/trace.jsonl \
|
||||
--use-trace-timestamps \
|
||||
--num-prompts 100 \
|
||||
--profile \
|
||||
--output-file "$SIMULATOR_OUTPUT_DIR/benchmark.json"
|
||||
```
|
||||
|
||||
The server and benchmark are separate processes, so exporting
|
||||
`SGLANG_SIMULATOR_OUTPUT_DIR` in the server terminal does not configure the
|
||||
benchmark terminal. The benchmark adapter reads `metrics.json` from this path
|
||||
after profiling and uses those server-side logical-time metrics for its serving
|
||||
table and output file. If the benchmark points at another directory, it may show
|
||||
unrelated stale metrics or client wall-clock values. Use the same fresh path in
|
||||
both terminals for every run.
|
||||
|
||||
The simulator always runs the SGLang runtime with `tp_size=ep_size=dp_size=pp_size=1`
|
||||
and both attention/decode context-parallel sizes set to `1`. Parallel CLI options
|
||||
accepted by SGLang are therefore ignored by this simulator entry point. This keeps
|
||||
simulator-only CPU work single-process; it does not change the modeled deployment.
|
||||
Set the real deployment topology under `scheduler` in `--sim-config-path`. That
|
||||
topology drives predictor and cache-resource modeling without launching physical
|
||||
parallel workers.
|
||||
|
||||
Other server options are normal SGLang command-line arguments. For direct Python
|
||||
integration, see
|
||||
[`test_simulation_sglang_runner.py`](test/test_simulation_sglang_runner.py); for the
|
||||
process/HTTP path, see
|
||||
[`test_simulation_sglang_serving.py`](test/test_simulation_sglang_serving.py).
|
||||
|
||||
## Simulation modes
|
||||
|
||||
| Mode | Behavior |
|
||||
|---|---|
|
||||
| `OFFLINE` | Advances the simulator's logical clock without sleeping. |
|
||||
| `BLOCKING` | Sleeps for predicted forward and visible L2-to-L1 load latency. |
|
||||
|
||||
Use server-side simulator metrics for comparisons. Client wall-clock duration is
|
||||
not the simulated timeline in `OFFLINE` mode. When using the benchmark adapter,
|
||||
make sure its `SGLANG_SIMULATOR_OUTPUT_DIR` matches the server's output directory
|
||||
so the printed table and `benchmark.json` are sourced from the current run's
|
||||
`metrics.json`.
|
||||
|
||||
## Configuration
|
||||
|
||||
A simulator configuration has three sections:
|
||||
|
||||
```json
|
||||
{
|
||||
"platform": {
|
||||
"accelerator": {"name": "h20_sxm"},
|
||||
"disk_read_bandwidth_gb": 8,
|
||||
"disk_write_bandwidth_gb": 8,
|
||||
"memory_read_bandwidth_gb": 64,
|
||||
"memory_write_bandwidth_gb": 64,
|
||||
"num_device_per_node": 1
|
||||
},
|
||||
"predictor": {
|
||||
"name": "replay",
|
||||
"database_path": "/absolute/path/to/replay_table.json"
|
||||
},
|
||||
"scheduler": {
|
||||
"tp_size": 4,
|
||||
"ep_size": 4,
|
||||
"dp_size": 1,
|
||||
"pp_size": 1,
|
||||
"cp_size": 1,
|
||||
"cp_style": "none",
|
||||
"data_type": "BF16",
|
||||
"kv_cache_data_type": "BF16",
|
||||
"backend_name": "sglang"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
- `platform` describes the simulated accelerator and storage bandwidth.
|
||||
- `predictor` selects forward-latency prediction.
|
||||
- `scheduler` describes the real target deployment topology and backend metadata.
|
||||
`tp_size`, `ep_size`, `dp_size`, `pp_size`, and `cp_size` are modeled values;
|
||||
they do not launch physical workers. For AIConfigurator, `tp_size` is converted
|
||||
to attention TP after removing modeled DP and CP, while `cp_size` is passed as
|
||||
AIConfigurator context parallelism. `cp_style` uses the AIConfigurator values
|
||||
such as `none`, `allgather`, `ulysses`, or `ring`. Decode-only `dcp_size` has no
|
||||
separate AIConfigurator field and is not modeled yet.
|
||||
|
||||
### Prefix-cache accuracy
|
||||
|
||||
Prefix-cache hit accuracy is highly sensitive to `max_total_tokens`. It controls
|
||||
the simulated device KV-cache capacity and participates in hierarchical host-cache
|
||||
sizing, so a mismatch changes eviction timing and device, host, and storage hit
|
||||
attribution. For deployment-faithful results, copy `max_total_num_tokens=N` from
|
||||
the real SGLang server startup log and launch the simulator with
|
||||
`--max-total-tokens N`. Avoid relying on a separately estimated capacity when
|
||||
comparing the simulator with production traces.
|
||||
|
||||
Supported predictors:
|
||||
|
||||
| Predictor | Purpose |
|
||||
|---|---|
|
||||
| `aiconfigurator` | Operator and module performance-database estimation. |
|
||||
| `ml` | A trained sklearn-compatible 18-feature latency model. |
|
||||
| `replay` | Exact or nearest-neighbor batch-composition replay. |
|
||||
|
||||
Relative predictor paths are resolved from the simulator configuration location.
|
||||
Environment variables in paths use `${NAME}` syntax.
|
||||
|
||||
## Workload formats
|
||||
|
||||
The Autobench trace format uses timestamps in milliseconds:
|
||||
|
||||
```json
|
||||
{"prompt":[1,2,3],"prompt_len":3,"output_len":1,"timestamp":200}
|
||||
```
|
||||
|
||||
Random and ShareGPT workloads are also supported by the runner API and serving
|
||||
benchmark paths.
|
||||
|
||||
## Validation
|
||||
|
||||
Run the CPU compatibility and unit tests from the repository root:
|
||||
|
||||
```bash
|
||||
pip install -e tools/sglang-simulator
|
||||
python3 -m pytest -q tools/sglang-simulator/test/test_simulation_sglang_runner.py
|
||||
python3 -m pytest -q tools/sglang-simulator/test/test_simulation_sglang_serving.py
|
||||
```
|
||||
|
||||
Run the two files as separate pytest commands because the runner test installs
|
||||
process-global simulator hooks and state.
|
||||
|
||||
Run repository checks before submitting:
|
||||
|
||||
```bash
|
||||
git ls-files -z tools/sglang-simulator | \
|
||||
xargs -0 env SKIP=no-commit-to-branch pre-commit run --files
|
||||
```
|
||||
|
||||
Runtime changes should also be validated in a matching official SGLang image
|
||||
with both `OFFLINE` and `BLOCKING` modes. Predictor changes should report
|
||||
step-level error, and scheduler or cache changes should compare request latency,
|
||||
throughput, and prefix-cache reuse against measured traces.
|
||||
Reference in New Issue
Block a user