The v2 spec workers got a published `ServerArgs` copy carrying two values: the target's context length and `--speculative-draft-load-format`. Neither is a process-wide config change — each is consumed by exactly one constructor — so the copy, the publish switch around the draft build, and the replay of the target's resolved overrides onto it all go away, and the values travel to the runner that owns them: - **Context length.** `TpModelWorker` already takes it (`context_length=None` keeps `server_args.context_length`); the four v2 draft workers and `build_draft_tp_worker` pass the target's, which every one of them has in scope as `target_worker` / `target_model_config`. - **Load format.** `ModelRunner._draft_load_format()` resolves it for a draft runner and `build_load_config` takes it, so the `LoadConfig` is per-runner. Model code also reads it off the bag while it builds — Inkling replaces per-element noise in its shared-expert scales under dummy loading — so the load is wrapped in a scoped bag override that puts the target's value back. - `skip_tokenizer_init` was on the copy for nobody: `TpModelWorker` already short-circuits the tokenizer for a draft worker (`or self.is_draft_worker`). `PrefillCudaGraphRunner._max_addressable_prefix_len` capped the prefix by `server_args.context_length`, which the copy used to carry for the draft; it now reads the runner's own `model_config.context_len`. That is also more accurate for the target, whose `--context-length` may be unset while the resolved context is shorter than the token table. What stays a variant is the dflash/dspark path's attention backend: backend selection reads it off the config object the draft runner holds, and the resolved gate has to survive the variant's publish. `draft_server_args_overrides` now carries only those fields and says why.
Test and Continuous Integration (CI) System in SGLang
This page covers principles and essentials: folder layout, how to run tests, registration, and suite selection. For complete references, see the skill guides:
- Writing tests — templates, fixtures, model selection, complete suite tables, checklist:
.claude/skills/write-sglang-test/SKILL.md - CI pipeline internals — stage flow diagrams, fast-fail layers, gating, partitioning, execution modes, debugging failures:
.claude/skills/ci-workflow-guide/SKILL.md
CI Pipeline Overview
The CI pipeline runs in three sequential stages: A (pre-flight, ~3 min) → B (basic, ~30 min) → C (advanced, ~30 min). Kernel and multimodal-gen tests run in parallel with stage B. For details on stage gating, fast-fail mechanisms, execution modes (PR vs scheduled vs /rerun-stage), and debugging CI failures, see the CI workflow guide.
Folder Organization
registered/: CI test files, auto-discovered byrun_suite.py. Most tests live here. JIT kernel tests are an exception (see below).manual/: Non-CI tests for local debugging or special setups.run_suite.py: CI runner — scansregistered/and JIT kernel directories.
The system supports both unittest and pytest. The launcher runs python filename.py -f with failfast enabled by default.
Make sure your file ends with exactly one of:
# for unittest
if __name__ == "__main__":
unittest.main()
# for pytest
if __name__ == "__main__":
import sys
sys.exit(pytest.main([__file__]))
Do not add custom argparse or modify sys.argv before these calls — the CI runner appends -f for failfast.
Run Tests Locally
# Single file
python3 test/registered/core/test_srt_endpoint.py
# Single test method
python3 test/registered/core/test_srt_endpoint.py TestSRTEndpoint.test_simple_decode
# Single JIT kernel test
python3 test/registered/jit/test_add_constant.py
# Run a suite
python3 test/run_suite.py --hw cpu --suite base-a-test-cpu
python3 test/run_suite.py --hw cuda --suite base-a-test-1-gpu-small
# Nightly tests
python3 test/run_suite.py --hw cuda --suite nightly-1-gpu --nightly
# With auto-partitioning (for parallel CI jobs)
python3 test/run_suite.py --hw cuda --suite base-b-test-1-gpu-small \
--auto-partition-id 0 --auto-partition-size 4
CI Registration
Every CI-discovered test file must call a registration function at module level:
from sglang.test.ci.ci_register import register_cuda_ci
register_cuda_ci(est_time=80, stage="base-b", runner_config="1-gpu-small")
Parameters: est_time (seconds), stage + runner_config (target stage and runner pool from scripts/ci/runner_configs.yml), nightly=True (nightly-only), disabled="reason" (temporarily disable).
Keep est_time, stage, runner_config as literal values — run_suite.py collects them by AST parsing.
JIT kernel correctness tests and benchmarks live under test/registered/jit/, same as other registered tests (their helpers stay alongside the kernel source under python/sglang/kernels/jit/ and are imported by absolute path):
- Correctness tests:
test/registered/jit/test_*.py→base-b-kernel-unit-test-1-gpu-large - Benchmarks:
test/registered/jit/benchmark/bench_*.py→base-b-kernel-benchmark-test-1-gpu-large
Choosing a Suite
Use the lightest suite that meets your test's needs. Full suite tables are in the write-sglang-test skill.
| Need | Suite |
|---|---|
| No GPU required | base-a-test-cpu |
| Small GPU (fits 5090, 32GB) | base-b-test-1-gpu-small (most tests go here) |
| Large GPU memory or Hopper features | base-b-test-1-gpu-large |
| JIT kernel correctness | base-b-kernel-unit-test-1-gpu-large |
| JIT kernel benchmarks | base-b-kernel-benchmark-test-1-gpu-large |
| Multi-GPU (2/4/8) | base-b-test-2-gpu-large, base-c-test-* |
| Long-running or experimental | nightly-* suites |
Steps for Adding a Test
See the write-sglang-test skill for templates, fixtures, model selection, and a complete checklist.
Multi-Hardware Backends
This README mostly describes the NVIDIA GPU CI pipeline. Other hardware backends (AMD, NPU) follow the same practices and use the multi-backend registry system. A scheduled job summarizes test coverage across all backends; here is an example run.
Tips
- Learn from existing examples in test/registered.
- Reuse servers — launching is expensive. Share one server across many test methods via
setUpClass. - Use as few GPUs as possible. Prefer 1-GPU runners.
- Each test file should take < 500 seconds; split if longer.
- Each GitHub Actions job should take < 30 minutes; split if longer.
- If tests are too slow for per-commit, consider nightly suites.
Other Notes
Adding New Models to Nightly CI
- Text models: Extend the global model list variables in
test_utils.py. - VLMs: Extend the
MODEL_THRESHOLDSdictionary intest/registered/eval/test_vlms_mmmu_eval.py.