Files
sglang/test
Cheng Wan d13d5c03ab config: decisions keyed on the attention backend read the configured pair
`--attention-backend` is one field of three: a launch that sets only
`--prefill-attention-backend` or `--decode-attention-backend` leaves the base
field at `None`. Seven decisions read that base field alone and therefore
answered from a field the operator never set. `attention_backends()` is the
pair with the base-field fallback already applied, so each site now asks it for
the half it actually needs:

- `inkling_common/attn` assembles backend-specific kwargs (rel_bias / score
  mods) and gates its fused prologue; the backend those describe is the one
  `self.attn` dispatches to, so `serving_attention_backend()` selects the pair
  member by `forward_batch.forward_mode`, mirroring
  `HybridAttnBackend._select_backend` exactly -- draft-extend routes through
  the prefill branch like the dispatcher does -- and preferring the
  runner-stamped pair, so a draft runner answers with its own backend. That
  preference only works if every backend that can enter a ForwardContext
  carries the stamp, so `DraftBackendFactory._create_backend` now stamps its
  products with the backend it resolved (draft override first), and the
  draft-extend conv-sidecar wrapper copies the wrapped backend's stamp -- the
  replacement backends the spec workers install had no stamp at all and fell
  back to the target's configured pair.
- The chunked-prefix-cache gate is a *prefill* feature -> prefill half. Reading
  the base field switched the feature off for every prefill-only configuration.
- `init_deterministic_inference_config` maps *prefill* knobs
  (SPLIT_TILE / PREFILL_TRUNCATION_ALIGN) -> prefill half; the map missed and
  left truncation unset.
- `two_batch_overlap` computes extend positions -> prefill half.
- mrope's interleaved-rope kernel runs in both phases -> both halves must
  support triton. This one is not conservative when it misreads:
  `support_triton(None)` answers **True**, so a `--prefill-attention-backend
  torch_native` launch took the triton path.
- The req-to-token writer has one caller, `alloc_for_extend` -> prefill half;
  its fallback pays several `.item()` syncs per request, so gating it on the
  decode half too would send every extend of a mixed launch through the slow
  path. `get_last_loc` (the spec-decode allocator's helper) keeps the
  both-halves reading: verify tokens are served by either half depending on
  `speculative_attention_mode`.
- The flashinfer version floor is a guard; it never fired for a launch that
  pinned flashinfer through a split field.

One more site the census found is not converted here: `gpt_oss` derives its
`sinks` parameter dtype from the backend, and a single parameter dtype cannot
serve a split pair (FA4 asserts bfloat16, trtllm_mha consumes float32), so
that one is a behaviour question rather than a config-source one and is fixed
in its own PR.

`test_split_attention_backend_decisions.py` pins the callable decisions by
calling them under a split-only publish, and pins the remaining ones
statically -- the file/why map fails if any of them goes back to the base field
(reverse-verified). It also asserts the `support_triton(None) is True` trap the
sweep exists for.

The stamp comes from the constructor, not the request: every factory leaf
answers ("effective_name", backend), because several map entries do not build
what their key says -- cutedsl_mla draft-extend builds the trtllm-mla backend,
"nsa" is a deprecated alias building dsa, and the hybrid-linear entries pick
fa3/intel_amx/triton by host, which no static rename table can express (a
review catch: on Blackwell the alias stamp reached Inkling's per-forward
kwargs assembly, which asserts a concrete kernel name, and crashed the first
draft-extend forward). The stamping is pinned by unit tests, not only by a
spec e2e: removing the child-stamping loop, stamping an alias from a leaf, or
dropping the wrapper copy goes red (reverse-verified), and a static guard
walks the factory source asserting no leaf answers an alias name. The child loop states its contract explicitly --
`create_decode_backend` passes `stamps_children=True` because its products
are per-step containers by construction, so a container without
`attn_backends` raises instead of being silently skipped by a defensive
probe. The `_version` invalidation names its contract (autograd's in-place
counter: private, chosen because it is the only per-tensor signal that ticks
on copy_-style updates; removal fails loudly). The version-floor guard's file
joins the pair-reader ratchet, and the one runner-seed chain read sharing the
backend's __init__ (`speculative_eagle_topk`) reads the spec bag.
2026-08-15 00:38:00 -07:00
..

Test and Continuous Integration (CI) System in SGLang

This page covers principles and essentials: folder layout, how to run tests, registration, and suite selection. For complete references, see the skill guides:

CI Pipeline Overview

The CI pipeline runs in three sequential stages: A (pre-flight, ~3 min) → B (basic, ~30 min) → C (advanced, ~30 min). Kernel and multimodal-gen tests run in parallel with stage B. For details on stage gating, fast-fail mechanisms, execution modes (PR vs scheduled vs /rerun-stage), and debugging CI failures, see the CI workflow guide.

Folder Organization

  • registered/: CI test files, auto-discovered by run_suite.py. Most tests live here. JIT kernel tests are an exception (see below).
  • manual/: Non-CI tests for local debugging or special setups.
  • run_suite.py: CI runner — scans registered/ and JIT kernel directories.

The system supports both unittest and pytest. The launcher runs python filename.py -f with failfast enabled by default.

Make sure your file ends with exactly one of:

# for unittest
if __name__ == "__main__":
    unittest.main()
# for pytest
if __name__ == "__main__":
    import sys
    sys.exit(pytest.main([__file__]))

Do not add custom argparse or modify sys.argv before these calls — the CI runner appends -f for failfast.

Run Tests Locally

# Single file
python3 test/registered/core/test_srt_endpoint.py

# Single test method
python3 test/registered/core/test_srt_endpoint.py TestSRTEndpoint.test_simple_decode

# Single JIT kernel test
python3 test/registered/jit/test_add_constant.py

# Run a suite
python3 test/run_suite.py --hw cpu --suite base-a-test-cpu
python3 test/run_suite.py --hw cuda --suite base-a-test-1-gpu-small

# Nightly tests (CUDA nightly suites take no --nightly; the stage is in the name)
python3 test/run_suite.py --hw cuda --suite nightly-test-1-gpu-large

# With auto-partitioning (for parallel CI jobs)
python3 test/run_suite.py --hw cuda --suite base-b-test-1-gpu-small \
    --auto-partition-id 0 --auto-partition-size 4

CI Registration

Every CI-discovered test file must call a registration function at module level:

from sglang.test.ci.ci_register import register_cuda_ci

register_cuda_ci(est_time=80, stage="base-b", runner_config="1-gpu-small")

Parameters: est_time (seconds), stage + runner_config (target stage and runner pool from scripts/ci/runner_configs.yml), nightly=True (nightly-only), disabled="reason" (temporarily disable).

Keep est_time, stage, runner_config as literal valuesrun_suite.py collects them by AST parsing.

JIT kernel correctness tests and benchmarks live under test/registered/jit/, same as other registered tests (their helpers stay alongside the kernel source under python/sglang/kernels/jit/ and are imported by absolute path):

  • Correctness tests: test/registered/jit/test_*.pybase-b-kernel-unit-test-1-gpu-large
  • Benchmarks: test/registered/jit/benchmark/bench_*.pybase-b-kernel-benchmark-test-1-gpu-large

Choosing a Suite

Use the lightest suite that meets your test's needs. Full suite tables are in the write-sglang-test skill.

Need Suite
No GPU required base-a-test-cpu
Small GPU (fits 5090, 32GB) base-b-test-1-gpu-small (most tests go here)
Large GPU memory or Hopper features base-b-test-1-gpu-large
JIT kernel correctness base-b-kernel-unit-test-1-gpu-large
JIT kernel benchmarks base-b-kernel-benchmark-test-1-gpu-large
Multi-GPU (2/4/8) base-b-test-2-gpu-large, base-c-test-*
Long-running or experimental nightly-* suites

Steps for Adding a Test

See the write-sglang-test skill for templates, fixtures, model selection, and a complete checklist.

Multi-Hardware Backends

This README mostly describes the NVIDIA GPU CI pipeline. Other hardware backends (AMD, NPU) follow the same practices and use the multi-backend registry system. A scheduled job summarizes test coverage across all backends; here is an example run.

Tips

  • Learn from existing examples in test/registered.
  • Reuse servers — launching is expensive. Share one server across many test methods via setUpClass.
  • Use as few GPUs as possible. Prefer 1-GPU runners.
  • Each test file should take < 500 seconds; split if longer.
  • Each GitHub Actions job should take < 30 minutes; split if longer.
  • If tests are too slow for per-commit, consider nightly suites.

Other Notes

Adding New Models to Nightly CI

  • Text models: Extend the global model list variables in test_utils.py.
  • VLMs: Extend the MODEL_THRESHOLDS dictionary in test/registered/eval/test_vlms_mmmu_eval.py.