`--attention-backend` is one field of three: a launch that sets only
`--prefill-attention-backend` or `--decode-attention-backend` leaves the base
field at `None`. Seven decisions read that base field alone and therefore
answered from a field the operator never set. `attention_backends()` is the
pair with the base-field fallback already applied, so each site now asks it for
the half it actually needs:
- `inkling_common/attn` assembles backend-specific kwargs (rel_bias / score
mods) and gates its fused prologue; the backend those describe is the one
`self.attn` dispatches to, so `serving_attention_backend()` selects the pair
member by `forward_batch.forward_mode`, mirroring
`HybridAttnBackend._select_backend` exactly -- draft-extend routes through
the prefill branch like the dispatcher does -- and preferring the
runner-stamped pair, so a draft runner answers with its own backend. That
preference only works if every backend that can enter a ForwardContext
carries the stamp, so `DraftBackendFactory._create_backend` now stamps its
products with the backend it resolved (draft override first), and the
draft-extend conv-sidecar wrapper copies the wrapped backend's stamp -- the
replacement backends the spec workers install had no stamp at all and fell
back to the target's configured pair.
- The chunked-prefix-cache gate is a *prefill* feature -> prefill half. Reading
the base field switched the feature off for every prefill-only configuration.
- `init_deterministic_inference_config` maps *prefill* knobs
(SPLIT_TILE / PREFILL_TRUNCATION_ALIGN) -> prefill half; the map missed and
left truncation unset.
- `two_batch_overlap` computes extend positions -> prefill half.
- mrope's interleaved-rope kernel runs in both phases -> both halves must
support triton. This one is not conservative when it misreads:
`support_triton(None)` answers **True**, so a `--prefill-attention-backend
torch_native` launch took the triton path.
- The req-to-token writer has one caller, `alloc_for_extend` -> prefill half;
its fallback pays several `.item()` syncs per request, so gating it on the
decode half too would send every extend of a mixed launch through the slow
path. `get_last_loc` (the spec-decode allocator's helper) keeps the
both-halves reading: verify tokens are served by either half depending on
`speculative_attention_mode`.
- The flashinfer version floor is a guard; it never fired for a launch that
pinned flashinfer through a split field.
One more site the census found is not converted here: `gpt_oss` derives its
`sinks` parameter dtype from the backend, and a single parameter dtype cannot
serve a split pair (FA4 asserts bfloat16, trtllm_mha consumes float32), so
that one is a behaviour question rather than a config-source one and is fixed
in its own PR.
`test_split_attention_backend_decisions.py` pins the callable decisions by
calling them under a split-only publish, and pins the remaining ones
statically -- the file/why map fails if any of them goes back to the base field
(reverse-verified). It also asserts the `support_triton(None) is True` trap the
sweep exists for.
The stamp comes from the constructor, not the request: every factory leaf
answers ("effective_name", backend), because several map entries do not build
what their key says -- cutedsl_mla draft-extend builds the trtllm-mla backend,
"nsa" is a deprecated alias building dsa, and the hybrid-linear entries pick
fa3/intel_amx/triton by host, which no static rename table can express (a
review catch: on Blackwell the alias stamp reached Inkling's per-forward
kwargs assembly, which asserts a concrete kernel name, and crashed the first
draft-extend forward). The stamping is pinned by unit tests, not only by a
spec e2e: removing the child-stamping loop, stamping an alias from a leaf, or
dropping the wrapper copy goes red (reverse-verified), and a static guard
walks the factory source asserting no leaf answers an alias name. The child loop states its contract explicitly --
`create_decode_backend` passes `stamps_children=True` because its products
are per-step containers by construction, so a container without
`attn_backends` raises instead of being silently skipped by a defensive
probe. The `_version` invalidation names its contract (autograd's in-place
counter: private, chosen because it is the only per-tensor signal that ticks
on copy_-style updates; removal fails loudly). The version-floor guard's file
joins the pair-reader ratchet, and the one runner-seed chain read sharing the
backend's __init__ (`speculative_eagle_topk`) reads the spec bag.
Test and Continuous Integration (CI) System in SGLang
This page covers principles and essentials: folder layout, how to run tests, registration, and suite selection. For complete references, see the skill guides:
- Writing tests — templates, fixtures, model selection, complete suite tables, checklist:
.claude/skills/write-sglang-test/SKILL.md - CI pipeline internals — stage flow diagrams, fast-fail layers, gating, partitioning, execution modes, debugging failures:
.claude/skills/ci-workflow-guide/SKILL.md
CI Pipeline Overview
The CI pipeline runs in three sequential stages: A (pre-flight, ~3 min) → B (basic, ~30 min) → C (advanced, ~30 min). Kernel and multimodal-gen tests run in parallel with stage B. For details on stage gating, fast-fail mechanisms, execution modes (PR vs scheduled vs /rerun-stage), and debugging CI failures, see the CI workflow guide.
Folder Organization
registered/: CI test files, auto-discovered byrun_suite.py. Most tests live here. JIT kernel tests are an exception (see below).manual/: Non-CI tests for local debugging or special setups.run_suite.py: CI runner — scansregistered/and JIT kernel directories.
The system supports both unittest and pytest. The launcher runs python filename.py -f with failfast enabled by default.
Make sure your file ends with exactly one of:
# for unittest
if __name__ == "__main__":
unittest.main()
# for pytest
if __name__ == "__main__":
import sys
sys.exit(pytest.main([__file__]))
Do not add custom argparse or modify sys.argv before these calls — the CI runner appends -f for failfast.
Run Tests Locally
# Single file
python3 test/registered/core/test_srt_endpoint.py
# Single test method
python3 test/registered/core/test_srt_endpoint.py TestSRTEndpoint.test_simple_decode
# Single JIT kernel test
python3 test/registered/jit/test_add_constant.py
# Run a suite
python3 test/run_suite.py --hw cpu --suite base-a-test-cpu
python3 test/run_suite.py --hw cuda --suite base-a-test-1-gpu-small
# Nightly tests (CUDA nightly suites take no --nightly; the stage is in the name)
python3 test/run_suite.py --hw cuda --suite nightly-test-1-gpu-large
# With auto-partitioning (for parallel CI jobs)
python3 test/run_suite.py --hw cuda --suite base-b-test-1-gpu-small \
--auto-partition-id 0 --auto-partition-size 4
CI Registration
Every CI-discovered test file must call a registration function at module level:
from sglang.test.ci.ci_register import register_cuda_ci
register_cuda_ci(est_time=80, stage="base-b", runner_config="1-gpu-small")
Parameters: est_time (seconds), stage + runner_config (target stage and runner pool from scripts/ci/runner_configs.yml), nightly=True (nightly-only), disabled="reason" (temporarily disable).
Keep est_time, stage, runner_config as literal values — run_suite.py collects them by AST parsing.
JIT kernel correctness tests and benchmarks live under test/registered/jit/, same as other registered tests (their helpers stay alongside the kernel source under python/sglang/kernels/jit/ and are imported by absolute path):
- Correctness tests:
test/registered/jit/test_*.py→base-b-kernel-unit-test-1-gpu-large - Benchmarks:
test/registered/jit/benchmark/bench_*.py→base-b-kernel-benchmark-test-1-gpu-large
Choosing a Suite
Use the lightest suite that meets your test's needs. Full suite tables are in the write-sglang-test skill.
| Need | Suite |
|---|---|
| No GPU required | base-a-test-cpu |
| Small GPU (fits 5090, 32GB) | base-b-test-1-gpu-small (most tests go here) |
| Large GPU memory or Hopper features | base-b-test-1-gpu-large |
| JIT kernel correctness | base-b-kernel-unit-test-1-gpu-large |
| JIT kernel benchmarks | base-b-kernel-benchmark-test-1-gpu-large |
| Multi-GPU (2/4/8) | base-b-test-2-gpu-large, base-c-test-* |
| Long-running or experimental | nightly-* suites |
Steps for Adding a Test
See the write-sglang-test skill for templates, fixtures, model selection, and a complete checklist.
Multi-Hardware Backends
This README mostly describes the NVIDIA GPU CI pipeline. Other hardware backends (AMD, NPU) follow the same practices and use the multi-backend registry system. A scheduled job summarizes test coverage across all backends; here is an example run.
Tips
- Learn from existing examples in test/registered.
- Reuse servers — launching is expensive. Share one server across many test methods via
setUpClass. - Use as few GPUs as possible. Prefer 1-GPU runners.
- Each test file should take < 500 seconds; split if longer.
- Each GitHub Actions job should take < 30 minutes; split if longer.
- If tests are too slow for per-commit, consider nightly suites.
Other Notes
Adding New Models to Nightly CI
- Text models: Extend the global model list variables in
test_utils.py. - VLMs: Extend the
MODEL_THRESHOLDSdictionary intest/registered/eval/test_vlms_mmmu_eval.py.