[CI] Key scheduled CUDA suites by runner_config instead of hand-written jobs (#34186)
This commit is contained in:
@@ -489,7 +489,7 @@ if torch.cuda.get_device_capability()[0] < 9:
|
||||
JIT kernel correctness tests and benchmarks live under `test/registered/kernels/ops/<group>/` and `test/registered/kernels/benchmark/<group>/`, mirroring the wrapper's group under `python/sglang/kernels/ops/` (NOT inside the `sglang` package -- a `register_*_ci(...)` call anywhere under `python/sglang/` is rejected by the `check-no-registered-tests-in-package` pre-commit hook). Only their test-only helpers (e.g. `benchmark/marker.py`) stay alongside the kernel source under `python/sglang/kernels/jit/` and are imported by absolute path. **CI does not run `pytest` in those directories directly.** The unified runner `test/run_suite.py` discovers every `test_*.py` and `bench_*.py` under `test/registered/`, collects `register_*_ci(...)` calls by **statically parsing each file's AST**, and executes the selected suite. Every test file must register at least one CUDA entry or the collector fails its sanity check.
|
||||
|
||||
- **PR / per-commit CUDA suites** (see `test/run_suite.py` → `PER_COMMIT_SUITES`): JIT unit tests use `base-b-kernel-unit-test-1-gpu-large` on H100 and `base-b-kernel-unit-test-4-gpu-b200` on B200/SM100 paths (see `.github/workflows/pr-test-jit-kernel.yml`). Multi-GPU JIT tests use `base-b-kernel-unit-test-8-gpu-h200`.
|
||||
- **Nightly kernel suite**: `nightly-kernel-1-gpu` with `--nightly` — typically used with `SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1` in CI for expanded parameter grids (see `python/sglang/kernels/jit/utils/common.py` → `should_run_full_tests` / `get_ci_test_range`). Wired in `.github/workflows/nightly-test-nvidia.yml` (e.g. `python3 run_suite.py --hw cuda --suite nightly-kernel-1-gpu --nightly --continue-on-error`).
|
||||
- **Nightly kernel suite**: register with `stage="nightly"` plus the `runner_config` of the machine it needs (e.g. `1-gpu-large`), giving the `nightly-test-1-gpu-large` suite. `.github/workflows/nightly-test-nvidia.yml` sets `SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1` for the whole nightly run, so the expanded parameter grids apply automatically (see `python/sglang/kernels/jit/utils/common.py` → `should_run_full_tests` / `get_ci_test_range`). There is no separate kernel-only nightly job: every nightly test on one machine type shares that machine's suite.
|
||||
|
||||
Registration pattern (module level, **literal** `est_time`, `stage`, and `runner_config` values — required for AST parsing):
|
||||
|
||||
@@ -499,12 +499,12 @@ from sglang.test.ci.ci_register import register_cuda_ci
|
||||
register_cuda_ci(est_time=30, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Optional B200/SM100 registration for tests that cover Blackwell-specific code paths
|
||||
# register_cuda_ci(est_time=30, stage="base-b-kernel-unit", runner_config="4-gpu-b200")
|
||||
# Optional second registration: same file also listed under the nightly kernel suite
|
||||
# (nightly suites use the legacy single-string suite=, not stage/runner_config)
|
||||
# register_cuda_ci(est_time=120, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
# Optional second registration: same file also runs nightly, same form,
|
||||
# stage is just "nightly" there (and no `nightly=True`)
|
||||
# register_cuda_ci(est_time=120, stage="nightly", runner_config="1-gpu-large")
|
||||
```
|
||||
|
||||
CI generates the suite name as `{stage}-test-{runner_config}`, so `stage="base-b-kernel-unit", runner_config="1-gpu-large"` becomes the `base-b-kernel-unit-test-1-gpu-large` suite you pass to `run_suite.py` below — don't put the `-test-` infix in `register_cuda_ci`. The single-string `suite=` form is only for nightly/stress/weekly suites.
|
||||
CI generates the suite name as `{stage}-test-{runner_config}`, so `stage="base-b-kernel-unit", runner_config="1-gpu-large"` becomes the `base-b-kernel-unit-test-1-gpu-large` suite you pass to `run_suite.py` below — don't put the `-test-` infix in `register_cuda_ci`. Nightly uses the same shape with `stage="nightly"`; the single-string `suite=` form is left only for `stress` and non-CUDA pools.
|
||||
|
||||
Keep `est_time`, `stage`, `runner_config`, and `suite` as literal values. `run_suite.py` collects them from the file AST, so computed values and helper wrappers can break CI discovery.
|
||||
|
||||
@@ -661,7 +661,7 @@ cd test && python3 run_suite.py --hw cuda --suite base-b-kernel-benchmark-test-1
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
- **`No CI registry found in ...` from `run_suite.py`**: add a module-level `register_cuda_ci(...)` with literal `est_time`, `stage`, and `runner_config` (and optional `nightly=True`); starred args and non-literal values break AST collection
|
||||
- **`No CI registry found in ...` from `run_suite.py`**: add a module-level `register_cuda_ci(...)` with literal `est_time`, `stage`, and `runner_config`; starred args and non-literal values break AST collection
|
||||
- **JIT compilation fails**: ensure the `.cuh` file is under `python/sglang/kernels/jit/csrc/`; reduce template argument combinations
|
||||
- **CUDA crash / illegal memory access**: `CUDA_LAUNCH_BLOCKING=1`; `compute-sanitizer --tool memcheck python ...`
|
||||
- **Unstable benchmark results**: `marker.do_bench` uses CUDA-graph-based timing by default; set `use_cuda_graph=False` only if the kernel can't be captured. `graph_clone_args` defaults to `"all"`; if you narrow it, it must still cover every *read* tensor — reusing a single buffer keeps it L2-hot and skews results. Keep *write* tensors in it too: they are what sets the rotation count, and a shared output buffer stays L2-hot the same way.
|
||||
|
||||
@@ -56,7 +56,7 @@ A per-commit suite name is **generated** from registration metadata as `{stage}-
|
||||
- **`runner_config`** — a runner-pool key from `scripts/ci/runner_configs.yml`, which maps it to the physical runner label (so `1-gpu-large` runs on `1-gpu-h100`). AMD/NPU use their own keys (e.g. `amd`).
|
||||
- **Suite** — `register_cuda_ci(stage="base-b", runner_config="1-gpu-small")` → `base-b-test-1-gpu-small`, the name you pass to `run_suite.py --suite`. The `-test-` is just the connector; never put it in `register_*_ci`.
|
||||
|
||||
> Legacy single-string `suite=` is only for suites that don't fit that shape — nightly/stress/weekly and some AMD/CPU/NPU pools (e.g. `suite="nightly-kernel-1-gpu", nightly=True`). Per-commit tests always use `stage=` + `runner_config=`.
|
||||
> CUDA nightly uses the same shape with `stage="nightly"` (e.g. `stage="nightly", runner_config="1-gpu-large"` → `nightly-test-1-gpu-large`) and **no** `nightly=True` — the stage name carries the cadence, and setting the flag makes the test silently never run. Legacy single-string `suite=` is left only for `stress` and some AMD/CPU/NPU pools.
|
||||
|
||||
### All CI Suites
|
||||
|
||||
@@ -113,13 +113,14 @@ A per-commit suite name is **generated** from registration metadata as `{stage}-
|
||||
|
||||
#### Nightly
|
||||
|
||||
Nightly suites are listed in `NIGHTLY_SUITES` in [`test/run_suite.py`](../../../test/run_suite.py). They run via `nightly-test-nvidia.yml`, `nightly-test-amd.yml`, and `nightly-test-npu.yml`, not `pr-test.yml`. Examples:
|
||||
Nightly suites are listed in `NIGHTLY_SUITES` in [`test/run_suite.py`](../../../test/run_suite.py). They run via `nightly-test-nvidia.yml`, `nightly-test-amd.yml`, and `nightly-test-npu.yml`, not `pr-test.yml`.
|
||||
|
||||
- `nightly-1-gpu` (CUDA)
|
||||
- `nightly-kernel-1-gpu` (CUDA, JIT kernel full grids)
|
||||
- `nightly-kernel-8-gpu-h200` (CUDA, multi-GPU JIT kernel nightly)
|
||||
- `nightly-8-gpu-h200` (CUDA)
|
||||
- `nightly-eval-vlm-2-gpu` (CUDA)
|
||||
CUDA nightly suites are named `nightly-test-{runner_config}` — one per machine type, holding everything that runs nightly on it. There is no per-purpose split (kernel / eval / perf / precision all share their machine's suite); `auto_partition` splits the work. Examples:
|
||||
|
||||
- `nightly-test-1-gpu-large` (CUDA)
|
||||
- `nightly-test-2-gpu-large` (CUDA)
|
||||
- `nightly-test-8-gpu-h200` (CUDA)
|
||||
- `nightly-test-4-gpu-gb300` (CUDA)
|
||||
- `nightly-amd` (AMD)
|
||||
- `nightly-amd-8-gpu-mi35x` (AMD)
|
||||
- `nightly-1-npu-a3` (NPU)
|
||||
@@ -330,8 +331,8 @@ register_cuda_ci(est_time=80, suite="base-b-test-1-gpu-small")
|
||||
# Per-commit test (large 1-gpu, runs on H100)
|
||||
register_cuda_ci(est_time=120, suite="base-b-test-1-gpu-large")
|
||||
|
||||
# Nightly-only test
|
||||
register_cuda_ci(est_time=200, suite="nightly-1-gpu", nightly=True)
|
||||
# Nightly-only test (same shape as per-commit, stage is just "nightly")
|
||||
register_cuda_ci(est_time=200, stage="nightly", runner_config="1-gpu-large")
|
||||
|
||||
# Multi-backend test (only when testing backend-specific code paths)
|
||||
register_cuda_ci(est_time=80, suite="base-a-test-1-gpu-small")
|
||||
@@ -345,7 +346,7 @@ register_cuda_ci(est_time=80, suite="base-b-test-1-gpu-small", disabled="flaky -
|
||||
Parameters:
|
||||
- `est_time`: estimated runtime in seconds (used for CI partitioning)
|
||||
- `suite`: which CI suite to run in (see suite tables above)
|
||||
- `nightly=True`: for nightly-only tests (default `False` = per-commit)
|
||||
- `nightly=True`: legacy cadence flag, for non-CUDA nightly suites only. CUDA nightly uses `stage="nightly"` and must leave this unset
|
||||
- `disabled="reason"`: temporarily disable with explanation
|
||||
|
||||
**Key principle**: Only add `register_amd_ci` / `register_npu_ci` when the test exercises backend-specific code paths. Common E2E tests just need `register_cuda_ci` — duplicating across backends wastes CI time.
|
||||
@@ -365,12 +366,12 @@ register_cuda_ci(est_time=120, stage="base-b-kernel-unit", runner_config="8-gpu-
|
||||
# Benchmarks in test/registered/jit/benchmark/
|
||||
register_cuda_ci(est_time=6, stage="base-b-kernel-benchmark", runner_config="1-gpu-large")
|
||||
|
||||
# Optional nightly registration — nightly suites use the legacy single-string suite=
|
||||
register_cuda_ci(est_time=120, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=120, suite="nightly-kernel-8-gpu-h200", nightly=True)
|
||||
# Optional nightly registration — same form, stage is just "nightly"
|
||||
register_cuda_ci(est_time=120, stage="nightly", runner_config="1-gpu-large")
|
||||
register_cuda_ci(est_time=120, stage="nightly", runner_config="8-gpu-h200")
|
||||
```
|
||||
|
||||
The `stage` + `runner_config` calls generate suites like `base-b-kernel-unit-test-1-gpu-large`; nightly keeps the legacy `suite=` string. Keep `est_time`, `stage`, `runner_config`, and `suite` as **literal values** — `run_suite.py` collects them by AST parsing.
|
||||
Every call generates a suite named `{stage}-test-{runner_config}`, e.g. `base-b-kernel-unit-test-1-gpu-large` and `nightly-test-1-gpu-large`. Keep `est_time`, `stage`, `runner_config`, and `suite` as **literal values** — `run_suite.py` collects them by AST parsing.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -63,6 +63,15 @@ on:
|
||||
type: boolean
|
||||
default: false
|
||||
|
||||
scheduled:
|
||||
description: 'The caller is a nightly or weekly workflow, not a per-commit one. Switches on per-file timeouts derived from est_time, the full jit_kernel grids, metrics upload, and serial shards unless dispatched with full_parallel.'
|
||||
type: boolean
|
||||
default: false
|
||||
job_timeout_minutes:
|
||||
description: 'Job-level cap, covering install and warmup on top of run_timeout_minutes.'
|
||||
type: string
|
||||
default: '240'
|
||||
|
||||
# Mirror pr-test.yml top-level env. Reusable workflows do NOT inherit caller's
|
||||
# workflow-level env across the workflow_call boundary, so anything pr-test.yml
|
||||
# defines must be redeclared here for the called job to see the same context.
|
||||
@@ -74,6 +83,15 @@ env:
|
||||
SGLANG_ENABLE_ASYNC_ASSERT: ${{ startsWith(inputs.self_name, 'base-a') && 'false' || 'true' }}
|
||||
SGLANG_CUDA_COREDUMP: "1"
|
||||
SGLANG_JIT_DEEPGEMM_FAST_WARMUP: true
|
||||
SGLANG_JIT_KERNEL_RUN_FULL_TESTS: ${{ inputs.scheduled && '1' || '0' }}
|
||||
# is_h200_system() (test_utils.py) raises the server-launch ceiling from 600s
|
||||
# to 3600s. Only the scheduled 8-gpu-h200 suites load models that need it;
|
||||
# per-commit stages keep the short ceiling so a hung launch fails fast.
|
||||
IS_H200: ${{ (inputs.scheduled && inputs.runner_config == '8-gpu-h200') && '1' || '0' }}
|
||||
# Scheduled suites pull hundreds of GB of checkpoints; the hub's 10s etag
|
||||
# default times out on a cold cache.
|
||||
HF_HUB_DOWNLOAD_TIMEOUT: 300
|
||||
HF_HUB_ETAG_TIMEOUT: 300
|
||||
SKIP_PR_TEST_HEALTH_CHECK: ${{ (fromJson(inputs.caller_inputs).skip_pr_test_health_check || fromJson(inputs.caller_inputs).test_parallel_dispatch || fromJson(inputs.caller_inputs).run_all_tests) && 'true' || 'false' }}
|
||||
PR_TEST_BYPASS_MAINTENANCE_ON_MAIN: ${{ github.ref == 'refs/heads/main' && 'true' || 'false' }}
|
||||
USE_VENV: false
|
||||
@@ -90,10 +108,10 @@ jobs:
|
||||
# $b200_runner (see runner_configs.py --map). rdma_devices is exported
|
||||
# below in a setup step via $GITHUB_ENV.
|
||||
runs-on: ${{ fromJson(fromJson(inputs.check_changes).runs_on_map)[inputs.runner_config] }}
|
||||
timeout-minutes: 240
|
||||
timeout-minutes: ${{ fromJson(inputs.job_timeout_minutes) }}
|
||||
strategy:
|
||||
fail-fast: false
|
||||
max-parallel: ${{ fromJson(inputs.partitions)[inputs.self_name].max_parallel }}
|
||||
max-parallel: ${{ (inputs.scheduled && !fromJson(inputs.caller_inputs).full_parallel) && 1 || fromJson(inputs.partitions)[inputs.self_name].max_parallel }}
|
||||
matrix:
|
||||
partition: ${{ fromJson(inputs.partitions)[inputs.self_name].arr }}
|
||||
steps:
|
||||
@@ -165,16 +183,39 @@ jobs:
|
||||
curl --fail --silent --show-error --max-time 15 --retry 3 --retry-delay 2 \
|
||||
"$URL" -o /tmp/partition-model.json
|
||||
|
||||
# Only test_nightly_precision_regression.py reads these; the others are
|
||||
# inert wherever it does not run.
|
||||
- name: Export precision baseline env
|
||||
if: inputs.scheduled
|
||||
env:
|
||||
BASELINE_HF_TOKEN: ${{ secrets.HF_TOKEN_PRECISION_STORE }}
|
||||
run: |
|
||||
{
|
||||
echo "SGLANG_PRECISION_BASELINE_DIR=/tmp/sglang_precision_baselines"
|
||||
echo "SGLANG_PRECISION_HF_REPO=${{ vars.SGLANG_PRECISION_HF_REPO }}"
|
||||
echo "SGLANG_PRECISION_HF_REVISION=${{ vars.SGLANG_PRECISION_HF_REVISION || 'main' }}"
|
||||
echo "SGLANG_PRECISION_COMMIT=${{ github.sha }}"
|
||||
echo "SGLANG_PRECISION_FORCE_UPDATE=${{ fromJson(inputs.caller_inputs).force_baseline_update && '1' || '0' }}"
|
||||
echo "SGLANG_PRECISION_HF_TOKEN=${BASELINE_HF_TOKEN}"
|
||||
} >> "$GITHUB_ENV"
|
||||
|
||||
- name: Run test
|
||||
timeout-minutes: ${{ fromJson(inputs.run_timeout_minutes) }}
|
||||
env:
|
||||
CONTINUE_ON_ERROR_FLAG: ${{ fromJson(inputs.check_changes).continue_on_error == 'true' && '--continue-on-error' || '' }}
|
||||
RUNNER_LABELS: ${{ fromJson(fromJson(inputs.check_changes).runs_on_map)[inputs.runner_config] }}
|
||||
GPU_CONFIG: ${{ inputs.runner_config }}
|
||||
# Read only by test_runai_model_loader.py (nightly 1-gpu); inert elsewhere.
|
||||
# Left unconditional because an `env:` key cannot be omitted, and an empty
|
||||
# value would reach the runai streamer as a malformed limit.
|
||||
RUNAI_STREAMER_MEMORY_LIMIT: 0
|
||||
run: |
|
||||
cd test
|
||||
python3 run_suite.py --hw cuda --suite ${{ inputs.self_name }} \
|
||||
--auto-partition-id ${{ matrix.partition }} \
|
||||
--auto-partition-size ${{ fromJson(inputs.partitions)[inputs.self_name].size }} \
|
||||
--partition-model-file /tmp/partition-model.json \
|
||||
${{ inputs.scheduled && '--timeout-from-est-time' || '' }} \
|
||||
${{ inputs.timeout_per_file && format('--timeout-per-file {0}', inputs.timeout_per_file) || '' }} \
|
||||
$CONTINUE_ON_ERROR_FLAG
|
||||
|
||||
@@ -183,6 +224,26 @@ jobs:
|
||||
timeout-minutes: 10
|
||||
run: python3 -m pytest -q ${{ inputs.extra_pytest_path }}
|
||||
|
||||
- name: Collect performance metrics
|
||||
if: ${{ always() && inputs.scheduled }}
|
||||
run: |
|
||||
python3 scripts/ci/utils/save_metrics.py \
|
||||
--gpu-config ${{ inputs.runner_config }} \
|
||||
--partition ${{ matrix.partition }} \
|
||||
--run-id ${{ github.run_id }} \
|
||||
--output test/metrics-${{ inputs.runner_config }}-partition-${{ matrix.partition }}.json \
|
||||
--search-dir test/performance_results_8_gpu \
|
||||
--search-dir test
|
||||
|
||||
- name: Upload partition metrics
|
||||
if: ${{ always() && inputs.scheduled }}
|
||||
uses: actions/upload-artifact@v4
|
||||
with:
|
||||
name: metrics-${{ inputs.runner_config }}-partition-${{ matrix.partition }}
|
||||
path: test/metrics-${{ inputs.runner_config }}-partition-${{ matrix.partition }}.json
|
||||
retention-days: 5
|
||||
if-no-files-found: ignore
|
||||
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
with:
|
||||
|
||||
@@ -1,3 +1,10 @@
|
||||
# Nightly CUDA tests. One job per runner_config, running that machine's
|
||||
# `nightly-test-{runner_config}` suite -- a test reaches a machine by declaring
|
||||
# `runner_config=` on register_cuda_ci, never by editing this file.
|
||||
#
|
||||
# Jobs go through the same _pr-test-stage.yml as the per-commit suites; what a
|
||||
# scheduled run needs on top is its `scheduled` input, so runner resolution,
|
||||
# rust-ext reuse and shard sizing are not reimplemented here.
|
||||
name: Nightly Test (Nvidia)
|
||||
|
||||
on:
|
||||
@@ -5,32 +12,28 @@ on:
|
||||
- cron: '0 14 */2 * *'
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
job_filter:
|
||||
description: 'Select which job to run (leave empty or "all" to run all jobs)'
|
||||
runner_filter:
|
||||
description: 'Select which runner_config to run (leave empty or "all" to run all)'
|
||||
required: false
|
||||
type: choice
|
||||
default: 'all'
|
||||
options:
|
||||
- 'all'
|
||||
- 'nightly-test-general-1-gpu-h100'
|
||||
- 'nightly-test-general-4-gpu-h100'
|
||||
- 'nightly-test-general-8-gpu-h200'
|
||||
- 'nightly-test-general-8-gpu-h20'
|
||||
- 'nightly-test-general-8-gpu-b200'
|
||||
- 'nightly-test-text-accuracy-2-gpu-h100'
|
||||
- 'nightly-test-text-perf-2-gpu-h100'
|
||||
- 'nightly-test-vlm-accuracy-2-gpu-h100'
|
||||
- 'nightly-test-vlm-perf-2-gpu-h100'
|
||||
- 'nightly-test-perf-4-gpu-b200'
|
||||
- 'nightly-test-perf-8-gpu-b200'
|
||||
- 'nightly-test-specialized-8-gpu-b200'
|
||||
- 'nightly-test-perf-4-gpu-gb300'
|
||||
- 'nightly-test-kernel-1-gpu-h100'
|
||||
- 'nightly-test-diffusion'
|
||||
- 'nightly-test-kernel-8-gpu-h200'
|
||||
- 'nightly-test-precision-8-gpu-h200'
|
||||
- '1-gpu-large'
|
||||
- '2-gpu-large'
|
||||
- '4-gpu-h100'
|
||||
- '4-gpu-b200'
|
||||
- '4-gpu-gb300'
|
||||
- '8-gpu-h200'
|
||||
- '8-gpu-b200'
|
||||
- 'diffusion'
|
||||
full_parallel:
|
||||
description: 'Run all shards of a job at once (faster, but competes with per-commit CI for machines). Off by default: one shard at a time.'
|
||||
required: false
|
||||
type: boolean
|
||||
default: false
|
||||
force_baseline_update:
|
||||
description: 'precision job only: refresh the rolling baseline instead of comparing (sets SGLANG_PRECISION_FORCE_UPDATE=1). Dispatch once after an intentional forward-path dtype/precision change stales the baseline; later scheduled runs compare against it and go green.'
|
||||
description: 'Refresh the precision rolling baseline instead of comparing. Dispatch once after an intentional forward-path precision change stales it; later runs compare against the new one.'
|
||||
required: false
|
||||
type: boolean
|
||||
default: false
|
||||
@@ -41,499 +44,160 @@ on:
|
||||
required: false
|
||||
type: string
|
||||
default: ''
|
||||
job_filter:
|
||||
description: 'Select which job to run (leave empty or "all" to run all jobs)'
|
||||
runner_filter:
|
||||
description: 'Select which runner_config to run (leave empty or "all" to run all)'
|
||||
required: false
|
||||
type: string
|
||||
default: 'all'
|
||||
full_parallel:
|
||||
description: 'Run all shards of a job at once (faster, but competes with per-commit CI for machines). Off by default: one shard at a time.'
|
||||
required: false
|
||||
type: boolean
|
||||
default: false
|
||||
|
||||
concurrency:
|
||||
group: nightly-test-nvidia-${{ inputs.ref || github.ref }}
|
||||
cancel-in-progress: ${{ github.event_name != 'workflow_call' }}
|
||||
|
||||
env:
|
||||
SGLANG_IS_IN_CI: true
|
||||
SGLANG_ENABLE_ASYNC_ASSERT: true
|
||||
SGLANG_CUDA_COREDUMP: "1"
|
||||
HF_HUB_DOWNLOAD_TIMEOUT: 300
|
||||
HF_HUB_ETAG_TIMEOUT: 300
|
||||
permissions:
|
||||
actions: write
|
||||
contents: read
|
||||
issues: read
|
||||
pull-requests: read
|
||||
|
||||
jobs:
|
||||
# General tests - 1 GPU
|
||||
nightly-test-general-1-gpu-h100:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'nightly-test-general-1-gpu-h100')
|
||||
runs-on: 1-gpu-h100
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
|
||||
- uses: ./.github/actions/check-maintenance
|
||||
|
||||
- name: Install dependencies
|
||||
run: |
|
||||
bash scripts/ci/cuda/ci_install_dependency.sh
|
||||
|
||||
- name: Run test
|
||||
timeout-minutes: 60
|
||||
env:
|
||||
RUNAI_STREAMER_MEMORY_LIMIT: 0
|
||||
run: |
|
||||
cd test
|
||||
python3 run_suite.py --hw cuda --suite nightly-1-gpu --nightly --continue-on-error
|
||||
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
|
||||
# JIT kernel full unit tests (expanded parameter ranges via SGLANG_JIT_KERNEL_RUN_FULL_TESTS)
|
||||
nightly-test-kernel-1-gpu-h100:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'nightly-test-kernel-1-gpu-h100')
|
||||
runs-on: 1-gpu-h100
|
||||
timeout-minutes: 60
|
||||
env:
|
||||
# Full jit_kernel test grids (see sglang.kernels.jit.utils.should_run_full_tests)
|
||||
SGLANG_JIT_KERNEL_RUN_FULL_TESTS: "1"
|
||||
# Match pr-test-jit-kernel workflow for consistent JIT warmup behavior
|
||||
SGLANG_JIT_DEEPGEMM_FAST_WARMUP: true
|
||||
# Allow maintenance bypass on default branch (same semantics as PR JIT workflow)
|
||||
PR_TEST_BYPASS_MAINTENANCE_ON_MAIN: ${{ github.ref == 'refs/heads/main' && 'true' || 'false' }}
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
|
||||
- uses: ./.github/actions/check-maintenance
|
||||
|
||||
- name: Install dependencies
|
||||
timeout-minutes: 20
|
||||
run: |
|
||||
bash scripts/ci/cuda/ci_install_dependency.sh
|
||||
|
||||
- name: Run jit kernel nightly suite
|
||||
timeout-minutes: 60
|
||||
run: |
|
||||
cd test
|
||||
python3 run_suite.py --hw cuda --suite nightly-kernel-1-gpu --nightly --continue-on-error
|
||||
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
|
||||
nightly-test-kernel-8-gpu-h200:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'nightly-test-kernel-8-gpu-h200')
|
||||
runs-on: 8-gpu-h200
|
||||
timeout-minutes: 240
|
||||
env:
|
||||
SGLANG_JIT_KERNEL_RUN_FULL_TESTS: "1"
|
||||
SGLANG_JIT_DEEPGEMM_FAST_WARMUP: true
|
||||
PR_TEST_BYPASS_MAINTENANCE_ON_MAIN: ${{ github.ref == 'refs/heads/main' && 'true' || 'false' }}
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
|
||||
- uses: ./.github/actions/check-maintenance
|
||||
|
||||
- name: Install dependencies
|
||||
timeout-minutes: 20
|
||||
run: |
|
||||
bash scripts/ci/cuda/ci_install_dependency.sh
|
||||
|
||||
- name: Run multi-GPU jit kernel nightly suite
|
||||
timeout-minutes: 90
|
||||
run: |
|
||||
cd test
|
||||
# Full grids run ~7x the in-CI parametrizations per world size; the
|
||||
# default 1200s per-file budget only fits the reduced PR sweep.
|
||||
python3 run_suite.py --hw cuda --suite nightly-kernel-8-gpu-h200 --nightly --continue-on-error --timeout-per-file 3600
|
||||
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
|
||||
# General tests - 4 GPU H100
|
||||
nightly-test-general-4-gpu-h100:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'nightly-test-general-4-gpu-h100')
|
||||
runs-on: 4-gpu-h100
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
|
||||
- uses: ./.github/actions/check-maintenance
|
||||
|
||||
- name: Install dependencies
|
||||
run: |
|
||||
bash scripts/ci/cuda/ci_install_dependency.sh
|
||||
|
||||
- name: Run test
|
||||
timeout-minutes: 60
|
||||
run: |
|
||||
cd test
|
||||
python3 run_suite.py --hw cuda --suite nightly-4-gpu --nightly --continue-on-error
|
||||
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
|
||||
# General tests - 8 GPU H200
|
||||
nightly-test-general-8-gpu-h200:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'nightly-test-general-8-gpu-h200')
|
||||
runs-on: 8-gpu-h200
|
||||
strategy:
|
||||
fail-fast: false
|
||||
max-parallel: 2
|
||||
matrix:
|
||||
partition: [0, 1, 2, 3]
|
||||
env:
|
||||
RUNNER_LABELS: 8-gpu-h200
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
|
||||
- uses: ./.github/actions/check-maintenance
|
||||
|
||||
- name: Install dependencies
|
||||
run: |
|
||||
bash scripts/ci/cuda/ci_install_dependency.sh
|
||||
|
||||
- name: Run common 8-GPU model tests
|
||||
if: always()
|
||||
timeout-minutes: 300
|
||||
env:
|
||||
GPU_CONFIG: "8-gpu-h200"
|
||||
IS_H200: "1"
|
||||
run: |
|
||||
cd test
|
||||
python3 run_suite.py --hw cuda --suite nightly-8-gpu-common --nightly --timeout-per-file=18000 --continue-on-error --auto-partition-id=${{ matrix.partition }} --auto-partition-size=4
|
||||
|
||||
- name: Run test
|
||||
timeout-minutes: 30
|
||||
env:
|
||||
GPU_CONFIG: "8-gpu-h200"
|
||||
run: |
|
||||
cd test
|
||||
python3 run_suite.py --hw cuda --suite nightly-8-gpu-h200 --nightly --continue-on-error --auto-partition-id=${{ matrix.partition }} --auto-partition-size=4
|
||||
|
||||
- name: Collect performance metrics
|
||||
if: always()
|
||||
run: |
|
||||
python3 scripts/ci/utils/save_metrics.py \
|
||||
--gpu-config 8-gpu-h200 \
|
||||
--partition ${{ matrix.partition }} \
|
||||
--run-id ${{ github.run_id }} \
|
||||
--output test/metrics-8gpu-h200-partition-${{ matrix.partition }}.json \
|
||||
--search-dir test/performance_results_8_gpu \
|
||||
--search-dir test
|
||||
|
||||
- name: Upload partition metrics
|
||||
if: always()
|
||||
uses: actions/upload-artifact@v4
|
||||
with:
|
||||
name: metrics-8gpu-h200-partition-${{ matrix.partition }}
|
||||
path: test/metrics-8gpu-h200-partition-${{ matrix.partition }}.json
|
||||
retention-days: 5
|
||||
if-no-files-found: ignore
|
||||
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
with:
|
||||
artifact-suffix: ${{ matrix.partition }}
|
||||
|
||||
# General tests - 8 GPU H20
|
||||
nightly-test-general-8-gpu-h20:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'nightly-test-general-8-gpu-h20')
|
||||
runs-on: 8-gpu-h20
|
||||
env:
|
||||
SGLANG_CI_RDMA_ALL_DEVICES: "mlx5_1,mlx5_2,mlx5_3,mlx5_4"
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
|
||||
- uses: ./.github/actions/check-maintenance
|
||||
|
||||
- name: Install dependencies
|
||||
run: |
|
||||
bash scripts/ci/cuda/ci_install_dependency.sh
|
||||
|
||||
- name: Run test
|
||||
timeout-minutes: 30
|
||||
env:
|
||||
GPU_CONFIG: "8-gpu-h20"
|
||||
run: |
|
||||
cd test
|
||||
python3 run_suite.py --hw cuda --suite nightly-8-gpu-h20 --nightly --continue-on-error
|
||||
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
|
||||
# General tests - 8 GPU B200
|
||||
nightly-test-general-8-gpu-b200:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'nightly-test-general-8-gpu-b200')
|
||||
runs-on: 8-gpu-b200
|
||||
strategy:
|
||||
fail-fast: false
|
||||
max-parallel: 2
|
||||
matrix:
|
||||
partition: [0, 1, 2, 3]
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
|
||||
- uses: ./.github/actions/check-maintenance
|
||||
|
||||
- name: Install dependencies
|
||||
run: |
|
||||
bash scripts/ci/cuda/ci_install_dependency.sh
|
||||
|
||||
- name: Run common 8-GPU model tests
|
||||
if: always()
|
||||
timeout-minutes: 200
|
||||
env:
|
||||
GPU_CONFIG: "8-gpu-b200"
|
||||
run: |
|
||||
cd test
|
||||
python3 run_suite.py --hw cuda --suite nightly-8-gpu-common --nightly --timeout-per-file=12000 --continue-on-error --auto-partition-id=${{ matrix.partition }} --auto-partition-size=4
|
||||
|
||||
- name: Collect performance metrics
|
||||
if: always()
|
||||
run: |
|
||||
python3 scripts/ci/utils/save_metrics.py \
|
||||
--gpu-config 8-gpu-b200 \
|
||||
--partition ${{ matrix.partition }} \
|
||||
--run-id ${{ github.run_id }} \
|
||||
--output test/metrics-8gpu-b200-partition-${{ matrix.partition }}.json \
|
||||
--search-dir test/performance_results_8_gpu \
|
||||
--search-dir test
|
||||
|
||||
- name: Upload partition metrics
|
||||
if: always()
|
||||
uses: actions/upload-artifact@v4
|
||||
with:
|
||||
name: metrics-8gpu-b200-partition-${{ matrix.partition }}
|
||||
path: test/metrics-8gpu-b200-partition-${{ matrix.partition }}.json
|
||||
retention-days: 5
|
||||
if-no-files-found: ignore
|
||||
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
with:
|
||||
artifact-suffix: ${{ matrix.partition }}
|
||||
|
||||
# Text model accuracy tests
|
||||
nightly-test-text-accuracy-2-gpu-h100:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'nightly-test-text-accuracy-2-gpu-h100')
|
||||
runs-on: 2-gpu-h100
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
|
||||
- uses: ./.github/actions/check-maintenance
|
||||
|
||||
- name: Install dependencies
|
||||
run: |
|
||||
bash scripts/ci/cuda/ci_install_dependency.sh
|
||||
|
||||
- name: Run eval test for text models
|
||||
timeout-minutes: 120
|
||||
run: |
|
||||
cd test
|
||||
python3 run_suite.py --hw cuda --suite nightly-eval-text-2-gpu --nightly --continue-on-error --timeout-per-file 4500
|
||||
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
|
||||
# Text model performance tests
|
||||
nightly-test-text-perf-2-gpu-h100:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'nightly-test-text-perf-2-gpu-h100')
|
||||
runs-on: 2-gpu-h100
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
|
||||
- uses: ./.github/actions/check-maintenance
|
||||
|
||||
- name: Install dependencies
|
||||
run: |
|
||||
bash scripts/ci/cuda/ci_install_dependency.sh
|
||||
|
||||
- name: Run performance test for text models
|
||||
timeout-minutes: 30
|
||||
env:
|
||||
GPU_CONFIG: "2-gpu-h100"
|
||||
run: |
|
||||
cd test
|
||||
rm -rf performance_results_text_models/
|
||||
python3 run_suite.py --hw cuda --suite nightly-perf-text-2-gpu --nightly --continue-on-error --timeout-per-file 3600
|
||||
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
|
||||
# VLM accuracy tests
|
||||
nightly-test-vlm-accuracy-2-gpu-h100:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'nightly-test-vlm-accuracy-2-gpu-h100')
|
||||
runs-on: 2-gpu-h100
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
|
||||
- uses: ./.github/actions/check-maintenance
|
||||
|
||||
- name: Install dependencies
|
||||
run: |
|
||||
bash scripts/ci/cuda/ci_install_dependency.sh
|
||||
|
||||
- name: Run eval test for VLM models (fixed MMMU-100)
|
||||
timeout-minutes: 120
|
||||
run: |
|
||||
cd test
|
||||
python3 run_suite.py --hw cuda --suite nightly-eval-vlm-2-gpu --nightly --continue-on-error --timeout-per-file 9000
|
||||
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
|
||||
# VLM performance tests
|
||||
nightly-test-vlm-perf-2-gpu-h100:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'nightly-test-vlm-perf-2-gpu-h100')
|
||||
runs-on: 2-gpu-h100
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
|
||||
- uses: ./.github/actions/check-maintenance
|
||||
|
||||
- name: Install dependencies
|
||||
run: |
|
||||
bash scripts/ci/cuda/ci_install_dependency.sh
|
||||
|
||||
- name: Run perf test for VLM models (MMMU)
|
||||
timeout-minutes: 30
|
||||
env:
|
||||
GPU_CONFIG: "2-gpu-h100"
|
||||
run: |
|
||||
cd test
|
||||
rm -rf performance_results_vlms/
|
||||
python3 run_suite.py --hw cuda --suite nightly-perf-vlm-2-gpu --nightly --continue-on-error --timeout-per-file 3600
|
||||
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
|
||||
# B200 Performance tests - 4 GPU
|
||||
nightly-test-perf-4-gpu-b200:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'nightly-test-perf-4-gpu-b200')
|
||||
runs-on: 4-gpu-b200
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
|
||||
- uses: ./.github/actions/check-maintenance
|
||||
|
||||
- name: Install dependencies
|
||||
run: |
|
||||
bash scripts/ci/cuda/ci_install_dependency.sh
|
||||
|
||||
- name: Run test
|
||||
timeout-minutes: 200
|
||||
env:
|
||||
GPU_CONFIG: "4-gpu-b200"
|
||||
run: |
|
||||
cd test
|
||||
python3 run_suite.py --hw cuda --suite nightly-4-gpu-b200 --nightly --continue-on-error --timeout-per-file 12000
|
||||
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
|
||||
# GB300 (Grace-Blackwell NVL4) performance tests - 4 GPU (ARM64)
|
||||
nightly-test-perf-4-gpu-gb300:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'nightly-test-perf-4-gpu-gb300')
|
||||
name: nightly-test-perf-4-gpu-gb300 (${{ matrix.model }})
|
||||
runs-on: 4-gpu-gb300-nightly
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
include:
|
||||
- model: glm5-nvfp4
|
||||
suite: nightly-4-gpu-gb300-glm5-nvfp4
|
||||
- model: qwen35-fp8
|
||||
suite: nightly-4-gpu-gb300-qwen35-fp8
|
||||
- model: deepseek-v4-pro-fp4
|
||||
suite: nightly-4-gpu-gb300-deepseek-v4-pro-fp4
|
||||
- model: kimi-k25-nvfp4
|
||||
suite: nightly-4-gpu-gb300-kimi-k25-nvfp4
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
|
||||
- uses: ./.github/actions/check-maintenance
|
||||
|
||||
- name: Install dependencies
|
||||
run: |
|
||||
bash scripts/ci/cuda/ci_install_dependency.sh
|
||||
|
||||
- name: Run test
|
||||
timeout-minutes: 600
|
||||
env:
|
||||
GPU_CONFIG: "4-gpu-gb300"
|
||||
run: |
|
||||
cd test
|
||||
python3 run_suite.py --hw cuda --suite ${{ matrix.suite }} --nightly --continue-on-error --timeout-per-file 7200
|
||||
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
|
||||
# Specialized B200 tests - 8 GPU, for specific backends and configs
|
||||
nightly-test-specialized-8-gpu-b200:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'nightly-test-perf-8-gpu-b200' || inputs.job_filter == 'nightly-test-specialized-8-gpu-b200')
|
||||
runs-on: 8-gpu-b200
|
||||
env:
|
||||
RUNNER_LABELS: 8-gpu-b200
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
|
||||
- uses: ./.github/actions/check-maintenance
|
||||
|
||||
- name: Install dependencies
|
||||
run: |
|
||||
bash scripts/ci/cuda/ci_install_dependency.sh
|
||||
|
||||
- name: Run test
|
||||
timeout-minutes: 60
|
||||
env:
|
||||
GPU_CONFIG: "8-gpu-b200"
|
||||
run: |
|
||||
cd test
|
||||
python3 run_suite.py --hw cuda --suite nightly-8-gpu-b200 --nightly --continue-on-error --timeout-per-file 2400
|
||||
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
|
||||
# SGLang-Diffusion nightly benchmark
|
||||
# run_all_tests skips the paths-filter, so main_package is 'true' and
|
||||
# sgl_kernel stays empty: every test runs, and no job waits on a wheel this
|
||||
# workflow never builds. pr_test_yml points back here so shard sizing reads
|
||||
# this file's own run_timeout_minutes.
|
||||
check-changes:
|
||||
uses: ./.github/workflows/_pr-test-check-changes.yml
|
||||
with:
|
||||
git_ref: ${{ inputs.ref || '' }}
|
||||
pr_test_yml: '.github/workflows/nightly-test-nvidia.yml'
|
||||
run_all_tests: true
|
||||
force_continue_on_error: true
|
||||
secrets: inherit
|
||||
|
||||
nightly-1-gpu-large:
|
||||
needs: check-changes
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.runner_filter == '' || inputs.runner_filter == 'all' || inputs.runner_filter == '1-gpu-large')
|
||||
uses: ./.github/workflows/_pr-test-stage.yml
|
||||
with:
|
||||
self_name: nightly-test-1-gpu-large
|
||||
runner_config: 1-gpu-large
|
||||
check_changes: ${{ toJson(needs.check-changes.outputs) }}
|
||||
caller_inputs: ${{ toJson(inputs) }}
|
||||
partitions: ${{ needs.check-changes.outputs.partitions }}
|
||||
run_timeout_minutes: '120'
|
||||
job_timeout_minutes: '180'
|
||||
scheduled: true
|
||||
secrets: inherit
|
||||
|
||||
nightly-2-gpu-large:
|
||||
needs: check-changes
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.runner_filter == '' || inputs.runner_filter == 'all' || inputs.runner_filter == '2-gpu-large')
|
||||
uses: ./.github/workflows/_pr-test-stage.yml
|
||||
with:
|
||||
self_name: nightly-test-2-gpu-large
|
||||
runner_config: 2-gpu-large
|
||||
check_changes: ${{ toJson(needs.check-changes.outputs) }}
|
||||
caller_inputs: ${{ toJson(inputs) }}
|
||||
partitions: ${{ needs.check-changes.outputs.partitions }}
|
||||
run_timeout_minutes: '240'
|
||||
job_timeout_minutes: '300'
|
||||
scheduled: true
|
||||
secrets: inherit
|
||||
|
||||
nightly-4-gpu-h100:
|
||||
needs: check-changes
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.runner_filter == '' || inputs.runner_filter == 'all' || inputs.runner_filter == '4-gpu-h100')
|
||||
uses: ./.github/workflows/_pr-test-stage.yml
|
||||
with:
|
||||
self_name: nightly-test-4-gpu-h100
|
||||
runner_config: 4-gpu-h100
|
||||
check_changes: ${{ toJson(needs.check-changes.outputs) }}
|
||||
caller_inputs: ${{ toJson(inputs) }}
|
||||
partitions: ${{ needs.check-changes.outputs.partitions }}
|
||||
run_timeout_minutes: '120'
|
||||
job_timeout_minutes: '180'
|
||||
scheduled: true
|
||||
secrets: inherit
|
||||
|
||||
nightly-4-gpu-b200:
|
||||
needs: check-changes
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.runner_filter == '' || inputs.runner_filter == 'all' || inputs.runner_filter == '4-gpu-b200')
|
||||
uses: ./.github/workflows/_pr-test-stage.yml
|
||||
with:
|
||||
self_name: nightly-test-4-gpu-b200
|
||||
runner_config: 4-gpu-b200
|
||||
check_changes: ${{ toJson(needs.check-changes.outputs) }}
|
||||
caller_inputs: ${{ toJson(inputs) }}
|
||||
partitions: ${{ needs.check-changes.outputs.partitions }}
|
||||
run_timeout_minutes: '150'
|
||||
job_timeout_minutes: '210'
|
||||
scheduled: true
|
||||
secrets: inherit
|
||||
|
||||
nightly-4-gpu-gb300:
|
||||
needs: check-changes
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.runner_filter == '' || inputs.runner_filter == 'all' || inputs.runner_filter == '4-gpu-gb300')
|
||||
uses: ./.github/workflows/_pr-test-stage.yml
|
||||
with:
|
||||
self_name: nightly-test-4-gpu-gb300
|
||||
runner_config: 4-gpu-gb300
|
||||
check_changes: ${{ toJson(needs.check-changes.outputs) }}
|
||||
caller_inputs: ${{ toJson(inputs) }}
|
||||
partitions: ${{ needs.check-changes.outputs.partitions }}
|
||||
run_timeout_minutes: '360'
|
||||
job_timeout_minutes: '420'
|
||||
scheduled: true
|
||||
# aarch64: the rust-ext cache key is x86_64-only, same reason base-c skips it.
|
||||
skip_prebuilt_rust_ext: true
|
||||
secrets: inherit
|
||||
|
||||
nightly-8-gpu-h200:
|
||||
needs: check-changes
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.runner_filter == '' || inputs.runner_filter == 'all' || inputs.runner_filter == '8-gpu-h200')
|
||||
uses: ./.github/workflows/_pr-test-stage.yml
|
||||
with:
|
||||
self_name: nightly-test-8-gpu-h200
|
||||
runner_config: 8-gpu-h200
|
||||
check_changes: ${{ toJson(needs.check-changes.outputs) }}
|
||||
caller_inputs: ${{ toJson(inputs) }}
|
||||
partitions: ${{ needs.check-changes.outputs.partitions }}
|
||||
run_timeout_minutes: '300'
|
||||
job_timeout_minutes: '360'
|
||||
scheduled: true
|
||||
secrets: inherit
|
||||
|
||||
nightly-8-gpu-b200:
|
||||
needs: check-changes
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.runner_filter == '' || inputs.runner_filter == 'all' || inputs.runner_filter == '8-gpu-b200')
|
||||
uses: ./.github/workflows/_pr-test-stage.yml
|
||||
with:
|
||||
self_name: nightly-test-8-gpu-b200
|
||||
runner_config: 8-gpu-b200
|
||||
check_changes: ${{ toJson(needs.check-changes.outputs) }}
|
||||
caller_inputs: ${{ toJson(inputs) }}
|
||||
partitions: ${{ needs.check-changes.outputs.partitions }}
|
||||
run_timeout_minutes: '360'
|
||||
job_timeout_minutes: '420'
|
||||
scheduled: true
|
||||
secrets: inherit
|
||||
|
||||
# Hand-written because it is not a registry suite: it drives run_comparison.py
|
||||
# and publishes a dashboard rather than running test files.
|
||||
nightly-test-diffusion:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'nightly-test-diffusion')
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.runner_filter == '' || inputs.runner_filter == 'all' || inputs.runner_filter == 'diffusion')
|
||||
runs-on: 4-gpu-h100
|
||||
timeout-minutes: 300
|
||||
env:
|
||||
SGLANG_IS_IN_CI: true
|
||||
SGLANG_ENABLE_ASYNC_ASSERT: true
|
||||
SGLANG_CUDA_COREDUMP: "1"
|
||||
HF_HUB_DOWNLOAD_TIMEOUT: 300
|
||||
HF_HUB_ETAG_TIMEOUT: 300
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
@@ -593,47 +257,19 @@ jobs:
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
|
||||
# Nightly precision regression - per-layer hidden state comparison
|
||||
nightly-test-precision-8-gpu-h200:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'nightly-test-precision-8-gpu-h200')
|
||||
runs-on: 8-gpu-h200
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
ref: ${{ inputs.ref || github.ref }}
|
||||
|
||||
- uses: ./.github/actions/check-maintenance
|
||||
|
||||
- name: Install dependencies
|
||||
run: |
|
||||
bash scripts/ci/cuda/ci_install_dependency.sh
|
||||
|
||||
- name: Run precision regression test
|
||||
timeout-minutes: 120
|
||||
env:
|
||||
SGLANG_PRECISION_BASELINE_DIR: /tmp/sglang_precision_baselines
|
||||
# Required: the test errors if SGLANG_PRECISION_HF_REPO is unset (no
|
||||
# local-only mode). Set the var + the HF_TOKEN_PRECISION_STORE secret.
|
||||
SGLANG_PRECISION_HF_REPO: ${{ vars.SGLANG_PRECISION_HF_REPO }}
|
||||
SGLANG_PRECISION_HF_REVISION: ${{ vars.SGLANG_PRECISION_HF_REVISION || 'main' }}
|
||||
HF_TOKEN: ${{ secrets.HF_TOKEN_PRECISION_STORE }}
|
||||
SGLANG_PRECISION_COMMIT: ${{ github.sha }}
|
||||
# '0' on scheduled runs (input unset); '1' only on a manual baseline-refresh dispatch.
|
||||
SGLANG_PRECISION_FORCE_UPDATE: ${{ inputs.force_baseline_update && '1' || '0' }}
|
||||
run: |
|
||||
cd test
|
||||
python3 run_suite.py --hw cuda --suite nightly-precision-8-gpu-h200 --nightly --continue-on-error --timeout-per-file 3600
|
||||
|
||||
- uses: ./.github/actions/upload-cuda-coredumps
|
||||
if: failure()
|
||||
|
||||
# Consolidate performance metrics from all jobs
|
||||
consolidate-metrics:
|
||||
if: github.repository == 'sgl-project/sglang' && always()
|
||||
# Every scheduled stage uploads metrics now, so all of them must finish
|
||||
# before the download step globs `*metrics-*`.
|
||||
needs:
|
||||
- nightly-test-general-8-gpu-h200
|
||||
- nightly-test-general-8-gpu-b200
|
||||
- nightly-1-gpu-large
|
||||
- nightly-2-gpu-large
|
||||
- nightly-4-gpu-h100
|
||||
- nightly-4-gpu-b200
|
||||
- nightly-4-gpu-gb300
|
||||
- nightly-8-gpu-h200
|
||||
- nightly-8-gpu-b200
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout code
|
||||
@@ -674,20 +310,14 @@ jobs:
|
||||
check-all-jobs:
|
||||
if: github.repository == 'sgl-project/sglang' && always()
|
||||
needs:
|
||||
- nightly-test-general-1-gpu-h100
|
||||
- nightly-test-general-4-gpu-h100
|
||||
- nightly-test-general-8-gpu-h200
|
||||
- nightly-test-general-8-gpu-h20
|
||||
- nightly-test-general-8-gpu-b200
|
||||
- nightly-test-text-accuracy-2-gpu-h100
|
||||
- nightly-test-text-perf-2-gpu-h100
|
||||
- nightly-test-vlm-accuracy-2-gpu-h100
|
||||
- nightly-test-vlm-perf-2-gpu-h100
|
||||
- nightly-test-perf-4-gpu-b200
|
||||
- nightly-test-specialized-8-gpu-b200
|
||||
- nightly-test-perf-4-gpu-gb300
|
||||
- nightly-1-gpu-large
|
||||
- nightly-2-gpu-large
|
||||
- nightly-4-gpu-h100
|
||||
- nightly-4-gpu-b200
|
||||
- nightly-4-gpu-gb300
|
||||
- nightly-8-gpu-h200
|
||||
- nightly-8-gpu-b200
|
||||
- nightly-test-diffusion
|
||||
- nightly-test-precision-8-gpu-h200
|
||||
- consolidate-metrics
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
|
||||
@@ -1,3 +1,5 @@
|
||||
# Weekly CUDA tests. Same shape as nightly-test-nvidia.yml (see its header); a
|
||||
# test joins by declaring `stage="weekly"` plus a runner_config.
|
||||
name: Weekly Test (Nvidia)
|
||||
|
||||
on:
|
||||
@@ -5,46 +7,53 @@ on:
|
||||
- cron: '0 0 * * 0' # Run every Sunday at midnight UTC
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
job_filter:
|
||||
description: 'Select which job to run (leave empty or "all" to run all jobs)'
|
||||
runner_filter:
|
||||
description: 'Select which runner_config to run (leave empty or "all" to run all)'
|
||||
required: false
|
||||
type: choice
|
||||
default: 'all'
|
||||
options:
|
||||
- 'all'
|
||||
- 'weekly-test-8-gpu-h200'
|
||||
- '8-gpu-h200'
|
||||
full_parallel:
|
||||
description: 'Run all shards of a job at once (faster, but competes with per-commit CI for machines). Off by default: one shard at a time.'
|
||||
required: false
|
||||
type: boolean
|
||||
default: false
|
||||
|
||||
concurrency:
|
||||
group: weekly-test-nvidia-${{ github.ref }}
|
||||
cancel-in-progress: true
|
||||
|
||||
env:
|
||||
SGLANG_IS_IN_CI: true
|
||||
SGLANG_ENABLE_ASYNC_ASSERT: true
|
||||
HF_HUB_DOWNLOAD_TIMEOUT: 300
|
||||
HF_HUB_ETAG_TIMEOUT: 300
|
||||
permissions:
|
||||
actions: write
|
||||
contents: read
|
||||
issues: read
|
||||
pull-requests: read
|
||||
|
||||
jobs:
|
||||
# Weekly tests - 8 GPU H200
|
||||
# run_all_tests skips the paths-filter, so main_package is 'true' and
|
||||
# sgl_kernel stays empty: every test runs, and no job waits on a wheel this
|
||||
# workflow never builds.
|
||||
check-changes:
|
||||
uses: ./.github/workflows/_pr-test-check-changes.yml
|
||||
with:
|
||||
pr_test_yml: '.github/workflows/weekly-test-nvidia.yml'
|
||||
run_all_tests: true
|
||||
force_continue_on_error: true
|
||||
secrets: inherit
|
||||
|
||||
weekly-test-8-gpu-h200:
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.job_filter == '' || inputs.job_filter == 'all' || inputs.job_filter == 'weekly-test-8-gpu-h200')
|
||||
runs-on: 8-gpu-h200
|
||||
timeout-minutes: 120
|
||||
env:
|
||||
RUNNER_LABELS: 8-gpu-h200
|
||||
steps:
|
||||
- name: Checkout code
|
||||
uses: actions/checkout@v4
|
||||
|
||||
- name: Install dependencies
|
||||
run: |
|
||||
bash scripts/ci/cuda/ci_install_dependency.sh
|
||||
|
||||
- name: Run weekly 8-GPU H200 tests
|
||||
timeout-minutes: 120
|
||||
env:
|
||||
GPU_CONFIG: "8-gpu-h200"
|
||||
IS_H200: "1"
|
||||
run: |
|
||||
cd test
|
||||
python3 run_suite.py --hw cuda --suite weekly-8-gpu-h200 --nightly --continue-on-error --timeout-per-file 7200
|
||||
needs: check-changes
|
||||
if: github.repository == 'sgl-project/sglang' && (inputs.runner_filter == '' || inputs.runner_filter == 'all' || inputs.runner_filter == '8-gpu-h200')
|
||||
uses: ./.github/workflows/_pr-test-stage.yml
|
||||
with:
|
||||
self_name: weekly-test-8-gpu-h200
|
||||
runner_config: 8-gpu-h200
|
||||
check_changes: ${{ toJson(needs.check-changes.outputs) }}
|
||||
caller_inputs: ${{ toJson(inputs) }}
|
||||
partitions: ${{ needs.check-changes.outputs.partitions }}
|
||||
run_timeout_minutes: '240'
|
||||
job_timeout_minutes: '300'
|
||||
scheduled: true
|
||||
secrets: inherit
|
||||
|
||||
@@ -108,7 +108,7 @@ A `capture_signature` (SHA-1 hash of schema version, max_tokens, ignore_eos, TP
|
||||
| `SGLANG_PRECISION_COMMIT` | _(auto-detected from git)_ | Override the sglang commit SHA tagged on push |
|
||||
| `SGLANG_PRECISION_HF_REPO` | _(required)_ | HuggingFace dataset repo for cross-runner baseline storage |
|
||||
| `SGLANG_PRECISION_HF_REVISION` | `main` | Branch/revision of the HF dataset |
|
||||
| `HF_TOKEN` | _(required in CI)_ | HuggingFace token with write access to the dataset |
|
||||
| `SGLANG_PRECISION_HF_TOKEN` | _(required in CI)_ | HuggingFace token with write access to the dataset. Kept off `HF_TOKEN`, which already carries the runner's gated-model read token |
|
||||
|
||||
---
|
||||
|
||||
@@ -116,31 +116,34 @@ A `capture_signature` (SHA-1 hash of schema version, max_tokens, ignore_eos, TP
|
||||
|
||||
### Workflow job
|
||||
|
||||
The nightly job `nightly-test-precision-8-gpu-h200` is defined in `.github/workflows/nightly-test-nvidia.yml` and runs on an 8-GPU H200 runner. It is included in the nightly suite via `test/run_suite.py`.
|
||||
The test is registered on the `nightly-8-gpu-h200` stage in `.github/workflows/nightly-test-nvidia.yml`, which runs through `_pr-test-stage.yml` like every other CUDA stage. The baseline env below is exported on every scheduled stage; only this test reads it.
|
||||
|
||||
Key CI configuration:
|
||||
|
||||
```yaml
|
||||
- name: Run precision regression test
|
||||
timeout-minutes: 120
|
||||
- name: Export precision baseline env
|
||||
if: inputs.scheduled
|
||||
env:
|
||||
SGLANG_PRECISION_BASELINE_DIR: /tmp/sglang_precision_baselines
|
||||
SGLANG_PRECISION_HF_REPO: ${{ vars.SGLANG_PRECISION_HF_REPO }}
|
||||
SGLANG_PRECISION_HF_REVISION: ${{ vars.SGLANG_PRECISION_HF_REVISION || 'main' }}
|
||||
HF_TOKEN: ${{ secrets.HF_TOKEN_PRECISION_STORE }}
|
||||
SGLANG_PRECISION_COMMIT: ${{ github.sha }}
|
||||
BASELINE_HF_TOKEN: ${{ secrets.HF_TOKEN_PRECISION_STORE }}
|
||||
run: |
|
||||
cd test
|
||||
python3 run_suite.py --hw cuda --suite nightly-precision-8-gpu-h200 --nightly --continue-on-error --timeout-per-file 3600
|
||||
{
|
||||
echo "SGLANG_PRECISION_BASELINE_DIR=/tmp/sglang_precision_baselines"
|
||||
echo "SGLANG_PRECISION_HF_REPO=${{ vars.SGLANG_PRECISION_HF_REPO }}"
|
||||
echo "SGLANG_PRECISION_HF_REVISION=${{ vars.SGLANG_PRECISION_HF_REVISION || 'main' }}"
|
||||
echo "SGLANG_PRECISION_COMMIT=${{ github.sha }}"
|
||||
echo "SGLANG_PRECISION_HF_TOKEN=${BASELINE_HF_TOKEN}"
|
||||
} >> "$GITHUB_ENV"
|
||||
```
|
||||
|
||||
`SGLANG_PRECISION_HF_TOKEN` rather than `HF_TOKEN`: the latter already carries the runner's gated-model read token, and overwriting it would turn every gated model on the job into a 401.
|
||||
|
||||
### Required GitHub secrets/variables
|
||||
|
||||
| Name | Type | Purpose |
|
||||
|------|------|---------|
|
||||
| `SGLANG_PRECISION_HF_REPO` | Repository variable | HF dataset repo ID (e.g. `org/sglang-precision-baselines`) — **required**, the test errors if unset |
|
||||
| `SGLANG_PRECISION_HF_REVISION` | Repository variable (optional) | Dataset branch (defaults to `main`) |
|
||||
| `HF_TOKEN_PRECISION_STORE` | Repository secret | HF token with write access to the dataset |
|
||||
| `HF_TOKEN_PRECISION_STORE` | Repository secret | HF token with write access to the dataset; exported to the job as `SGLANG_PRECISION_HF_TOKEN` |
|
||||
|
||||
### GitHub Step Summary
|
||||
|
||||
@@ -187,7 +190,7 @@ export SGLANG_PRECISION_MODELS="your-org/your-model"
|
||||
export SGLANG_PRECISION_BASELINE_DIR="/tmp/my_precision_baselines"
|
||||
export SGLANG_PRECISION_DIFF_THRESHOLD="1e-3"
|
||||
export SGLANG_PRECISION_HF_REPO="your-org/sglang-precision-baselines"
|
||||
export HF_TOKEN="hf_..."
|
||||
export SGLANG_PRECISION_HF_TOKEN="hf_..."
|
||||
|
||||
cd test
|
||||
python3 -m pytest registered/debug_utils/test_nightly_precision_regression.py -v
|
||||
@@ -201,7 +204,7 @@ python3 -m pytest registered/debug_utils/test_nightly_precision_regression.py -v
|
||||
export SGLANG_PRECISION_MODELS="your-org/your-model"
|
||||
export SGLANG_PRECISION_BASELINE_DIR="/tmp/test_baselines"
|
||||
export SGLANG_PRECISION_HF_REPO="your-org/sglang-precision-baselines"
|
||||
export HF_TOKEN="hf_..."
|
||||
export SGLANG_PRECISION_HF_TOKEN="hf_..."
|
||||
export SGLANG_PRECISION_FORCE_UPDATE="1" # first run: establish baseline
|
||||
|
||||
cd test
|
||||
@@ -249,7 +252,7 @@ python3 -m pytest registered/debug_utils/test_nightly_precision_regression.py -v
|
||||
- SGLang installed in development mode
|
||||
- GPUs matching the model's requirements
|
||||
- `huggingface_hub` installed
|
||||
- A **HuggingFace dataset** for baseline storage and a write-capable `HF_TOKEN`. The HF store is **mandatory** — `SGLANG_PRECISION_HF_REPO` must be set or the test will error at startup. This is because the nightly CI runners are ephemeral (no persistent local disk), so baselines must survive across runs via the HF dataset. There is currently no local-only fallback.
|
||||
- A **HuggingFace dataset** for baseline storage and a write-capable `SGLANG_PRECISION_HF_TOKEN`. The HF store is **mandatory** — `SGLANG_PRECISION_HF_REPO` must be set or the test will error at startup. This is because the nightly CI runners are ephemeral (no persistent local disk), so baselines must survive across runs via the HF dataset. There is currently no local-only fallback.
|
||||
|
||||
### Quick local test
|
||||
|
||||
@@ -258,7 +261,7 @@ python3 -m pytest registered/debug_utils/test_nightly_precision_regression.py -v
|
||||
export SGLANG_PRECISION_MODELS="Qwen/Qwen2.5-0.5B-Instruct"
|
||||
export SGLANG_PRECISION_BASELINE_DIR="/tmp/precision_baselines"
|
||||
export SGLANG_PRECISION_HF_REPO="your-org/sglang-precision-baselines"
|
||||
export HF_TOKEN="hf_..."
|
||||
export SGLANG_PRECISION_HF_TOKEN="hf_..."
|
||||
|
||||
# First run: establish baseline
|
||||
cd test
|
||||
@@ -356,7 +359,7 @@ The framework uses a **rolling baseline**: every successful comparison updates t
|
||||
|
||||
### No local-only mode
|
||||
|
||||
The test requires a HuggingFace dataset (`SGLANG_PRECISION_HF_REPO`) and a write-capable `HF_TOKEN`. There is no local-only fallback. This is by design — CI runners have no persistent local disk, so the HF dataset is the only way to carry baselines across runs. If you need to run the test locally, you must set up a HF dataset (even a private one) and provide the corresponding token.
|
||||
The test requires a HuggingFace dataset (`SGLANG_PRECISION_HF_REPO`) and a write-capable `SGLANG_PRECISION_HF_TOKEN`. There is no local-only fallback. This is by design — CI runners have no persistent local disk, so the HF dataset is the only way to carry baselines across runs. If you need to run the test locally, you must set up a HF dataset (even a private one) and provide the corresponding token.
|
||||
|
||||
---
|
||||
|
||||
@@ -369,4 +372,4 @@ The test requires a HuggingFace dataset (`SGLANG_PRECISION_HF_REPO`) and a write
|
||||
| `python/sglang/srt/debug_utils/comparator/` | Tensor comparison engine |
|
||||
| `python/sglang/srt/debug_utils/dumper.py` | Runtime hidden-state capture |
|
||||
| `.github/workflows/nightly-test-nvidia.yml` | CI workflow definition |
|
||||
| `test/run_suite.py` | Test suite registration (includes `nightly-precision-8-gpu-h200`) |
|
||||
| `test/run_suite.py` | Test suite registration (includes `nightly-test-8-gpu-h200`) |
|
||||
|
||||
@@ -137,9 +137,23 @@ def _repo_relative_path(p: str) -> str:
|
||||
return p[idx + len(marker) :] if idx >= 0 else p
|
||||
|
||||
|
||||
# Slow-run variance is largely additive (cold HF cache, slow server launch), so
|
||||
# the multiplier alone under-provisions at both ends: test_encoder_dp runs
|
||||
# 200-426s but once took over 1185s against a 1.5x budget of 765s, and
|
||||
# test_lora_deepseek_v3_base_logprob_diff (est 1800) landed on exactly 1.5 * est.
|
||||
# Every file gets the same absolute slack on top of the proportional one.
|
||||
DERIVED_TIMEOUT_SLACK = 1800.0
|
||||
DERIVED_TIMEOUT_FACTOR = 1.5
|
||||
|
||||
|
||||
def derive_timeout_per_file(est_time: float) -> float:
|
||||
est = float(est_time)
|
||||
return max(est * DERIVED_TIMEOUT_FACTOR, est + DERIVED_TIMEOUT_SLACK)
|
||||
|
||||
|
||||
def run_unittest_files(
|
||||
files: Union[List[TestFile], List[CIRegistry]],
|
||||
timeout_per_file: float,
|
||||
timeout_per_file: Optional[float] = None,
|
||||
continue_on_error: bool = False,
|
||||
enable_retry: bool = False,
|
||||
max_attempts: int = 2,
|
||||
@@ -150,7 +164,8 @@ def run_unittest_files(
|
||||
|
||||
Args:
|
||||
files: List of TestFile objects to run
|
||||
timeout_per_file: Timeout in seconds for each test file
|
||||
timeout_per_file: Fixed timeout in seconds for every test file, or None
|
||||
to derive each file's budget from its own est_time.
|
||||
continue_on_error: If True, continue running remaining tests even if one fails.
|
||||
If False, stop at first failure (default behavior for PR tests).
|
||||
enable_retry: If True, retry failed tests that appear to be accuracy/performance
|
||||
@@ -178,6 +193,12 @@ def run_unittest_files(
|
||||
# FIXME: remove this branch after migrating all tests to use CIRegistry
|
||||
filename, estimated_time = file.name, file.estimated_time
|
||||
|
||||
file_timeout = (
|
||||
timeout_per_file
|
||||
if timeout_per_file is not None
|
||||
else derive_timeout_per_file(estimated_time)
|
||||
)
|
||||
|
||||
process = None
|
||||
output_lines = []
|
||||
|
||||
@@ -235,7 +256,7 @@ def run_unittest_files(
|
||||
run_one_file,
|
||||
args=(filename,),
|
||||
kwargs={"capture_output": enable_retry},
|
||||
timeout=timeout_per_file,
|
||||
timeout=file_timeout,
|
||||
)
|
||||
|
||||
if ret_code == 0:
|
||||
@@ -281,24 +302,22 @@ def run_unittest_files(
|
||||
# TimeoutError aborts run_one_file before its elapsed write;
|
||||
# record the timeout cap as an upper bound so the file still
|
||||
# appears in the TIMINGS block below.
|
||||
file_elapsed[filename] = float(timeout_per_file)
|
||||
file_elapsed[filename] = float(file_timeout)
|
||||
# Retry once on timeout: usually a stuck server / hung device.
|
||||
# A real hang times out again and is reported.
|
||||
if enable_retry and attempt < max_attempts:
|
||||
logger.info(
|
||||
f"\n[CI Retry] {filename} timed out after "
|
||||
f"{timeout_per_file}s; waiting {retry_wait_seconds}s "
|
||||
f"{file_timeout}s; waiting {retry_wait_seconds}s "
|
||||
f"before retry (attempt {attempt + 1}/{max_attempts})\n"
|
||||
)
|
||||
time.sleep(retry_wait_seconds)
|
||||
attempt += 1
|
||||
continue
|
||||
logger.info(
|
||||
f"\n✗ TIMEOUT: {filename} after {timeout_per_file} seconds\n"
|
||||
)
|
||||
logger.info(f"\n✗ TIMEOUT: {filename} after {file_timeout} seconds\n")
|
||||
if was_retried:
|
||||
retried_tests.append((filename, attempt, "timeout"))
|
||||
failed_tests.append((filename, f"timeout after {timeout_per_file}s"))
|
||||
failed_tests.append((filename, f"timeout after {file_timeout}s"))
|
||||
break
|
||||
|
||||
if not file_passed:
|
||||
|
||||
@@ -26,6 +26,16 @@ from huggingface_hub.errors import (
|
||||
)
|
||||
|
||||
|
||||
def _store_token() -> Optional[str]:
|
||||
"""Write token for the baseline dataset repo.
|
||||
|
||||
Deliberately not HF_TOKEN: that name already carries the runner's
|
||||
gated-model read token, so writing the store token there would shadow it
|
||||
and turn every gated model on the job into a 401.
|
||||
"""
|
||||
return os.environ.get("SGLANG_PRECISION_HF_TOKEN") or None
|
||||
|
||||
|
||||
@dataclass
|
||||
class HfStoreConfig:
|
||||
repo: str
|
||||
@@ -38,7 +48,7 @@ class HfStoreConfig:
|
||||
raise RuntimeError(
|
||||
"SGLANG_PRECISION_HF_REPO is not set. The precision baseline "
|
||||
"store is required (there is no local-only mode); set the repo "
|
||||
"and HF_TOKEN_PRECISION_STORE."
|
||||
"and SGLANG_PRECISION_HF_TOKEN."
|
||||
)
|
||||
revision = os.environ.get("SGLANG_PRECISION_HF_REVISION", "main")
|
||||
return cls(repo=repo, revision=revision)
|
||||
@@ -148,6 +158,7 @@ def fetch_latest_baseline(
|
||||
repo_type="dataset",
|
||||
revision=config.revision,
|
||||
allow_patterns=[f"{run_path}/tensors/*"],
|
||||
token=_store_token(),
|
||||
),
|
||||
what="snapshot download",
|
||||
)
|
||||
@@ -171,6 +182,7 @@ def _read_manifest(config: HfStoreConfig) -> tuple[list[dict[str, Any]], str]:
|
||||
repo_type="dataset",
|
||||
filename="manifest.jsonl",
|
||||
revision=config.revision,
|
||||
token=_store_token(),
|
||||
),
|
||||
what="manifest fetch",
|
||||
)
|
||||
@@ -215,7 +227,7 @@ def push_run(
|
||||
# Dedup: same model+date+sha → skip tensor upload but still refresh meta
|
||||
# + comparator_report + append a new manifest row, so pass-1 baseline and
|
||||
# pass-2 stats both land. force=True re-uploads tensors too.
|
||||
api = HfApi()
|
||||
api = HfApi(token=_store_token())
|
||||
date_str, date_path = _today_path()
|
||||
model_sanitized = _sanitize_model_name(model)
|
||||
sha7 = (
|
||||
@@ -306,7 +318,7 @@ def prune_old_runs(
|
||||
# dry_run defaults True because model=None+keep_days=0 would wipe the
|
||||
# store. Live mode rewrites the manifest before deleting folders so a
|
||||
# mid-run failure leaves manifest pointing at the kept rows only.
|
||||
api = HfApi()
|
||||
api = HfApi(token=_store_token())
|
||||
rows, _ = _read_manifest(config)
|
||||
if not rows:
|
||||
return {"kept": [], "pruned": []}
|
||||
|
||||
@@ -4,11 +4,11 @@ Pre-commit hook: validate CI registry calls under test/registered/.
|
||||
|
||||
1. Every test file must contain a CI registry call (register_cuda_ci,
|
||||
register_amd_ci, etc.).
|
||||
2. A CUDA test must register its PR-test suite via the modern
|
||||
2. A CUDA test must register its suite via the modern
|
||||
`stage=`/`runner_config=` form. The legacy single-string `suite=` is reserved
|
||||
for the nightly/stress/weekly families (and for AMD/CPU/NPU suites); any other
|
||||
CUDA `suite=` resolves to a name no PR-test workflow invokes, so the test
|
||||
silently never runs. Two shapes are rejected:
|
||||
for the stress family (and for AMD/CPU/NPU suites); any other CUDA `suite=`
|
||||
resolves to a name no workflow invokes, so the test silently never runs.
|
||||
Two shapes are rejected:
|
||||
a. `{stage}-test-{runner_config}` -- the modern name stuffed back into the
|
||||
legacy form. Reported with the exact stage/runner split to use.
|
||||
b. an older `{stage}-{runner_config}` PR-test name (e.g. the pre-migration
|
||||
@@ -33,11 +33,10 @@ import sys
|
||||
# shape is always expressible (and should be expressed) the modern way.
|
||||
_MODERN_SHAPE = re.compile(r"^(.+)-test-(.+)$")
|
||||
|
||||
# The only suite families a CUDA registry may keep on the legacy single-string
|
||||
# `suite=` form. Everything else is a PR-test/base stage that must use the
|
||||
# modern stage=/runner_config= form (otherwise its effective_suite matches no
|
||||
# suite the PR-test workflows invoke, and the test silently never runs).
|
||||
_LEGACY_CUDA_PREFIXES = ("nightly", "stress", "weekly")
|
||||
# The only CUDA suite family still allowed on the legacy single-string `suite=`
|
||||
# form. Anything else needs stage=/runner_config=, or its effective_suite matches
|
||||
# no suite any workflow invokes and the test silently never runs.
|
||||
_LEGACY_CUDA_PREFIXES = ("stress",)
|
||||
|
||||
|
||||
def _defines_testcase(tree: ast.AST) -> bool:
|
||||
@@ -118,8 +117,6 @@ def main() -> int:
|
||||
and r.runner_config is None
|
||||
):
|
||||
continue
|
||||
# nightly/stress/weekly are the only CUDA suites allowed to stay on
|
||||
# the legacy single-string form.
|
||||
if r.suite.split("-", 1)[0] in _LEGACY_CUDA_PREFIXES:
|
||||
continue
|
||||
m = _MODERN_SHAPE.match(r.suite)
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
{
|
||||
"_comment": "Manual overrides for list_stage_models.py. by_file/by_suite ADD models the static scan cannot see (models built dynamically, read from configs, passed via CLI args). deny REMOVES false-positive ids the heuristic mistakes for models. Keys in by_file are repo-relative test paths (e.g. test/registered/foo/test_bar.py). suite_labels maps legacy suite= registrations (no runner_config) to the GH runner label(s) their dispatching workflow hardcodes in runs-on -- a list, because one suite can run on several labels (nightly-8-gpu-common). $b200_runner is the dynamic-b200 placeholder from runner_configs.yml. Deliberately absent: nightly-4-gpu-gb300-* (run as k8s pods, not GHA runners) and nightly-2-gpu (registered but dispatched by no workflow); both stay visible in unmapped_suites.",
|
||||
"_comment": "Manual overrides for list_stage_models.py. by_file/by_suite ADD models the static scan cannot see (models built dynamically, read from configs, passed via CLI args). deny REMOVES false-positive ids the heuristic mistakes for models. Keys in by_file are repo-relative test paths (e.g. test/registered/foo/test_bar.py). suite_labels maps legacy suite= registrations (no runner_config) to the GH runner label(s) their dispatching workflow hardcodes in runs-on -- a list, because one suite can run on several labels. $b200_runner is the dynamic-b200 placeholder from runner_configs.yml. Suites registered with stage=/runner_config= need no entry: their label resolves through runner_configs.yml. Anything unmappable stays visible in unmapped_suites.",
|
||||
"by_file": {},
|
||||
"by_suite": {},
|
||||
"suite_labels": {
|
||||
@@ -7,21 +7,7 @@
|
||||
"base-b-kernel-unit-1-gpu-b200": ["$b200_runner"],
|
||||
"base-b-kernel-unit-1-gpu-large": ["1-gpu-h100"],
|
||||
"base-b-kernel-unit-8-gpu-h200": ["8-gpu-h200"],
|
||||
"nightly-1-gpu": ["1-gpu-h100"],
|
||||
"nightly-4-gpu": ["4-gpu-h100"],
|
||||
"nightly-4-gpu-b200": ["$b200_runner"],
|
||||
"nightly-8-gpu-b200": ["8-gpu-b200"],
|
||||
"nightly-8-gpu-common": ["8-gpu-h200", "8-gpu-b200"],
|
||||
"nightly-8-gpu-h200": ["8-gpu-h200"],
|
||||
"nightly-eval-text-2-gpu": ["2-gpu-h100"],
|
||||
"nightly-eval-vlm-2-gpu": ["2-gpu-h100"],
|
||||
"nightly-kernel-1-gpu": ["1-gpu-h100"],
|
||||
"nightly-kernel-8-gpu-h200": ["8-gpu-h200"],
|
||||
"nightly-perf-text-2-gpu": ["2-gpu-h100"],
|
||||
"nightly-perf-vlm-2-gpu": ["2-gpu-h100"],
|
||||
"nightly-precision-8-gpu-h200": ["8-gpu-h200"],
|
||||
"stress": ["8-gpu-h200"],
|
||||
"weekly-8-gpu-h200": ["8-gpu-h200"]
|
||||
"stress": ["8-gpu-h200"]
|
||||
},
|
||||
"deny": [
|
||||
"tok/req",
|
||||
|
||||
@@ -128,8 +128,11 @@ def compute_partitions(
|
||||
in-source `est_time` / `(1.0, 0.0)`.
|
||||
`full_parallel=True` lifts the matrix-fanout throttle.
|
||||
"""
|
||||
# Allowlist: stages pr-test.yml dispatches. Stress / weekly /
|
||||
# nightly-* live in test/registered/ but pr-test doesn't run them.
|
||||
# Allowlist of the stages this workflow dispatches -- what keeps stress /
|
||||
# weekly / nightly out, since CUDA scheduled suites no longer carry
|
||||
# `nightly=True`. The nightly filter still matters for CPU: some tests sit on
|
||||
# a dispatched suite with the flag set, so run_suite.py skips them and their
|
||||
# est_time must not inflate the shard count.
|
||||
dispatched_suites = set(run_timeouts) | set(_BASE_A_OVERRIDES)
|
||||
suite_tests = defaultdict(list)
|
||||
for t in tests:
|
||||
|
||||
@@ -702,11 +702,8 @@ def _extract_runner_configs(content):
|
||||
|
||||
|
||||
def _extract_legacy_suites(content):
|
||||
"""Pull every legacy single-string `suite=` from `register_cuda_ci(...)` calls.
|
||||
|
||||
Mirrors _extract_runner_configs for the legacy nightly/weekly shape: a file
|
||||
may register on multiple pools, so collect all of them rather than the first.
|
||||
"""
|
||||
"""Pull every legacy single-string `suite=` from `register_cuda_ci(...)`
|
||||
calls. Used only to report why such a file is not dispatchable."""
|
||||
out = []
|
||||
for args in re.finditer(
|
||||
r"^[^#\n]*register_cuda_ci\s*\(([^)]*)\)", content, re.MULTILINE
|
||||
@@ -717,38 +714,6 @@ def _extract_legacy_suites(content):
|
||||
return out
|
||||
|
||||
|
||||
# Legacy nightly/weekly CUDA suites register with a single-string `suite=`
|
||||
# instead of `runner_config=`, so they carry no runner metadata of their own.
|
||||
# Map each to the runner_config in scripts/ci/runner_configs.yml whose hardware
|
||||
# matches the runner the nightly/weekly pipeline actually uses (see
|
||||
# .github/workflows/{nightly,weekly}-test-nvidia.yml), so /rerun-test can still
|
||||
# dispatch a single nightly/weekly test. The runner label, install script,
|
||||
# timeout and rdma_devices are then resolved from
|
||||
# runner_configs.yml as usual, keeping that file the single source of truth for
|
||||
# runner details.
|
||||
#
|
||||
# Suites on hardware with no matching runner_config (e.g. nightly-4-gpu-gb300)
|
||||
# and non-CUDA suites (npu/amd) are intentionally absent and stay
|
||||
# non-dispatchable until a matching runner_config exists.
|
||||
_LEGACY_SUITE_TO_RUNNER_CONFIG = {
|
||||
"nightly-1-gpu": "1-gpu-large",
|
||||
"nightly-kernel-1-gpu": "1-gpu-large",
|
||||
"nightly-eval-text-2-gpu": "2-gpu-large",
|
||||
"nightly-perf-text-2-gpu": "2-gpu-large",
|
||||
"nightly-eval-vlm-2-gpu": "2-gpu-large",
|
||||
"nightly-perf-vlm-2-gpu": "2-gpu-large",
|
||||
"nightly-4-gpu": "4-gpu-h100",
|
||||
"nightly-4-gpu-b200": "4-gpu-b200",
|
||||
"nightly-8-gpu-common": ["8-gpu-h200", "8-gpu-b200"],
|
||||
"nightly-8-gpu-h200": "8-gpu-h200",
|
||||
"nightly-kernel-8-gpu-h200": "8-gpu-h200",
|
||||
"nightly-precision-8-gpu-h200": "8-gpu-h200",
|
||||
"nightly-8-gpu-h20": "8-gpu-h20",
|
||||
"nightly-8-gpu-b200": "8-gpu-b200",
|
||||
"weekly-8-gpu-h200": "8-gpu-h200",
|
||||
}
|
||||
|
||||
|
||||
def _dispatch_err(suite, msg):
|
||||
"""Build a detect_suite error result for the given suite."""
|
||||
return {
|
||||
@@ -811,11 +776,10 @@ def detect_suite(file_path_from_test):
|
||||
pool it should run on — so this returns a *list* of dispatch dicts, one
|
||||
per registration. Runner label, install script, timeout, and rdma_devices
|
||||
are all resolved from scripts/ci/runner_configs.yml — the
|
||||
same single source of truth that drives the main PR test pipeline.
|
||||
|
||||
Legacy nightly/weekly CUDA suites (single-string `suite=`) are dispatchable
|
||||
too: each suite name is mapped to the matching runner_config via
|
||||
_LEGACY_SUITE_TO_RUNNER_CONFIG, then resolved the same way.
|
||||
same single source of truth that drives the main PR test pipeline. Every
|
||||
dispatchable CUDA suite, per-commit and scheduled alike, goes through that
|
||||
one path; the legacy single-string `suite=` carries no runner_config and is
|
||||
reported as non-dispatchable.
|
||||
|
||||
CPU files yield a single-element list. A file with no recognised (or no
|
||||
dispatchable) registration yields a one-element list whose dict has an
|
||||
@@ -837,19 +801,7 @@ def detect_suite(file_path_from_test):
|
||||
results.append(_resolve_runner_config(rc, full_path, suite))
|
||||
return results
|
||||
|
||||
# Legacy nightly/weekly CUDA suites: single-string `suite=`, no
|
||||
# runner_config. Map each mappable suite to its runner_config and resolve.
|
||||
legacy_suites = _extract_legacy_suites(content)
|
||||
mappable = [s for s in legacy_suites if s in _LEGACY_SUITE_TO_RUNNER_CONFIG]
|
||||
if mappable:
|
||||
results = []
|
||||
for s in mappable:
|
||||
rcs = _LEGACY_SUITE_TO_RUNNER_CONFIG[s]
|
||||
if isinstance(rcs, str):
|
||||
rcs = [rcs]
|
||||
for rc in rcs:
|
||||
results.append(_resolve_runner_config(rc, full_path, s))
|
||||
return results
|
||||
|
||||
if re.search(r"^[^#\n]*register_cpu_ci\s*\(", content, re.MULTILINE):
|
||||
return [
|
||||
@@ -869,11 +821,11 @@ def detect_suite(file_path_from_test):
|
||||
return [
|
||||
_dispatch_err(
|
||||
suite,
|
||||
f"Suite `{suite}` in `{full_path}` is not dispatchable via "
|
||||
f"/rerun-test. It has no entry in _LEGACY_SUITE_TO_RUNNER_CONFIG "
|
||||
f"— either it is a non-CUDA suite (npu/amd) or it runs on "
|
||||
f"hardware with no matching runner_config in "
|
||||
f"scripts/ci/runner_configs.yml.",
|
||||
f"Suite `{suite}` in `{full_path}` is registered with the legacy "
|
||||
f"single-string `suite=`, which carries no runner_config and so "
|
||||
f"is not dispatchable via /rerun-test. Re-register it with "
|
||||
f"`stage=`/`runner_config=` (CUDA), or dispatch its own "
|
||||
f"workflow (npu/amd).",
|
||||
)
|
||||
]
|
||||
|
||||
|
||||
+2
-2
@@ -50,8 +50,8 @@ python3 test/registered/jit/test_add_constant.py
|
||||
python3 test/run_suite.py --hw cpu --suite base-a-test-cpu
|
||||
python3 test/run_suite.py --hw cuda --suite base-a-test-1-gpu-small
|
||||
|
||||
# Nightly tests
|
||||
python3 test/run_suite.py --hw cuda --suite nightly-1-gpu --nightly
|
||||
# Nightly tests (CUDA nightly suites take no --nightly; the stage is in the name)
|
||||
python3 test/run_suite.py --hw cuda --suite nightly-test-1-gpu-large
|
||||
|
||||
# With auto-partitioning (for parallel CI jobs)
|
||||
python3 test/run_suite.py --hw cuda --suite base-b-test-1-gpu-small \
|
||||
|
||||
@@ -12,7 +12,7 @@ from sglang.test.test_utils import (
|
||||
popen_launch_server,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=810, suite="nightly-4-gpu-b200", nightly=True)
|
||||
register_cuda_ci(est_time=1200, stage="nightly", runner_config="4-gpu-b200")
|
||||
|
||||
NEMOTRON_3_SUPER_NVFP4_MODEL = "nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4"
|
||||
|
||||
|
||||
@@ -6,8 +6,9 @@ from sglang.test.performance_test_runner import PerformanceTestParams
|
||||
from sglang.test.run_combined_tests import run_combined_tests
|
||||
from sglang.test.test_utils import ModelLaunchSettings
|
||||
|
||||
# Runs on both H200 and B200 via nightly-8-gpu-common suite
|
||||
register_cuda_ci(est_time=1800, suite="nightly-8-gpu-common", nightly=True)
|
||||
# Runs on both H200 and B200: registered once per runner_config below
|
||||
register_cuda_ci(est_time=2880, stage="nightly", runner_config="8-gpu-h200")
|
||||
register_cuda_ci(est_time=2880, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
GLM_52_FP8_MODEL_PATH = "zai-org/GLM-5.2-FP8"
|
||||
|
||||
|
||||
@@ -6,8 +6,9 @@ from sglang.test.performance_test_runner import PerformanceTestParams
|
||||
from sglang.test.run_combined_tests import run_combined_tests
|
||||
from sglang.test.test_utils import ModelLaunchSettings
|
||||
|
||||
# Runs on both H200 and B200 via nightly-8-gpu-common suite
|
||||
register_cuda_ci(est_time=1800, suite="nightly-8-gpu-common", nightly=True)
|
||||
# Runs on both H200 and B200: registered once per runner_config below
|
||||
register_cuda_ci(est_time=1320, stage="nightly", runner_config="8-gpu-h200")
|
||||
register_cuda_ci(est_time=1320, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
GLM_4_6_MODEL_PATH = "zai-org/GLM-4.6"
|
||||
|
||||
|
||||
@@ -5,9 +5,10 @@ from sglang.test.performance_test_runner import PerformanceTestParams
|
||||
from sglang.test.run_combined_tests import run_combined_tests
|
||||
from sglang.test.test_utils import ModelLaunchSettings
|
||||
|
||||
# Runs on both H200 and B200 via nightly-8-gpu-common suite
|
||||
# Runs on both H200 and B200: registered once per runner_config below
|
||||
# Higher est_time due to 6 variants with both performance and accuracy tests
|
||||
register_cuda_ci(est_time=1800, suite="nightly-8-gpu-common", nightly=True)
|
||||
register_cuda_ci(est_time=690, stage="nightly", runner_config="8-gpu-h200")
|
||||
register_cuda_ci(est_time=690, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
GPT_OSS_120B_MXFP4_MODEL_PATH = "openai/gpt-oss-120b"
|
||||
GPT_OSS_120B_EAGLE3_DRAFT_MODEL_PATH = "lmsys/EAGLE3-gpt-oss-120b-bf16"
|
||||
|
||||
@@ -10,7 +10,8 @@ from sglang.test.test_utils import ModelLaunchSettings, is_blackwell_system
|
||||
|
||||
# NVFP4 needs Blackwell FP4 kernels, so this runs on the Blackwell leg of the
|
||||
# common 8-GPU suite (Hopper is skipped below).
|
||||
register_cuda_ci(est_time=3600, suite="nightly-8-gpu-common", nightly=True)
|
||||
register_cuda_ci(est_time=3600, stage="nightly", runner_config="8-gpu-h200")
|
||||
register_cuda_ci(est_time=3600, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
INKLING_NVFP4_MODEL = "thinkingmachines/Inkling-NVFP4"
|
||||
INKLING_SMALL_NVFP4_MODEL = "thinkingmachines/Inkling-Small-NVFP4"
|
||||
|
||||
@@ -6,8 +6,9 @@ from sglang.test.performance_test_runner import PerformanceTestParams
|
||||
from sglang.test.run_combined_tests import run_combined_tests
|
||||
from sglang.test.test_utils import ModelLaunchSettings
|
||||
|
||||
# Runs on both H200 and B200 via nightly-8-gpu-common suite
|
||||
register_cuda_ci(est_time=3600, suite="nightly-8-gpu-common", nightly=True)
|
||||
# Runs on both H200 and B200: registered once per runner_config below
|
||||
register_cuda_ci(est_time=2820, stage="nightly", runner_config="8-gpu-h200")
|
||||
register_cuda_ci(est_time=2820, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
KIMI_K25_MODEL_PATH = "moonshotai/Kimi-K2.5"
|
||||
|
||||
|
||||
@@ -12,7 +12,8 @@ from sglang.test.ci.ci_register import register_cuda_ci
|
||||
from sglang.test.kits.eval_accuracy_kit import GSM8KMixin
|
||||
from sglang.test.server_fixtures.default_fixture import DefaultServerBase
|
||||
|
||||
register_cuda_ci(est_time=600, suite="nightly-8-gpu-common", nightly=True)
|
||||
register_cuda_ci(est_time=300, stage="nightly", runner_config="8-gpu-h200")
|
||||
register_cuda_ci(est_time=300, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
|
||||
class TestLing26Flash(GSM8KMixin, DefaultServerBase):
|
||||
|
||||
@@ -6,8 +6,9 @@ from sglang.test.performance_test_runner import PerformanceTestParams
|
||||
from sglang.test.run_combined_tests import run_combined_tests
|
||||
from sglang.test.test_utils import ModelLaunchSettings
|
||||
|
||||
# Runs on both H200 and B200 via nightly-8-gpu-common suite
|
||||
register_cuda_ci(est_time=1800, suite="nightly-8-gpu-common", nightly=True)
|
||||
# Runs on both H200 and B200: registered once per runner_config below
|
||||
register_cuda_ci(est_time=1800, stage="nightly", runner_config="8-gpu-h200")
|
||||
register_cuda_ci(est_time=1800, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
LLAMA4_MODEL_PATH = "meta-llama/Llama-4-Scout-17B-16E-Instruct"
|
||||
|
||||
|
||||
@@ -6,7 +6,7 @@ from sglang.test.performance_test_runner import PerformanceTestParams
|
||||
from sglang.test.run_combined_tests import run_combined_tests
|
||||
from sglang.test.test_utils import ModelLaunchSettings
|
||||
|
||||
register_cuda_ci(est_time=1200, suite="nightly-8-gpu-h200", nightly=True)
|
||||
register_cuda_ci(est_time=1200, stage="nightly", runner_config="8-gpu-h200")
|
||||
|
||||
# LongCat-Flash-Lite-FP8 is the smallest member of the LongCat family
|
||||
# (~138 GB FP8 weights, hidden=3072, 14 layers, 256 routed + 128 zero
|
||||
|
||||
@@ -6,8 +6,9 @@ from sglang.test.performance_test_runner import PerformanceTestParams
|
||||
from sglang.test.run_combined_tests import run_combined_tests
|
||||
from sglang.test.test_utils import ModelLaunchSettings
|
||||
|
||||
# Runs on both H200 and B200 via nightly-8-gpu-common suite
|
||||
register_cuda_ci(est_time=1800, suite="nightly-8-gpu-common", nightly=True)
|
||||
# Runs on both H200 and B200: registered once per runner_config below
|
||||
register_cuda_ci(est_time=1860, stage="nightly", runner_config="8-gpu-h200")
|
||||
register_cuda_ci(est_time=1860, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
MINIMAX_M25_MODEL_PATH = "MiniMaxAI/MiniMax-M2.5"
|
||||
|
||||
|
||||
@@ -8,9 +8,10 @@ from sglang.test.performance_test_runner import PerformanceTestParams
|
||||
from sglang.test.run_combined_tests import run_combined_tests
|
||||
from sglang.test.test_utils import ModelLaunchSettings, is_blackwell_system
|
||||
|
||||
# Runs on both H200 and B200 via nightly-8-gpu-common suite
|
||||
# Runs on both H200 and B200: registered once per runner_config below
|
||||
# Note: trtllm_mla backend may have hardware-specific behavior
|
||||
register_cuda_ci(est_time=3000, suite="nightly-8-gpu-common", nightly=True)
|
||||
register_cuda_ci(est_time=3000, stage="nightly", runner_config="8-gpu-h200")
|
||||
register_cuda_ci(est_time=3000, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
MISTRAL_LARGE3_FP8_MODEL_PATH = "mistralai/Mistral-Large-3-675B-Instruct-2512"
|
||||
MISTRAL_LARGE3_NVFP4_MODEL_PATH = "mistralai/Mistral-Large-3-675B-Instruct-2512-NVFP4"
|
||||
|
||||
@@ -7,8 +7,9 @@ from sglang.test.performance_test_runner import PerformanceTestParams
|
||||
from sglang.test.run_combined_tests import run_combined_tests
|
||||
from sglang.test.test_utils import ModelLaunchSettings, is_blackwell_system
|
||||
|
||||
# Runs on both Hopper and Blackwell via nightly-8-gpu-common suite
|
||||
register_cuda_ci(est_time=5400, suite="nightly-8-gpu-common", nightly=True)
|
||||
# Runs on both Hopper and Blackwell: registered once per runner_config below
|
||||
register_cuda_ci(est_time=3360, stage="nightly", runner_config="8-gpu-h200")
|
||||
register_cuda_ci(est_time=3360, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
NEMOTRON_3_SUPER_BF16_MODEL = "nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16"
|
||||
NEMOTRON_3_SUPER_NVFP4_MODEL = "nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4"
|
||||
|
||||
@@ -6,8 +6,9 @@ from sglang.test.performance_test_runner import PerformanceTestParams
|
||||
from sglang.test.run_combined_tests import run_combined_tests
|
||||
from sglang.test.test_utils import ModelLaunchSettings
|
||||
|
||||
# Runs on both H200 and B200 via nightly-8-gpu-common suite
|
||||
register_cuda_ci(est_time=1800, suite="nightly-8-gpu-common", nightly=True)
|
||||
# Runs on both H200 and B200: registered once per runner_config below
|
||||
register_cuda_ci(est_time=3000, stage="nightly", runner_config="8-gpu-h200")
|
||||
register_cuda_ci(est_time=3000, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
QWEN35_MODEL_PATH = "Qwen/Qwen3.5-397B-A17B-FP8"
|
||||
|
||||
|
||||
@@ -5,7 +5,8 @@ from sglang.test.ci.ci_register import register_cuda_ci
|
||||
from sglang.test.run_combined_tests import run_combined_tests
|
||||
from sglang.test.test_utils import ModelLaunchSettings
|
||||
|
||||
register_cuda_ci(est_time=510, suite="nightly-8-gpu-common", nightly=True)
|
||||
register_cuda_ci(est_time=1680, stage="nightly", runner_config="8-gpu-h200")
|
||||
register_cuda_ci(est_time=1680, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
RING_2_5_1T_MODEL_PATH = "inclusionAI/Ring-2.5-1T"
|
||||
|
||||
|
||||
@@ -12,7 +12,7 @@ from sglang.test.test_deterministic_utils import (
|
||||
TestDeterministicBase,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=240, suite="nightly-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=240, stage="nightly", runner_config="1-gpu-large")
|
||||
|
||||
DEEPSEEK_MODEL = "lmsys/sglang-ci-dsv3-test"
|
||||
|
||||
|
||||
@@ -12,7 +12,7 @@ from sglang.test.test_deterministic_utils import (
|
||||
TestDeterministicBase,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=200, suite="nightly-4-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=200, stage="nightly", runner_config="4-gpu-h100")
|
||||
|
||||
QWEN3_NEXT = "Qwen/Qwen3-Next-80B-A3B-Instruct"
|
||||
|
||||
|
||||
@@ -17,7 +17,7 @@ from sglang.test.test_utils import (
|
||||
write_github_step_summary,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=450, suite="nightly-4-gpu-b200", nightly=True)
|
||||
register_cuda_ci(est_time=900, stage="nightly", runner_config="4-gpu-b200")
|
||||
|
||||
FULL_DEEPSEEK_V3_FP4_MODEL_PATH = "nvidia/DeepSeek-V3-0324-FP4"
|
||||
SERVER_LAUNCH_TIMEOUT = 1000
|
||||
|
||||
@@ -12,7 +12,7 @@ from sglang.test.test_utils import (
|
||||
popen_launch_server,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=800, suite="nightly-4-gpu-b200", nightly=True)
|
||||
register_cuda_ci(est_time=1770, stage="nightly", runner_config="4-gpu-b200")
|
||||
|
||||
|
||||
class FlashinferTrtllmGenMoeBackendFP8Base:
|
||||
|
||||
@@ -11,7 +11,7 @@ import unittest
|
||||
from sglang.test.ci.ci_register import register_cuda_ci
|
||||
from sglang.test.server_fixtures.pcg_spec_fixture import PCGSpecBase
|
||||
|
||||
register_cuda_ci(est_time=531, suite="nightly-4-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=130, stage="nightly", runner_config="4-gpu-h100")
|
||||
|
||||
|
||||
class TestPCGWithEAGLE3(PCGSpecBase, unittest.TestCase):
|
||||
|
||||
@@ -15,7 +15,7 @@ from sglang.test.test_utils import (
|
||||
CustomTestCase,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=531, suite="nightly-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=110, stage="nightly", runner_config="1-gpu-large")
|
||||
|
||||
|
||||
class TestPCGWithDFlash(PCGSpecBase, CustomTestCase):
|
||||
|
||||
@@ -8,7 +8,7 @@ import unittest
|
||||
from sglang.test.ci.ci_register import register_cuda_ci
|
||||
from sglang.test.server_fixtures.pcg_spec_fixture import PCGSpecBase
|
||||
|
||||
register_cuda_ci(est_time=531, suite="nightly-4-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=450, stage="nightly", runner_config="4-gpu-h100")
|
||||
|
||||
|
||||
class TestPCGWithMTP(PCGSpecBase, unittest.TestCase):
|
||||
|
||||
@@ -19,7 +19,7 @@ from sglang.test.test_utils import (
|
||||
)
|
||||
|
||||
# CI Registration
|
||||
register_cuda_ci(est_time=180, suite="nightly-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=250, stage="nightly", runner_config="1-gpu-large")
|
||||
register_amd_ci(est_time=180, suite="stage-b-test-1-gpu-large-amd")
|
||||
|
||||
|
||||
|
||||
@@ -20,7 +20,7 @@ register_amd_ci(
|
||||
suite="nightly-amd-accuracy-8-gpu-mi35x-qwen35-triton-dcp",
|
||||
nightly=True,
|
||||
)
|
||||
register_cuda_ci(est_time=4800, suite="nightly-4-gpu-b200", nightly=True)
|
||||
register_cuda_ci(est_time=4800, stage="nightly", runner_config="4-gpu-b200")
|
||||
|
||||
QWEN35_MODEL_PATH = os.environ.get("QWEN3_5_MODEL_PATH", "Qwen/Qwen3.5-397B-A17B-FP8")
|
||||
SERVER_LAUNCH_TIMEOUT = 4800
|
||||
|
||||
@@ -21,7 +21,7 @@ from sglang.test.test_utils import (
|
||||
popen_launch_server,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=40, suite="nightly-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=70, stage="nightly", runner_config="1-gpu-large")
|
||||
register_amd_ci(est_time=40, suite="nightly-amd-1-gpu", nightly=True)
|
||||
register_cpu_ci(est_time=225, suite="base-c-test-cpu")
|
||||
|
||||
|
||||
@@ -50,7 +50,7 @@ from sglang.test.test_utils import (
|
||||
run_distributed_test,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=30, suite="nightly-2-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=30, stage="nightly", runner_config="2-gpu-large")
|
||||
register_amd_ci(est_time=60, suite="nightly-amd", nightly=True)
|
||||
|
||||
|
||||
|
||||
@@ -35,7 +35,7 @@ from sglang.test.test_utils import (
|
||||
popen_launch_server,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=300, suite="nightly-4-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=290, stage="nightly", runner_config="4-gpu-h100")
|
||||
register_amd_ci(
|
||||
est_time=300,
|
||||
suite="nightly-amd-4-gpu",
|
||||
|
||||
@@ -9,6 +9,8 @@ Env knobs:
|
||||
SGLANG_PRECISION_COMMIT override sglang sha (7-40 hex) tagged on push
|
||||
SGLANG_PRECISION_HF_REPO required HF dataset repo for cross-runner
|
||||
baseline storage; see precision_baseline_store
|
||||
SGLANG_PRECISION_HF_TOKEN write token for that repo (not HF_TOKEN, which
|
||||
carries the runner's gated-model read token)
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -47,7 +49,7 @@ try:
|
||||
except Exception: # pragma: no cover
|
||||
_hfs = None
|
||||
|
||||
register_cuda_ci(est_time=3600, suite="nightly-precision-8-gpu-h200", nightly=True)
|
||||
register_cuda_ci(est_time=3600, stage="nightly", runner_config="8-gpu-h200")
|
||||
|
||||
DEFAULT_MODELS_FOR_NIGHTLY_PRECISION = "zai-org/GLM-5.2-FP8"
|
||||
DEFAULT_DIFF_THRESHOLD = 1e-3
|
||||
|
||||
@@ -17,7 +17,7 @@ from sglang.test.test_utils import (
|
||||
popen_launch_server,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=120, suite="nightly-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=240, stage="nightly", runner_config="1-gpu-large")
|
||||
register_amd_ci(est_time=120, suite="nightly-amd-1-gpu", nightly=True)
|
||||
register_cpu_ci(est_time=622, suite="base-c-test-cpu")
|
||||
|
||||
|
||||
@@ -19,7 +19,7 @@ from sglang.test.test_utils import (
|
||||
popen_launch_server,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=1200, suite="nightly-8-gpu-b200", nightly=True)
|
||||
register_cuda_ci(est_time=450, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
KIMI_LINEAR_MODEL = "moonshotai/Kimi-Linear-48B-A3B-Instruct"
|
||||
PHYSICAL_PAGE_SIZE = 64
|
||||
|
||||
@@ -23,7 +23,7 @@ from sglang.test.test_utils import (
|
||||
try_cached_model,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=420, suite="nightly-eval-text-2-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=200, stage="nightly", runner_config="2-gpu-large")
|
||||
|
||||
# 72 routed experts + 48 replicas = 120 physical, 60 per rank, so two thirds of
|
||||
# the routed (token, expert) pairs get double-counted when ranks disagree. At 24
|
||||
|
||||
@@ -23,7 +23,7 @@ from sglang.test.test_utils import (
|
||||
# downloading on cache miss. Use a longer timeout than the default 600s.
|
||||
NIGHTLY_EVAL_SERVER_TIMEOUT = 1800
|
||||
|
||||
register_cuda_ci(est_time=3600, suite="nightly-eval-text-2-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=2880, stage="nightly", runner_config="2-gpu-large")
|
||||
|
||||
MODEL_SCORE_THRESHOLDS = {
|
||||
# sgl-eval (zero-shot chat, \boxed{}, math_verify grading). Thresholds are
|
||||
|
||||
@@ -19,7 +19,7 @@ from sglang.test.test_utils import (
|
||||
# Use a longer timeout than the default 600s.
|
||||
NIGHTLY_EVAL_SERVER_TIMEOUT = 1800
|
||||
|
||||
register_cuda_ci(est_time=7200, suite="nightly-eval-vlm-2-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=7200, stage="nightly", runner_config="2-gpu-large")
|
||||
|
||||
MODEL_THRESHOLDS = {
|
||||
# Conservative thresholds on 100 MMMU samples, especially for latency thresholds
|
||||
|
||||
@@ -6,9 +6,7 @@ from sglang.test.performance_test_runner import PerformanceTestParams
|
||||
from sglang.test.run_combined_tests import run_combined_tests
|
||||
from sglang.test.test_utils import ModelLaunchSettings
|
||||
|
||||
register_cuda_ci(
|
||||
est_time=7200, suite="nightly-4-gpu-gb300-deepseek-v4-pro-fp4", nightly=True
|
||||
)
|
||||
register_cuda_ci(est_time=7200, stage="nightly", runner_config="4-gpu-gb300")
|
||||
|
||||
MODEL_PATH = "deepseek-ai/DeepSeek-V4-Pro"
|
||||
SERVER_LAUNCH_TIMEOUT = 3600
|
||||
|
||||
@@ -6,7 +6,7 @@ from sglang.test.performance_test_runner import PerformanceTestParams
|
||||
from sglang.test.run_combined_tests import run_combined_tests
|
||||
from sglang.test.test_utils import ModelLaunchSettings
|
||||
|
||||
register_cuda_ci(est_time=7200, suite="nightly-4-gpu-gb300-glm5-nvfp4", nightly=True)
|
||||
register_cuda_ci(est_time=2280, stage="nightly", runner_config="4-gpu-gb300")
|
||||
|
||||
MODEL_PATH = "nvidia/GLM-5.2-NVFP4"
|
||||
|
||||
|
||||
@@ -6,9 +6,7 @@ from sglang.test.performance_test_runner import PerformanceTestParams
|
||||
from sglang.test.run_combined_tests import run_combined_tests
|
||||
from sglang.test.test_utils import ModelLaunchSettings
|
||||
|
||||
register_cuda_ci(
|
||||
est_time=7200, suite="nightly-4-gpu-gb300-kimi-k25-nvfp4", nightly=True
|
||||
)
|
||||
register_cuda_ci(est_time=7200, stage="nightly", runner_config="4-gpu-gb300")
|
||||
|
||||
MODEL_PATH = "nvidia/Kimi-K2.5-NVFP4"
|
||||
DRAFT_MODEL_PATH = "lightseekorg/kimi-k2.5-eagle3-mla"
|
||||
|
||||
@@ -6,7 +6,7 @@ from sglang.test.performance_test_runner import PerformanceTestParams
|
||||
from sglang.test.run_combined_tests import run_combined_tests
|
||||
from sglang.test.test_utils import ModelLaunchSettings
|
||||
|
||||
register_cuda_ci(est_time=7200, suite="nightly-4-gpu-gb300-qwen35-fp8", nightly=True)
|
||||
register_cuda_ci(est_time=7200, stage="nightly", runner_config="4-gpu-gb300")
|
||||
|
||||
MODEL_PATH = "Qwen/Qwen3.5-397B-A17B-FP8"
|
||||
|
||||
|
||||
@@ -25,7 +25,7 @@ from sglang.kernels.ops.kv_canary.verify import VerifyPlan
|
||||
from sglang.kernels.ops.kv_canary.write import WritePlan
|
||||
from sglang.test.ci.ci_register import register_amd_ci, register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=900, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=40, stage="nightly", runner_config="1-gpu-large")
|
||||
# AMD mirrors the CUDA nightly registration (nightly-only, no per-PR suite).
|
||||
register_amd_ci(est_time=900, suite="nightly-amd-kernel-1-gpu", nightly=True)
|
||||
|
||||
|
||||
@@ -16,7 +16,7 @@ from sglang.kernels.ops.kv_canary.scatter_req_token_ids import (
|
||||
)
|
||||
from sglang.test.ci.ci_register import register_amd_ci, register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=180, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=20, stage="nightly", runner_config="1-gpu-large")
|
||||
# AMD mirrors the CUDA nightly registration (nightly-only, no per-PR suite).
|
||||
# Note: amd_ci_exec.sh sets SGLANG_IS_IN_CI, so this runs the CI-reduced range
|
||||
# (_BS_AXIS_CI/_SEQ_LEN_AXIS_CI via get_benchmark_range), same as CUDA nightly.
|
||||
|
||||
@@ -32,7 +32,7 @@ from sglang.kernels.ops.kv_canary.verify import (
|
||||
)
|
||||
from sglang.test.ci.ci_register import register_amd_ci, register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=900, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=20, stage="nightly", runner_config="1-gpu-large")
|
||||
# AMD mirrors the CUDA nightly registration (nightly-only, no per-PR suite).
|
||||
# Note: amd_ci_exec.sh sets SGLANG_IS_IN_CI, so this runs the CI-reduced range
|
||||
# (build_fast_matrix_cases via get_benchmark_range), same as CUDA nightly.
|
||||
|
||||
@@ -30,7 +30,7 @@ from sglang.kernels.ops.kv_canary.verify import (
|
||||
from sglang.kernels.ops.kv_canary.write import WritePlan, launch_canary_write_kernel
|
||||
from sglang.test.ci.ci_register import register_amd_ci, register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=900, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=20, stage="nightly", runner_config="1-gpu-large")
|
||||
# AMD mirrors the CUDA nightly registration (nightly-only, no per-PR suite).
|
||||
# Note: amd_ci_exec.sh sets SGLANG_IS_IN_CI, so this runs the CI-reduced range
|
||||
# (build_fast_matrix_cases via get_benchmark_range), same as CUDA nightly.
|
||||
|
||||
@@ -14,7 +14,7 @@ from sglang.test.ci.ci_register import register_amd_ci, register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=20, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=30, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=20, stage="nightly", runner_config="1-gpu-large")
|
||||
register_amd_ci(est_time=20, stage="jit-kernel-unit", runner_config="amd")
|
||||
|
||||
|
||||
|
||||
@@ -14,7 +14,7 @@ from sglang.srt.layers.attention.dsa.utils import (
|
||||
from sglang.srt.utils import is_sm100_supported
|
||||
from sglang.test.ci.ci_register import register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=180, suite="nightly-4-gpu-b200", nightly=True)
|
||||
register_cuda_ci(est_time=180, stage="nightly", runner_config="4-gpu-b200")
|
||||
|
||||
BLOCK_KV = 64
|
||||
HEAD_DIM = 128
|
||||
|
||||
@@ -17,7 +17,7 @@ from sglang.srt.layers.attention.dsa.utils import (
|
||||
from sglang.srt.utils import is_sm90_supported, is_sm100_supported
|
||||
from sglang.test.ci.ci_register import register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=40, suite="nightly-4-gpu-b200", nightly=True)
|
||||
register_cuda_ci(est_time=40, stage="nightly", runner_config="4-gpu-b200")
|
||||
|
||||
BLOCK_KV = 64
|
||||
HEAD_DIM = 128
|
||||
|
||||
@@ -10,7 +10,7 @@ from sglang.test.ci.ci_register import register_amd_ci, register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=64, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=256, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=390, stage="nightly", runner_config="1-gpu-large")
|
||||
register_amd_ci(est_time=64, suite="jit-kernel-unit-test-amd")
|
||||
|
||||
DEVICE = "cuda"
|
||||
|
||||
@@ -46,11 +46,7 @@ register_cuda_ci(
|
||||
runner_config="8-gpu-h200",
|
||||
)
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(
|
||||
est_time=300,
|
||||
suite="nightly-kernel-8-gpu-h200",
|
||||
nightly=True,
|
||||
)
|
||||
register_cuda_ci(est_time=110, stage="nightly", runner_config="8-gpu-h200")
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Test parameters
|
||||
|
||||
@@ -36,7 +36,7 @@ from sglang.test.kernels.utils import multigpu_pytest_main
|
||||
|
||||
register_cuda_ci(est_time=240, stage="extra-b", runner_config="8-gpu-h200")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=240, suite="nightly-kernel-8-gpu-h200", nightly=True)
|
||||
register_cuda_ci(est_time=70, stage="nightly", runner_config="8-gpu-h200")
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Test parameters
|
||||
|
||||
@@ -10,7 +10,7 @@ from sglang.test.ci.ci_register import register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=44, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=176, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=220, stage="nightly", runner_config="1-gpu-large")
|
||||
|
||||
DEVICE = "cuda"
|
||||
DTYPE = torch.bfloat16
|
||||
|
||||
@@ -14,7 +14,7 @@ from sglang.test.ci.ci_register import register_amd_ci, register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=15, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=120, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=30, stage="nightly", runner_config="1-gpu-large")
|
||||
register_amd_ci(est_time=30, suite="nightly-amd-kernel-1-gpu", nightly=True)
|
||||
|
||||
DEVICE = "cuda"
|
||||
|
||||
@@ -18,7 +18,7 @@ from sglang.test.ci.ci_register import register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=16, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=120, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=20, stage="nightly", runner_config="1-gpu-large")
|
||||
|
||||
CORRECTNESS_BATCH_SIZES = get_ci_test_range(
|
||||
[1, 2, 8, 128, 256, 512, 1536, 2048, 4096, 11008, 16384],
|
||||
|
||||
@@ -17,7 +17,7 @@ from sglang.test.ci.ci_register import register_amd_ci, register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=10, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=60, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=20, stage="nightly", runner_config="1-gpu-large")
|
||||
register_amd_ci(est_time=15, suite="nightly-amd-kernel-1-gpu", nightly=True)
|
||||
|
||||
DEVICE = "cuda"
|
||||
|
||||
@@ -31,7 +31,7 @@ from sglang.test.ci.ci_register import register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=15, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=60, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=20, stage="nightly", runner_config="1-gpu-large")
|
||||
|
||||
DEVICE = "cuda"
|
||||
DTYPES = get_ci_test_range([torch.bfloat16, torch.float16], [torch.bfloat16])
|
||||
|
||||
@@ -16,7 +16,7 @@ from sglang.test.ci.ci_register import register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=37, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=148, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=110, stage="nightly", runner_config="1-gpu-large")
|
||||
|
||||
HIDDEN_DIMS = [1024, 4096, 5120, 6144, 7168]
|
||||
ROUTER_GEMM_CASES = get_ci_test_range(
|
||||
|
||||
@@ -24,7 +24,7 @@ from sglang.test.ci.ci_register import register_cuda_ci
|
||||
from sglang.test.kernels.utils import multigpu_pytest_main
|
||||
|
||||
register_cuda_ci(est_time=240, stage="base-c", runner_config="4-gpu-b200")
|
||||
register_cuda_ci(est_time=480, suite="nightly-8-gpu-b200", nightly=True)
|
||||
register_cuda_ci(est_time=480, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
_HIDDEN_SIZE = 7168
|
||||
_GEMM_AR_K_TOTAL = 12288
|
||||
|
||||
@@ -10,7 +10,7 @@ from sglang.test.ci.ci_register import register_amd_ci, register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=28, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=120, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=40, stage="nightly", runner_config="1-gpu-large")
|
||||
register_amd_ci(est_time=55, stage="jit-kernel-unit", runner_config="amd")
|
||||
|
||||
BS_LIST = [2**n for n in range(0, 15)]
|
||||
|
||||
@@ -9,7 +9,7 @@ from sglang.test.ci.ci_register import register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=10, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=120, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=20, stage="nightly", runner_config="1-gpu-large")
|
||||
|
||||
|
||||
def sglang_jit_fused_add_rmsnorm(
|
||||
|
||||
@@ -10,7 +10,7 @@ from sglang.test.ci.ci_register import register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=37, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=148, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=130, stage="nightly", runner_config="1-gpu-large")
|
||||
|
||||
|
||||
def sglang_aot_qknorm(
|
||||
|
||||
@@ -10,7 +10,7 @@ from sglang.test.ci.ci_register import register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=15, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=120, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=20, stage="nightly", runner_config="1-gpu-large")
|
||||
|
||||
|
||||
def sglang_jit_qknorm_across_heads(
|
||||
|
||||
@@ -10,7 +10,7 @@ from sglang.test.ci.ci_register import register_amd_ci, register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=45, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=240, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=160, stage="nightly", runner_config="1-gpu-large")
|
||||
register_amd_ci(est_time=45, suite="jit-kernel-unit-test-amd")
|
||||
|
||||
|
||||
|
||||
@@ -15,7 +15,7 @@ from sglang.test.ci.ci_register import register_amd_ci, register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=30, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=120, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=100, stage="nightly", runner_config="1-gpu-large")
|
||||
register_amd_ci(est_time=30, stage="jit-kernel-unit", runner_config="amd")
|
||||
|
||||
EPS = 1e-5
|
||||
|
||||
@@ -16,7 +16,7 @@ from sglang.test.quant_ref_utils import (
|
||||
dequantize_nvfp4_to_dtype,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=300, suite="nightly-4-gpu-b200", nightly=True)
|
||||
register_cuda_ci(est_time=300, stage="nightly", runner_config="4-gpu-b200")
|
||||
|
||||
if torch.cuda.get_device_capability() < (10, 0):
|
||||
pytest.skip(
|
||||
|
||||
@@ -7,7 +7,7 @@ from sglang.srt.layers.moe.topk import biased_grouped_topk_gpu, biased_grouped_t
|
||||
from sglang.srt.utils import get_device
|
||||
from sglang.test.ci.ci_register import register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=2, suite="nightly-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=40, stage="nightly", runner_config="1-gpu-large")
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
|
||||
@@ -12,7 +12,7 @@ from sglang.test.ci.ci_register import register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=28, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=120, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=50, stage="nightly", runner_config="1-gpu-large")
|
||||
|
||||
|
||||
def ceil_div(a, b):
|
||||
|
||||
@@ -13,7 +13,7 @@ from sglang.test.ci.ci_register import register_cuda_ci
|
||||
|
||||
register_cuda_ci(est_time=16, stage="base-b-kernel-unit", runner_config="1-gpu-large")
|
||||
# Nightly is not redundant here: it sets SGLANG_JIT_KERNEL_RUN_FULL_TESTS=1 to expand get_ci_test_range sweeps.
|
||||
register_cuda_ci(est_time=120, suite="nightly-kernel-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=20, stage="nightly", runner_config="1-gpu-large")
|
||||
|
||||
try:
|
||||
from sglang.srt.utils import is_hip
|
||||
|
||||
@@ -34,10 +34,7 @@ import sglang as sgl
|
||||
from sglang.test.ci.ci_register import register_cuda_ci
|
||||
from sglang.test.test_utils import CustomTestCase
|
||||
|
||||
register_cuda_ci(
|
||||
est_time=300,
|
||||
suite="nightly-8-gpu-b200",
|
||||
)
|
||||
register_cuda_ci(est_time=1800, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
BASE_MODEL = "deepseek-ai/DeepSeek-V3.1-Base"
|
||||
LORA_HF_REPO = "yushengsu/lora-diff-DeepSeek-V3.1-Base"
|
||||
|
||||
@@ -27,7 +27,7 @@ from sglang.test.ci.ci_register import (
|
||||
register_xpu_ci,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=200, suite="nightly-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=10, stage="nightly", runner_config="1-gpu-large")
|
||||
register_amd_ci(est_time=200, suite="nightly-amd-1-gpu", nightly=True)
|
||||
register_cpu_ci(est_time=6, suite="base-c-test-cpu")
|
||||
register_xpu_ci(est_time=10, suite="stage-a-test-1-gpu-xpu")
|
||||
|
||||
@@ -34,10 +34,7 @@ import sglang as sgl
|
||||
from sglang.test.ci.ci_register import register_cuda_ci
|
||||
from sglang.test.test_utils import CustomTestCase
|
||||
|
||||
register_cuda_ci(
|
||||
est_time=360,
|
||||
suite="nightly-8-gpu-b200",
|
||||
)
|
||||
register_cuda_ci(est_time=420, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
BASE_MODEL = "moonshotai/Kimi-K2.5"
|
||||
LORA_HF_REPO = "yushengsu/lora-diff-Kimi-K2.5"
|
||||
|
||||
@@ -27,7 +27,7 @@ from sglang.test.test_utils import (
|
||||
popen_launch_server,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=150, suite="nightly-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=180, stage="nightly", runner_config="1-gpu-large")
|
||||
register_amd_ci(est_time=150, suite="nightly-amd-1-gpu", nightly=True)
|
||||
|
||||
|
||||
|
||||
@@ -4,7 +4,7 @@ import sglang as sgl
|
||||
from sglang.test.ci.ci_register import register_cuda_ci
|
||||
from sglang.test.test_utils import CustomTestCase
|
||||
|
||||
register_cuda_ci(est_time=300, suite="nightly-4-gpu")
|
||||
register_cuda_ci(est_time=300, stage="nightly", runner_config="4-gpu-h100")
|
||||
|
||||
PROMPTS = [
|
||||
"Hello, my name is",
|
||||
|
||||
@@ -5,7 +5,7 @@ from sglang.srt.utils.common import temp_set_env
|
||||
from sglang.test.ci.ci_register import register_cuda_ci
|
||||
from sglang.test.test_utils import CustomTestCase
|
||||
|
||||
register_cuda_ci(est_time=380, suite="nightly-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=50, stage="nightly", runner_config="1-gpu-large")
|
||||
|
||||
TEST_GCS_MODEL = "gs://vertex-model-garden-public-us/codegemma/codegemma-2b/"
|
||||
|
||||
|
||||
@@ -20,11 +20,7 @@ from sglang.test.server_fixtures.disaggregation_fixture import (
|
||||
PDDisaggregationServerBase,
|
||||
)
|
||||
|
||||
register_cuda_ci(
|
||||
est_time=450,
|
||||
suite="nightly-8-gpu-b200",
|
||||
nightly=True,
|
||||
)
|
||||
register_cuda_ci(est_time=750, stage="nightly", runner_config="8-gpu-b200")
|
||||
|
||||
|
||||
class TestGLM52DSACacheLayerSplit(PDDisaggregationServerBase, GSM8KMixin):
|
||||
|
||||
@@ -32,7 +32,7 @@ from sglang.test.kits.eval_accuracy_kit import GSM8KMixin
|
||||
from sglang.test.kits.prefix_cache_branching_kit import PrefixCacheBranchingMixin
|
||||
from sglang.test.server_fixtures.default_fixture import DefaultServerBase
|
||||
|
||||
register_cuda_ci(est_time=600, suite="nightly-4-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=570, stage="nightly", runner_config="4-gpu-h100")
|
||||
|
||||
KIMI_LINEAR_MODEL = "moonshotai/Kimi-Linear-48B-A3B-Instruct"
|
||||
|
||||
|
||||
@@ -17,7 +17,7 @@ from sglang.test.test_utils import (
|
||||
|
||||
# 60 test classes testing hybrid parallelism configurations
|
||||
# Each test launches server + runs MMLU eval (~90s per test)
|
||||
register_cuda_ci(est_time=5400, suite="weekly-8-gpu-h200", nightly=True)
|
||||
register_cuda_ci(est_time=8160, stage="weekly", runner_config="8-gpu-h200")
|
||||
|
||||
|
||||
class Test00(CustomTestCase):
|
||||
|
||||
@@ -9,7 +9,7 @@ from sglang.test.test_utils import (
|
||||
parse_models,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=3600, suite="nightly-perf-text-2-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=450, stage="nightly", runner_config="2-gpu-large")
|
||||
|
||||
RESULT_DIR = "performance_results_text_models"
|
||||
|
||||
|
||||
@@ -11,7 +11,7 @@ from sglang.test.test_utils import (
|
||||
parse_models,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=7200, suite="nightly-perf-vlm-2-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=750, stage="nightly", runner_config="2-gpu-large")
|
||||
|
||||
RESULT_DIR = "performance_results_vlms"
|
||||
|
||||
|
||||
@@ -23,7 +23,7 @@ MIMO_LAUNCH_TIMEOUT = 3600
|
||||
|
||||
# MiMo V2.5 is pre-cached on the eight-H200 runner. The H200-only nightly suite
|
||||
# exercises the asymmetric MHA host pool end to end without adding PR CI cost.
|
||||
register_cuda_ci(est_time=1200, suite="nightly-8-gpu-h200", nightly=True)
|
||||
register_cuda_ci(est_time=270, stage="nightly", runner_config="8-gpu-h200")
|
||||
|
||||
|
||||
class TestUnifiedMiMoHiCacheLoadBackKL(CustomTestCase):
|
||||
|
||||
+1
-1
@@ -26,7 +26,7 @@ from sglang.test.test_utils import (
|
||||
GLM5_MODEL = "zai-org/GLM-5.2-FP8"
|
||||
GLM5_LAUNCH_TIMEOUT = 3600
|
||||
|
||||
register_cuda_ci(est_time=900, suite="nightly-8-gpu-h200", nightly=True)
|
||||
register_cuda_ci(est_time=690, stage="nightly", runner_config="8-gpu-h200")
|
||||
|
||||
|
||||
class AccuracyTwoPassMixin:
|
||||
|
||||
@@ -34,7 +34,7 @@ from sglang.test.test_utils import (
|
||||
popen_launch_server,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=150, suite="nightly-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=60, stage="nightly", runner_config="1-gpu-large")
|
||||
|
||||
_MODEL_NAME = "Qwen/Qwen3-0.6B"
|
||||
# We address the up half via the HF-style unfused name "up_proj.weight". sglang's
|
||||
|
||||
@@ -20,7 +20,7 @@ from sglang.test.test_utils import (
|
||||
popen_launch_server,
|
||||
)
|
||||
|
||||
register_cuda_ci(est_time=120, suite="nightly-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=50, stage="nightly", runner_config="1-gpu-large")
|
||||
register_amd_ci(est_time=120, suite="nightly-amd-1-gpu", nightly=True)
|
||||
register_cpu_ci(est_time=184, suite="base-c-test-cpu")
|
||||
|
||||
|
||||
@@ -11,7 +11,7 @@ from sglang.test.test_utils import CustomTestCase
|
||||
|
||||
# Note: MI300 (gfx942) has 64KB shared memory limit but kernel needs 66KB
|
||||
# MI35x (gfx950/CDNA4) may have different limits - testing on MI35x only
|
||||
register_cuda_ci(est_time=10, suite="nightly-1-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=20, stage="nightly", runner_config="1-gpu-large")
|
||||
register_amd_ci(est_time=10, suite="nightly-amd-1-gpu-mi35x", nightly=True)
|
||||
|
||||
device_type = getattr(torch.accelerator.current_accelerator(), "type", "cpu")
|
||||
|
||||
@@ -7,7 +7,7 @@ from sglang.test.ci.ci_register import register_amd_ci, register_cuda_ci
|
||||
from sglang.test.kits.mmmu_vlm_kit import MMMUMultiModelTestBase
|
||||
from sglang.test.test_utils import is_in_ci
|
||||
|
||||
register_cuda_ci(est_time=500, suite="nightly-4-gpu", nightly=True)
|
||||
register_cuda_ci(est_time=510, stage="nightly", runner_config="4-gpu-h100")
|
||||
register_amd_ci(est_time=500, suite="nightly-amd-4-gpu", nightly=True)
|
||||
|
||||
MODELS = [
|
||||
|
||||
+33
-40
@@ -121,34 +121,16 @@ PER_COMMIT_SUITES = {
|
||||
# Nightly test suites (run nightly, organized by GPU configuration)
|
||||
NIGHTLY_SUITES = {
|
||||
HWBackend.CUDA: [
|
||||
"nightly-1-gpu",
|
||||
"nightly-2-gpu",
|
||||
"nightly-4-gpu",
|
||||
"nightly-4-gpu-b200",
|
||||
"nightly-8-gpu",
|
||||
"nightly-8-gpu-h200",
|
||||
"nightly-8-gpu-h20",
|
||||
"nightly-8-gpu-b200",
|
||||
"nightly-8-gpu-h200-basic", # Basic tests for large models on H200
|
||||
"nightly-8-gpu-b200-basic", # Basic tests for large models on B200
|
||||
"nightly-8-gpu-common", # Common tests that run on both H200 and B200
|
||||
"nightly-kernel-1-gpu",
|
||||
"nightly-kernel-8-gpu-h200",
|
||||
# Eval and perf suites (2-gpu)
|
||||
"nightly-eval-text-2-gpu",
|
||||
"nightly-eval-vlm-2-gpu",
|
||||
"nightly-perf-text-2-gpu",
|
||||
"nightly-perf-vlm-2-gpu",
|
||||
# GB300 (4x GB300 NVL4) nightly suites
|
||||
"nightly-4-gpu-gb300",
|
||||
"nightly-4-gpu-gb300-deepseek-v4-pro-fp4",
|
||||
"nightly-4-gpu-gb300-glm5-nvfp4",
|
||||
"nightly-4-gpu-gb300-kimi-k25",
|
||||
"nightly-4-gpu-gb300-kimi-k25-nvfp4",
|
||||
"nightly-4-gpu-gb300-qwen35-fp8",
|
||||
"nightly-4-gpu-gb300-qwen35-nvfp4",
|
||||
# Nightly precision regression (per-layer hidden state comparison)
|
||||
"nightly-precision-8-gpu-h200",
|
||||
# `stage="nightly"` + a runner_config, same `{stage}-test-{runner_config}`
|
||||
# shape as the per-commit suites. No `nightly=True`: the stage name
|
||||
# carries the cadence; only the legacy suites below still need the flag.
|
||||
"nightly-test-1-gpu-large",
|
||||
"nightly-test-2-gpu-large",
|
||||
"nightly-test-4-gpu-h100",
|
||||
"nightly-test-4-gpu-b200",
|
||||
"nightly-test-4-gpu-gb300",
|
||||
"nightly-test-8-gpu-h200",
|
||||
"nightly-test-8-gpu-b200",
|
||||
],
|
||||
HWBackend.AMD: [
|
||||
"nightly-amd",
|
||||
@@ -194,7 +176,9 @@ OTHER_SUITES = {
|
||||
],
|
||||
HWBackend.CUDA: [
|
||||
"stress",
|
||||
"weekly-8-gpu-h200",
|
||||
# `stage="weekly"` -- same shape. The three dicts group names for
|
||||
# readability only; validation reads their union.
|
||||
"weekly-test-8-gpu-h200",
|
||||
],
|
||||
}
|
||||
|
||||
@@ -244,14 +228,10 @@ def filter_tests(
|
||||
if t.backend == hw and t.effective_suite == suite and t.nightly == nightly
|
||||
]
|
||||
|
||||
valid_suites = (
|
||||
NIGHTLY_SUITES.get(hw, []) if nightly else PER_COMMIT_SUITES.get(hw, [])
|
||||
)
|
||||
|
||||
if suite not in valid_suites:
|
||||
print(
|
||||
f"Warning: Unknown suite {suite} for backend {hw.name}, nightly={nightly}"
|
||||
)
|
||||
# Union of all three dicts, not just the per-commit or nightly half:
|
||||
# CUDA nightly suites are selected by name alone, without --nightly.
|
||||
if suite not in _valid_suites_by_backend().get(hw, set()):
|
||||
print(f"Warning: Unknown suite {suite} for backend {hw.name}")
|
||||
|
||||
enabled_tests = [t for t in ci_tests if t.disabled is None]
|
||||
skipped_tests = [t for t in ci_tests if t.disabled is not None]
|
||||
@@ -369,9 +349,11 @@ def run_a_suite(args):
|
||||
|
||||
pretty_print_tests(args, ci_tests, skipped_tests)
|
||||
|
||||
# None hands the per-file budget over to est_time (see run_unittest_files).
|
||||
timeout = None if args.timeout_from_est_time else args.timeout_per_file
|
||||
|
||||
# Add extra timeout when retry is enabled
|
||||
timeout = args.timeout_per_file
|
||||
if args.enable_retry:
|
||||
if timeout is not None and args.enable_retry:
|
||||
timeout += args.retry_timeout_increase
|
||||
|
||||
return run_unittest_files(
|
||||
@@ -399,7 +381,10 @@ def main():
|
||||
parser.add_argument(
|
||||
"--nightly",
|
||||
action="store_true",
|
||||
help="Run nightly tests instead of per-commit tests.",
|
||||
help=(
|
||||
"Include tests registered with nightly=True (AMD/CPU/NPU). CUDA "
|
||||
"scheduled suites are selected by name and take no flag."
|
||||
),
|
||||
)
|
||||
parser.add_argument(
|
||||
"--timeout-per-file",
|
||||
@@ -407,6 +392,14 @@ def main():
|
||||
default=1200,
|
||||
help="The time limit for running one file in seconds (default: 1200).",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--timeout-from-est-time",
|
||||
action="store_true",
|
||||
help=(
|
||||
"Derive each file's time limit from its own est_time instead of "
|
||||
"the flat --timeout-per-file, for suites mixing fast and slow tests."
|
||||
),
|
||||
)
|
||||
parser.add_argument(
|
||||
"--continue-on-error",
|
||||
action="store_true",
|
||||
|
||||
Reference in New Issue
Block a user