ci: add issue filing and suspect PR identification to log analyzer (#22899)

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
ishandhanani
2026-04-15 14:27:14 -07:00
committed by GitHub
co-authored by Claude Opus 4.6
parent efc267ca29
commit 9497001b0c
3 changed files with 61 additions and 19 deletions
View File
+41
View File
@@ -94,9 +94,50 @@ Write markdown with this structure:
### Noise
- warnings that were not causal
### Suspect PRs (sglang)
- PR #NNNN: "<title>" — reason this could be related
(only if the failure may originate in sglang)
### Recommended Fix
...
```
Keep the report concrete. Avoid generic summaries. If you are unsure, say so
and explain what evidence is missing.
## Filing Issues
After completing your analysis, if the root cause is actionable and clearly
attributable to a specific repo, open a GitHub issue using `gh issue create`.
**Rules:**
- Only file an issue if you have concrete evidence (specific error, file, line).
Do NOT file issues for flaky infra, transient timeouts, or unclear failures.
- One issue per root cause. Do not create duplicates — search existing issues
first with `gh issue list --repo <repo> --search "<keywords>"`.
- File against the correct repo:
- **`NVIDIA/srt-slurm`**: orchestration bugs, config handling, srtctl behavior,
recipe/YAML issues, incorrect flags or environment variables being passed to
workers, worker launch failures caused by srt-slurm itself.
- **`sgl-project/sglang`**: Do NOT auto-file issues here. Instead, use `gh` to
review recent commits from the past day on the sglang repo:
```
gh api repos/sgl-project/sglang/commits?since=$(date -u -d '24 hours ago' +%Y-%m-%dT%H:%M:%SZ)&per_page=50
```
Identify any commits/PRs that could plausibly have caused the failure based
on the files changed and the error you found. List these as "Suspect PRs" in
the report with links and a brief explanation of why each is suspicious. Let
the human decide whether to follow up.
- Use this format for srt-slurm issues:
```
gh issue create --repo NVIDIA/srt-slurm \
--title "<concise title>" \
--body "<body>"
```
- The issue body should include:
- A short summary of the failure
- The exact error message and which log file it came from
- The job ID and relevant config (model, flags, etc.)
- Do NOT include API keys, tokens, or secrets in the issue.
- If you are unsure which repo to file against, or if the failure is ambiguous,
do NOT file an issue. Just note it in the report.
+20 -19
View File
@@ -9,25 +9,26 @@
# nightly-test-<runner>.yml workflow.
# Never edit workflow YAML files directly for these changes.
dsr1-fp8-gb200-dynamo-sglang:
model: deepseek-ai/DeepSeek-R1-0528
model-prefix: dsr1
runner: gb200
precision: fp8
framework: dynamo-sglang
multinode: true
disagg: true
seq-len-configs:
- isl: 1024
osl: 1024
search-space:
- conc-list: [1024, 2048, 4096, 6144]
# https://github.com/NVIDIA/srt-slurm/blob/sglang-nightly-regression/recipes/gb200-fp8/1k1k/max-tpt.yaml
config_file: recipes/gb200-fp8/1k1k/max-tpt.yaml
- conc-list: [4096]
# https://github.com/NVIDIA/srt-slurm/blob/sglang-nightly-regression/recipes/gb200-fp8/1k1k/ultra-tpt.yaml
config_file: recipes/gb200-fp8/1k1k/ultra-tpt.yaml
# TODO: re-enable after testing log analyzer (see follow-up PR)
# dsr1-fp8-gb200-dynamo-sglang:
# model: deepseek-ai/DeepSeek-R1-0528
# model-prefix: dsr1
# runner: gb200
# precision: fp8
# framework: dynamo-sglang
# multinode: true
# disagg: true
# seq-len-configs:
# - isl: 1024
# osl: 1024
# search-space:
# - conc-list: [1024, 2048, 4096, 6144]
# # https://github.com/NVIDIA/srt-slurm/blob/sglang-nightly-regression/recipes/gb200-fp8/1k1k/max-tpt.yaml
# config_file: recipes/gb200-fp8/1k1k/max-tpt.yaml
#
# - conc-list: [4096]
# # https://github.com/NVIDIA/srt-slurm/blob/sglang-nightly-regression/recipes/gb200-fp8/1k1k/ultra-tpt.yaml
# config_file: recipes/gb200-fp8/1k1k/ultra-tpt.yaml
dsr1-fp4-gb200-dynamo-sglang:
model: nvidia/DeepSeek-R1-0528-NVFP4-v2