Fix CI slack reporting: surface crashed/missing model jobs in model_results.json (#49427)
* notification_service: populate job_link for no-artifact jobs
Add MACHINE_TYPE_TO_GPU mapping and matrix_name_to_github_jobs lookup to
recover GitHub job URLs for jobs that produced no artifact (OOM kills,
silent exits, container init failures).
Two-tier lookup in the no-artifact branch:
1. artifact_name_to_job_map (keyed by "single-gpu_/multi-gpu_ + artifact_key"):
covers OOM-killed jobs where the "Test suite reports artifacts" step still
appeared in the GitHub API (conclusion=failure/skipped).
2. matrix_name_to_github_jobs (parsed from job name + runner_group_name):
fallback for early container failures where that step never ran at all.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fixup: correct regex for parsing matrix folder from job name
Job name format is:
"Model CI / run_models_gpu (aws-g5-12xlarge-cache, 0) / run_models_gpu (models/mistral4)"
The folder is in the LAST parenthetical group, not the first.
Use r"\(([^)]+)\)\s*$" to anchor to end of string.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* notification_service: detect partial GPU slot crash (one slot missing artifact)
When one GPU slot uploaded its artifact but the other did not (e.g. OCI
exec error on the upload step), the inner loop only processes the slot
that succeeded. The missing slot stays invisible: error=False, no job_link.
Fix: collect actual_gpus from paths before the loop, then after the loop
diff against expected_gpus from matrix_name_to_github_jobs. Any slot in
the diff gets error=True and a job_link populated via the same two-tier
lookup used for the all-crash case.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* notification_service: add per-GPU error_type field to model_results.json
error_type is a dict keyed by GPU slot ("single"/"multi"), mirroring the
structure of job_link and failures. Two values for now:
- "no_artifact": slot was scheduled but uploaded nothing
- "incomplete_artifact": artifact exists but summary_short is missing
Keeps error (bool) unchanged for backwards compatibility.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* fixup: use "crashed" instead of "no_artifact" for error_type
More honest about what we know — something crashed, cause TBD.
Will be refined after discussion with Tarek.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* style: ruff format notification_service.py
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* trigger model CI for bert/gpt2/vit with intentional crash and failure to test PR #49427
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* use single-GPU-only transformers-ci branch for test run
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* trigger CI with multi-gpu restored
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* increase bert OOM test to 400 GB to crash both single and multi GPU runners
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
* revert debugging changes used to validate PR #49427
Remove the intentional OOM crash in bert, the intentional assert failure
in gpt2, and restore the full self-scheduled-caller.yml (all CI jobs,
correct Slack channel, main branch of transformers-ci reusable workflow,
no subdirs restriction).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
---------
Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>