transformers
65dc402e - Fix CI slack reporting: surface crashed/missing model jobs in model_results.json (#49427)

Commit
2 days ago
Fix CI slack reporting: surface crashed/missing model jobs in model_results.json (#49427) * notification_service: populate job_link for no-artifact jobs Add MACHINE_TYPE_TO_GPU mapping and matrix_name_to_github_jobs lookup to recover GitHub job URLs for jobs that produced no artifact (OOM kills, silent exits, container init failures). Two-tier lookup in the no-artifact branch: 1. artifact_name_to_job_map (keyed by "single-gpu_/multi-gpu_ + artifact_key"): covers OOM-killed jobs where the "Test suite reports artifacts" step still appeared in the GitHub API (conclusion=failure/skipped). 2. matrix_name_to_github_jobs (parsed from job name + runner_group_name): fallback for early container failures where that step never ran at all. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fixup: correct regex for parsing matrix folder from job name Job name format is: "Model CI / run_models_gpu (aws-g5-12xlarge-cache, 0) / run_models_gpu (models/mistral4)" The folder is in the LAST parenthetical group, not the first. Use r"\(([^)]+)\)\s*$" to anchor to end of string. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * notification_service: detect partial GPU slot crash (one slot missing artifact) When one GPU slot uploaded its artifact but the other did not (e.g. OCI exec error on the upload step), the inner loop only processes the slot that succeeded. The missing slot stays invisible: error=False, no job_link. Fix: collect actual_gpus from paths before the loop, then after the loop diff against expected_gpus from matrix_name_to_github_jobs. Any slot in the diff gets error=True and a job_link populated via the same two-tier lookup used for the all-crash case. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * notification_service: add per-GPU error_type field to model_results.json error_type is a dict keyed by GPU slot ("single"/"multi"), mirroring the structure of job_link and failures. Two values for now: - "no_artifact": slot was scheduled but uploaded nothing - "incomplete_artifact": artifact exists but summary_short is missing Keeps error (bool) unchanged for backwards compatibility. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fixup: use "crashed" instead of "no_artifact" for error_type More honest about what we know — something crashed, cause TBD. Will be refined after discussion with Tarek. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * style: ruff format notification_service.py Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * trigger model CI for bert/gpt2/vit with intentional crash and failure to test PR #49427 Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * use single-GPU-only transformers-ci branch for test run Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * trigger CI with multi-gpu restored Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * increase bert OOM test to 400 GB to crash both single and multi GPU runners Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * revert debugging changes used to validate PR #49427 Remove the intentional OOM crash in bert, the intentional assert failure in gpt2, and restore the full self-scheduled-caller.yml (all CI jobs, correct Slack channel, main branch of transformers-ci reusable workflow, no subdirs restriction). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
Author
Parents
Loading