transformers
1e718279 - [CacheHardIntegrationTest] use dedicated safetensors repo to fix Xet bucket cache corruption (#49182)

Commit
8 days ago
[CacheHardIntegrationTest] use dedicated safetensors repo to fix Xet bucket cache corruption (#49182) * [CacheHardIntegrationTest] use dedicated tiny repo with safetensors to avoid Xet corruption The original `hf-internal-testing/tiny-random-GPTJForCausalLM` repo only has `pytorch_model.bin`, which the Xet/bucket cache serves with corrupted wte.weight bytes (same file size, wrong content) — causing all 5 test_cache_gptj_model_* tests to fail on bucket cache CI runners since Sep 26. Switch to a new dedicated repo that stores the same weights as safetensors, uploaded directly from the EFS runner where the correct model was confirmed. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * [CacheHardIntegrationTest] revert to original model, keep backup repo note The Xet chunk corruption that caused bucket-cache failures (Sep 25-26) was resolved on the HF infrastructure side. Reverting to the original tiny model; the safetensors-based backup repo is noted in a comment for future reference. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * [CacheHardIntegrationTest] add @slow to test_cache_gptj_model Ensures these tests are included in slow CI runs (PR comment CI uses -m slow but the test only had @require_torch_accelerator, so it was being skipped there). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * [CacheHardIntegrationTest] fix @slow placement for @parameterized.expand @slow must be below @parameterized.expand (not above) so the mark is inherited by the generated test methods — matching the pattern used by test_static_cache_greedy_decoding_pad_left. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * [CacheHardIntegrationTest] strip file to only gptj test for CI verification Temporary: keep only test_cache_gptj_model to verify @slow + @parameterized.expand works correctly in CI slow runs. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * [CacheHardIntegrationTest] use safetensors repo to fix intermittent Xet corruption The Xet bucket cache intermittently serves corrupted pytorch_model.bin for tiny-random-GPTJForCausalLM (wrong wte.weight → wrong golden outputs). Using the dedicated safetensors-based repo avoids this entirely. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * [CacheHardIntegrationTest] use original model_id for CI verification Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * [CacheHardIntegrationTest] use safetensors repo to fix Xet FUSE corruption Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * [CacheHardIntegrationTest] fix Xet FUSE corruption: use safetensors repo The original `hf-internal-testing/tiny-random-GPTJForCausalLM` repo uses pytorch_model.bin. On Nvidia bucket CI runners, /mnt/cache is mounted via Xet FUSE, and the Xet chunks for that blob are corrupted — causing wrong `wte.weight` bytes and thus wrong golden outputs for all 5 parameterized `test_cache_gptj_model` variants. Fix: use `hf-internal-testing/tiny-random-GPTJForCausalLM-for-CacheHardIntegrationTest` which stores weights in safetensors format (unaffected by Xet chunk corruption). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> --------- Co-authored-by: ydshieh <ydshieh@users.noreply.github.com>
Author
Parents
Loading