DeepSpeed
a4490b2a - Loosen the late-iteration gradient tolerance in the checkpointing test (#8648)

Commit
15 days ago
Loosen the late-iteration gradient tolerance in the checkpointing test (#8648) ## Description `TestZeroUserBackwardWithCheckpointing::test_checkpointed_multiple_backward` compares DDP and engine gradients over three optimizer iterations at dtype-default tolerances, and flips between pass and fail across torch versions on cpu (fails on 2.10/2.11 kernels, passes on 2.13) without any behavioral change. A minimal two-rank repro isolates the mechanism: ``` torch 2.13 torch 2.11 (= CI family) iter 0 grads exact match exact match post-step — linear2.weight differs by 1.7e-6 <- the seed iter 1 grads exact match exact match (quantization absorbs it) iter 2 grads exact match 7.8e-3 <- exceeds rtol final weights 1 bf16 ulp 1 bf16 ulp ``` The engine keeps fp32 accounting while the DDP reference steps the bf16 weights directly, so the two mathematically-equivalent paths diverge by ~one bf16 ulp of weight per optimizer step even from bit-identical gradients; by iteration 2 that crosses a reduction/relu rounding boundary. Whether the difference survives bf16 quantization depends on the CPU bf16 kernel rounding — hence the version dependence. This PR gives iterations >= 2 an absolute tolerance floor (`atol=1e-2` alongside the bf16-default `rtol=1.6e-2`). The test already documents "small differences at later iterations are expected due to bfloat16 precision"; the floor absorbs quantization-boundary chaos while still catching real gradient errors. ## Validation (executed on real hardware) - 20-core x86_64 CPU, gloo, 2 ranks (`LOCAL_SIZE=2`) - torch 2.11 (the failing kernel family): 6/6 parametrizations **fail -> pass** - torch 2.13 (the passing family): 6/6 pass before and after (no regression) - pre-commit passes on the changed file Exposed by the `LOCAL_SIZE=4` multi-rank CPU run in #8381. Signed-off-by: Guokai Ma <guokai.ma@intel.com>
Author
Parents
Loading