Loosen the late-iteration gradient tolerance in the checkpointing test (#8648)
## Description
`TestZeroUserBackwardWithCheckpointing::test_checkpointed_multiple_backward`
compares DDP and engine gradients over three optimizer iterations at
dtype-default tolerances, and flips between pass and fail across torch
versions on cpu (fails on 2.10/2.11 kernels, passes on 2.13) without any
behavioral change.
A minimal two-rank repro isolates the mechanism:
```
torch 2.13 torch 2.11 (= CI family)
iter 0 grads exact match exact match
post-step — linear2.weight differs by 1.7e-6 <- the seed
iter 1 grads exact match exact match (quantization absorbs it)
iter 2 grads exact match 7.8e-3 <- exceeds rtol
final weights 1 bf16 ulp 1 bf16 ulp
```
The engine keeps fp32 accounting while the DDP reference steps the bf16
weights directly, so the two mathematically-equivalent paths diverge by
~one bf16 ulp of weight per optimizer step even from bit-identical
gradients; by iteration 2 that crosses a reduction/relu rounding
boundary. Whether the difference survives bf16 quantization depends on
the CPU bf16 kernel rounding — hence the version dependence.
This PR gives iterations >= 2 an absolute tolerance floor (`atol=1e-2`
alongside the bf16-default `rtol=1.6e-2`). The test already documents
"small differences at later iterations are expected due to bfloat16
precision"; the floor absorbs quantization-boundary chaos while still
catching real gradient errors.
## Validation (executed on real hardware)
- 20-core x86_64 CPU, gloo, 2 ranks (`LOCAL_SIZE=2`)
- torch 2.11 (the failing kernel family): 6/6 parametrizations **fail ->
pass**
- torch 2.13 (the passing family): 6/6 pass before and after (no
regression)
- pre-commit passes on the changed file
Exposed by the `LOCAL_SIZE=4` multi-rank CPU run in #8381.
Signed-off-by: Guokai Ma <guokai.ma@intel.com>