Delegate normalization layers to NNlib functional operators (#2701)
BatchNorm, InstanceNorm, GroupNorm and LayerNorm now wrap the functional
normalization operators added in NNlib v0.9.41 (batchnorm, instancenorm,
groupnorm, normalise), dropping Flux's own _norm_layer_forward /
_track_stats! implementation. Flux.normalise becomes a thin wrapper over
NNlib.normalise, kept for backward compatibility.
cuDNN-accelerated BatchNorm now lives in NNlib (NNlibCUDACUDNNExt) and is
selected automatically for CuArrays, so FluxCUDAcuDNNExt and the cuDNN
dependency are removed. The AMDGPU MIOpen batchnorm fast path is removed
as well, falling back to NNlib's generic path pending integration into
NNlib (FluxML/NNlib.jl#752).
As a side effect, LayerNorm and Flux.normalise now add eps (not eps^2)
to the variance, matching the other normalization layers. Half-precision
(fp16/bf16) normalization parameter handling is deferred to a follow-up;
those tests are temporarily disabled.
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>