[mlir][xegpu] Support N-D block transfers in VectorToXeGPU (#210527)
Extend the vector.transfer_read/transfer_write lowerings so they can
produce N-D xegpu.load_nd/store_nd, not just 1D/2D, and relax the
out-of-bounds handling to match load_nd's implicit-zero padding.
Restructure both patterns as "block first, then scatter as fallback.
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>