Bound total output allocation size in Tile kernel (#28070)
### Description
The per-axis SafeInt multiplication added in #27566 detects overflow
when computing an individual output dimension, but combinations of
per-axis repeats can still request an int64-representable total that is
unreasonably large. This PR adds a 4 GiB upper bound on the total tiled
byte count in the CPU, CUDA, and WebGPU Tile kernels and extends
validation/tests for the new behavior.
### Changes
- `onnxruntime/core/providers/cpu/tensor/tile.cc`: compute the output
shape with division-based checks that reject negative repeats, int64
overflow, and total tiled byte counts above the supported maximum before
allocation. The maximum is clamped to `size_t::max()` for 32-bit builds,
and the bound applies to `std::string` tensors as well because their
output buffers still allocate per-element backing storage.
- `onnxruntime/core/providers/cuda/tensor/tile.cc`: same output-size
bound applied to keep CPU and CUDA behavior consistent.
- `onnxruntime/core/providers/webgpu/tensor/tile.cc`: same output-size
bound applied to keep WebGPU behavior consistent, plus repeats
rank/length validation matching CPU/CUDA.
- `onnxruntime/test/providers/cpu/tensor/tile_op_test.cc`: tests cover
malformed repeats rank/length, 1-D, multi-axis, double (8-byte element),
and string cases that exceed the bound, plus a positive test confirming
a moderate (4 MB) output is still accepted.
### Motivation and Context
Follow-up to #27566, which fixed per-axis overflow but did not bound
total allocation size.
---------
Co-authored-by: Gopalakrishnan Nallasamy <gopalakrishnan.nallasamy@microsoft.com>
Co-authored-by: Gopalakrishnan Nallasamy <gnallasamy@microsoft.com>