fix(quantization): avoid intermediate overflow when calculating quantization ranges (#33104)
### Description
Promote both range endpoints to float64 before subtracting them in
`compute_scale_zp`. The scale is still returned in the original
floating-point dtype, and the zero point keeps the quantized dtype.
Add coverage for FP16 and FP32 ranges with signed and unsigned 8-bit and
16-bit quantization, including symmetric ranges. Add QDQ MatMul tests
that quantize large finite weights and run CPU inference with per-tensor
and per-channel weight quantization.
The FP16 inference cases first check that the current CPU build can
initialize the original FP16 MatMul. They skip only when that kernel is
unavailable; the FP32 inference cases and dtype-level regression tests
remain enabled.
### Motivation and Context
The current subtraction runs in the input dtype before its result is
converted to float64. For example, FP16 endpoints of -40000 and 40000
are both finite, but their difference overflows FP16. This produces an
infinite scale even though dividing the range by the quantized range
would produce a representable scale. FP32 inputs can encounter the same
problem.
Doing the subtraction in float64 also changes rounding for some finite,
non-overflowing ranges. This change therefore preserves the output
dtypes, but does not promise identical scale or zero-point values for
every previously finite input.
Validation on macOS arm64, Python 3.12.3, NumPy 2.5.3:
- All 20 new regression subtests fail against the original function and
pass with the change.
- The complete `test_quant_util.py` and `test_op_matmul.py` modules
report 15 passed, 8 skipped, and 35 subtests passed. These totals
include the new regressions. The skips cover missing SciPy, existing
FP16 calibration failures, and unsupported FP16 QLinearMatMul cases; the
four new QDQ inference cases run successfully.
- Repository-pinned `lintrunner` and `git diff --check` pass.
After adding the CPU kernel check, both modules pass again with the same
counts on this arm64 machine. A separate injected-error probe confirms
that a missing FP16 MatMul skips only its two subcases while both FP32
cases still run, and an unrelated missing-operator error is not skipped.
This probe does not replace a native x86/x64 test run.
The tests import the checkout's Python quantization code and use the
ONNX Runtime 1.30.0 CPU wheel for native execution. This validation does
not include a native build from the current checkout, GPU execution, or
the full repository test suite.