onnxruntime
b40377d9 - fix(quantization): avoid intermediate overflow when calculating quantization ranges (#33104)

Commit
3 days ago
fix(quantization): avoid intermediate overflow when calculating quantization ranges (#33104) ### Description Promote both range endpoints to float64 before subtracting them in `compute_scale_zp`. The scale is still returned in the original floating-point dtype, and the zero point keeps the quantized dtype. Add coverage for FP16 and FP32 ranges with signed and unsigned 8-bit and 16-bit quantization, including symmetric ranges. Add QDQ MatMul tests that quantize large finite weights and run CPU inference with per-tensor and per-channel weight quantization. The FP16 inference cases first check that the current CPU build can initialize the original FP16 MatMul. They skip only when that kernel is unavailable; the FP32 inference cases and dtype-level regression tests remain enabled. ### Motivation and Context The current subtraction runs in the input dtype before its result is converted to float64. For example, FP16 endpoints of -40000 and 40000 are both finite, but their difference overflows FP16. This produces an infinite scale even though dividing the range by the quantized range would produce a representable scale. FP32 inputs can encounter the same problem. Doing the subtraction in float64 also changes rounding for some finite, non-overflowing ranges. This change therefore preserves the output dtypes, but does not promise identical scale or zero-point values for every previously finite input. Validation on macOS arm64, Python 3.12.3, NumPy 2.5.3: - All 20 new regression subtests fail against the original function and pass with the change. - The complete `test_quant_util.py` and `test_op_matmul.py` modules report 15 passed, 8 skipped, and 35 subtests passed. These totals include the new regressions. The skips cover missing SciPy, existing FP16 calibration failures, and unsupported FP16 QLinearMatMul cases; the four new QDQ inference cases run successfully. - Repository-pinned `lintrunner` and `git diff --check` pass. After adding the CPU kernel check, both modules pass again with the same counts on this arm64 machine. A separate injected-error probe confirms that a missing FP16 MatMul skips only its two subcases while both FP32 cases still run, and an unrelated missing-operator error is not skipped. This probe does not replace a native x86/x64 test run. The tests import the checkout's Python quantization code and use the ONNX Runtime 1.30.0 CPU wheel for native execution. This validation does not include a native build from the current checkout, GPU execution, or the full repository test suite.
Author
Parents
Loading