[CUDA] Follow up INT2 grouped GEMM test coverage and plugin build (PR4a2) (#33045)
## Summary
PR4a2 follows up the five non-blocking comments from the final review of
#32963 at `d3d60acf26`. Based on main after #32963 was squash-merged as
`a4d55e2b6f`. No production kernel math, alignment contract, or runtime
eligibility changes.
## Review Follow-ups
- **Coverage**
([comment](https://github.com/microsoft/onnxruntime/pull/32963#discussion_r4158213954)):
allow compute capability major 8 in the direct-kernel fixture instead of
exactly 8.0. Test parameters continue to select the SM80 implementation.
Legacy W4/W8 alignment regressions remain enabled, removing their
fixture-level skip on A10/SM86.
- **Plugin build footprint**
([comment](https://github.com/microsoft/onnxruntime/pull/32963#discussion_r4158213967)):
restrict dense/W4/W8 baselines and disabled benchmarks to legacy tests;
add source-specific plugin/adapter definitions. Reduce plugin
internal-test contrib implementation sources from 16 to 4, retaining the
adaptor needed for INT2 unpacking. Plugin retains INT2 FP16/BF16
numerical, metadata, valid-offset and misalignment coverage. No measured
CI wall-time improvement is claimed.
- **Diagnostic options**
([comment](https://github.com/microsoft/onnxruntime/pull/32963#discussion_r4158213972)):
remove unconditional duplicate suppressions. CUDA 13+ inherits 970/2189
from shared options. CUDA <13 retains target-local suppressions with an
explanatory comment because the internal INT2 TU exposes third-party
Abseil/Protobuf header diagnostics under warning-as-error. Removing them
reproduced a CUDA 12.8 compile failure; the explicit version guard fixes
it.
- **Rebase**
([comment](https://github.com/microsoft/onnxruntime/pull/32963#discussion_r4158213979)):
#32932 is already in the main base. This PR does not replay that fix and
has no `moe_quantization.cc` diff.
- **Headers**
([comment](https://github.com/microsoft/onnxruntime/pull/32963#discussion_r4158214003)):
add the requested Microsoft copyright/MIT license headers to the helper
header, implementation CU and test TU.
## Validation
All compilation, formatting, CMake checks and GPU tests ran inside
Docker `jiafa-dev` on an A100-SXM4-80GB.
| Check | Result |
| --- | --- |
| CUDA 12.8 plugin standalone harness | 16 passed, zero skipped |
| CUDA 12.8 legacy standalone harness | 23 enabled passed, zero skipped;
3 benchmarks disabled |
| CUDA 13.0 plugin standalone harness | 16 passed, zero skipped |
| Fresh legacy/plugin test TU compilation with CUDA 12.8 and 13.0
headers | Passed |
| Fresh four plugin implementation CUs with CUDA 12.8 and 13.0,
warning-as-error | Passed |
| Fresh legacy INT2 implementation CU, CUDA 12.8 | Passed |
| Scoped execution of actual CMake source/definition and option blocks |
Passed; four contrib sources enabled / zero disabled, exactly one
suppression per diagnostic per toolkit |
| Symbol audit | Plugin baseline dependencies/benchmarks absent; legacy
coverage retained |
| Production body comparison | Unchanged apart from requested headers |
| clang-format dry run and git diff --check | Passed |
### Limitations
- Focused standalone validation, not a full top-level CMake
configure/build or provider integration run. Existing Ninja recipes were
adapted to compile current sources and new options into fresh objects;
actual changed CMake blocks were extracted and executed with capture
functions. Full target integration remains for CI.
- Plugin harnesses freshly compile the test and all four contrib
implementation CUs, reusing unchanged core/static libraries and CUDA
support objects. Legacy freshly compiles the test and INT2 CU, reusing
unchanged baseline kernel objects and standalone plugin support shims.
- Only SM80 hardware was available. SM86/SM89 execution, Windows, full
regression and sanitizer runs were not exercised for this follow-up.
CUDA 13 legacy test compilation passed; legacy runtime execution used
CUDA 12.8.
- Exact local commands and compiler/test logs are recorded under
`build/pr4a2-review-validation/` on the validation machine; these
ignored scratch artifacts are not committed.