Enable fpA_intB GEMM in CUDA builds and add configurable options #29622
[CUDA] In-memory tactic autotuning for fpA_intB MatMulNBits
e4e013d6
Enable onnxruntime_USE_FPA_INTB_GEMM by default when CUDA is enabled
bc32e55a
Add ep.cuda.fpa_intb_gemm / fpa_intb_profile_m session-config keys
8a2ee181
Test ep.cuda.fpa_intb_gemm / fpa_intb_profile_m session-config keys
faf42962
Simplify fpA_intB enable option to boolean; prepacked forces the path
c1cb3320
Fix fpA_intB MatMulNBits CI failures and address review feedback
0ae5e83f
Clarify fpA_intB config-toggle test comment (numeric-equivalence scope)
8ad2e0a3
refactoring
5a8a30df
fix(cuda): use ORT env helper instead of std::getenv in cuda_runtime_…
d72d7a6b
fix Windows CI
eab48120
fix(cuda): use _dupenv_s on Windows to avoid C4996 in cuda_runtime_utils
abc2588c
fix build
6f40ffa2
fix lint warnings
774906f2
address feedbacks
fcf29207
remove stale test
ab7df6ec
tianleiwu
force pushed
from
ad6c9168
to
ab7df6ec
69 days ago
Assignees
No one assigned
Login to write a write a comment.
Login via GitHub