[CANN] Fix deadlock during parallel execution (#29740)
### Motivation and Context
During parallel execution (`ORT_PARALLEL`), graph partitioning and
compilation may be distributed across different worker threads. However,
the GraphEngine function `aclgrphBuildFinalize` is thread-sensitive and
must be executed by the exact same thread that initialized the engine
(`aclgrphBuildInitialize`) - otherwise, this cross-thread mismatch
results in a deadlock.
### Description
- The initialization and finalization calls are now executed inside a
dedicated background thread (`g_ge_thread`), synchronized via
`std::promise` and `std::future` to ensure proper execution flow.
- Any exceptions thrown inside this background thread are safely caught
and rethrown in the parent thread to ensure proper error propagation.
### Reproduce the Issue
Below is the simplified C++ reproduction code. The `inference()`
function is the C/C++ sample taken directly from the [ONNX Runtime
documentation](https://onnxruntime.ai/docs/execution-providers/community-maintained/CANN-ExecutionProvider.html#cc),
with `ORT_PARALLEL` mode enabled.
```cpp
#include <sys/wait.h>
#include <thread>
#include <unistd.h>
#include <vector>
extern void inference();
void run(int num_proc, int num_threads) {
std::vector<pid_t> child_pids;
child_pids.reserve(num_proc);
for (int i = 0; i < num_proc; ++i) {
auto pid = fork();
if (pid == 0) {
std::vector<std::thread> threads;
threads.reserve(num_threads);
for (int j = 0; j < num_threads; ++j) {
threads.emplace_back(std::thread([] {
// The program deadlocks inside here
// during aclgrphBuildFinalize
inference();
}));
}
for (auto& thread : threads) {
thread.join();
}
exit(0);
}
else if (pid > 0) {
child_pids.push_back(pid);
}
}
for (auto pid : child_pids) {
int status;
waitpid(pid, &status, 0);
}
}
int main() {
run(10, 10); // stress test, deadlock occurs even with run(1, 1) or during a direct inference() call
return 0;
}
```
---------
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>