jaxlib/gpu: pre-allocate rocBLAS workspace for rocsolver potrf on ROCm
Without a pre-allocated workspace, rocBLAS internally calls hipMalloc on
each potrf invocation. On ROCm, hipMalloc is synchronous and can add
1-5ms of overhead per call, which explains the ~1.7x gap between JAX
Cholesky and rocsolver-bench (which pre-allocates its own workspace).
Use the rocBLAS two-phase workspace query API
(rocblas_start_device_memory_size_query / rocblas_stop_device_memory_size_query)
to determine the exact workspace size needed for rocsolver_spotrf, then
allocate from the XLA scratch allocator and set it via rocblas_set_workspace
before calling rocsolver_spotrf. This follows the same pattern as GesddImpl.
After the kernel completes, clear the workspace pointer with
rocblas_set_workspace(handle, nullptr, 0) to revert to the default memory
model for other operations sharing the handle.